Data creation system, learning system, estimation system, processing device, evaluation system, data creation method and program

By superimposing the offset amplitude in the image generation system to generate the second image data, the problem of reduced recognition performance of the learning model in the prior art is solved, and higher recognition accuracy and reliability are achieved.

CN116438569BActive Publication Date: 2025-09-09PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180072094.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-10
Filing Date
2021-11-05
Publication Date
2025-09-09
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

In the prior art, when an image generation system generates learning data, the recognition performance of the learned model decreases. In particular, in an X-ray image object recognition system, the replacement of pixel values ​​in a predetermined area leads to information loss, affecting recognition accuracy.

Method used

By obtaining the pixel values ​​of the feature area and superimposing the offset amplitude in a predetermined area of ​​the first image, the second image data is generated, ensuring that the predetermined area has a peripheral shape corresponding to the feature area, and expanding the learning data to generate an image that is closer to the real world.

Benefits of technology

The possibility of degradation of the recognition performance of the learned model is reduced, and the recognition accuracy and reliability of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116438569B_ABST
    Figure CN116438569B_ABST
Patent Text Reader

Abstract

The problem to be overcome by the present invention is to reduce the possibility of causing a decrease in the recognition performance of a learned model. A data creation system (1) creates second image data (D12) for use as learning data for generating a learned model (M1) based on first image data (D11). The data creation system (1) includes an acquisition unit (102) and a superposition unit (105). The acquisition unit (102) acquires feature image data indicating the individual pixel values ​​of a plurality of pixels included in a feature area. The superposition unit (105) creates the second image data (D12) by superimposing an offset amplitude on the individual pixel values ​​of a plurality of pixels included in a predetermined area of ​​a first image represented by the first image data (D11). The predetermined area has an outer peripheral shape corresponding to the feature area. The offset amplitude is determined based on the individual pixel values ​​of the plurality of pixels included in the feature area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to data creation systems, learning systems, estimation systems, processing devices, evaluation systems, data creation methods, and programs. More particularly, the present invention relates to a data creation system, data creation method, and program for creating image data for use as learning data for generating a learned model. The present invention also relates to a processing device for use in the data creation system, an evaluation system including the processing device, a learning system for generating a learned model, and an estimation system using the learned model. Background Art

[0002] Patent Document 1 discloses an image generation system. The image generation system includes an acquisition unit, a calculation unit, a transformation unit, and an image generation unit. The acquisition unit acquires an image in a first region included in a first image and an image in a second region included in a second image. The calculation unit calculates transformation parameters for transforming the image in the first region so that the color information of the image in the first region is similar to the color information of the image in the second region. The transformation unit transforms the first image using the transformation parameters.

[0003] The image generation unit generates a third image by combining the transformed first image with the second image. Specifically, assuming that the width and height of the first image are (width, height) and the specific coordinates of the second image are (x, y), the image generation unit superimposes the first image on the second image so that the specific coordinates (x, y) of the second image are located at the upper left corner of the first image, and replaces the pixel values ​​of the second image with the pixel values ​​of the first image within the range from (x, y) to (x + width, y + height).

[0004] For example, the X-ray image object recognition system disclosed in Patent Document 1 replaces the pixel values ​​of pixels in a predetermined region of the second image, ranging from (x, y) to (x + width, y + height), with those of the first image. Consequently, information related to the pixel values ​​of pixels in the predetermined region of the second image disappears from the third image, potentially generating an unrealistic image. Consequently, using the third image as learning data to generate a learned model can lead to a decrease in the learned model's recognition performance during the inference phase.

[0005] Prior art literature

[0006] Patent Literature

[0007] Patent Document 1: Japanese Patent Application Laid-Open No. 2017-45441 Summary of the Invention

[0008] In view of the foregoing background, an object of the present invention is to provide a data creation system, a learning system, an estimation system, a processing device, an evaluation system, a data creation method and a program, all of which are configured or designed to reduce the possibility of causing a degradation in the recognition performance of the learned model.

[0009] According to one aspect of the present invention, a data creation system creates second image data for use as learning data for generating a learned model based on first image data. The data creation system includes an acquisition unit and an overlay unit. The acquisition unit acquires feature image data indicating the individual pixel values ​​of a plurality of pixels included in a feature area. The overlay unit creates the second image data by overlaying an offset amplitude on the individual pixel values ​​of a plurality of pixels included in a predetermined area of ​​a first image represented by the first image data. The predetermined area has an outer peripheral shape corresponding to the feature area. The offset amplitude is determined based on the individual pixel values ​​of the plurality of pixels included in the feature area.

[0010] A learning system according to another aspect of the present invention generates a learned model using a learning dataset including learning data as second image data created by the above-described data creation system.

[0011] An estimation system according to yet another aspect of the present invention uses the learned model generated by the above-mentioned learning system to perform estimation related to an object to be recognized.

[0012] A processing device according to another aspect of the present invention is used as the first processing device in a data creation system including a first processing device and a second processing device. The first processing device includes an extraction unit that extracts extracted image data including pixel values ​​of a plurality of pixels included in a predetermined extraction area from third image data used as learning data. The second processing device includes the acquisition unit and the superposition unit.

[0013] A processing device according to another aspect of the present invention is used as the second processing device in a data creation system including a first processing device and a second processing device. The second processing device includes an extraction unit that extracts extracted image data including pixel values ​​of a plurality of pixels included in a predetermined extraction area from third image data used as learning data. The second processing device includes the acquisition unit and the superposition unit.

[0014] According to another aspect of the present invention, an evaluation system includes a processing device and a learning system. The processing device extracts extracted image data including individual pixel values ​​of a plurality of pixels included in a predetermined extraction area from third image data representing a third image including a pixel area indicating an object to be identified. The processing device outputs the extracted image data thus extracted. The learning system generates a learned model. The learned model outputs an estimation result similar to a case where the third image data is the object to be identified in response to a second image represented by second image data or a predetermined area in the second image. The predetermined area is an area included in the first image and having an outer peripheral shape corresponding to the extraction area. The first image includes a pixel area indicating the object to be identified and is represented by first image data. The second image is created by superimposing an offset amplitude on the individual pixel values ​​of the plurality of pixels included in the predetermined area of ​​the first image. The offset amplitude is determined based on the individual pixel values ​​of the plurality of pixels included in the extraction area.

[0015] Another processing device according to yet another aspect of the present invention is used as the processing device of the above-mentioned evaluation system.

[0016] Another learning system according to still another aspect of the present invention is used as a learning system of the above-mentioned evaluation system.

[0017] According to another aspect of the present invention, an evaluation system includes a processing device and an estimation system. The processing device extracts extracted image data including individual pixel values ​​of a plurality of pixels included in a predetermined extraction area from third image data representing a third image including a pixel area indicating an object to be identified. The processing device outputs the extracted image data thus extracted. The estimation system uses a learned model to perform an estimation related to the object to be identified. The learned model outputs an estimation result similar to a case where the third image data is the object to be identified, in response to a second image represented by second image data or a predetermined area in the second image. The predetermined area is an area included in the first image and having an outer peripheral shape corresponding to the extraction area. The first image includes a pixel area indicating the object to be identified and is represented by first image data. The second image is created by superimposing an offset amplitude on the individual pixel values ​​of the plurality of pixels included in the predetermined area of ​​the first image. The offset amplitude is determined based on the individual pixel values ​​of the plurality of pixels included in the extraction area.

[0018] Another processing device according to yet another aspect of the present invention is used as the processing device of the above-mentioned evaluation system.

[0019] Another estimation system according to still another aspect of the present invention is used as the estimation system of the above-mentioned evaluation system.

[0020] According to another aspect of the present invention, a data creation method is a data creation method for creating second image data for use as learning data for generating a learned model based on first image data. The data creation method includes an acquisition step and a superposition step. The acquisition step includes acquiring feature image data indicating the respective pixel values ​​of a plurality of pixels included in a feature area. The superposition step includes creating the second image data by superimposing an offset amplitude on the respective pixel values ​​of a plurality of pixels included in a predetermined area of ​​a first image represented by the first image data. The predetermined area has an outer peripheral shape corresponding to the feature area. The offset amplitude is determined based on the respective pixel values ​​of the plurality of pixels included in the feature area.

[0021] A program according to yet another aspect of the present invention is designed to cause one or more processors to perform the above-mentioned data creation method. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a block diagram illustrating a schematic structure of an overall evaluation system including a data creation system according to a first embodiment;

[0023] Figure 2 illustrating an exemplary first image represented by first image data input to the data creation system;

[0024] Figure 3 illustrating an exemplary defective product as an object to be identified by the learned model in the data creation system;

[0025] Figure 4 illustrating another exemplary defective product as another object to be identified by the learned model in the data creation system;

[0026] Figure 5 illustrating another exemplary defective product as another object to be identified by the learned model in the data creation system;

[0027] Figure 6 showing an exemplary third image represented by third image data to be input to the data creation system;

[0028] Figure 7 illustrating an exemplary image represented by the feature image data acquired by the data creation system;

[0029] Figure 8 showing an exemplary first image represented by first image data input to the data creation system;

[0030] Figure 9 showing an exemplary second image represented by second image data created by the data creation system;

[0031] Figure 10 An exemplary image represented by the feature image data acquired by the data creation system is shown;

[0032] Figure 11 illustrating a cross section of an object represented by feature image data acquired by the data creation system;

[0033] Figure 12 Illustrate how the data creation system can determine the offset magnitude;

[0034] Figure 13 Illustrate how the data creation system can determine the offset magnitude;

[0035] Figure 14 A to C illustrate how the data creation system can superimpose offset amplitudes;

[0036] Figure 15 Illustrate how the data creation system can superimpose offset amplitudes;

[0037] Figure 16 Illustrate how the data creation system can superimpose offset amplitudes;

[0038] Figure 17 is a flowchart showing the process of operation of the data creation system;

[0039] Figure 18 Illustrate how the data creation system of the comparative example superimposes the offset amplitude;

[0040] Figure 19 is a block diagram illustrating a schematic structure of an overall evaluation system including a data creation system according to a second embodiment;

[0041] Figure 20 Illustrate how the data creation system may define a maintenance area;

[0042] Figure 21 Illustrate how the data creation system can superimpose offset amplitudes;

[0043] Figure 22 A and B illustrate how the data creation system can determine the offset magnitude;

[0044] Figure 23 Illustrate how the data creation system can superimpose offset amplitudes;

[0045] Figure 24 is a block diagram illustrating a schematic configuration of an overall evaluation system including a data creation system according to a first modification;

[0046] Figure 25Illustrate how the data creation system may define a maintenance area;

[0047] Figure 26 shows an exemplary image represented by feature image data acquired by the data creation system according to the second modification; and

[0048] Figure 27 is a block diagram illustrating a schematic configuration of a data creation system according to a third modification. DETAILED DESCRIPTION

[0049] The drawings referred to in the following description of the embodiments are all schematic representations. Therefore, the ratios of the sizes (including thicknesses) of the various components shown in the drawings do not always reflect their actual size ratios.

[0050] (1) First embodiment

[0051] (1.1) Overview

[0052] like Figure 1 As shown, the data creation system 1 according to the exemplary embodiment creates the second image data D12 based on the first image data D11. The first image data D11 is a diagram showing the image (the first image Im11, reference Figure 8 ) is data showing information related to another image (second image Im12, reference Figure 9 The information related to each image includes coordinates (X coordinates and Y coordinates) of a plurality of pixels forming the image and pixel values ​​corresponding to the respective coordinates.

[0053] The second image data D12 is used as learning data for generating the learned model M1. In other words, the second image data D12 is learning data for generating a model through machine learning. As used herein, "model" refers to a program designed to estimate the condition of an object to be identified in response to the input of data related to the object to be identified (object), and output an estimation result (recognition result). In addition, as used herein, "learned model" refers to a model related to the completion of machine learning using learning data. In addition, "learning data" refers to a data set that combines input information (image data D1) to be input for the model and labels attached to the input information, i.e., the so-called "training data". That is, in this embodiment, the learned model M1 is a model for which machine learning is performed through supervised learning.

[0054] In this embodiment, if Figure 2As shown, the object to be identified may be, for example, a weld bead B10. The weld bead B10 is formed at a boundary B14 (welding point) between the first base metal B11 and the second base metal B12 when two or more welding base materials (for example, a first base metal B11 and a second base metal B12 in this example) are welded together via a welding material B13. Figure 2 , the first base metal B11 and the second base metal B12 are arranged along the Y-axis (vertically), and the weld bead B10 is formed to be elongated along the X-axis (transversely). The size and shape of the weld bead B10 mainly depend on the welding material B13. Therefore, when the image data D3 of the object to be identified covering the weld bead B10 is input, the learned model M1 estimates the condition of the weld bead B10 and outputs the estimation result. Specifically, the learned model M1 outputs information indicating whether the weld bead B10 is a defective product or a non-defective (i.e., good) product, as well as information related to the type of defect in the case where the weld bead B10 is a defective product, as an estimation result. That is, the learned model M1 is used to determine whether the weld bead B10 is a good product. In other words, the learned model M1 is used to perform a welding appearance test to determine whether the welding is performed correctly.

[0055] The decision on whether the weld bead B10 is good or defective can be made, for example, based on whether the length of the weld bead B10, the height of the weld bead B10, the elevation angle of the weld bead B10, the weld depth of the weld bead B10, the excess metal of the weld bead B10, and the misalignment of the weld point of the weld bead B10 (including the degree of offset of the starting end of the weld bead B10) fall within their respective tolerance ranges. For example, if at least one of these parameters listed above does not fall within its tolerance range, the weld bead B10 is determined to be a defective product. Alternatively, the decision on whether the weld bead B10 is good or defective can also be made, for example, based on whether the weld bead B10 has any undercut B2 (refer to FIG. 2 ). Figure 3 ), whether the weld bead B10 has any pits B3 (reference Figure 4 ), whether the weld bead B10 has any spatter B4 (reference Figure 5 ), and whether weld bead B10 has any protrusions. For example, if at least one of the defects listed above occurs, weld bead B10 is determined to be a defective product. In the following description, such defects are sometimes referred to as "defects."

[0056] In order to perform model-related machine learning, it is necessary to collect a large amount of image data related to the object to be identified, including defective products, as learning data. However, in the case where the frequency of occurrence of the object to be identified being proven to be defective is low, the learning data required to generate a learned model M1 with high identifiability is often insufficient. Therefore, in order to overcome this problem, by performing data enhancement processing on the learning data (original learning data) obtained by actually photographing the weld bead B10 using a camera device, model-related machine learning can be performed when the amount of learning data increases. As used herein, data enhancement processing refers to processing for expanding learning data by performing various types of processing on the learning data, such as translation, enlargement or reduction (expansion or contraction), rotation, inversion, and noise or defect application.

[0057] The data creation system 1 according to this embodiment creates second image data D12 by, for example, superimposing (applying) offset amplitude data indicating a defect onto first image data D11, which is original learning data. In this way, the learning data (i.e., image data of the object to be identified, including defective products) is expanded.

[0058] like Figure 1 As shown, the data creation system 1 includes an acquisition unit 102 and a superimposition unit 105 .

[0059] The acquisition unit 102 acquires the indication feature region R0 (reference Figure 7 ). The feature image data includes the coordinates (X coordinates and Y coordinates) of each pixel included in the image Im20 representing the predetermined feature region R0, and the pixel value (data value) of each pixel. The feature region R0 may be, for example, a region having a defect. As will be described later, the feature image data can be extracted, for example, from original learning data (third image data D13) representing an image (including a defect) of an object to be identified as data indicating a region having a defect.

[0060] The superimposition section 105 superimposes the offset amplitude on a predetermined region R10 (reference region R11) of the first image Im11 represented by the first image data D11. Figure 8 ) is created based on the respective pixel values ​​of the plurality of pixels included in the predetermined region R10. The predetermined region R10 is a region having an outer peripheral shape corresponding to the feature region R0. The offset amplitude is determined based on the respective pixel values ​​of the plurality of pixels included in the feature region R0. The offset amplitude can be determined, for example, for each of the pixels forming the feature region R0. The offset amplitude can be determined, for example, by the determination unit 104. How the determination unit 104 determines the offset amplitude will be described later. In addition, as used herein, "superposition" means adding the offset amplitude to the pixel value of the pixel included in the first image Im11.

[0061] As can be seen, according to this embodiment, the second image Im12 is generated by superimposing the offset amplitude on the individual pixel values ​​of a plurality of pixels included in the predetermined region R10 of the first image Im11. Therefore, this embodiment enables the pixel values ​​of the pixels of the first image Im11 (i.e., the image represented by the first image data D11 as the original learning data) to be reflected in the pixel values ​​of the individual pixels in the region included in the second image Im12 and corresponding to the predetermined region R10. This makes it possible to generate learning data that represents an image that is closer to an image that can exist in the real world, thereby helping to reduce the possibility of causing a decline in the performance of the learned model M1 generated based on the learning data.

[0062] Furthermore, according to the learning system 2 of this embodiment (see Figure 1 ) A learned model M1 is generated using a learning data set that includes learning data of the second image data D12 created by the data creation system 1. This helps to reduce the possibility of causing a decline in the performance of identifying the learned model M1. The learning data used to generate the learned model M1 may include not only the second image data D12 (extended data) but also the original first image data D11. In other words, the image data D1 according to the present embodiment includes at least the second image data D12, and may include both the first image data D11 and the second image data D12. In addition, the learning data used to generate the learned model M1 may include original learning data (third image data D13) representing an image (including defects) of an object to be identified.

[0063] According to the estimation system 3 of this embodiment (refer to Figure 1 ) uses the learned model M1 generated by the learning system 2 to make an estimation related to the object to be recognized (for example, the welding bead B10). This helps reduce the possibility of causing a decrease in the performance of the learned model M1.

[0064] The data creation method according to the present embodiment is designed to create second image data D12 for use as learning data for generating a learned model M1 based on first image data D11. The data creation method includes an acquisition step and a superposition step. The acquisition step includes acquiring feature image data indicating the individual pixel values ​​of a plurality of pixels included in a feature region R0. The superposition step includes creating second image data D12 by superimposing an offset amplitude on the individual pixel values ​​of a plurality of pixels included in a predetermined region R10 of a first image Im11 represented by the first image data D11. The predetermined region R10 has an outer peripheral shape corresponding to the feature region R0. The offset amplitude is determined based on the individual pixel values ​​of the plurality of pixels included in the feature region R0. This helps to reduce the possibility of causing a decrease in the performance of identifying the learned model M1. The data creation method is used on a computer system (data creation system 1). That is, the data creation method can also be implemented as a program. The program according to the present embodiment is designed to cause one or more processors to perform the data creation method according to the present embodiment. The program can be distributed after being recorded in some non-transitory storage medium.

[0065] (1.2) Details

[0066] Next, an overall system (hereinafter referred to as “evaluation system 100 ”) including the data creation system 1 according to the present embodiment will be described in detail with reference to the drawings.

[0067] (1.2.1) Overall structure

[0068] like Figure 1 As shown, the evaluation system 100 includes a data creation system 1, a learning system 2, an estimation system 3, and one or more camera devices 6 (in Figure 1 Only one camera device 6 is shown in FIG.

[0069] It is assumed that the data creation system 1, learning system 2, and estimation system 3 are implemented as servers, for example. Assume that the server used herein is implemented as a single server device. That is, it is assumed that the main functions of the data creation system 1, learning system 2, and estimation system 3 are provided for a single server device.

[0070] Alternatively, the server may be implemented as multiple server devices. Specifically, the functions of data creation system 1, learning system 2, and estimation system 3 may be provided by three different server devices. Alternatively, two of these three systems may be provided by a single server device. Alternatively, for example, these server devices may form a cloud computing system.

[0071] Furthermore, the server device may be installed inside the factory where welding is performed or outside the factory (e.g., at a service headquarters), whichever is appropriate. If the respective functions of the data creation system 1, the learning system 2, and the estimation system 3 are provided for three different server devices, each of these server devices is preferably connected to the other server devices to be ready for communication with the other server devices.

[0072] The data creation system 1 is configured to create image data D1 for use as learning data for generating the learned model M1. As used herein, "creating learning data" may refer not only to generating new learning data separately from the original learning data, but also to generating new learning data by updating the original learning data.

[0073] As used herein, the learned model M1 may, for example, include a model using a neural network, or a model generated by deep learning using a multi-layer neural network. Examples of neural networks may include convolutional neural networks (CNNs) and Bayesian neural networks (BNNs). The learned model M1 may, for example, be implemented by installing the learned neural network into an integrated circuit such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA). However, the learned model M1 does not have to be a model generated by deep learning. Alternatively, for example, the learned model M1 may also be a model generated by a support vector machine or a decision tree.

[0074] In this embodiment, as described above, the data creation system 1 has a function for expanding learning data by performing data enhancement processing on the original learning data (first image data D11). In the following description, a person using the evaluation system 100 including the data creation system 1 will be referred to as a "user." For example, the user may be an operator monitoring a manufacturing process such as a welding process in a factory, or a manager in charge.

[0075] like Figure 1 As shown, the data creation system 1 includes a processor 10 , a communication interface 11 , a display device 12 , and an operating member 13 .

[0076] exist Figure 1 In the example shown, a storage device 14 for storing learning data (image data D1) is provided outside the data creation system 1. However, this is merely an example and should not be construed as limiting. Alternatively, the data creation system 1 may further include a storage device 14. In this case, the storage device 14 may also be a memory built into the processor 10. The storage device 14 for storing the image data D1 includes a programmable nonvolatile memory such as an electrically erasable programmable read-only memory (EEPROM).

[0077] Alternatively, some functions of the data creation system 1 may be distributed to a telecommunications device capable of communicating with a server. Examples of telecommunications devices, as used herein, include personal computers (including laptops and desktop computers) and mobile devices such as smartphones and tablet computers. In this embodiment, the functions of the display device 12 and operating member 13 are provided for the telecommunications device to be used by the user. A dedicated application software program is pre-installed in the telecommunications device that enables the telecommunications device to communicate with the server.

[0078] The processor 10 can be implemented as a computer system including one or more processors (microprocessors) and one or more memories. That is, the one or more processors can perform the functions of the processor 10 by executing one or more programs (applications) stored in one or more memories. In this embodiment, the program is pre-stored in the memory of the processor 10. Alternatively, the program can be downloaded via a telecommunication line such as the Internet, or distributed after being stored in a non-transitory storage medium such as a memory card.

[0079] Processor 10 performs processing for controlling communication interface 11, display device 12, and operating member 13. It is assumed that the functions of processor 10 are performed by a server. Processor 10 also performs image processing. In particular, processor 10 performs data enhancement processing to create second image data D12 based on first image data D11. Processor 10 will be described in detail in the next section.

[0080] The display device 12 can be implemented as a liquid crystal display or an organic electroluminescent (EL) display. As described above, the display device 12 is provided for the telecommunications device. Alternatively, the display device 12 can also be a touch screen panel display. The display device 12 displays (outputs) information related to the first image data D11, the second image data D12, and the third image data D13. In addition to the first image data D11, the second image data D12, and the third image data D13, the display device 12 also displays various types of information related to the generation of learning data.

[0081] Examples of the operating member 13 include a mouse, a keyboard, and a pointing device. As described above, the operating member 13 may be provided for the telecommunications device to be used by the user. If the display device 12 is a touch screen panel display of the telecommunications device, the display device 12 may also have the function of the operating member 13.

[0082] The communication interface 11 is a communication interface for communicating with one or more camera devices 6 directly or indirectly, for example, via another server having the function of a production management system. In the present embodiment, it is assumed that the function of the communication interface 11 and the function of the processor 10 are provided for the same server. However, this is only an example and should not be interpreted as limiting. Alternatively, for example, the function of the communication interface 11 may also be provided for a telecommunications device. The communication interface 11 receives first image data D11 as original learning data from the camera device 6. In addition, the communication interface 11 also receives third image data D13 as original learning data representing an image of an object to be identified (including defects) from the camera device 6.

[0083] In this embodiment, the imaging device 6 is a distance image sensor for measuring the distance to an object. The imaging device 6 measures the distance to an object using, for example, a time-of-flight (TOF) method. Therefore, the image data captured by the imaging device 6 is distance image data in which each of the plurality of pixels in the image is assigned a value indicating the distance from the imaging device 6 to the object. In short, each of the first image data D11 and the third image data D13 is distance image data in which the pixel value of each pixel is a distance value.

[0084] like Figure 2 As shown, the first image data D11 is data representing a first image Im11 encompassing an object to be identified. As described above, the object to be identified may be, for example, a weld bead B10 formed at a boundary B14 between first base metal B11 and second base metal B12 when the first base metal B11 and the second base metal B12 are welded together via welding material B13.

[0085] For example, the first image data D11 is selected as a target of data enhancement processing from a large amount of image data related to the object to be recognized captured by the camera 6 according to a user's command. The evaluation system 100 preferably includes a user interface (which may be an operating member 13) that accepts a user's command related to his or her selection.

[0086] The learning system 2 generates a learned model M1 using a learning data set including a plurality of image data D1 (including a plurality of second image data D12) created by the data creation system 1. The learning data set is generated by attaching a label indicating a good product or a defective product or a label indicating the type and location of a defect regarding a defective product to each image data in the plurality of image data D1. Examples of types of defects include undercuts, pits, and spatters. The operation of attaching labels is performed by a user on the evaluation system 100 via a user interface such as an operating member 13. In one variation, the operation of attaching labels can also be performed by a learned model having a function for attaching labels to the image data D1. The learning system 2 generates a learned model M1 by performing machine learning related to the condition (including good condition, bad condition, type of defect, and location of defect) of the object to be identified (e.g., weld bead B10) using the learning data set.

[0087] Alternatively, the learning system 2 may attempt to improve the performance of the learned model M1 by relearning using a learning dataset that includes newly acquired learning data. For example, if a new type of defect is discovered in the object to be identified (e.g., weld bead B10), the learning system 2 may be relearned with respect to the new type of defect.

[0088] The estimation system 3 estimates the condition of the object to be identified (including good condition, bad condition, type of defect, and location of defect) using the learned model M1 generated by the learning system 2. The estimation system 3 is configured to be ready to communicate with one or more camera devices 6 directly or indirectly, for example, via another server having the function of a production management system. The estimation system 3 receives image data D3 of the object to be identified generated by using the camera device 6 to capture a weld bead B10 that has been formed through actual welding process steps.

[0089] Based on the learned model M1, the estimation system 3 determines whether weld bead B10 captured in the image data D3 of the object to be identified is a good product or a defective product. If weld bead B10 is a defective product, the estimation system 3 estimates the type and location of the defect. The estimation system 3 outputs the recognition result (i.e., the estimation result) related to the image data D3 of the object to be identified to, for example, a telecommunications device or a production management system used by a user. This allows the user to check the estimation result via the telecommunications device. Optionally, the production management system can control the production facility to discard welded parts that are determined to be defective products based on the estimation results obtained by the production management system before they are transported and undergo the next process step.

[0090] (1.2.2) Data enhancement processing

[0091] The processor 10 has the function of performing data enhancement processing. Specifically, Figure 1As shown, the processor 10 includes an extraction unit 101, an acquisition unit 102, a setting unit 103, a determination unit 104, and an overlay unit 105. Note that the extraction unit 101, the acquisition unit 102, the setting unit 103, the determination unit 104, and the overlay unit 105 do not have substantial structures but merely represent functions to be performed by the processor 10.

[0092] The extraction section 101 extracts extracted image data including respective pixel values ​​of a plurality of pixels included in a predetermined extraction region R1 from the third image data D13 used as learning data.

[0093] As described above, in this embodiment, the third image data D13 is a data representing the object to be identified (including defects (reference Figures 3 to 5 The extraction unit 101 extracts, for example, a region having a defect as the extraction region R1.

[0094] The extraction region R1 may be defined, for example, by a user using a telecommunication device.

[0095] The extraction unit 101 causes the display device 12 to display a three-dimensional image (third image) Im13 represented by the third image data D13. As described above, the third image data D13 is distance image data. Therefore, the processor (extraction unit 101) determines the three-dimensional shape of the object by setting the coordinates of each pixel and the pixel value as a constituent point of the coordinate value of the three-dimensional coordinate system (i.e., X, Y, and Z coordinates) for each pixel, and connecting a plurality of such constituent points together. The processor 10 then projects the three-dimensional shape of the object thus determined onto a two-dimensional plane, thereby displaying the three-dimensional image Im13 (projected image) on the display device 12. The projected image displayed on the display device 12 is preferably capable of changing the viewpoint position or any other parameter thereof according to an operation command input by the user via the operating member 13.

[0096] Figure 6 An exemplary three-dimensional image Im13 (projected image) is shown. Figure 6 In the example shown, a pit B3 has been formed as a defect in a portion of the weld bead B10. Figure 6 In the example shown, a plurality of protrusions B30 each protruding away from the dimple B3 are shown on the outer peripheral portion of the dimple B3. These protrusions B30 are caused by noise and do not actually exist in the actual weld bead B10.

[0097] The user specifies the extraction region R1 using the operation member 13 while viewing the three-dimensional image Im13 displayed on the display device 12. Figure 6In the example shown, the user can specify the extraction region R1 by, for example, selecting a plurality of pixels included in the region of the pit B3 using a pointer. Alternatively, the user can specify the extraction region R1 by selecting a plurality of pixels forming the periphery of the pit B3. Still further, the user can specify the extraction region R1 by specifying an area having an arbitrary shape (e.g., a rectangular shape) such that the pit B3 is included therein.

[0098] In this specific example, the extraction unit 101 displays a three-dimensional image Im13 and an end button on the display device 12. The user uses a mouse as an operating member 13 to specify several points (as contour points) along the outline of the extraction region R1 (e.g., along its circumference), and then presses the end button displayed on the display device 12. The extraction unit 101 defines the extraction region R1 as a range formed by connecting the specified contour points with each other via a line, a curve, or a combination thereof. It can be seen that the data creation system 1 may further include a specifying unit 15 for specifying a predetermined extraction region R1 based on the third image data D13 and in accordance with a user's operating command (the function of which is performed by the operating member 13 and the extraction unit 101 in combination).

[0099] The extraction unit 101 extracts data related to the extraction region R1 specified by the user from the third image data D13 as extracted image data.

[0100] The acquisition unit 102 acquires characteristic image data indicating the pixel values ​​of the plurality of pixels included in the characteristic region R0. In this example, the acquisition unit 102 acquires the extracted image data extracted by the extraction unit 101 as the characteristic image data. Therefore, the shape (peripheral shape) of the characteristic region R0 corresponds to (i.e., coincides with) the shape of the extraction region R1. Furthermore, the characteristic image data includes the pixel values ​​of the plurality of pixels included in the extraction region R1 as the pixel values ​​of the plurality of pixels in the characteristic region R0.

[0101] Figure 7 An exemplary three-dimensional image Im20 (projected image) represented by the feature image data is illustrated. Figure 7 In the exemplary three-dimensional image Im20 shown, the Figure 6 The third image Im13 shown extracts the extracted image data related to the extraction region R1 as the feature image data. Since the extraction region R1, which is the source of the extracted image data, includes the pit B3, the three-dimensional shape object represented by the feature image data has a bottom portion that is recessed from the outer periphery of the feature region R0 (i.e., the bottom surface of the pit B3). Figure 7 , the bottom surface of the pit B3 is schematically indicated by a two-dot chain line.

[0102] The determination unit 104 determines the offset width based on the pixel values ​​of the plurality of pixels in the characteristic region R0. In this embodiment, the determination unit 104 converts the pixel values ​​of the plurality of pixels in the characteristic region R0 into the offset width. In this embodiment, the determination unit 104 functions as a conversion unit for converting the pixel values ​​of the plurality of pixels in the characteristic region R0 into the offset width.

[0103] The determination unit 104 determines the offset width based on the virtual cross section or virtual line. Specifically, the determination unit 104 converts the pixel values ​​of the plurality of pixels in the characteristic region R0 into the offset width based on the virtual cross section or virtual line. The virtual cross section or virtual line is set by the setting unit 103 using the pixel values ​​of two or more pixels forming the periphery (contour) of the characteristic region R0.

[0104] More specifically, the setting unit 103 sets, for each of two or more pixels forming the periphery of the characteristic region R0, a constituent point P1 whose pixel coordinates and pixel values ​​are defined as coordinate values ​​(X coordinate, Y coordinate, Z coordinate) of the three-dimensional coordinate system. The setting unit 103 uses the coordinate values ​​of a plurality of such constituent points P1 to set a virtual cross section or a virtual line within the three-dimensional coordinate system.

[0105] In this embodiment, the setting unit 103 sets a plurality of virtual lines as virtual cross sections or virtual lines. Each of the plurality of virtual lines is set as a line segment A1 that connects two constituent points P1 arranged side by side in one direction among constituent points P1 of two or more pixels forming the periphery of the characteristic region R0.

[0106] For example, Figure 10 As shown, the setting unit 103 selects any one pixel among a plurality of pixels forming the periphery of the characteristic region R0, thereby selecting a constituent point P1 corresponding to the pixel. Figure 10 When the constituent point P11 shown in FIG. 1 is set, the setting unit 103 sets the line segment A1 (for example, Figure 10 The starting point of the line segment A1 is the constituent point P11 and the end point thereof is in one direction (e.g., Figure 10 The constituent point P1 (such as the constituent point P11) adjacent to the constituent point P11 in the X-axis direction (shown in FIG. 1 ) and forming a portion of the periphery of the characteristic region R0 Figure 10 In this way, one line segment A1 is set as a virtual line. The setting unit 103 sets the line segment A1 in the same manner for a plurality of pixels forming the periphery of the characteristic region R0 (i.e., a plurality of constituent points P1). In this way, a plurality of line segments A1 are set as a plurality of virtual lines (see Figure 10 ).

[0107] It can be said that the plurality of line segments A1 (virtual lines) thus provided represent a virtual cross section defined by the periphery of the characteristic region R0 (for example, the opening of the dimple B3 when the characteristic region R0 is the dimple B3 ).

[0108] In this case, the periphery of the characteristic region R0 is not necessarily flat (ie, the Z coordinate value is not always constant). Therefore, the line segment A1 connecting the two constituent points P1 to each other may be inclined with respect to the XY plane. Figure 11 The surface of the object (ie, the inner surface of the pit B3) is shown and the Figure 10 The line segment A10 shown in FIG. 1 and the line L1 intercepted by the plane of the Z axis. Figure 11 As shown, the constituent point P12 has a larger Z coordinate value than the constituent point P11 , thereby causing the line segment A10 to be tilted relative to the X axis.

[0109] The determination unit 104 converts the pixel values ​​(i.e., Z coordinate values) of the plurality of pixels included in the characteristic region R0 into offset widths based on the plurality of virtual lines (line segments A1) thus set. Specifically, the determination unit 104 converts the pixel values ​​of the plurality of pixels corresponding to each of the plurality of line segments A1 into offset widths such that the offset widths are equal to zero at the two constituent points P1 at both ends of the line segment A1 of interest.

[0110] The determination unit 104 sets a constituent point whose coordinates and pixel values ​​of the pixel of interest are defined as coordinate values ​​(X coordinate, Y coordinate, and Z coordinate) of a three-dimensional coordinate system for each of the plurality of pixels included in the characteristic region R0. The determination unit 104 then converts each pixel value of the plurality of pixels included in the characteristic region R0 into an offset width based on the distance between each coordinate value of the constituent point for the plurality of pixels and a virtual cross section or virtual line (i.e., the distance in the Z-axis direction).

[0111] Next, refer to Figure 12 and Figure 13 In one example, how the determination unit 104 may determine the offset magnitude is described.

[0112] like Figure 12 As shown, the determination unit 104 locates a constituent point P2 having a smaller pixel value (Z coordinate value) than any other constituent point among the constituent points of the plurality of pixels corresponding to the line segment A1 (i.e., the constituent points forming the line L1 representing the surface of the object), and sets a straight line C1 passing through the constituent point P2 and parallel to the X-axis. The determination unit 104 determines the distance D0 in the Z-axis direction between the constituent point P2 and the line segment A1.

[0113] In addition, if Figure 13As shown, the determination unit 104 also determines, for each constituent point of the plurality of pixels corresponding to line segment A1, the distance E1 from line segment A1 to line C1 in the Z-axis direction, the distance E2 from the constituent point to line segment A1 in the Z-axis direction, and the distance E3 from the constituent point to line C1 in the Z-axis direction. As is clear from this definition, E1 = E2 + E3 is satisfied for each constituent point.

[0114] For each component point, the determination unit 104 calculates the ratio of distance E2 to distance E1 (E2 / E1) and multiplies this ratio by distance D0 (i.e., D0×E2 / E1) to set the offset width. In this example, at component points P1 for pixels forming the periphery of characteristic region R0 (i.e., component points P1 at both ends of line segment A1), distance E2 is zero. Therefore, at these component points P1, the offset width is zero. On the other hand, for component point P2, since distance E2 = distance E1, the offset width is equal to D0.

[0115] The superimposing section 105 creates the second image data D12 by superimposing the offset amplitude on the respective pixel values ​​of a plurality of pixels included in the predetermined region R10 of the first image Im11 represented by the first image data D11 .

[0116] The first image data D11 is image data for use as learning data, and may be, for example, original learning data regarding a non-defective object to be identified (ie, the object to be identified is a good product). Figure 8 An exemplary three-dimensional image Im11 (projected image) is shown as a first image Im11 represented by the first image data D11.

[0117] The predetermined region R10 may be any region included in the first image Im11 and having an outer peripheral shape corresponding to the characteristic region R0 .

[0118] The predetermined area R10 can be designated by the user using the above-mentioned telecommunication device, for example. The user can designate the predetermined area R10 using, for example, an operating member while looking at the three-dimensional image Im11 displayed on the display device 12. Figure 8In the example shown, the user specifies the predetermined region R10 by, for example, selecting the coordinates of an arbitrary point included in the predetermined region R10 using a pointer on the display device 12 (for example, if the predetermined region R10 has a circular shape on the XY plane, the user selects the coordinates of the center of the circle). As described above, the outer peripheral shape of the predetermined region R10 is determined to correspond to the outer peripheral shape of the feature region R0. Therefore, specifying the coordinates of a point included in the predetermined region R10 enables the position of the entire predetermined region R10 to be specified. Note that if rotation of the image represented by the feature image data is allowed, not only the coordinates of an arbitrary point included in the predetermined region R10 but also the rotation angle can be specified.

[0119] The superimposition unit 105 superimposes the offset width determined by the determination unit 104 on the pixel values ​​of the plurality of pixels included in the predetermined region R10 thus specified. That is, the superimposition unit 105 adds the offset width determined by the determination unit 104 for each pixel of each line segment A1 to the pixel value of the pixel corresponding to the former pixel in the predetermined region R10.

[0120] Next, refer to Figure 14 A to C are used to describe a specific exemplary superimposition process of the superimposition unit 105.

[0121] The superimposition section 105 locates two pixels corresponding to the two pixels of the constituent points P1 at both ends of the line segment A1 serving as the superimposition source, among the plurality of pixels included in the predetermined region R10 of the first image Im11, and locates the constituent points P100 for the located two pixels. In addition, the superimposition section 105 also determines a line segment R101 (which forms a part of the line L100 passing through the surface of the object) connecting the two constituent points P1 and defined by the plurality of constituent points (see Figure 14 A). Line L100 and line segment R101 are lines representing the outline (surface) shape of the object taken along the XZ plane. Note that although Figure 14 The center line L100 of A is illustrated as a straight line, but as Figure 15 As shown, line L100 (line segment R101 ) actually has a concavo-convex shape corresponding to the concavo-convexity of the surface of the object.

[0122] The superimposing section 105 also sets a virtual line C101 by offsetting the line L100 by a distance D0 in the Z-axis direction. This virtual line C101, like the line L100 (line segment R101), actually has a concavoconvex shape corresponding to the concavoconvexity of the surface of the object.

[0123] Next, the superimposing unit 105 replaces the pixel value (Z coordinate value) of each of the plurality of constituent points forming the line segment R101 (i.e., the portion of the line L100 between the constituent points P100, P100) with the value of the virtual line C101 (see Figure 14 B).

[0124] Finally, the superimposing unit 105 shifts the pixel value (Z coordinate value) of each of the plurality of constituent points forming the line segment R101 backward toward the line segment R101 by the ratio of the distance E3 to the distance E1 multiplied by the distance D0 (i.e., D0×E3 / E1) (see Figure 14 C) to determine the line L200 passing through the surface of the object after superposition. In this way, a new pixel value can be determined by adding an offset amplitude (D0-D0×E3 / E1)=D0×E2 / E1 to the original pixel value for each of the multiple pixels corresponding to the multiple constituent points forming line L200.

[0125] From a different perspective, the determination unit 104 transforms Figure 12 The image within the quadrilateral frame F1 shown is transformed into Figure 14 The image within the quadrilateral frame F100 shown in C is converted into an offset amplitude for each of the multiple pixels corresponding to the line segment A1. The quadrilateral frame F1 is defined by two constituent points P1 and two virtual points IP1 as follows. These two virtual points IP1 are defined as the intersections of two straight lines extending from the two constituent points P1 along the Z-axis direction and the line C1. The quadrilateral frame F100 is defined by two constituent points P100 and two virtual points IP100 as the intersections of two straight lines extending from the two constituent points P100 along the Z-axis direction and the virtual line C101. The superimposition unit 105 then adds the offset amplitude determined by the determination unit 104 for each pixel on the line segment A1 to the pixel value of the pixel corresponding to the former pixel on the line segment R101. In short, the determination unit 104 determines the offset amplitude through projective transformation.

[0126] Figure 16 The shape of the line L200 is shown, which passes through the surface of the object after superposition and is formed by passing through the Figure 11 The offset amplitude determined by transforming the cross-sectional shape shown (i.e., line L1 through the surface of the object) is superimposed on the Figure 15 The pixel value of the pixel corresponding to the line segment R101 shown in FIG. Figure 16 , a line L100 (line segment R101 ) passing through the surface of the object before the offset amplitude is superimposed is also shown as an imaginary line.

[0127] The determination section 104 and the superposition section 105 perform this conversion and superposition processing on the constituent points P100 corresponding to all pixels forming the periphery of the predetermined region R10 in the same manner. As a result, the second image data D12 (reference image data D12) is created in which the offset amplitude is superimposed (i.e., added) on the respective pixel values ​​of the plurality of pixels included in the predetermined region R10. Figure 9 ).exist Figure 9 In the example shown, the offset amplitude data indicating the offset amplitude is superimposed on two different regions R11 and R12 (respectively Figure 8 The offset amplitude data superimposed on the region R12 is obtained by inverting the offset amplitude data superimposed on the region R11.

[0128] (1.2.3) Operation

[0129] Next, refer to Figure 17 Hereinafter, an exemplary operation of the data creation system 1 will be described. Note that the procedure of the operation to be described below is merely an example and should not be construed as limiting.

[0130] First, in order to acquire feature image data, processor 10 of data creation system 1 acquires third image data D13 as original learning data representing an image with defects regarding an object to be recognized (in ST1 ).

[0131] The processor 10 defines an extraction region R1 having a defect in the third image data D13. The processor 10 extracts extraction image data including the pixel values ​​of the plurality of pixels included in the extraction region R1 (in ST2) and acquires the extracted image data as feature image data (in ST3). The processor 10 then converts the feature image data into offset amplitude data, thereby determining the offset amplitude (in ST4).

[0132] In addition, the processor 10 also acquires first image data D11 as original learning data (in ST5 ), and defines a predetermined region R10 in the first image Im11 represented by the first image data D11 (in ST6 ).

[0133] The processor 10 creates second image data D12 by superimposing the offset amplitude data on the predetermined region R10 (in ST7). The processor 10 then outputs the second image data D12 thus created (in ST8). The second image data D12 is stored in the storage device 14 as learning data (image data D1) with the label "defective" attached, similarly to the third image data D13.

[0134] (1.2.4) Advantages

[0135] As described above, the data creation system 1 according to the present embodiment creates the second image data D12 by superimposing the offset amplitude on the respective pixel values ​​of the plurality of pixels included in the predetermined region R10 of the first image Im11 represented by the first image data D11. The offset amplitude data is generated so that the offset amplitude becomes equal to zero at each of the plurality of pixels forming the periphery of the predetermined region R10.

[0136] In this case, weld bead B10 has an arbitrary height shape in the Z-axis direction. Therefore, the height (i.e., Z coordinate value) of the peripheral portion of extraction region R1 of third image data D13, which is the superimposition source, does not always coincide with the height (i.e., Z coordinate value) of the peripheral portion of predetermined region R10 of first image data D11, which is the superimposition destination. Therefore, simply replacing the individual pixel values ​​of the plurality of pixels included in predetermined region R10 with the individual pixel values ​​of the plurality of pixels included in extraction region R1 may result in a height difference in the Z-axis direction at the boundary portion of predetermined region R10 in the image created after the replacement.

[0137] On the other hand, it is conceivable that the data creation system as a comparative example creates offset amplitude data so that at one of the plurality of pixels forming the periphery of the predetermined region R10 (for example, at the position corresponding to the pixel value of the feature region R0), the offset amplitude data is generated while maintaining the correlation (i.e., the difference) between the pixel values ​​(Z coordinate values) of all pixels included in the feature region R0. Figure 15 The offset amplitude is equal to zero at the pixel corresponding to the constituent point P101 shown. This makes it possible to Figure 18 As shown, the pixel portion corresponding to the constituent point P101 is connected without causing any height difference. However, in this case, a height difference may still be caused in the Z-axis direction at the pixel portion corresponding to a different constituent point P102 other than the constituent point P101 among the plurality of pixels forming the periphery of the predetermined region R10.

[0138] In contrast, the data creation system 1 according to this embodiment enables data to be superimposed without causing any height differences in the boundary portion, thereby enabling the creation of pseudo-data that is even closer to image data that could exist in the real world. Then, based on the learned model M1 generated using the second image data D12 thus obtained as learning data, the condition of the object to be recognized represented by the object to be recognized image data D3 is estimated. This reduces the possibility of erroneous recognition of the condition of the object to be recognized due to the presence of height differences. Consequently, this helps reduce the possibility of a decline in the performance of the learned model M1 in recognizing objects.

[0139] (2) Second embodiment

[0140] In the data creation system 1 according to the second embodiment, the processor 10 maintains the correlation between the respective pixel values ​​of adjacent pixels for pixels falling within a predetermined range (hereinafter referred to as the "maintenance region") of the feature region R0 when converting the feature image data into the offset amplitude data. This is different from the data creation system 1 according to the first embodiment described above. In the following description, any constituent elements of the second embodiment having substantially the same functions as the corresponding portions of the data creation system 1 according to the first embodiment described above will be denoted by the same reference numerals as those of the corresponding portions, and their description will be omitted as appropriate.

[0141] like Figure 19 As shown, the processor 10 according to this embodiment further includes a holding region defining section 106 and a threshold specifying section 107. Note that the holding region defining section 106 and the threshold specifying section 107 do not have substantial structures but merely represent functions to be performed by the processor 10.

[0142] The threshold setting unit 107 sets a threshold. In this embodiment, the threshold setting unit 107 sets the threshold in response to a user command. The threshold is a value to be compared with the pixel values ​​(Z coordinate values) of the plurality of pixels included in the extraction region R1 (feature region R0). The user sets the threshold using the operating member 13 while viewing the image displayed on the display device 12.

[0143] For example, in the three-dimensional image Im13 (refer to Figure 6 ) is displayed on the display device 12, the threshold value specifying unit 107 displays the pixel area whose pixel values ​​(i.e., distance values) are equal to or greater than the threshold value and the pixel area whose pixel values ​​(distance values) are less than the threshold value in different modes (e.g., in two different colors). When the user changes the threshold value by operating the telecommunication device, the display content on the display device 12 (i.e., the range of the pixel area whose pixel values ​​are equal to or greater than the threshold value and the pixel area whose pixel values ​​are less than the threshold value) also changes accordingly. This enables the user to specify any desired threshold value while viewing the three-dimensional image Im13 displayed on the display device 12.

[0144] The holding region defining unit 106 defines a holding region based on a comparison result between a threshold and each pixel value of a plurality of pixels included in the characteristic region R0. For example, the holding region defining unit 106 sets a pixel region having pixel values ​​less than the threshold as a holding region.

[0145] Figure 20 Shown in the Figure 10 The exemplary results of the maintenance area defined on the line segment A10 and the cross section of the Z axis are shown. Figure 20In the illustrated example, a region R2 between two intersection points P20 , P20 where a line L1 passing through the surface of the object and a line Th1 indicating a threshold value intersect each other is defined as a maintenance region.

[0146] It can be seen that the holding area definition section 106 defines the holding area based on the threshold value specified by the threshold value specifying section 107 .

[0147] The determination unit 104 determines the offset amplitude so as to maintain the correlation between the pixel values ​​of the plurality of pixels included in the maintained region. In other words, with respect to the remaining portion of feature region R0 (hereinafter referred to as the "modified region") other than the maintained region, the determination unit 104 changes the correlation between the pixel values ​​(Z coordinate values) of the plurality of pixels included in the modified region. On the other hand, with respect to the maintained region of feature region R0, the determination unit 104 maintains the correlation between the pixel values ​​of the plurality of pixels included in the maintained region.

[0148] Specifically, for the change area, the determination unit 104 sets a straight line C1 and a distance D0 in the same manner as in the first embodiment, and sets D0×E2 / E1 as the offset width for each pixel. Note that the straight line C1 is parallel to the straight line Th1 representing the threshold value.

[0149] On the other hand, for the maintenance area, the determination section 104 determines the offset width so as to maintain the correlation (ie, the difference) between the pixel values ​​(Z coordinate values) of adjacent pixels.

[0150] Will refer to Figures 12 to 14 and Figure 21 First, for each of the plurality of pixels included in the characteristic region R0 (including both the conversion region and the maintenance region), the determination unit 104 sets D0×E2 / E1 as the offset width (reference Figure 12 and Figure 13 The superposition unit 105 determines a new pixel value for each of the plurality of pixels included in the characteristic region R0 based on the offset width determined by the determination unit 104 (refer to Figure 14 C).

[0151] Finally, for a plurality of pixels included in the maintenance area (i.e., pixels whose pixel values ​​are less than the threshold), the superposition unit 105 replaces their pixel values ​​(Z coordinate values) using the pixel value at the constituent point P200 at the boundary of the maintenance area as a reference value to maintain the correlation (difference) between the pixel values ​​of adjacent pixels. Figure 21 In the illustrated example, the pixel values ​​between two constituent points P200 are changed from the values ​​indicated by the two-dot chain line to the values ​​indicated by the solid line by replacing the pixel values ​​to maintain the correlation between the pixel values ​​of adjacent pixels.

[0152] From a different perspective, the determination unit 104 transforms Figure 22 The image within the quadrilateral frame F11 shown in A is transformed into Figure 22 On the other hand, the determination unit 104 Figure 22 The image within the quadrilateral frame F12 shown in A is moved (i.e., translated) as it is to Figure 22 In this manner, the determination unit 104 converts the pixel values ​​of the plurality of pixels corresponding to the line segment A1 into offset widths.

[0153] A quadrilateral frame F11 is defined by two constituent points P1 and two virtual points IP11 located at the intersection of two straight lines extending from the two constituent points P1 in the Z-axis direction and a line Th1 indicating a threshold value. A quadrilateral frame F12 is defined by two virtual points IP11 and two virtual points IP12 located at the intersection of two straight lines extending from the two virtual points IP11 in the Z-axis direction and a straight line C1.

[0154] The quadrilateral frame F102 is defined by two virtual points IP102 and IP101. The two virtual points IP102 are located at the intersection of two straight lines extending from the two constituent points P100 in the Z-axis direction and the virtual line C101. The two virtual points IP101 are determined by offsetting the Z coordinate value of the virtual point IP102 by the distance between the virtual points IP11 and IP12 of the quadrilateral frame F12. The quadrilateral frame F101 is defined by the two constituent points P100 and the two virtual points IP101.

[0155] In short, the determination unit 104 determines the offset width using projective transformation.

[0156] Figure 23 The shape of the line L200 is shown, which is obtained by the processor 10 according to the present embodiment and is obtained by comparing with Figure 15 The line segment R101 shown corresponds to the pixel portion of the surface of the object after superposition.

[0157] According to this embodiment, the correlation between the individual pixel values ​​of the plurality of pixels included in the maintenance area is maintained before and after the transformation performed by the determination unit 104 (transformation unit). This reduces the likelihood of a transformation that tilts the bottom surface of the dimple B3 relative to the horizontal plane, for example. Therefore, the data creation system 1 according to this embodiment enables the generation of learning data that represents an image closer to that which would exist in the real world. Consequently, this helps reduce the likelihood of a decline in the recognition performance of the learned model M1 generated using the learning data.

[0158] (3) Modification

[0159] Note that the above-described embodiments are merely typical of various embodiments of the present invention and should not be construed as limiting. Rather, the exemplary embodiments can be readily modified in various ways based on design choices or any other factors without departing from the scope of the present invention. Furthermore, the functions of the data creation system 1 according to the above-described exemplary embodiments may also be implemented as a data creation method, a computer program, or a non-transitory storage medium storing a computer program.

[0160] Next, modifications of the exemplary embodiment will be listed one by one. Note that the modifications described below can be appropriately combined and adopted.

[0161] The data creation system 1 according to the present invention includes a computer system. The computer system may include a processor and a memory as its main hardware components. The functions of the data creation system 1 according to the present invention can be performed by causing the processor to execute a program stored in the memory of the computer system. The program can be pre-stored in the memory of the computer system. Alternatively, the program can also be downloaded via a telecommunications line, or distributed after being recorded in some non-transitory storage medium such as a memory card, an optical disk, or a hard disk drive (any of which is readable by the computer system). The processor of the computer system can be composed of a single or multiple electronic circuits including a semiconductor integrated circuit (IC) or a large-scale integrated circuit (LSI). As used herein, "integrated circuits" such as IC or LSI are referred to by different names depending on the degree of their integration. Examples of integrated circuits include system LSI, very large-scale integrated circuit (VLSI) and ultra-large-scale integrated circuit (ULSI). Alternatively, a field programmable gate array (FPGA) to be programmed after the LSI is manufactured or a logic device that allows reconfiguration of connections or circuit sections within the LSI can also be used as the processor. These electronic circuits can be integrated together on a single chip or distributed on multiple chips, whichever is appropriate. These multiple chips can be aggregated together in a single device or distributed across multiple devices, without limitation. As used herein, a "computer system" includes a microcontroller comprising one or more processors and one or more memories. Therefore, a microcontroller can also be implemented as a single or multiple electronic circuits comprising semiconductor integrated circuits or large-scale integrated circuits.

[0162] Furthermore, in the above embodiment, the multiple functions of the data creation system 1 are integrated into a single housing. However, this is not an essential structure of the data creation system 1. Alternatively, the constituent elements of the data creation system 1 may be distributed across multiple different housings.

[0163] Instead, multiple functions of the data creation system 1 may be gathered together in a single housing. Alternatively, at least some functions of the data creation system 1 (eg, some functions of the data creation system 1 ) may also be implemented as a cloud computing system, for example.

[0164] (3.1) First Modification

[0165] Will refer to Figure 24 and Figure 25 A data creation system 1 according to a first modification will be described. In the data creation system 1 according to this modification, a holding region definition unit 106 defines a holding region based on designated pixels, which is a difference from the data creation system 1 according to the second embodiment. In the following description, any constituent elements of this first modification that have substantially the same functions as their counterparts in the data creation system 1 according to the second embodiment described above will be denoted by the same reference numerals as those of the corresponding components, and their description will be omitted as appropriate.

[0166] like Figure 24 As shown, the processor 10 according to this modification includes a range specifying unit 108 instead of the threshold specifying unit 107 .

[0167] The range specifying unit 108 specifies a range covering at least one pixel (hereinafter referred to as a “specific pixel”) among the pixels included in the characteristic region R0 . The holding region defining unit 106 defines a holding region covering the at least one pixel (specific pixel) specified by the range specifying unit 108 .

[0168] In this case, the range specifying section 108 specifies a specific pixel according to a user's command. For example, the user can specify a specific pixel using the operating member 13 while looking at an image displayed on the display device 12.

[0169] For example, in the cross section of the characteristic region R0 (refer to Figure 25 ) is displayed on the display device 12, the user specifies a pixel corresponding to an arbitrary constituent point forming a part of the surface of the characteristic region R0 as a specific pixel. Figure 25 In the example shown, two pixels corresponding to constituent points P31 and P32 are designated as specific pixels. The range designation unit 108 sets a virtual cross section including a straight line C30 passing through these two constituent points P31 and P32. The maintenance region definition unit 106 defines a maintenance region that includes pixels whose pixel values ​​(Z coordinate values) are smaller than those of the virtual cross section.

[0170] In one example, the virtual cross section including the straight line C30 may be a virtual cross section including the straight line C30 and another straight line defined by translating the straight line C30 in the Y-axis direction. In another example, the virtual cross section including the straight line C30 may be a virtual cross section including two constituent points P31 and P32 and another constituent point specified separately from the constituent points P31 and P32 (i.e., a total of three constituent points). Note that the three constituent points defining the virtual cross section do not necessarily include both the two constituent points P31 and P32. Alternatively, the three constituent points may be specified on three different cross sections, respectively.

[0171] As can be seen, the processor 10 according to this modification defines a maintenance region that includes the specific pixels specified by the range specification unit 108. Furthermore, the determination unit 104 converts the pixel values ​​of the plurality of pixels included in the characteristic region R0 into offset amplitudes in order to maintain the correlation between the pixel values ​​of the plurality of pixels included in the maintenance region. As with the second embodiment described above, this reduces the likelihood of a transformation that would tilt the bottom surface of the dimple B3 relative to the horizontal plane, for example. Consequently, the data creation system 1 according to this modification enables the creation of learning data representing an image that more closely resembles an image that would exist in the real world.

[0172] In a modified example, the maintenance region definition unit 106 may define the maintenance region based solely on a specific pixel specified by the user. For example, the maintenance region definition unit 106 may set the pixel value (Z coordinate value) of the specific pixel specified by the user as a threshold. The maintenance region definition unit 106 may then define the maintenance region based on a comparison result between the threshold and the pixel values ​​(Z coordinate values) of the plurality of pixels included in the feature region R0 (i.e., based on their respective magnitudes).

[0173] In another variation, the processor 10 may include both the threshold specification unit 107 and the range specification unit 108. In this case, the maintenance region definition unit 106 may define the maintenance region based solely on the threshold specified by the threshold specification unit 107. Alternatively, the maintenance region definition unit 106 may define the maintenance region based solely on the specific pixels specified by the range specification unit 108. Still further alternatively, the maintenance region definition unit 106 may define the maintenance region based on both the threshold and the specific pixels (e.g., based on an AND (logical product) or an OR (logical sum) of these two results).

[0174] (3.2) Second Modification

[0175] Will refer to Figure 26A data creation system 1 according to a second modification will be described. In the data creation system 1 according to this modification, a setting unit 103 sets a virtual plane based on the coordinate values ​​of constituent points P1 of two or more pixels forming the periphery of feature region R0, and a determination unit 104 performs transformation based on the virtual plane. This is a difference from the data creation system 1 according to the first embodiment described above. In the following description, any constituent elements of this second modification that have substantially the same functions as their counterparts in the data creation system 1 according to the first embodiment described above will be denoted by the same reference numerals as those of the corresponding components, and their description will be omitted as appropriate.

[0176] In this modification, the setting unit 103 sets a virtual plane as a virtual cross section or a virtual line so that the average distance between the virtual plane and the coordinate values ​​of the constituent points P1 of two or more pixels forming the periphery of the characteristic region R0 becomes minimum.

[0177] For example, Figure 26 As shown in FIG. 1 , the setting unit 103 selects a plurality of pixels forming the periphery of the characteristic region R0, thereby selecting a plurality of corresponding constituent points P1. When the plurality of constituent points P1 are selected, the setting unit 103 determines a virtual plane based on the respective coordinate values ​​(X, Y, and Z coordinates) of the constituent points P1, which minimizes the average distance between the virtual plane and the constituent points P1 (i.e., the distance in the Z-axis direction).

[0178] The determination unit 104 determines the offset width based on the virtual plane set by the setting unit 103. The determination unit 104 can determine the offset width by, for example, three-dimensional projective transformation. In this case, the superposition unit 105 sets a virtual plane for the predetermined region R10 in the same manner as in the case of the characteristic region R0, and can superimpose the offset width determined by the determination unit 104 on the pixel values ​​of the plurality of pixels included in the virtual plane.

[0179] The data creation system 1 according to this modification enables the determination unit 104 to convert the pixel values ​​of multiple pixels in the characteristic region R0 into offset widths based on a virtual plane, which is a virtual cross section or virtual line. This enables the creation of pseudo data that is closer to image data that can exist in the real world.

[0180] The configuration of the setting section 103 according to this modification example can be applied to the processor 10 according to the second embodiment.

[0181] (3.3) Third Modification

[0182] In the data creation system 1 , the processing device (hereinafter referred to as “first processing device”) 110 including the extraction section 101 and the processing device (hereinafter referred to as “second processing device”) 120 including the acquisition section 102 and the superposition section 105 may be two different devices.

[0183] For example, Figure 27 As shown, the first processing device 110 includes a processor (hereinafter referred to as "first processor") 1001, a communication interface (hereinafter referred to as "first communication interface") 111, a display device 12, and an operating member 13. The first processor 1001 of the first processing device 110 includes an extraction unit 101. The first processing device 110 includes a designation unit 15 (which includes the operating member 13 and the extraction unit 101).

[0184] The first communication interface 111 receives third image data D13 as raw learning data representing an image (including defects) of an object to be recognized from the imaging device 6 .

[0185] The extraction section 101 (specifying section 15 ) extracts extracted image data including respective pixel values ​​of a plurality of pixels included in a predetermined extraction region R1 from the third image data D13 .

[0186] The first communication interface 111 (transmitting unit) outputs (transmits) the extracted image data D20 extracted by the extracting unit 101 to the second processing device 120 .

[0187] The second processing device 120 includes a processor (hereinafter referred to as “second processor”) 1002 and a communication interface (hereinafter referred to as “second communication interface”) 112. The second processor 1002 of the second processing device 120 includes an acquisition unit 102, a setting unit 103, a determination unit 104, and an overlay unit 105.

[0188] The second communication interface 112 receives the first image data D11 as raw learning data from the imaging device 6. In addition, the second communication interface 112 also receives the extracted image data D20 from the first processing device 110.

[0189] The acquisition unit 102 acquires the extracted image data D20 received by the second communication interface 112 as feature image data indicating the pixel values ​​of a plurality of pixels included in the feature region R0. The setting unit 103 sets a virtual cross section or virtual line based on the pixel values ​​of two or more pixels forming the periphery (contour) of the feature region R0. The determination unit 104 determines an offset amplitude based on the virtual cross section or virtual line. The superposition unit 105 superimposes the offset amplitude determined by the determination unit 104 on the respective pixel values ​​of a plurality of pixels included in the predetermined region R10 of the first image Im11 represented by the first image data D11, thereby creating second image data D12.

[0190] The second processing device 120, for example, can cause the second communication interface 112 to transmit the second image data D12 thus created to the first processing device 110. In this case, the user can cause the learning system 2 to generate the learned model M1 using the second image data D12 thus received.

[0191] The second processing device 120 can transmit the second image data D12 thus generated to an external server including a learning system. The learning system of the external server uses a learning dataset including the learning data of the second image data D12 to generate a learned model M1. This learned model M1 outputs an estimation result similar to that in the case where the third image data D13 is an object to be identified, in response to the second image Im12 represented by the second image data D12 or a predetermined region R10 within the second image Im12. In this case, the predetermined region R10 is an area included in the first image Im11 and having a peripheral shape corresponding to the extraction region R1. The first image Im11 includes a pixel region indicating the object to be identified and is represented by the first image data D11. The second image Im12 is generated by superimposing an offset amplitude determined based on the individual pixel values ​​of the multiple pixels included in the extraction region R1 on the individual pixel values ​​of the multiple pixels included in the predetermined region R10. An exemplary estimation result may indicate whether the object to be identified is a good product or a defective product. If the object to be identified is a defective product, the exemplary estimation result may include the type of defect. An exemplary estimation result may include the size of the defect if the object to be identified is a defective product. If the object to be identified is a defective product and the defect type is a pit B3, an exemplary estimation result may include the depth of the pit B3. The user may receive the learned model M1 generated in this manner from an external server.

[0192] (3.4) Other variations

[0193] In a modified example, the object to be identified does not necessarily have to be the weld bead B10. That is, the learned model M1 does not necessarily have to be used in a welding appearance test to check whether welding is performed correctly.

[0194] In another variation, the feature image data need not be extracted image data extracted from the third image data D13 used as learning data. Alternatively, the feature image data may be data arbitrarily created by a user, for example. Furthermore, the feature image data acquired by acquisition unit 102 need not be the extracted image data extracted by extraction unit 101. Alternatively, for example, acquisition unit 102 may acquire the feature image data from another device via communication. Still further alternatively, acquisition unit 102 may acquire feature image data pre-stored in a storage device of processor 10 from the storage device.

[0195] In yet another modification, the feature image data is not necessarily data of an image of a defective product, but may be data of an image of a good product.

[0196] In yet another modification, the first image data D11 may be data representing an image showing a defect of the object to be identified. In yet another modification, the first image data D11 may be the same as the third image data D13.

[0197] In still another modification, the processor 10 performs the processes for defining the extraction region R1 , defining the predetermined region R10 , and defining the maintenance region according to an appropriately set reference instead of following a user's command.

[0198] In another variation, for example, the plurality of line segments A1 need not be line segments aligned with the X axis, but may be line segments aligned with the Y axis. However, when viewed along the Z axis (i.e., when viewed from the perspective of the line segment A1 drawn with Figure 10 The plurality of line segments A1 are preferably parallel to each other (when viewed from the front of the paper).

[0199] In still another modification, the image represented by the characteristic image data may be deformed (for example, rotated, reversed, or enlarged or reduced).

[0200] In still another modification, the first image data D11 and the third image data D13 do not need to be distance image data, but may also be luminance image data.

[0201] As used herein, "image data" does not necessarily refer to image data captured by an image sensor. Instead, it may refer to two-dimensional data such as CG images, or, as described in the basic example, two-dimensional data formed by arranging multiple one-dimensional data captured by a range image sensor. Alternatively, "image data" may refer to three-dimensional or higher-dimensional data. Furthermore, "pixel" as used herein does not necessarily refer to the pixels of an image actually captured by an image sensor, but may also refer to individual elements of two-dimensional data.

[0202] The evaluation system 100 may include only some of the constituent elements of the data creation system 1. For example, the evaluation system 100 may include only the first processing device 110 and the second processing device 120 of the data creation system 1 (see Figure 13 ) and the learning system 2. The functions of the first processing device 110 and the functions of the learning system 2 can be provided by a single device. Alternatively, the evaluation system 100 can include, for example, only the first processing device 110 of the first processing device 110 and the second processing device 120 of the data creation system 1 and the estimation system 3. The functions of the first processing device 110 and the functions of the estimation system 3 can be provided by a single device.

[0203] (4) Various aspects

[0204] As can be seen from the foregoing description, the above-mentioned embodiments and their variations may be specific implementations of the following aspects of the present invention.

[0205] According to a first aspect, a data creation system (1) creates second image data (D12) for use as learning data for generating a learned model (M1) based on first image data (D11). The data creation system (1) includes an acquisition unit (102) and a superposition unit (105). The acquisition unit (102) acquires feature image data indicating the respective pixel values ​​of a plurality of pixels included in a feature region (R0). The superposition unit (105) creates the second image data (D12) by superimposing an offset amplitude on the respective pixel values ​​of a plurality of pixels included in a predetermined region (R10) of a first image (Im11) represented by the first image data (D11). The predetermined region (R10) has an outer peripheral shape corresponding to the feature region (R0). The offset amplitude is determined based on the respective pixel values ​​of the plurality of pixels included in the feature region (R0).

[0206] This aspect enables the generation of learning data that represents images that are closer to images that may exist in the real world, thereby helping to reduce the possibility of causing a decline in the performance of the learned model (M1) generated based on the learning data.

[0207] In the data creation system (1) according to the second aspect which can be implemented in combination with the first aspect, the first image data (D11) is distance image data including pixel values ​​expressed as distance values.

[0208] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0209] The data creation system (1) according to the third aspect, which can be implemented in combination with the second aspect, further includes a determination unit (104) and a setting unit (103). The determination unit (104) determines the offset amplitude based on the respective pixel values ​​of a plurality of pixels in the characteristic region (R0). The setting unit (103) sets a virtual cross section or a virtual line based on the respective pixel values ​​of two or more pixels forming the periphery of the characteristic region (R0). The determination unit (104) determines the offset amplitude by transformation based on the virtual cross section or the virtual line.

[0210] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0211] In the data creation system (1) according to the fourth aspect that can be implemented in combination with the third aspect, the setting unit (103) sets, for each of two or more pixels forming the periphery of the feature region (R0), a constituent point (P1) using the coordinates and pixel values ​​of each pixel as coordinate values ​​of a three-dimensional coordinate system. The setting unit (103) sets a virtual section or a virtual line within the three-dimensional coordinate system using the coordinate values ​​of the constituent point (P1) for each of the two or more pixels.

[0212] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0213] In the data creation system (1) according to the fifth aspect that can be implemented in combination with the fourth aspect, the virtual section or virtual line includes a virtual plane. The setting unit (103) sets the virtual plane so that the average distance between the coordinate values ​​of the constituent points (P1) of two or more pixels and the virtual plane is minimized.

[0214] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0215] In the data creation system (1) according to the sixth aspect that can be implemented in combination with the fourth aspect, the virtual section or virtual line includes a plurality of virtual lines. The setting unit (103) sets a line segment (A1, A10) as each of the plurality of virtual lines, the line segment connecting two constituent points (P11, P12) arranged side by side in one direction and selected from constituent points (P1) for two or more pixels.

[0216] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0217] The data creation system (1) according to the seventh aspect, which can be implemented in combination with any one of the first to sixth aspects, further includes a determination unit (104). The determination unit (104) determines an offset width based on respective pixel values ​​of a plurality of pixels in the feature region (R0). The determination unit (104) determines the offset width using a projective transformation.

[0218] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0219] The data creation system (1) according to the eighth aspect, which can be implemented in combination with any one of the first to seventh aspects, further includes a determination unit (104) and a maintenance region definition unit (106). The determination unit (104) determines an offset amplitude based on the respective pixel values ​​of a plurality of pixels in the feature region (R0). The maintenance region definition unit (106) defines a maintenance region in the feature region (R0) that maintains the correlation between the respective pixel values ​​of adjacent pixels. The determination unit (104) determines the offset amplitude by transformation so as to maintain the correlation between the respective pixel values ​​of the plurality of pixels included in the maintenance region.

[0220] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0221] The data creation system (1) according to the ninth aspect, which can be implemented in combination with the eighth aspect, further includes a threshold value specifying unit (107) for specifying a threshold value. The maintenance region defining unit (106) defines the maintenance region based on a comparison result between the threshold value and each pixel value of a plurality of pixels included in the feature region (R0).

[0222] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0223] The data creation system (1) according to the tenth aspect, which can be implemented in combination with the eighth aspect or the ninth aspect, further includes a range specifying unit (108). The range specifying unit (108) specifies a range covering at least one pixel among a plurality of pixels included in the feature region (R0). The maintenance region defining unit (106) defines the maintenance region so that the maintenance region covers the at least one pixel specified by the range specifying unit (108).

[0224] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0225] The data creation system (1) according to the eleventh aspect, which can be implemented in combination with any one of the first to tenth aspects, further includes an extraction unit (101). The extraction unit (101) extracts extracted image data including pixel values ​​of a plurality of pixels included in a predetermined extraction region (R1) from third image data (D13) used as learning data. The acquisition unit (102) acquires the extracted image data as feature image data.

[0226] This aspect enables the generation of learning data that represents images that are closer to images that can exist in the real world.

[0227] The learning system according to the twelfth aspect generates a learned model (M1) using a learning data set including second image data (D12) as learning data created by the data creation system (1) according to any one of the first to eleventh aspects.

[0228] This aspect helps to reduce the possibility of causing degradation in the performance of the learned model (M1).

[0229] The estimation system according to the thirteenth aspect uses the learned model (M1) generated by the learning system according to the twelfth aspect to perform estimation regarding the object to be recognized.

[0230] According to this aspect, the object to be recognized is estimated using the learned model while reducing the possibility of causing degradation in recognition performance, thereby enabling an appropriate estimation result to be obtained.

[0231] The data creation method according to the fourteenth aspect is designed to create second image data (D12) for use as learning data for generating a learned model (M1) based on first image data (D11). The data creation method includes an acquisition step and a superposition step. The acquisition step includes acquiring feature image data indicating the respective pixel values ​​of a plurality of pixels included in a feature region (R0). The superposition step includes creating the second image data (D12) by superimposing an offset amplitude on the respective pixel values ​​of a plurality of pixels included in a predetermined region (R10) of a first image (Im11) represented by the first image data (D11). The predetermined region (R10) has an outer peripheral shape corresponding to the feature region (R0). The offset amplitude is determined based on the respective pixel values ​​of the plurality of pixels included in the feature region (R0).

[0232] This aspect enables the generation of learning data that represents images that are closer to images that may exist in the real world, thereby helping to reduce the possibility of causing a decline in the performance of the learned model (M1) generated based on the learning data.

[0233] The program according to the fifteenth aspect is designed to cause one or more processors to perform the data creation method according to the fourteenth aspect.

[0234] This aspect enables the generation of learning data that represents images that are closer to images that may exist in the real world, thereby helping to reduce the possibility of causing a decline in the performance of the learned model (M1) generated based on the learning data.

[0235] The data creation system (1) according to the sixteenth aspect, which can be implemented in combination with the eleventh aspect, includes a first processing device (110) and a second processing device (120). The first processing device (110) includes an extraction unit (101). The second processing device (120) includes an acquisition unit (102) and a superposition unit (105). The first processing device (110) sends the extracted image data (D20) to the second processing device (120). The second processing device (120) receives the extracted image data (D20) from the first processing device (110). The acquisition unit (102) of the second processing device (120) acquires the extracted image data (D20) as feature image data.

[0236] This aspect helps to reduce the possibility of causing degradation in the performance of the learned model (M1).

[0237] In the data creation system (1) according to the seventeenth aspect, which can be implemented in combination with the sixteenth aspect, the first processing device (110) further includes a designating unit (15). The designating unit (15) designates a predetermined extraction region (R1) based on the third image data (D13) according to an operation command input by a user.

[0238] The processing device according to the eighteenth aspect is used as the first processing device (110) of the data creation system (1) according to the sixteenth aspect or the seventeenth aspect.

[0239] The processing device according to the nineteenth aspect is used as the second processing device (120) of the data creation system (1) according to the sixteenth aspect or the seventeenth aspect.

[0240] The evaluation system (100) of the twentieth aspect includes a processing device (110) and a learning system (2). The processing device (110) extracts extracted image data including the respective pixel values ​​of a plurality of pixels included in a predetermined extraction region (R1) from third image data (D13) representing a third image (Im13) including a pixel region indicating an object to be identified. The processing device (110) outputs the extracted image data (D20) thus extracted. The learning system (2) generates a learned model (M1). The learned model (M1) outputs an estimation result similar to the case where the third image data (D13) is an object to be identified in response to the second image (Im12) represented by the second image data (D12) or the predetermined region (R10) in the second image (Im12). The predetermined region (R10) is an area included in the first image (Im11) and having an outer peripheral shape corresponding to the extraction region (R1). The first image (Im11) includes a pixel region indicating an object to be identified and is represented by the first image data (D11). The second image (Im12) is generated by superimposing an offset amplitude on each pixel value of a plurality of pixels included in a predetermined region (R10) of the first image (Im11). The offset amplitude is determined based on each pixel value of a plurality of pixels included in the extraction region (R1).

[0241] This aspect helps to reduce the possibility of causing degradation in the performance of the learned model (M1).

[0242] The processing device (110) according to the twenty-first aspect is used as the processing device (110) of the evaluation system (100) according to the twentieth aspect.

[0243] The learning system (2) according to the twenty-second aspect is used as the learning system (2) of the evaluation system (100) according to the twentieth aspect.

[0244] The evaluation system (100) according to the twenty-third aspect includes a processing device (110) and an estimation system (3). The processing device (110) extracts extracted image data including the respective pixel values ​​of a plurality of pixels included in a predetermined extraction region (R1) from third image data (D13) representing a third image (Im13) including a pixel region indicating an object to be identified. The processing device (110) outputs the extracted image data (D20) thus extracted. The estimation system (3) uses a learned model (M1) to perform an estimation related to the object to be identified. The learned model (M1) outputs an estimation result similar to the case where the third image data (D13) is the object to be identified in response to a predetermined region (R10) in the second image (Im12) or the second image (Im2) represented by the second image data (D12). The predetermined region (R10) is an area included in the first image (Im11) and having an outer peripheral shape corresponding to the extraction region (R1). A first image (Im11) includes a pixel region indicating an object to be recognized and is represented by first image data (D11). A second image (Im12) is generated by superimposing an offset magnitude on the respective pixel values ​​of a plurality of pixels included in a predetermined region (R10) of the first image (Im11). The offset magnitude is determined based on the respective pixel values ​​of the plurality of pixels included in the extraction region (R1).

[0245] This aspect helps to reduce the possibility of causing degradation in the performance of the learned model (M1).

[0246] The processing device (110) of the twenty-fourth aspect is used as the processing device (110) of the evaluation system (100) of the twenty-third aspect.

[0247] The learning system (2) according to the twenty-fifth aspect is used as the estimation system (3) of the evaluation system (100) according to the twenty-third aspect.

[0248] Note that the constituent elements according to the second to eleventh aspects and the sixteenth and seventeenth aspects are not essential constituent elements of the data creation system (1) but may be omitted as appropriate.

[0249] Description of Reference Numerals

[0250] 1Data creation system

[0251] 101 Extraction Department

[0252] 102 Acquisition Department

[0253] 103 Settings Department

[0254] 104 Determination Department

[0255] 105 superposition part

[0256] 106 Maintaining Area Definition Department

[0257] 107 Threshold setting unit

[0258] 108 Scope Designation Department

[0259] 15 Designated Department

[0260] 100 Evaluation System

[0261] 110 First processing device (processing device)

[0262] 120 Second processing device

[0263] 2Learning System

[0264] 3 Estimation System

[0265] D11 First image data

[0266] D12 Second image data

[0267] D13 Third image data

[0268] D20 Extract image data

[0269] Im11 first image

[0270] Im12 Second image

[0271] Im13 third image

[0272] R0 characteristic area

[0273] R1 Extraction Area

[0274] R10 reserved area

[0275] P1, P11, P12 constitute the points

[0276] A1, A10 line segment

[0277] M1 Learning Model

Claims

1. A data creation system configured to create second image data for use as learning data for generating a learned model based on first image data, the data creation system comprising: an acquisition section configured to acquire feature image data indicating respective pixel values ​​of a plurality of pixels included in the feature area; as well as A superposition unit is configured to create the second image data by superimposing an offset amplitude on the respective pixel values ​​of a plurality of pixels included in a predetermined area of ​​the first image represented by the first image data, wherein the predetermined area has a peripheral shape corresponding to the feature area, and the offset amplitude is determined based on the respective pixel values ​​of the plurality of pixels included in the feature area.

2. The data creation system according to claim 1, wherein: The first image data is range image data including pixel values ​​expressed as distance values.

3. The data creation system according to claim 2, further comprising: a determination unit configured to determine the offset amplitude based on respective pixel values ​​of a plurality of pixels in the feature area; as well as a setting section configured to set a virtual cross section or a virtual line based on respective pixel values ​​of two or more pixels forming a periphery of the characteristic region, The determining unit is configured to determine the offset amplitude based on the virtual cross section or the virtual line.

4. The data creation system according to claim 3, wherein: The setting section is configured to set, for each of the two or more pixels, a constituent point using the coordinates and pixel value of the pixel as a coordinate value of a three-dimensional coordinate system, and The setting section is configured to set the virtual cross section or the virtual line within the three-dimensional coordinate system using coordinate values ​​of constituent points for each of the two or more pixels.

5. The data creation system according to claim 4, wherein: The virtual section or the virtual line includes a virtual plane, and The setting section is configured to set the virtual plane so that an average distance between the coordinate values ​​of the constituent points for the two or more pixels and the virtual plane is minimized.

6. The data creation system according to claim 4, wherein: The virtual section or the virtual line includes a plurality of virtual lines, and The setting section is configured to set, as each of the plurality of virtual lines, a line segment connecting two constituent points arranged side by side in one direction and selected from constituent points for the two or more pixels.

7. The data creation system according to any one of claims 1 to 6, further comprising a determination unit configured to determine the offset width based on respective pixel values ​​of a plurality of pixels in the feature area, in, The determination section is configured to determine the offset magnitude using a projective transformation.

8. The data creation system according to any one of claims 1 to 7, further comprising: a determination unit configured to determine the offset amplitude based on respective pixel values ​​of a plurality of pixels in the feature area; as well as a maintenance region defining section configured to define a maintenance region in the feature region, wherein correlation between respective pixel values ​​of adjacent pixels is maintained in the maintenance region, The determining unit is configured to determine the offset amplitude so as to maintain the correlation between pixel values ​​of the plurality of pixels included in the maintenance area.

9. The data creation system according to claim 8, further comprising a threshold value specifying unit configured to specify a threshold value, in, The holding region defining section is configured to define the holding region based on a comparison result between the threshold value and respective pixel values ​​of a plurality of pixels included in the feature region.

10. The data creation system according to claim 8 or 9, further comprising a range specifying section configured to specify a range covering at least one pixel among a plurality of pixels included in the feature area, in, The holding region defining section is configured to define the holding region so that the holding region covers at least one pixel specified by the range specifying section.

11. The data creation system according to any one of claims 1 to 10, further comprising an extraction unit configured to extract extracted image data including respective pixel values ​​of a plurality of pixels included in a predetermined extraction area from the third image data to be used as the learning data, in, The acquisition section is configured to acquire the extracted image data as the feature image data.

12. The data creation system according to claim 11, comprising a first processing device and a second processing device, in, The first processing device includes the extraction unit, The second processing device includes the acquisition unit and the superposition unit, The first processing device is configured to send the extracted image data to the second processing device, The second processing device is configured to receive the extracted image data from the first processing device, and The acquisition section of the second processing device is configured to acquire the extracted image data as the feature image data.

13. The data creation system according to claim 12, wherein: The first processing device further includes a designating section configured to designate the predetermined extraction area based on the third image data according to an operation command input by a user.

14. A learning system configured to generate a learned model using a learning dataset, the learning dataset including learning data as second image data, the second image data being created by the data creation system according to any one of claims 1 to 11. 15 . An estimation system configured to perform estimation related to an object to be recognized using the learned model generated by the learning system according to claim 14 .

16. A data creation method for creating second image data for use as learning data for generating a learned model based on first image data, the data creation method comprising: an acquiring step for acquiring feature image data indicating respective pixel values ​​of a plurality of pixels included in the feature area; as well as A superimposition step for creating the second image data by superimposing an offset amplitude on the respective pixel values ​​of a plurality of pixels included in a predetermined area of ​​the first image represented by the first image data, wherein the predetermined area has a peripheral shape corresponding to the feature area, and the offset amplitude is determined based on the respective pixel values ​​of the plurality of pixels included in the feature area. 17 . A non-transitory storage medium storing a program designed to cause one or more processors to perform the data creation method according to claim 16 .

18. A processing device used as the first processing device of the data creation system according to claim 12 or 13.

19. A processing device used as the second processing device of the data creation system according to claim 12 or 13.

20. An evaluation system comprising a processing device and a learning system, the processing device being configured to extract, from third image data representing a third image including a pixel area indicating an object to be recognized, extracted image data including respective pixel values ​​of a plurality of pixels included in a predetermined extraction area, and output the extracted image data thus extracted, The learning system is configured to generate a learned model, which is configured to output an estimation result similar to a case where the third image data is the object to be identified in response to a second image represented by second image data or a predetermined area in the second image, wherein the second image data is generated by superimposing an offset amplitude on the individual pixel values ​​of a plurality of pixels included in the predetermined area in the first image represented by the first image data, including a pixel area indicating the object to be identified, the predetermined area having an outer peripheral shape corresponding to the extraction area, and the offset amplitude is determined based on the individual pixel values ​​of the plurality of pixels included in the extraction area.

21. A processing device for use as the processing device of the evaluation system according to claim 20.

22. A learning system for use as a learning system of the evaluation system according to claim 20.

23. An evaluation system comprising a processing device and an estimation system, the processing device being configured to extract, from third image data representing a third image including a pixel area indicating an object to be recognized, extracted image data including respective pixel values ​​of a plurality of pixels included in a predetermined extraction area, and output the extracted image data thus extracted, The estimation system is configured to use the learned model to make an estimation related to the object to be identified, The learned model is configured to output an estimation result similar to an estimation result for a case where the third image data is the object to be identified, in response to a second image represented by second image data or a predetermined area in the second image, wherein the second image data is created by superimposing an offset amplitude on the individual pixel values ​​of a plurality of pixels included in a predetermined area in a first image represented by first image data, including a pixel area indicating the object to be identified, the predetermined area having an outer peripheral shape corresponding to the extraction area, and the offset amplitude is determined based on the individual pixel values ​​of the plurality of pixels included in the extraction area.

24. A processing device for use as the processing device of the evaluation system according to claim 23.

25. An estimation system for use as the estimation system according to claim 23.

Citation Information

Patent Citations

  • Image generation method and image generation system

    JP2017045441A

  • Image processing device and method which use two images

    CN1964668A

  • Data generation device, data generation method, and data generation program

    JP2019114116A