Image processing device, image processing method, and image processing program
The image processing device facilitates on-site visual inspection and re-learning using a head-mounted display, addressing misidentifications by adapting the model to the current environment through direct interaction and data generation, enhancing detection accuracy and efficiency.
Patent Information
- Application Number
- JP2024010605
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-08-07
AI Technical Summary
Existing visual inspection systems using trained inference models face misidentifications due to limited training data and environmental differences, requiring complex re-learning processes that increase on-site workload.
An image processing device and method that allows users to perform visual inspections and re-learning while wearing a head-mounted display, utilizing an image receiving unit, region estimation, learning data generation, and learning unit to adapt the model to the current environment.
Enables visual confirmation of test results and efficient learning processes directly on-site, improving detection accuracy by adapting the model to the actual inspection environment.
Smart Images

Figure 2025115899000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, and an image processing program that support learning processing such as machine learning. [Background technology]
[0002] It is said that 20-30% of labor costs are spent on inspection work at manufacturing sites such as factories. Visual inspections, in particular, involve visual inspections to determine whether a product is defective, and include sensory inspections where standardization of standards is difficult, placing a heavy burden on the physical and mental health of inspectors. To address these issues, visual inspections using AI (Artificial Intelligence) and other technologies are increasingly being used.
[0003] For example, Patent Document 1 discloses an apparatus for performing visual inspection using a neural network model that has been trained in advance by machine learning. Based on image data of an object to be inspected, the trained inference model is used to estimate visual defects. The estimated defects are then displayed superimposed on the object to be inspected on a see-through head-mounted display. This allows inspectors to easily visually identify objects with visual defects. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Utility Model Registration No. 3223248 Summary of the Invention [Problem to be solved by the invention]
[0005] For example, when performing visual inspection using a trained inference model, there are cases where a non-defective part is mistakenly identified as defective, or where a defective part is mistakenly identified as not defective. Because there is a limit to the number and type of training data that can be trained in advance, for example, incorrect identification is likely to occur for unknown defects. Another factor that can cause incorrect identification is that the environment in which the visual inspection is performed is different from the environment during training.
[0006] To reduce such misidentifications, it is desirable to retrain an existing trained model to adapt it to the new inspection environment. For example, by retraining the model by adding new training data to the misidentified objects, it is expected that the detection accuracy will improve.
[0007] However, re-learning requires the complex work of taking a new image of the object to be learned using a camera or other device, importing the image data into a terminal such as a computer, identifying the area to be learned within the imported image, creating learning data, and then performing the learning process again, which increases the workload on-site.
[0008] Therefore, the present invention provides an image processing technique that allows the user to visually check the test results and perform learning and relearning processes while wearing a head-mounted display. [Means for solving the problem]
[0009] The image processing device of the present invention is an image processing device that displays an image on a head-mounted display, and includes an image receiving unit that receives image data captured of the foreground of the head-mounted display, a region estimation unit that estimates a predetermined feature region from the image data using an inference model of a neural network, an image control unit that outputs information related to the feature region, a region selection unit that selects an arbitrary region in the image data, a learning data generation unit that generates an image included in the region selected by the region selection unit as learning data, and a learning unit that performs learning using at least the generated learning data and generates the inference model.
[0010] In addition, the image processing method of the present invention is an image processing method for displaying an image on a head-mounted display, and includes the steps of receiving image data captured of the foreground of the head-mounted display, estimating a predetermined feature region from the image data using a neural network inference model, outputting information about the feature region, selecting an arbitrary region in the image data, generating an image included in the selected region as training data, and performing training using at least the generated training data to generate an inference model. [Effects of the Invention]
[0011] According to the present invention, the test results can be visually confirmed and learning and relearning processes can be performed while the head-mounted display is still attached. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a block diagram showing the overall configuration of an image processing system according to an embodiment. [Figure 2] FIG. 1 is a diagram showing a configuration of a transparent head-mounted display according to an embodiment. [Figure 3] 1 is a flowchart showing the flow of processing in an inspection mode of an image processing apparatus according to an embodiment; [Figure 4] FIG. 1 is a diagram showing a display example of a transparent head-mounted display according to an embodiment; [Figure 5] 1 is a flowchart showing the flow of processing in the re-learning mode of the image processing device according to an embodiment; [Figure 6] FIG. 10 is a diagram showing a display example in a re-learning mode of the transparent head mounted display according to the embodiment; [Figure 7] FIG. 1 is a diagram schematically illustrating learning data according to an embodiment. [Figure 8] 1 is a flowchart showing the flow of processing in a learning mode of an image processing device according to an embodiment; [Figure 9]FIG. 10 is a diagram showing a display example in a learning mode of the transparent head mounted display according to the embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present invention will be described, but the present invention is not limited to the following embodiments.
[0014] (Embodiment) The image processing device according to this embodiment will be described with reference to the drawings.
[0015] [1. Configuration] 1 shows the overall configuration of an image processing system according to this embodiment. A transparent head-mounted display 101, a camera 102, and an input device 103 are connected to an image processing device 100. A user, who is an examiner, wears the transparent head-mounted display 101 and performs an examination.
[0016] The image processing device 100 includes an information receiving unit 104 that receives position information and rotation information from the transparent head-mounted display 101, an image receiving unit 105 that receives image data from the camera 102, an operation control unit 106 that outputs a control signal for controlling the operation of the image processing device 100 based on input from the input device 103, a region estimation unit 107 that estimates a predetermined feature region from the image data using a learned inference model 114, and an image control unit 108 that displays the feature region estimated by the region estimation unit 107 on the transparent head-mounted display 101 via an image output unit 109.
[0017] The image processing device 100 also includes a region selection unit 110 that selects an arbitrary region of image data, a training data generation unit 111 that generates training data, a training data storage unit 112 that stores the training data, and a training unit 113 that performs training processing using a neural network. The training data is data used for training the neural network, and in the case of supervised learning, for example, it is so-called teacher data.
[0018] The image processing device 100 is composed of a memory and a processor, and each block shown inside the image processing device 100 in Fig. 1 is realized by software that operates in cooperation with the memory and processor. It may also be realized as a program that runs on a computer. Note that it is not necessary to realize all of the blocks by software, and some may be realized by hardware.
[0019] The transparent head mounted display 101 is connected to the image processing device 100 by wire or wirelessly. The camera 102 is configured integrally with the transparent head mounted display 101 and captures a scene (foreground) in real space that the user visually recognizes through the transparent head mounted display 101. Note that the camera 102 only needs to be able to capture a scene that the user can visually recognize, and may be configured separately from the transparent head mounted display 101, for example, using a wide-angle camera or an omnidirectional camera.
[0020] The input device 103 is, for example, an input device such as a keyboard or a mouse. Of course, any other input device may be used as long as it can be used in combination with the transparent head-mounted display 101. For example, the user's hand movements may be detected and input may be performed using hand gestures. In this case, the user's hand movements may be detected by a sensor or the like attached to the hand, or the user's hand movements may be photographed and analyzed using a camera.
[0021] When capturing images of the user's hand movements, the camera 102 may be used for capturing the images. In this case, image data captured by the camera 102 is received by the image receiving unit 105, and the hand movements are analyzed by the operation control unit 106 to generate a control signal, thereby substituting for the operation of the input device 103. Furthermore, the operation of the image processing device 100 can be controlled in combination with the input device 103.
[0022] Furthermore, if the transparent head mounted display 101 has a line of sight detection function, input based on the line of sight may also be used. Line of sight information detected by the transparent head mounted display 101 is sent to the operation control unit 106 via the information receiving unit 104. The operation control unit 106 then analyzes the movement of the user's line of sight and converts it into a control signal for controlling the operation of the image processing device 100. This makes it possible to control the operation of the image processing device 100 using the line of sight movement. Furthermore, the operation of the image processing device 100 can also be controlled by combining it with the input device 103, hand gestures, etc.
[0023] 2 shows an example of the configuration of the transmissive head mounted display 101. The transmissive head mounted display 101 includes a half mirror 210, a liquid crystal display unit 211, a gyro sensor 212 that detects the position and rotation of the transmissive head mounted display 101, and an eye gaze sensor 213 that detects the movement of the user's eye 201.
[0024] The user can view the real space through the half mirror 210. Specifically, the user views a foreground 200, which is the scenery in front of the user's eyes. The foreground 200 is also photographed by the camera 102 and input to the image processing device 100. The image processing device 100 acquires position information and rotation information of the transparent head-mounted display 101 from the gyro sensor 212. Based on this position information and rotation information, the image processing device 100 corrects the display position and display angle of the processed image and displays it on the liquid crystal display unit 211. This correction is performed so that the output image of the image processing device 100 is displayed superimposed on the foreground 200. This allows the user to view the image output from the image processing device 100 superimposed on the foreground 200.
[0025] Although the configuration of the see-through head mounted display 101 has been described using an example, the configuration is not limited to this. Any other configuration may be used as long as it has a function of superimposing an image output from the image processing device 100 on the scenery of the real world visually recognized by the user.
[0026] [2. Operation] The operation of the image processing device 100 will be described with reference to the drawings. In this embodiment, image processing for visual inspection of parts used in factories, etc. will be described as an example. The operation of the image processing device 100 can be broadly divided into three modes: inspection mode, re-learning mode, and learning mode. The operation of each mode will be described below.
[0027] [2-1. Inspection mode] The inspection mode is a mode for visually inspecting a part. A pre-trained inference model is used to identify areas containing predetermined features (feature areas). In this embodiment, areas containing defects in the appearance of the part (hereinafter referred to as abnormal areas) are treated as feature areas. Here, defects refer to defects in appearance, such as scratches, dents, hair or other adhesions, distortions, etc.
[0028] 3 is a flowchart showing the processing in the inspection mode of the image processing device 100. First, when a user wears the transparent head mounted display 101, a foreground 200 comes into the user's field of vision. Then, the camera 102 attached to the transparent head mounted display 101 captures an image of the foreground 200, and image data of the captured foreground 200 is input to the image receiving unit 105 (step S301).
[0029] Next, the region estimation unit 107 estimates a feature region using the trained inference model 114 of the neural network (step S302).
[0030] Generally, when a feature region is estimated using a neural network, an inference model is used that has been trained so that the output value of the neural network is high when an image having a predetermined feature is input. By using this inference model, a relatively high numerical value is output when an image having the predetermined feature is input, and a relatively low numerical value is output otherwise. Based on this output value, it is estimated whether the input image contains the predetermined feature.
[0031] In this embodiment, an inference model that has been trained to have a high output value when an image containing a defect is input and a low output value when an image containing no defect is input is used as the trained inference model 114. The output value of the trained inference model 114 is used as an evaluation value, and if the evaluation value exceeds a predetermined value, it is estimated to be an abnormal area containing a defect. The evaluation value is, for example, a value between 0 and 1, and the predetermined value in this case is, for example, 0.5. The evaluation value can also be expressed as a percentage, and the predetermined value in this case is, for example, 50%.
[0032] Learning can also be performed for each pixel. For example, if a neural network is configured to input a 512 x 512 pixel image and output 512 x 512 pixels, learning will be performed so that when an image with a predetermined feature is input, the output value of the area corresponding to the pixel in that feature region will be high. Using an inference model trained in this way, an evaluation value can be calculated for each pixel in the input image. A group of pixels with an evaluation value exceeding a predetermined value is then estimated to be an abnormal region. This makes it possible to estimate the size and shape of the abnormal region.
[0033] The region estimated to be an abnormal region is displayed on the transparent head mounted display 101 by the image control unit 108 (step S303). In FIG. 3, HMD (Head Mounted Display) refers to the transparent head mounted display 101. Here, the image control unit 108 displays the estimated abnormal region on the liquid crystal display unit 211 so as to be superimposed on the foreground 200 on the transparent head mounted display 101. Specifically, the superimposed display is performed by correcting the display position, display angle, etc. of the image indicating the abnormal region based on position information and rotation information of the transparent head mounted display 101. An example display on the transparent head mounted display 101 is shown in FIG. 4.
[0034] 4 shows an example of a display on the transmissive head mounted display 101. In Fig. 4(a) to (d), reference numeral 400 schematically shows an image presented to the user via the half mirror 210 of the transmissive head mounted display 101. That is, it shows a state in which the image transmitted through the half mirror 210 and the image displayed on the liquid crystal display unit 211 are combined.
[0035] 4(a) is a display example when there is no defect in the part 402. Because the part 402 does not contain any defect, there is no output from the image processing device 100, and the user views the foreground 200 including the platform 401 and the part 402 through the half mirror.
[0036] FIG. 4(b) is a display example in which a part 402 contains a defect. Marks 411 and 412 indicate estimated abnormal areas. In this way, the abnormal area is displayed superimposed on the part in real space and can be visually recognized, so the user can clearly identify the location of the abnormal area in the part 402. The marks 411 and 412 are displayed on the liquid crystal display unit 211 of the transmissive head-mounted display 101 by the image processing device 100. These marks are superimposed on the foreground 200 by the half mirror 210 and are visually recognized by the user.
[0037] Mark 413 is a mark indicating that an abnormal area exists on the back side of the component. An abnormal area on the back side of the component 402 cannot be detected from an image captured from only one direction, but the user can rotate the stage 401 while wearing the see-through head-mounted display 101 or move around the periphery of the stage 401 to capture images of the component 402 from various angles. This makes it possible to estimate the abnormal area on the back side of the component 402. Mark 413 is displayed in a different manner from marks 411 and 412 to indicate that an abnormal area exists on the back side of the component.
[0038] Next, the user visually checks whether or not there is a defect in the part 402 using the marks 411 to 413 displayed on the transparent head-mounted display 101 as clues. Since the marks 411 to 413 are displayed to indicate areas that may contain defects, the user can focus on checking the areas indicated by the marks 411 to 413. This allows for more efficient inspection than when all areas are visually checked. If a defect is confirmed visually, the part is determined to be defective and the inspection is terminated.
[0039] If the abnormal region is estimated correctly, the user can make a final judgment simply by visually inspecting it. However, depending on the inspection environment, good estimation results may not be obtained. For example, this may occur when inspecting a large number of unlearned parts or when the work environment has changed. In such cases, the reliability of the evaluation value indicating the abnormal region may decrease, and an area may be estimated as abnormal even though it does not contain any defects. In such a situation, the user must carefully visually inspect the area indicated by the marks 411 to 413 to determine whether it is a true abnormal region, which reduces the efficiency of the inspection work.
[0040] Therefore, in such a case, the user can display auxiliary information on the see-through head-mounted display 101. The auxiliary information is information that serves as a reference for the user when performing a visual inspection.
[0041] If auxiliary information is necessary, the user instructs the image processing device 100 to display the auxiliary information (step S304). In response, the image control unit 108 outputs the auxiliary information together (step S305). This state is shown in FIG. 4(c).
[0042] Figure 4(c) is an example of a display showing auxiliary information about a part. 421 to 423 are evaluation values obtained by converting the output values of the trained inference model 114 into percentages. As mentioned above, the higher the evaluation value, the more likely the part is to contain a defect. In this example, the evaluation values for the areas marked 411 and 413 are high at "90%" and "80%, respectively, so these areas are likely to contain defects. Therefore, the user can make a final determination that a defect is present simply by casually checking visually.
[0043] On the other hand, if the evaluation value is low, there is a possibility that the estimation of the abnormal area is incorrect. In this case, the evaluation value of mark 412 is low at "51%, so there is a possibility that the estimation of this area is incorrect. Therefore, the user carefully checks the area of mark 412 visually. If a defect is confirmed visually, the estimation result is maintained. On the other hand, if no defect is confirmed, the estimation result is corrected. In this way, by referring to the evaluation value, it is possible to prioritize and focus on checking areas that are likely to have been misjudged, allowing the user to perform inspections efficiently.
[0044] Another example of auxiliary information related to parts is shown in 424. Here, the part model number and the number of abnormal areas estimated by the area estimation unit 117 are displayed. By displaying the model number, it is possible to prevent misidentification of parts. Furthermore, if there are many abnormal areas, it is possible to prevent oversight of confirmation by displaying the total number of abnormal areas. By performing an inspection with reference to such auxiliary information, it is possible to improve the efficiency of the work. When the inspection is completed, the process ends (step S306).
[0045] FIG. 4(d) shows an example of an enlarged display of a component. If an enlarged display is required, the user instructs the image processing device 100 to perform an enlarged display (step S304). The image control unit 108 then outputs the enlarged image (step S305). The enlarged image 430 enlarged by the image processing device 100 is displayed on the LCD display unit 211 and superimposed on the half mirror 210 so that it is visible to the user. The magnification ratio of the enlarged image 430 is controlled by the image control unit 108, allowing it to be enlarged and displayed at a desired magnification. For example, as shown in FIG. 4(d), the entire component can be enlarged, or by specifying only the area of the mark 412, only the area of the mark 412 can be enlarged and displayed. This is an effective function when a more thorough inspection is required, such as when the defective part is small. Of course, a reduced display is also possible. This is effective when a bird's-eye view of the entire component is required. The user visually checks the enlarged or reduced image, and when all inspections are complete, the process ends (step S306).
[0046] The auxiliary information described in FIG. 4(c) is not limited to this, and other information such as information necessary for inspection can also be displayed. Although not shown here, the aggregated results according to the level of the evaluation value may be displayed. For example, evaluation value ranks may be defined, and evaluation values over 90% may be classified as Rank A, evaluation values over 70 to 90% as Rank B, and evaluation values over 50 to 70% as Rank C, and the number of abnormal areas for each rank may be displayed. For example, areas with a rank A are extremely likely to contain defects, so parts with a large number of rank A areas are presumed to have a large number of defects. By displaying the number of areas for each rank, the condition of the parts can be confirmed at a glance.
[0047] Although the numerical evaluation value itself is presented to the user as an example of auxiliary information, this is not limiting. It is sufficient to present the user with an approximate magnitude of the evaluation value, and for example, instead of displaying a numerical value, the marks 411 to 413 may be displayed in different colors or shapes for each rank.
[0048] Note that instructions for auxiliary information can be given by any means, such as the input device 103, hand gestures, or gaze input. Although not shown here, when displaying auxiliary information such as that shown in FIG. 4(c), for example, a button labeled "Display auxiliary information" may be displayed on the screen, and the button may be operated by an input device such as a mouse, hand gestures, gaze input, or the like. When wearing the see-through head-mounted display 101, it may be difficult to use input devices such as a mouse or keyboard, so it is desirable to be able to give instructions by hand gestures, gaze input, or the like.
[0049] Similarly, enlarged display can be performed by any means, such as the input device 103, hand gestures, or eye gaze input. When using hand gestures, enlargement or reduction can be achieved by pinching in or out using two fingers, for example. This allows for intuitive and easy-to-understand operations.
[0050] [2-2. Re-learning mode] Next, the re-learning mode will be explained. If the area estimation unit 107 can stably and satisfactorily estimate the abnormal area, there is no need to re-learn the inference model. However, an inference model trained in advance under specified conditions is not omnipotent, and as mentioned above, there are cases where the abnormal area is not estimated correctly when the inspection environment changes, for example. In this case, it is desirable to re-learn the trained inference model. This is because by customizing the trained inference model trained under specified conditions based on the actual environment, it is possible to construct an inference model that is more suited to the actual environment.
[0051] 5 is a flowchart showing the processing of the image processing device 100 in the relearning mode. The processing of the relearning mode is executed following the processing of the inspection mode. This will be described with reference to FIG. 6 as well.
[0052] 6 shows an example of a display on the transmissive head mounted display 101. In Fig. 6(a) to (d), reference numeral 600 schematically shows an image presented to the user via the half mirror 210 of the transmissive head mounted display 101. That is, it shows a state in which the image transmitted through the half mirror 210 and the image displayed on the liquid crystal display unit 211 are combined.
[0053] FIG. 6(a) shows the inspection mode, with abnormal areas displayed by marks 601 and 602. Here, mark 602 has an underline under it because its evaluation value is relatively low at "60%." This is to alert the user to the fact that the evaluation value is lower than a predetermined value (e.g., 70%) and encourage visual confirmation or relearning. FIG. 6(b) shows the state in which part 402 is enlarged. If necessary, the user can display an enlarged image of the part and visually check the area of mark 602 to determine whether or not it contains any defects.
[0054] If it is determined that the area of the mark 602 does not contain any defects, i.e., if the estimation by the area estimation unit 107 is incorrect, the user can switch to re-learning mode and re-learn the trained inference model 114. The switch to re-learning mode is performed based on an instruction from the user. For example, the switch to re-learning mode is performed based on an instruction from the input device 103 or a hand gesture.
[0055] When the system enters the relearning mode, the process shown in Fig. 5 begins. First, the user specifies a target area for relearning using the area selection unit 110 (step S501). The area marked with mark 602 is determined to be an abnormal area, even though it is not actually a defect-free area. Therefore, as shown in Fig. 6(c), the area marked with mark 602 is specified as a target area for relearning, i.e., a learning area. Then, as shown in Fig. 6(d), the specified area is displayed by mark 605.
[0056] The user can specify an area by directly pointing at a component in real space. As described above, the movement of the user's finger 604 is captured by the camera 102 and analyzed by the operation control unit 106 to detect the location where the user is pointing. Since an area can be specified by directly pointing at a component in real space, an intuitive and easy-to-use operation can be realized. The area can also be selected using the gaze sensor 213. For example, the area the user is gazing at may be detected and that area may be selected.
[0057] Once the region is selected, learning data for re-learning is generated (step S502). FIG. 7 is a diagram showing a schematic diagram of the learning data. Reference numeral 700 denotes learning data when the entire component is the learning target. Region 701 is the region selected by region selection unit 110. Region 701 may be designated by the region indicated by mark 602 in inspection mode, or the user may designate the region by drawing an arbitrary rectangle. Alternatively, the user may designate only one point and create a rectangle of a predetermined width and height centered on that point to designate the region. Learning data 700 is then generated by associating region 701 with information indicating that it is a "normal region."
[0058] The training data 700 includes an area 702 that is estimated to be an abnormal area. This abnormal area has been estimated well by the existing trained inference model 114, so there is basically no need to re-learn it. However, if re-learning is performed with a focus on only the area 701, the accuracy of the area estimation that was previously performed well may decrease. In such a case, the area that has been estimated well may also be included in the targets of re-learning. This is expected to suppress the decrease in accuracy. In this case, the training data 700 may be generated by associating area 701 with information indicating that it is a "normal area" and area 702 with information indicating that it is an "abnormal area."
[0059] Although the regions 701 and 702 have been described as rectangular regions, the target regions may be specified in more detail, and the defective regions themselves may be specified. For example, the regions 703 and 704 in FIG. 7(b) may be specified as regions to be re-learned.
[0060] In this embodiment, the case where the entire part is the learning target has been described, but for example, only the defective part may be the learning target. In this case, learning data can be generated by using areas 701 and 702 in Figure 7(b) as learning data, associating area 703 with information indicating that it is a "normal area," and associating area 704 with information indicating that it is an "abnormal area."
[0061] When the generation of the learning data is completed, the learning unit 113 performs a re-learning process (step S503). Re-learning is started by pressing a button 606 on the screen. The button 606 may be pressed using the input device 103, or may be operated by hand gestures, eye gaze input, or the like.
[0062] The re-learning process is performed using the trained inference model 114 as the initial value and the new learning data generated in step S502. The trained inference model 114 is then updated with the inference model after re-learning (step S504). Once all re-learning is complete, the re-learning mode process ends (step S505). By using the updated trained inference model 114 in the inspection mode, it is expected that the area indicated by the mark 602 will be estimated as a normal area.
[0063] In this embodiment, it has been described that re-learning is performed using the new training data generated in step S502, but re-learning may also be performed using training data that has already been trained. This is effective when it is not desired to make major changes to the existing inference model. In this case, this can be achieved by reading out the existing training data from the training data storage unit 112 and performing re-learning together with the new training data.
[0064] [2-3. Learning Mode] Next, we will explain the learning mode. Up until now, we have been assuming that there is an already trained inference model 114. The learning mode is a mode in which new learning data is created and an inference model is generated based on this new learning data.
[0065] 8 is a flowchart showing the processing in the learning mode of the image processing device 100. The processing in the learning mode is performed with the user wearing the transparent head mounted display 101. When the processing in the learning mode starts, first, the area selection unit 110 specifies an area to which a learning tag is to be attached (step S801). Next, the learning data generation unit 111 associates tag information (step S802). Then, learning data is generated (step S803). This process will be described with reference to FIG. 9.
[0066] Fig. 9 shows an example of a display on the transparent head mounted display 101. Reference numeral 900 schematically shows an image presented to the user via the half mirror 210 of the transparent head mounted display 101. Fig. 9(a) shows a state in which a part 901 in real space is visible. Here, an example will be described in which learning data is generated by associating a defective area 902 with tag information indicating that the area is "defective."
[0067] First, the area designation (step S801) will be described. FIG. 9(b) shows how an area is designated by a hand gesture. Finger 904 is the user's finger in real space, and shows how finger 904 is pointing near defect 902. When the operation control unit 106 analyzes the operation of the user's finger, area selection unit 110 designates an area within a predetermined range from the tip of finger 904 as a nearby area. Then, image control unit 108 displays an image showing this nearby area 903.
[0068] Next, a defective area 905 is identified from this nearby area 903, as shown in FIG. 9(c). The defective area 905 can be identified by applying a technique such as area division to the nearby area 903. This allows the user to automatically perform the process up to identifying the defective area 905 by simply pointing to the area near the defect 902. Note that although the nearby area 903 has been described as the target of area division here, the entire screen may be the target of area division without providing a nearby area. The defective area 905 can also be manually identified by finger movement, the input device 103, or the like.
[0069] When the defective area 905 is identified, it is tagged for learning (step S802). In this embodiment, tag information indicating "defective" is associated with the defective area 905. Specifically, a "class label" that classifies the type of defect is associated with the defective area. The "class label" is expressed as a numerical value according to the type of defect, such as "0" for normal, "1" for a scratch, or "2" for a dent.
[0070] Once tagging is complete, learning data is generated (step S803). Here, the explanation has been given based on the image displayed on the transparent head-mounted display 101, but in reality, the identification of the defective area 905 and the association of tag information are performed on the image data captured by the camera 102. The tag information is input, for example, from the input device 103. Then, the image data associated with the tag information is converted to a predetermined size and generated as learning data. Here, an image corresponding to 900 in FIG. 9(c) is generated as learning data.
[0071] Although the entire component is the learning target here, if only the defective region is the learning target, for example, only the neighboring region 903 or the defective region 905 may be output as learning data. The generated learning data is held in the learning data holding unit 112.
[0072] When the creation of the learning data is completed (step S804), the learning unit 113 performs a learning process for the neural network using the learning data stored in the learning data storage unit 112 (step S805). Then, when the learning is completed, a trained inference model 114 is created (step S806), and the learning mode ends.
[0073] 9(a) to 9(c), the vicinity area 903 and the defect area 905 are superimposed images output from the image processing device 100, and the rest are transparently displayed images in real space. However, since the images in real space are captured by the camera 102, some or all of the objects in real space can also be output from the image processing device 100.
[0074] Furthermore, although the learning mode has been described as a mode in which learning data is created and new learning is performed, it can also be applied, for example, when performing re-learning. In the above-mentioned re-learning mode, re-learning is performed on the area identified in the inspection mode. However, if a defect is found in an area where no abnormal area has been estimated, for example, in the state of FIG. 4(b), learning data can be generated using steps S801 to S803 of this mode. This makes it possible to re-learn not only for areas where a location that is not actually defective is estimated to be defective, but also for areas that are actually defective but are estimated to be non-defective.
[0075] As described above, the embodiments have been described as examples of the present invention. However, the present invention is not limited to these, and can be applied to embodiments in which modifications, substitutions, additions, omissions, etc. are made. For example, the image processing device of the present invention can achieve similar operations even if the transmissive head-mounted display 101 in FIG. 1 is replaced with a non-transmissive head-mounted display. In the embodiments described above, information about the characteristic region and the foreground 200 are combined using the half mirror 210. In contrast, when a non-transmissive head-mounted display is used, the image control unit 108 combines the information about the characteristic region onto the image of the foreground 200 so as to superimpose it. This operation will be described with reference to FIG. 1.
[0076] Image receiving unit 105 receives an image of foreground 200 from camera 102 mounted on head-mounted display 101. Image control unit 108 combines information about the characteristic region estimated by region estimation unit 107 with the image of foreground 200, and outputs the combined image to head-mounted display 101 via image output unit 109. This allows the user to visually recognize the image in which the information about the characteristic region is superimposed on foreground 200. Other operations are the same as those in the embodiments described above, and can be similarly implemented for any of the operations in the inspection mode, relearning mode, and learning mode described with reference to FIGS. 3 to 9.
[0077] With this configuration, the user can perform the inspection and learning process while wearing the head-mounted display.
[0078] [3. Effects, etc.] The image processing device 100 of the present invention is an image processing device that displays an image on a head-mounted display 101, and includes an image receiving unit 105 that receives image data captured of a scene in front of the head-mounted display 101, a region estimation unit 107 that estimates a predetermined feature region from the image data using a neural network inference model 114, an image control unit 108 that outputs information regarding the feature region, a region selection unit 110 that selects an arbitrary region in the image data, a training data generation unit 111 that generates images included in the region selected by the region selection unit 110 as training data, and a learning unit 113 that performs training using at least the training data generated by the training data generation unit 111 to generate an inference model. Since inspection and training can be performed while referring to the image displayed on the head-mounted display 101, work efficiency can be improved. Furthermore, by adding objects in an actual work site to the training data, an inference model adapted to the real environment can be generated.
[0079] The image control unit 108 may output display data in which information about the characteristic region is superimposed on the image data received by the image receiving unit 105. This causes information about the estimated characteristic region to be superimposed on the foreground, allowing tasks such as testing and learning to be performed while wearing the head-mounted display.
[0080] The head-mounted display 101 may also be a see-through head-mounted display. In this case, the image control unit 108 outputs display data for superimposing information about the characteristic region on a foreground that is transparently displayed on the see-through head-mounted display. Since the user can directly view the foreground, the user can perform work in a state closer to reality.
[0081] The system may further include a learning data holding unit 112 that holds learned learning data, and the learning unit 113 may perform learning on a trained inference model 114 using the learning data held in the learning data holding unit 112 and the learning data generated by the learning data generation unit 111. By performing learning based on an existing inference model, it is possible to perform existing region estimation well while preventing erroneous determination in specific situations.
[0082] The image control unit 108 may output the display data including the area selected by the area selection unit 110. The selected area is displayed as mark 605 in FIG. 6(d), so that the selected area can be confirmed.
[0083] The image control unit 108 may enlarge or reduce the display data before outputting it. If the defect is small or if you want to check the condition of the part more carefully, you can enlarge the image and display it. You can also get an overview of the entire part by reducing the image and displaying it.
[0084] The image control unit 108 may use the output of the inference model 114 as an evaluation value and output this evaluation value as part of the display data. By referring to the evaluation value of the inference model, it is possible to confirm the accuracy of the estimation in each estimated area. Furthermore, by referring to this evaluation value, it is possible to determine whether re-learning is necessary.
[0085] The image receiving unit 105 may further receive image data of the user's actions, and may further include an action control unit 106 that analyzes the image data of the user's actions and outputs control information, and the area selecting unit 110 may select an area based on the control information. For example, since the user can operate the image processing device 100 by hand gestures, a series of operations can be performed while wearing the head-mounted display 101.
[0086] The information receiving unit 104 may further receive gaze information of the user from the head mounted display 101, and the area selecting unit 110 may select an area based on the gaze information. Since gaze input can be used as part of the interface, area selection operation becomes easier while wearing the head mounted display 101.
[0087] The image processing method of this embodiment is an image processing method for displaying an image on a head-mounted display, and includes the steps of receiving image data of a scene in front of the head-mounted display, estimating a predetermined feature region from the image data using a neural network inference model, outputting information about the feature region, selecting an arbitrary region in the image data, generating an image included in the selected region as training data, and performing training using at least the generated training data to generate an inference model. Since inspection and training can be performed while referring to the image displayed on the head-mounted display 101, work efficiency can be improved. Furthermore, by adding objects in an actual work site to the training data, an inference model adapted to the real environment can be generated.
[0088] The image processing program of this embodiment is an image processing program that causes a computer to execute an image processing method for displaying an image on a head-mounted display, and executes the following steps: receiving image data of a scene in front of the head-mounted display; estimating a predetermined feature region from the image data using a neural network inference model; outputting information about the feature region; selecting an arbitrary region from the image data; generating an image included in the selected region as training data; and performing training using at least the generated training data to generate an inference model. Since inspection and training can be performed while referring to the image displayed on the head-mounted display 101, work efficiency can be improved. Furthermore, by adding objects in an actual work site to the training data, an inference model adapted to the real environment can be generated. [Explanation of symbols]
[0089] 100 Image processing device 101 Head-mounted display (transparent or non-transparent) 102 Camera 103 Input Devices 104 Information Receiving Unit 105 Image receiving unit 106 Motion control section 107 Area estimation part 108 Image control unit 109 Image output unit 110 Area selection section 111 Learning data generation unit 112 Learning data storage unit 113 Learning Department 114 Trained Inference Models 200 Foreground 201 Eye 210 Half Mirror 211 LCD display section 212 Gyro Sensor 213 Eye Sensor 400, 600, 700, 900 statues 401 units 402, 901 parts 411~413, 601, 602 marks 421~423 rating 424 Supplementary Information 430, 603 Enlarged image 604, 904 fingers 605 Nearby Area 606 Button 701~704 area 902 Bad 903 Nearby Area 905 Bad area
Claims
1. An image processing device for displaying an image on a head-mounted display, an image receiving unit that receives image data of a scene in front of the head mounted display; and an area estimating unit that estimates a predetermined feature area from the image data using an inference model of a neural network; an image control unit that outputs information about the characteristic region; an area selection unit for selecting an arbitrary area in the image data; a learning data generation unit that generates an image included in the selected region as learning data; An image processing device comprising: a learning unit that performs learning using at least the learning data and generates the inference model.
2. The image processing device according to claim 1 , wherein the image control unit outputs display data in which information about the characteristic region is superimposed on the image data received by the image receiving unit.
3. the head-mounted display is a see-through head-mounted display, The image processing device according to claim 1 , wherein the image control unit outputs display data for superimposing information about the characteristic region on the foreground that is transparently displayed on the transparent head-mounted display.
4. further comprising a learning data storage unit for storing learned learning data; the inference model is a trained inference model trained using the training data stored in the training data storage unit, The image processing device according to claim 2 or 3, wherein the learning unit performs learning on the trained inference model using the learning data stored in the learning data storage unit and the learning data generated by the learning data generation unit.
5. The image processing device according to claim 4 , wherein the image control unit outputs the display data including the area selected by the area selection unit.
6. The image processing device according to claim 5 , wherein the image control unit enlarges or reduces the display data before outputting it.
7. The image processing device according to claim 5 , wherein the image control unit uses the output of the inference model as an evaluation value and outputs the evaluation value by including it in the display data.
8. The image receiving unit further receives image data of a user's action, further comprising an operation control unit that analyzes image data obtained by capturing the user's operation and outputs control information; The image processing device according to claim 4 , wherein the region selection unit selects the region based on the control information.
9. the information receiving unit further receives gaze information of the user from the head-mounted display; The image processing device according to claim 4 , wherein the area selection unit selects the area based on the line-of-sight information.
10. An image processing method for displaying an image on a head-mounted display, comprising: receiving image data capturing a scene in front of the head-mounted display; estimating a predetermined feature region from the image data using a neural network inference model; outputting information about the feature region; selecting an arbitrary region in the image data; generating an image included in the selected region as learning data; A step of performing learning using at least the learning data and generating the inference model. An image processing method comprising:
11. An image processing program that causes a computer to execute an image processing method for displaying an image on a head-mounted display, receiving image data capturing a scene in front of the head-mounted display; estimating a predetermined feature region from the image data using a neural network inference model; outputting information about the feature region; selecting an arbitrary region in the image data; generating an image included in the selected region as learning data; A step of performing learning using at least the learning data and generating the inference model; An image processing program that executes the following.
Citation Information
Patent Citations
Visual defect inspection device
JP3223248U