Image filter generation system, image filter generation device, learning device, and learning method
By generating an image filter generation system, a well-trained model is generated using machine learning to infer suitable combinations of image filters and parameters. This solves the problem of text recognition errors caused by environmental fluctuations in OCR and improves recognition accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MITSUBISHI ELECTRIC CORP
- Filing Date
- 2022-02-15
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies in OCR struggle to cope with fluctuations in the factory environment, such as light exposure, shooting position and angle deviations, and individual differences in workpieces, resulting in a high text recognition error rate.
An image filter generation system utilizes machine learning to generate a trained model that infers suitable combinations of image filters and parameters, thereby reducing text recognition errors. The system includes a visual sensor, a learning device, and an inference device. It uses machine learning to generate a trained model and infer suitable combinations of image filters and parameters.
It improves the accuracy of text recognition in OCR and reduces recognition errors caused by environmental fluctuations.
Smart Images

Figure CN118451479B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an image filter generation system, an image filter generation device, a learning device, a learning method, and a program. Background Technology
[0002] Previously, it was known that in so-called OCR (Optical Character Recognition / Reader), which recognizes text recorded on a workpiece by a camera, a device that has been learned by machine learning during image processing of the captured image of the workpiece was used.
[0003] Patent Document 1 discloses an image processing apparatus that uses a neural network for a sequential planning unit, which outputs a sequential plan of image transformation filters applied in image processing. In Patent Document 1, a learning control unit enables the neural network to learn by learning pairs of learning data pairs, consisting of learning images and sequential patterns that can combine image transformation filter sets. Specifically, the learning control unit enables the neural network to learn by feeding back an error, i.e., a loss, calculated based on the following information: the sequential plan output by inputting the learning images contained in the learning data pairs into the neural network; and the sequential patterns contained in the learning data pairs.
[0004] Patent Document 2 discloses an image correction apparatus for generating an appropriate image for an input captured image. In Patent Document 2, a statistical learning rule is constructed by learning the parameters of a spatial filter that sets a small region image segmented from a sample image as the appropriate image, using these parameters as teaching values. Furthermore, in Patent Document 2, the small region image is corrected using a spatial filter created based on parameters output from the statistical learning rule by inputting the pixel values of the pixels contained in the small region image. The small region image is obtained by segmenting the input captured image.
[0005] Patent Document 1: Japanese Patent Application Publication No. 2020-154600
[0006] Patent Document 2: Japanese Patent Application Publication No. 2009-10853 Summary of the Invention
[0007] In the devices described in Patent Documents 1 and 2, only predetermined, highly reliable image filters are prepared. After learning their combinations and parameters, the most suitable combination and parameters of image filters for image processing of the input image are inferred. Therefore, the devices described in Patent Documents 1 and 2 cannot cope with environmental fluctuations that occur in the actual application of OCR, such as light entering from factory windows at different times of day, deviations in the position, orientation, and rotation angle of the workpiece being photographed, and individual differences in the workpiece, which may lead to misrecognition of text.
[0008] This invention is proposed in view of the above-mentioned actual situation, and its purpose is to reduce the misrecognition of characters.
[0009] To achieve the above objectives, the present invention relates to an image filter generation system for generating image filters used in image processing of object image data prior to OCR, the object image data being image data of an object captured by a shooting unit. The image filter generation system includes: an image filter generation device that generates image filters; a learning device that learns the correlation between pre-acquired object image data and image filters used in image processing of the object image data; and an inference device that infers suitable image filters for image processing of the object image data for OCR. The learning device includes: a learning data acquisition unit that acquires learning data including object image data, image filter correlation data, and OCR score data, the image filter correlation data being data representing combinations of image filters used in image processing of the object image data and the values of parameters for each image filter, and the OCR score data being data representing the score of text recognition output by OCR when image processing of the object image data is performed using image filters based on the image filter correlation data; a trained model generation unit that generates a trained model representing the correlation between the object image data, image filter correlation data, and OCR score data using machine learning with the learning data; and a trained model output unit that outputs the trained model. The inference apparatus includes: an object image data acquisition unit that acquires object image data for OCR; an inference result data generation unit that inputs the object image data for OCR into a trained model and generates inference result data, representing a suitable combination of image filters for image processing of the object image data for OCR and the values of parameters for each image filter; and an inference result data output unit that outputs the inference result data. The image filter generation apparatus includes: an image filter generation unit that generates image filters based on the inference result data; and an image filter output unit that outputs the image filters.
[0010] The effects of the invention
[0011] According to the present invention, the learning device generates a well-trained model representing the correlation between object image data, image filter correlation data, and OCR score data. Therefore, the inference device, by inputting the object image data subjected to OCR into the well-trained model, can generate and output inference result data that is inferred to be the highest-scoring text recognition output during OCR. As a result, compared to an image filter generation system that does not generate a well-trained model representing the correlation between object image data, image filter correlation data, and OCR score data, the image filter generation system of the present invention can reduce text misrecognition. Attached Figure Description
[0012] Figure 1 This is an overall illustration of the image filter generation system according to Implementation Method 1.
[0013] Figure 2 This is a diagram illustrating the functional structure of the image filter generation system according to Implementation Method 1.
[0014] Figure 3 This is a block diagram illustrating the hardware structure of each device involved in Implementation 1.
[0015] Figure 4 This is an explanatory diagram of the learning data involved in Implementation Method 1.
[0016] Figure 5 This is a diagram illustrating an overview of the processing of the output inference result data involved in Implementation 1.
[0017] Figure 6 This is a flowchart of the trained model generation process involved in Implementation Method 1.
[0018] Figure 7 This is a flowchart of the inference result data generation process involved in Implementation Method 1.
[0019] Figure 8 This is a flowchart of the image filter generation process involved in Implementation Method 1.
[0020] Figure 9 This is a diagram illustrating the function of the image filter generation system involved in Implementation Method 1.
[0021] Figure 10 This is a flowchart of the inference result data generation process involved in Implementation Method 2.
[0022] Figure 11 This is a flowchart of the image filter generation process involved in Implementation Method 2.
[0023] Figure 12 This is a diagram illustrating the functional structure of the vision sensor involved in Implementation Method 3. Detailed Implementation
[0024] Hereinafter, the image filter generation system, image filter generation apparatus, inference apparatus, inference method, and program related to the implementation of the present invention will be described in detail with reference to the accompanying drawings. Furthermore, the same or corresponding parts in the drawings are labeled with the same reference numerals.
[0025] [Implementation Method 1]
[0026] (Regarding the image filter generation system 1 involved in Implementation Method 1)
[0027] The image filter generation system 1 according to Embodiment 1 of the present invention is, for example, a system for generating image filters used for image processing before performing OCR (Optical Character Recognition / Reader) on image data, wherein the image data is obtained by photographing so-called workpieces such as products or parts produced in a factory.
[0028] like Figure 1 As shown, the image filter generation system 1 includes a vision sensor 100, which is an example of an imaging device and also an example of an image filter generation device that generates image filters used for image processing of the captured image data. Furthermore, the image filter generation system 1 includes a learning device 200 that learns the correlation between pre-acquired image data of an object (i.e., object image data) and the image filters used for image processing of the object image data before OCR. Additionally, the image filter generation system 1 includes an inference device 300 that infers image filters adapted to the image processing of the object image data for OCR. Finally, the image filter generation system 1 includes a storage device 400 for storing data. The vision sensor 100, learning device 200, inference device 300, and storage device 400 can transmit and receive data via a LAN (Local Area Network) not shown.
[0029] In the image filter generation system 1, firstly, the vision sensor 100 performs image processing on pre-generated image filters using pre-captured object image data, attempting OCR. Secondly, the vision sensor 100 outputs learning data based on the attempt results of the OCR-generated object image data to the learning device 200, which generates a trained model using machine learning with the acquired learning data. Thirdly, the learning device 200 outputs the generated trained model and stores it in the storage device 400, and the inference device 300 retrieves the trained model stored in the storage device 400.
[0030] Furthermore, when the vision sensor 100 actually captures an object for OCR, it outputs object image data to the inference device 300. The inference device 300 then inputs the acquired object image data into a training model to generate inference result data, representing the inference result of an image filter adapted to the image processing of the object image data, and outputs this data to the vision sensor 100. The vision sensor 100 then generates an image filter based on the acquired inference result data, performs image processing on the object image data using the generated image filter, and then performs OCR.
[0031] (Regarding the visual sensor 100 involved in Implementation Method 1)
[0032] like Figure 2 As shown, the vision sensor 100 includes a camera 110, which is an example of a capturing component for photographing objects. Additionally, the vision sensor 100 includes an image filter association data generation unit 120, which generates image filter association data representing combinations of multiple types of image filters and the parameters of each image filter. Furthermore, the vision sensor 100 includes an image filter generation unit 130 for generating image filters, an image filter output unit 140 for outputting image filters, an image processing unit 150 for performing image processing, and an OCR unit 160 for performing OCR. Additionally, the vision sensor 100 includes an object image data output unit 170 for outputting object image data, a learning data output unit 180 for outputting learning data, and an inference result data acquisition unit 190 for acquiring inference result data.
[0033] (Regarding the learning device 200 involved in Implementation Method 1)
[0034] The learning device 200 is, for example, a computer device such as a personal computer, a server computer, or a supercomputer. The learning device 200 includes a learning data acquisition unit 210 for acquiring learning data, a trained model generation unit 220 for generating a trained model, and a trained model output unit 230 for outputting the trained model. The trained model generation unit 220 includes a reward calculation unit 221 for calculating the reward (described later) and a value function update unit 222 for updating the value function (described later).
[0035] (Regarding the inference device 300 involved in Embodiment 1)
[0036] The inference device 300 is the same computer device as the learning device 200. The inference device 300 includes a trained model acquisition unit 310 for acquiring a trained model, an object image data acquisition unit 320 for acquiring object image data, an inference result data generation unit 330 for generating inference result data, and an inference result data output unit 340 for outputting inference result data.
[0037] (Regarding the storage device 400 involved in Embodiment 1)
[0038] Storage device 400 is, for example, an HDD (Hard Disk Drive) or NAS (Network Attached Storage) connected to a communication network via a LAN. Storage device 400 includes a trained model storage unit 410 for storing trained models.
[0039] (Regarding the hardware structure of the learning device 200 according to Embodiment 1)
[0040] like Figure 3 As shown, the learning device 200 has a control unit 51 that executes processing according to the control program 59. The control unit 51 has a CPU (Central Processing Unit). The control unit 51 performs processing according to the control program 59 as... Figure 2 The trained model generation unit 220, the reward calculation unit 221, and the value function update unit 222 shown in the diagram are in operation.
[0041] return Figure 3 The learning device 200 has a main storage unit 52 used as the working area of the control unit 51 for loading control program 59. The main storage unit 52 has RAM (Random Access Memory).
[0042] In addition, the learning device 200 has an external storage unit 53 that stores the control program 59 in advance. The external storage unit 53 supplies the data stored in the program to the control unit 51 according to the instructions of the control unit 51, and stores the data supplied from the control unit 51. The external storage unit 53 may be a non-volatile memory such as flash memory, HDD (Hard Disk Drive), or SSD (Solid State Drive).
[0043] In addition, the learning device 200 has an operation unit 54 operated by the user. Input information is supplied to the control unit 51 via the operation unit 54. The operation unit 54 includes information input components such as a keyboard, mouse, and touch panel.
[0044] In addition, the learning device 200 includes a display unit 55 that displays information input via the operation unit 54 and information output via the control unit 51. The display unit 55 includes display devices such as LCD (Liquid Crystal Display) and organic EL (Electro-Luminescence) displays.
[0045] return Figure 3 The learning device 200 includes a transceiver unit 56 for transmitting and receiving information. The transceiver unit 56 includes information communication components such as a communication network terminal device and a wireless communication device that are connected to a network. The transceiver unit 56 serves as… Figure 2 The data acquisition unit 210 and the trained model output unit 230 shown in the diagram are in operation.
[0046] return Figure 3 In the learning device 200, the main storage unit 52, the external storage unit 53, the operation unit 54, the display unit 55, and the transceiver unit 56 are all connected to the control unit 51 via the internal bus 50.
[0047] The learning device 200 utilizes the main storage unit 52, external storage unit 53, operation unit 54, display unit 55, and transceiver unit 56 as resources through the control unit 51, thereby achieving... Figure 2 The functions of the aforementioned parts 210, 220-222, and 230 are as follows: For example, the learning device 200 performs the learning data acquisition step performed by the learning data acquisition unit 210. Additionally, for example, the learning device 200 performs the trained model generation step performed by the trained model generation unit 220, the reward calculation step performed by the reward calculation unit 221, and the value function update step performed by the value function update unit 222. Furthermore, for example, the learning device 200 performs the trained model output step performed by the trained model output unit 230.
[0048] (Regarding the hardware structure of the inference device 300 involved in Implementation 1)
[0049] In addition, such as Figure 3 As shown, the inference device 300, like the learning device 200, also includes a control unit 51, a main storage unit 52, an external storage unit 53, an operation unit 54, a display unit 55, and a transceiver unit 56. The control unit 51 operates according to the control program 59. Figure 2 The inference result data generation unit 330 shown is functioning. Additionally, the transceiver unit 56 functions as... Figure 2 The trained model acquisition unit 310, the object image data acquisition unit 320, and the inference result data output unit 340 shown are in operation.
[0050] return Figure 3The inference device 300 utilizes the main storage unit 52, external storage unit 53, operation unit 54, display unit 55, and transceiver unit 56 as resources through the control unit 51, thereby achieving... Figure 2 The functions of the aforementioned parts 310 to 330 are as follows. For example, the inference device 300 performs the trained model acquisition step performed by the trained model acquisition unit 310, the item image data acquisition step performed by the item image data acquisition unit 320, the inference result data generation step performed by the inference result data generation unit 330, and the inference result data output step performed by the inference result data output unit 340.
[0051] (Regarding the hardware structure of the vision sensor 100 involved in Embodiment 1)
[0052] Additionally, although not illustrated, the vision sensor 100 includes a control unit 51, a main storage unit 52, an external storage unit 53, an operation unit 54, and a transceiver unit 56. The control unit 51 operates according to the control program 59. Figure 2 The image filter associated data generation unit 120, image filter generation unit 130, image filter output unit 140, image processing unit 150, and OCR unit 160 shown in the diagram are functional. Additionally, the transceiver unit 56 functions as... Figure 2 The item image data output unit 170, the learning data output unit 180, and the inference result data acquisition unit 190 are all in operation.
[0053] return Figure 3 The vision sensor 100 utilizes the main storage unit 52, external storage unit 53, operation unit 54, and transceiver unit 56 as resources through the control unit 51, thereby achieving... Figure 2 The functions of the aforementioned parts 120 to 190 are as follows. For example, the vision sensor 100 performs the image filter association data generation step performed by the image filter association data generation unit 120, the image filter generation step performed by the image filter generation unit 130, and the image filter output step performed by the image filter output unit 140. Additionally, for example, the vision sensor 100 performs the image processing step performed by the image processing unit 150 and the OCR step performed by the OCR unit 160. Furthermore, for example, the vision sensor 100 performs the object image data output step performed by the object image data output unit 170, the learning data output step performed by the learning data output unit 180, and the inference result data acquisition step performed by the inference result data acquisition unit 190.
[0054] (Details regarding the functional structure of the vision sensor 100 according to Embodiment 1)
[0055] return Figure 2Camera 110 photographs a workpiece that is within the allowable range of predetermined design values, serving as an example of an article, thus generating article image data. Here, camera 110 can photograph workpieces transported during manufacturing via an actual manufacturing production line, workpieces transported via a manufacturing production line similar to the actual production line, or workpieces photographed in a simulated manufacturing environment. Furthermore, when photographing workpieces in a simulated environment, camera 110 can, for example, simulate fluctuations in the environment envisioned during manufacturing—specifically, time periods such as morning, noon, and evening, the orientation of the transported workpiece, and the angle of rotation—and thus generate multiple types of article image data.
[0056] When attempting to perform OCR on object image data pre-acquired from camera 110, image filter association data generation unit 120 generates image filter association data for the image filters used in image processing. Furthermore, the object image data for which OCR is attempted includes object image data of workpieces actually captured during past manufacturing processes and object image data of workpieces captured in a simulated environment. Additionally, the combination of multiple types of image filters represented by the image filter association data is, for example, a combination of multiple types of image filters selected from known image filters such as binarization, dilation, shrinkage, smoothing filters, noise removal filters, contour extraction filters, high-pass filters, low-pass filters, cropping, and edge enhancement filters. Furthermore, the parameters of each image filter represented by the image filter association data are, for example, a combination of values of multiple types of parameters selected from known parameters such as threshold, kernel size, gain, maximum value, and minimum value.
[0057] Furthermore, the image filter association data generation unit 120 may generate image filter association data based on image filters actually used in past manufacturing processes. Alternatively, the image filter association data generation unit 120 may use random numbers to select combinations of image filters and parameters of each image filter to generate image filter association data.
[0058] The image filter generation unit 130 generates an image filter based on image filter association data. Here, for example, we will discuss the following case: the combination of image filters shown in the image filter association data is a combination of a noise removal filter and a contour extraction filter, where the parameter of the noise removal filter is a first parameter and the parameter of the contour extraction filter is a second parameter. In this case, the image filter generation unit 130 generates an image filter that combines the noise removal filter with the first parameter set and the contour extraction filter with the second parameter set.
[0059] When the image filter generation unit 130 generates an image filter, the image filter output unit 140 outputs the generated image filter to the image processing unit 150.
[0060] The image processing unit 150 uses the image filter obtained from the image filter output unit 140 to perform image processing on the item image data.
[0061] The OCR unit 160 performs OCR on the image data of the object after image processing and outputs a score that represents the reliability of character recognition.
[0062] The object image data output unit 170 outputs the object image data obtained by the camera 110 and subjected to OCR to the inference device 300.
[0063] The learning data output unit 180 outputs learning data to the learning device 200. Here, the learning data includes object image data for which OCR was attempted, and image filter association data that determines the image filter used in the image processing of the object image data. Furthermore, the learning data also includes OCR score data, representing the score output after attempting OCR through image processing of the object image data using an image filter based on the image filter association data. Therefore, the learning data comprises object image data related to the workpiece during past manufacturing, image filter association data, and OCR score data.
[0064] Here, the processing of learning data generated by the vision sensor 100 in order for the learning data output unit 180 to output learning data to the learning device 200 will be explained. First, as Figure 4 As shown, when the natural number is set to m, the m types of object image data generated by the camera 110 are set as IMG-1, IMG-2, ..., IMG-m. Similarly, when the natural number is set to n, the n types of image filter association data generated by the image filter association data generation unit 120 are set as F / P-001, F / P-002, ..., F / P-00n. The image filter generation unit 130 generates n types of image filters based on the n types of image filter association data F / P-001, F / P-002, ..., F / P-00n. Furthermore, the image processing unit 150 performs image processing on each object image data IMG-1, IMG-2, ..., IMG-m using the n types of image filters, and the OCR unit 160 performs OCR on the processed object image data of m×n types, outputting m×n types of OCR score data.
[0065] Here, the OCR score data of the first item image data IMG-1, which underwent image processing and OCR based on the image filters associated with each image filter F / P-001, F / P-002, ..., F / P-00n, is designated as IMG-1_F / P-001, IMG-1_F / P-002, ..., IMG-1_F / P-00n. Similarly, the OCR score data of the second item image data IMG-2, which underwent image processing and OCR based on the image filters associated with each image filter F / P-001, F / P-002, ..., F / P-00n, is designated as IMG-2_F / P-001, IMG-2_F / P-002, ..., IMG-2_F / P-00n. In addition, the OCR score data of the m-th item image data IMG-m, which has undergone image processing and OCR based on the image filters associated with each image filter F / P-001, F / P-002, ..., F / P-00n, is set as IMG-m_F / P-001, IMG-m_F / P-002, ..., IMG-m_F / P-00n.
[0066] As a result, the learning data output unit 180 outputs data containing m types of item image data IMG-1, IMG-2, ..., IMG-m, n types of image filter association data F / P-001, F / P-002, ..., F / P-00n, m×n types of OCR score data IMG-1_F / P-001, IMG-1_F / P-002, ..., IMG-1_F / P-00n, IMG-2_F / P-001, IMG-2_F / P-002, ..., IMG-2_F / P-00n, ..., IMG-m_F / P-001, IMG-m_F / P-002, ..., IMG-m_F / P-00n as learning data.
[0067] return Figure 2 The inference result data acquisition unit 190 acquires the inference result data output from the inference device 300. Meanwhile, the image filter generation unit 130 generates an image filter based on the inference result data, and the image filter output unit 140 outputs the image filter to the image processing unit 150. Furthermore, the image processing unit 150 performs image processing on the acquired image filter, and the OCR unit 160 performs OCR on the processed image data.
[0068] (Details regarding the functional structure of the learning device 200 according to Embodiment 1)
[0069] The learning data acquisition unit 210 acquires learning data output from the vision sensor 100. For example, the learning data acquisition unit 210 acquires image data of m types of objects (IMG-1, IMG-2, ..., IMG-m, n types of image filter association data (F / P-001, F / P-002, ..., F / P-00n) and OCR score data of m×n types (IMG-1_F / P-001, IMG-1_F / P-002, ..., IMG-1_F / P-00n, IMG-2_F / P-001, IMG-2_F / P-002, ..., IMG-2_F / P-00n, ..., IMG-m_F / P-001, IMG-m_F / P-002, ..., IMG-m_F / P-00n) as learning data.
[0070] The trained model generation unit 220 generates a trained model representing the correlation between object image data, image filter association data, and OCR score data using machine learning with acquired learning data of multiple types. The trained model generation unit 220 generates the trained model using Q-learning, an example of a known reinforcement learning algorithm. Here, reinforcement learning refers to machine learning where an agent, acting within an environment, observes the parameters of the environment, i.e., the current state, and decides on the appropriate action. In reinforcement learning, the environment dynamically changes due to the agent's actions, and the agent is rewarded based on these changes. Furthermore, in reinforcement learning, the agent repeats these actions, learning the action strategy that yields the highest reward through a series of actions.
[0071] Furthermore, in Q-learning, the action value is calculated based on the action value function, which is an example of a value function, as the action strategy that yields the highest reward. Here, the state of the environment at time t is set as s. t Let the action at time t be a. t Due to action a t The state that has changed is denoted as s. t+1 , will be due to the state changing from s t Change to s t+1 The reward is set to r. t+1 Let the discount rate be γ, and the learning coefficient be α, where 0 < γ ≤ 1 and 0 < α ≤ 1 hold true. Furthermore, let the action value function be Q(s) t a t In the case of ), the action value function Q(s) t a t The general update formula for ) is represented by the following formula 1.
[0072] [Equation 1]
[0073]
[0074] Furthermore, in Q-learning, when the value of an action is set to Q, if the action with the highest value at time t+1 is a... t+1 The value of action Q is greater than the value of action a performed at time t. t If the action value Q is increased, then the action value Q increases if action a... t+1 The value of action Q is less than the value of action a. t If the action value Q is less than the action value Q, then the action value Q is reduced. In other words, in Q-learning, to make the action a at time t... t The action value Q is close to the optimal action value at time t+1, for the action value function Q(s) t a t The optimal action value Q in a given environment is then updated. As a result, the optimal action value Q in a given environment is propagated sequentially to the action values Q in previous environments.
[0075] The trained model generation unit 220 substitutes the values of the object image data contained in the training data into the state s. t And substitute the values of the image filter association data contained in the learning data into action a. t This allows Q-learning to generate a well-trained model. Furthermore, it relates the value-oriented state s based on the object image data. t The substitution can be arbitrary. For example, the numerical value representing the image data of the item can be set as x, and a predefined constant can be set as u. In this case, regarding state s t s t =u×x holds true.
[0076] Additionally, regarding the value-oriented action a based on image filter-related data... t Substituting this into the equation, as long as it is based on the action value function Q(s) t a t ) and state s t And regarding action a t Perform calculations and be able to base them on action a t The combination of image filters and the parameters of each filter are determined, allowing for arbitrary substitution. For example, the numerical value representing the image filter correlation data can be set as y, and a predefined constant can be set as v. In this case, regarding action a... t a t =v×y holds true.
[0077] The reward calculation unit 221 calculates the reward r based on the numerical values of the image data representing the items, the numerical values of the image filter association data, and the score based on the OCR score data contained in the learning data. t+1 Calculations are performed. For example, when the reward calculation unit 221 compares two types of learning data, if the score based on the OCR score data changes due to a change in the value of at least one of the values representing the item image data and the value representing the image filter association data, then the assigned reward r is also adjusted. t+1 Change. Specifically, if the score increases, the reward r increases. t+1 For example, the reward calculation unit 221 assigns a reward of +1; on the other hand, if the score value decreases, the reward r is reduced. t+1 For example, the return calculation unit 221 assigns a return of -1.
[0078] Here, for example, we will discuss the image data IMG-1 of the first item and the image filter association data F / P-001 and F / P-002 of the two types. In this case, since the values representing the image filter association data F / P-001 and F / P-002 are different, the scores based on the OCR score data IMG-1_F / P-001 and IMG-1_F / P-002 are also different. Therefore, when the reward calculation unit 221 sets the scores based on the OCR score data IMG-1_F / P-001 and IMG-1_F / P-002 to SC1 and SC2, when the image filter association data changes from F / P-001 to F / P-002, if (SC2-SC1)>0, a reward of +1 is given; on the other hand, if (SC2-SC1)≤0, a reward of -1 is given.
[0079] The value function update unit 222 updates the value based on the return r calculated by the return calculation unit 221. t+1 For the action value function Q(s) t a t The value function update unit 222 generates the action value function Q(s) and updates it accordingly. t a t The data is used as a training model.
[0080] Each time the learning data acquisition unit 210 acquires learning data from the visual sensor 100, the trained model generation unit 220 repeatedly performs the feedback process. t+1 The calculation and action value function Q(s) t a t The updated model generation unit 220 updates the action value function Q(s) using the update formula shown in Equation 1 above each time. t a tWhen updating, a function Q(s) representing the updated action value is generated. t a t The data is used as a training model.
[0081] The trained model output unit 230 will generate the trained model, which represents the action value function Q(s). t a t The data is output and stored in storage device 400.
[0082] (Details regarding the functional structure of the inference device 300 involved in Embodiment 1)
[0083] The trained model acquisition unit 310 acquires the trained model stored in the storage device 400.
[0084] The article image data acquisition unit 320 acquires the article image data for OCR output from the vision sensor 100. In this embodiment, the article image data for OCR acquired by the article image data acquisition unit 320 is article image data of the workpiece pre-captured on the actual manufacturing production line before the application of OCR in the vision sensor 100. Specifically, the article image data for OCR includes various types of article image data that require image processing, such as article image data with blurred text written on the workpiece, article image data of the workpiece captured in bright indoor conditions, and article image data of the workpiece captured in dark indoor conditions. Furthermore, the article image data that requires image processing may also include data indicating the probability of capturing images on the actual manufacturing production line.
[0085] The inference result data generation unit 330 inputs the object image data for OCR into the trained model and generates a first inference result data and a second inference result data that is different from the first inference result data.
[0086] Here, a summary of the processing of the first inference result data and the second inference result data output by the well-trained model that is input into OCR object image data will be explained. First, the learning data used by the learning device 200 for machine learning includes image data of 4 categories (IMG-1, IMG-2, IMG-3, IMG-4), image filter association data of 5 categories (F / P-001, F / P-002, F / P-003, F / P-004, F / P-005), and OCR score data of 20 categories (IMG-1_F / P-001, IMG-1_F / P-002, ..., IMG-1_F / P-005, IMG-2_F / P-001, IMG-2_F / P-002, ..., IMG-2_F / P-005, ..., IMG-4_F / P-001, IMG-4_F / P-002, ..., IMG-4_F / P-005).
[0087] In addition, such as Figure 5 As shown, the scores of the 20 OCR score data IMG-1_F / P-001, IMG-1_F / P-002, ..., IMG-1_F / P-005, IMG-2_F / P-001, IMG-2_F / P-002, ..., IMG-2_F / P-005, ..., IMG-4_F / P-001, IMG-4_F / P-002, ..., IMG-4_F / P-005 are 99, 60, ..., 0, 70, 10, ..., 11, ..., 20, 91, ..., 91.
[0088] In addition, such as Figure 5 As shown, the probability that the OCR-generated item image data is the same as the first item image data IMG-1 is 9%, the probability that it is the same as the second item image data IMG-2 is 60%, the probability that it is the same as the third item image data IMG-3 is 30%, and the probability that it is the same as the fourth item image data IMG-4 is 1%.
[0089] In this case, such as Figure 5As shown, the first item image data IMG-1 scored the highest (99 points) when processed and OCR was performed using an image filter based on the first image filter associated data F / P-001. The second item image data IMG-2 scored the highest (98 points) when processed and OCR was performed using an image filter based on the third image filter associated data F / P-003. The third item image data IMG-3 scored the highest (100 points) when processed and OCR was performed using an image filter based on the second image filter associated data F / P-002. The fourth item image data IMG-4 scored the highest (91 points) when processed and OCR was performed using either the second image filter associated data F / P-002 or the fifth image filter associated data F / P-005.
[0090] Here, for example, we will discuss the case where the inference result data generation unit 330 assigns the condition that the inference result data has two categories and the score value is greater than or equal to 90 points to the trained model. In this case, firstly, the trained model determines candidates for combinations of image filter association data of the two categories with the highest coverage of each item image data IMG-1 to IMG-4, which indicate the score value of each item image data having been processed and OCR has been greater than or equal to 90 points.
[0091] Specifically, the trained model did not have a 100% coverage combination in the two types of image filter association data. Therefore, the first image filter association data F / P-001 and the second image filter association data F / P-002, with a coverage of 75%, were calculated as the first candidate, and the first image filter association data F / P-001 and the third image filter association data F / P-003 were calculated as the second candidate. Furthermore, regarding the maximum score for image processing and OCR using image filters based on the first candidate image filter association data F / P-001 and F / P-002, the first item image data IMG-1 scored 99 points, the second item image data IMG-2 scored 70 points, the third item image data IMG-3 scored 100 points, and the fourth item image data IMG-4 scored 91 points. In addition, regarding the maximum score when using image filters based on the associated data F / P-001 and F / P-003 as the second candidates for image processing and OCR, the first item image data IMG-1 scored 99 points, the second item image data IMG-2 scored 98 points, the third item image data IMG-3 scored 91 points, and the fourth item image data IMG-4 scored 80 points.
[0092] Furthermore, the trained model calculates the expected scores of the first and second candidates based on the probabilities of obtaining image data IMG-1 to IMG-4 of each item from the actual manufacturing production line, and outputs the candidate with the higher expected score as the inference result data. Specifically, the expected score of the first candidate is 81.82 (99×0.09+70×0.60+100×0.30+91×0.01=81.82). On the other hand, the expected score of the second candidate is 95.81 (99×0.09+98×0.60+91×0.30+80×0.01=95.81). Therefore, the trained model outputs the second candidate, namely the first image filter association data F / P-001 and the third image filter association data F / P-003, as the first and second inference result data. As a result, the inference result data generation unit 330 generates first image filter association data F / P-001 and third image filter association data F / P-003 as first inference result data and second inference result data.
[0093] Furthermore, in this embodiment, the inference result data generation unit 330 generates two types of inference result data: first inference result data and second inference result data. However, it may also generate three or more types of inference result data. For example, the inference result data generation unit 330 may also generate three types of inference result data: first inference result data, second inference result data, and third inference result data.
[0094] In this case, the trained model uses the first image filter associated data F / P-001, the second image filter associated data F / P-002, and the third image filter associated data F / P-003, which cover 100% of the image data, as the first candidate, and calculates the first image filter associated data F / P-001, the third image filter associated data F / P-003, and the fifth image filter associated data F / P-005 as the second candidate. Furthermore, regarding the maximum score for image processing and OCR using the image filters based on the first candidate image filter associated data F / P-001, F / P-002, and F / P-003, the first item image data IMG-1 scores 99 points, the second item image data IMG-2 scores 98 points, the third item image data IMG-3 scores 100 points, and the fourth item image data IMG-4 scores 91 points. In addition, regarding the maximum score when using image filters based on the associated data F / P-001, F / P-003, and F / P-005 as the second candidates for image processing and OCR, the first item image data IMG-1 scored 99 points, the second item image data IMG-2 scored 98 points, and the third item image data IMG-3 and the fourth item image data IMG-4 scored 91 points.
[0095] Therefore, the expected score for candidate 1 is 98.62 (99×0.09+98×0.60+100×0.30+91×0.01=98.62). On the other hand, the expected score for candidate 2 is 95.92 (99×0.09+98×0.60+91×0.30+91×0.01=95.92). Furthermore, in this case, even if the probability of obtaining image data IMG-1 to IMG-4 of each item changes, the expected score for candidate 1 is still higher than the expected score for candidate 2. For example, if the probability of obtaining image data IMG-1 to IMG-4 for each item is 25%, then the expected score of the first candidate, 97 points ((99+98+100+91) / 4=97), is higher than the expected score of the second candidate, 94.75 points ((99+98+91+91) / 4=94.75). Therefore, the trained model outputs the first candidate, i.e., the first image filter associated data F / P-001, the second image filter associated data F / P-002, and the third image filter associated data F / P-003, as the first inference result data, the second inference result data, and the third inference result data. As a result, the inference result data generation unit 330 generates the first image filter associated data F / P-001, the second image filter associated data F / P-002, and the third image filter associated data F / P-003 as the first inference result data, the second inference result data, and the third inference result data.
[0096] Furthermore, the inference result data generation unit 330 may not assign the condition that the score value is greater than or equal to 90 points to the trained model. Even in this case, the trained model can still output the first inference result data and the second inference result data by determining the combination of image filter association data F / P-001 to F / P-005 with the highest expected score.
[0097] Furthermore, the OCR-processed article image data acquired by the article image data acquisition unit 320 may not include data indicating the probability of an image being captured on the actual manufacturing production line. In this case, the inference result data generation unit 330 may determine the combination of image filter association data F / P-001 to F / P-005 with the highest expected score, provided that the probabilities of obtaining the acquired article image data are exactly the same.
[0098] return Figure 2The inference result data output unit 340 outputs the first inference result data and the second inference result data to the vision sensor 100 as the generated inference result data. Therefore, in the vision sensor 100, the inference result data acquisition unit 190 acquires the first inference result data and the second inference result data. Furthermore, the image filter generation unit 130 generates a first image filter based on the first inference result data and a second image filter based on the second inference result data, and the image filter output unit 140 outputs the first image filter and the second image filter to the image processing unit 150. The image processing unit 150 performs image processing on the object image data using each image filter, and the OCR unit 160 performs OCR on the object image data after each image processing step.
[0099] (Regarding the well-trained model generation process involved in Implementation Method 1)
[0100] Next, a flowchart will be used to illustrate the actions generated and output by the learning device 200 of the trained model. When the power is turned on, the learning device 200 begins execution. Figure 6 The trained model generation process is shown. First, the learning data acquisition unit 210 acquires new learning data from the visual sensor 100 (step S101). For example, the learning data acquisition unit 210 acquires data containing... Figure 4 The data shown includes image data of m types of items (IMG-1, IMG-2, ..., IMG-m, n), image filter association data of types (F / P-001, F / P-002, ..., F / P-00n), and OCR score data of m×n types (IMG-1_F / P-001, IMG-1_F / P-002, ..., IMG-1_F / P-00n, IMG-2_F / P-001, IMG-2_F / P-002, ..., IMG-2_F / P-00n, ..., IMG-m_F / P-001, IMG-m_F / P-002, ..., IMG-m_F / P-00n).
[0101] Next, the trained model generation unit 220 generates a trained model using machine learning with multiple types of acquired learning data. Specifically, the reward calculation unit 221 calculates the reward r based on the item image data, image filter association data, and OCR score data included in the acquired learning data. t+1 Calculation is performed (step S102). For example, for the first item image data IMG-1, when the image filter associated data changes from F / P-001 to F / P-002, if (SC2-SC1)>0, the reward calculation unit 221 assigns a reward of +1; on the other hand, if (SC2-SC1)≤0, the reward calculation unit 221 assigns a reward of -1.
[0102] Next, the value function update unit 222 updates the value function based on the calculated return r. t+1 Action value function Q(s) t a t The value function update unit 222 updates the state s based on the numerical value x representing the item image data (step S103). t Perform calculations and apply the numerical value y, representing the associated data of the image filters, to the action a. t The calculation is performed. Then, the value function update unit 222 updates the action value function Q(s) using the update formula shown in Equation 1 above. t a t The updated action value function Q(s) is then performed. The trained model generator 220 will then represent the updated action value function. t a t The data of the trained model is output to the storage device 400, so that the trained model storage unit 410 stores it (step S104), and the processing ends.
[0103] (Regarding the inference result data generation and processing involved in Implementation Method 1)
[0104] Next, a flowchart will be used to describe the actions of the inference device 300 in generating and outputting inference result data. When the power is turned on, the inference device 300 begins execution. Figure 7 The inference result data generation process is as follows: First, the trained model acquisition unit 310 acquires the trained model stored in the storage device 400 (step S201). Next, the object image data acquisition unit 320 acquires the newly OCR-enabled object image data from the vision sensor 100 (step S202). Next, the inference result data generation unit 330 inputs the newly OCR-enabled object image data into the trained model and generates the first inference result data and the second inference result data (step S203). Then, the inference result data output unit 340 outputs the generated first inference result data and the second inference result data to the vision sensor 100 (step S204), ending the process.
[0105] (Regarding the image filter generation process involved in Implementation Method 1)
[0106] Next, a flowchart will be used to illustrate the actions of the image filter generated and output by the vision sensor 100. When the power is turned on, the vision sensor 100 begins to execute... Figure 8The image filter generation process is shown below. First, the object image data output unit 170 outputs the OCR-processed object image data to the inference device 300 (step S301). Next, the inference result data acquisition unit 190 acquires the first inference result data and the second inference result data output from the inference device 300 (step S302). Next, the image filter generation unit 130 generates a first image filter based on the first inference result data and a second image filter based on the second inference result data (step S303). Then, the image filter output unit 140 outputs the first image filter and the second image filter to the image processing unit 150 (step S304), and the processing ends.
[0107] As described above, according to the image filter generation system 1 of this embodiment, the vision sensor 100 generates an image filter used in image processing before performing OCR on the object image data, which is obtained by the camera 110 taking pictures of the object.
[0108] Here, for example, we discuss the case of OCR for image data of workpieces captured by vision sensors in a factory. In this case, sometimes due to environmental factors such as the workpiece not being positioned correctly, or the factory being too bright or too dark when the workpiece was photographed, it is impossible to obtain image data of items that are easy to recognize text, and text may be misrecognized. Therefore, in the past, technicians have relied on know-how to manually combine various image filters and set the parameters of each image filter, preparing multiple types of image filters with high OCR reliability in a specific environment, and using the image filter with the highest OCR score in the current environment. However, the prepared image filters need to take into account the type of workpiece, such as its material, color, and shape, as well as the types and parameters of the combined image filters, thus resulting in a problem that preparation and application by hand require a lot of time.
[0109] In contrast, in the image filter generation system 1 of this embodiment, the vision sensor 100 takes into account the type of workpiece shown in the object image data, the type of image filter combined shown in the image filter association data, and the parameters, and automatically generates an image filter.
[0110] By adopting the above-described manner, the image filter generation system 1 of this embodiment can shorten the time from obtaining the OCR image data of the item until image processing using the image filter is performed, compared to preparing and applying the image filter manually.
[0111] Furthermore, according to the image filter generation system 1 of this embodiment, in the learning device 200, the learning data acquisition unit 210 acquires learning data including object image data and image filter correlation data from the vision sensor 100. Additionally, the trained model generation unit 220 generates a trained model representing the correlation between the object image data and the image filter correlation data using machine learning with the learning data, and the trained model output unit 230 outputs the trained model and stores it in the storage device 400.
[0112] Furthermore, in the inference device 300, the object image data acquisition unit 320 acquires object image data for OCR. Additionally, the inference result data generation unit 330 inputs the OCR-processed object image data into the trained model acquired from the storage device 400 by the trained model acquisition unit 310, generating first inference result data and second inference result data. Furthermore, the inference result data output unit 340 outputs the first and second inference result data to the vision sensor 100. Moreover, in the vision sensor 100, the image filter generation unit 130 generates a first image filter based on the first inference result data and a second image filter based on the second inference result data, and the image filter output unit 140 outputs the first and second image filters.
[0113] Therefore, in the vision sensor 100, the image processing unit 150 can perform image processing on the object image data using the first image filter and on the object image data using the second image filter. Furthermore, the OCR unit 160 can perform OCR on the object image data after image processing using the first image filter and on the object image data after image processing using the second image filter.
[0114] Here, for example, such as Figure 9As shown, the newly OCR-processed object image data is designated as IMG-0, the first inference result data generated and output by the inference device 300 is designated as F / P-001, and the second inference result data is designated as F / P-002. In this case, the vision sensor 100 generates a first image filter based on the first inference result data F / P-001 and a second image filter based on the second inference result data F / P-002. Furthermore, the vision sensor 100 performs image processing and OCR on the object image data IMG-0 using each image filter. At this time, the output OCR score data is designated as IMG-0_F / P-001 and IMG-0_F / P-002, and the scores based on each OCR score data IMG-0_F / P-001 and IMG-0_F / P-002 are designated as SCA and SCB, respectively. In this case, if (SCA-SCB)>0, the visual sensor 100 uses the result of text recognition obtained by image processing and OCR through the first image filter; on the other hand, if (SCA-SCB)≤0, the visual sensor 100 uses the result of text recognition obtained by image processing and OCR through the second image filter.
[0115] By configuring it as described above, the vision sensor 100 can select the image filter with the highest score for character recognition during OCR from the first image filter and the second image filter. Therefore, within the so-called "takt time" from when the camera 110 takes a picture of the workpiece until OCR is performed, the vision sensor 100 can select the most suitable image filter for the object image data to be OCRed each time the camera 110 takes a picture of the workpiece, and perform image processing and OCR using that image filter. As a result, compared to an image filter generation system that does not generate first and second inference result data, the image filter generation system 1 according to this embodiment can reduce misrecognition of characters.
[0116] Furthermore, in this embodiment, the inference device 300 generates and outputs inference result data of two types, but it can also generate and output inference result data of three or more types. For example, when the inference device 300 generates and outputs inference result data of three types, the visual sensor 100 can select the image filter with the highest score for text recognition during OCR from the first image filter, the second image filter, and the third image filter. In this case, image processing and OCR using image filters based on the three types of inference result data need to be attempted within the aforementioned time frame. Therefore, when the inference device 300 generates and outputs inference result data of three or more types, the number of types of inference result data needs to be determined considering the time frame.
[0117] Furthermore, according to the image filter generation system 1 of this embodiment, the learning data acquired by the learning data acquisition unit 210 includes object image data, image filter association data, and OCR score data. Moreover, the trained model generation unit 220 generates a trained model representing the correlation between the object image data, image filter association data, and OCR score data using machine learning based on the learning data.
[0118] By configuring it as described above, the inference device 300 can generate and output inference result data that is inferred to be the highest-scoring text recognition output during OCR by inputting object image data into a trained model. As a result, compared to an image filter generation system that does not generate a trained model representing the correlation between object image data, image filter correlation data, and OCR score data, the image filter generation system 1 according to this embodiment can reduce text misrecognition.
[0119] Furthermore, according to the image filter generation system 1 of this embodiment, in the learning device 200, the machine learning performed by the trained model generation unit 220 uses the action value function Q(s). t a t Reinforcement learning. Furthermore, the trained model generation unit 220 increases the score represented by the OCR score data when at least one of the two types of learning data (item image data and image filter associated data) changes, thus increasing the reward r. t+1 Increase, on the other hand, make the reward r decrease when the score decreases. t+1 Reduce, thereby affecting the action value function Q(s) t a t The model is then updated. Furthermore, the trained model generator 220 generates the updated action value function Q(s). t a t The data is used as a training model.
[0120] By configuring it as described above, the inference device 300 can input the object image data to be OCRed by a well-trained model obtained through reinforcement learning based on the scores represented by the OCR score data, thereby generating and outputting the inference result data that is inferred to be the character recognition with the highest score when OCR was performed. As a result, compared to an image filter generation system that does not generate a well-trained model obtained through reinforcement learning based on the scores represented by the OCR score data, the image filter generation system 1 according to this embodiment can reduce the misrecognition of characters.
[0121] Furthermore, according to the image filter generation system 1 of this embodiment, the article image data included in the learning data includes article image data of the workpiece actually photographed during past manufacturing.
[0122] By adopting the above-described manner, compared to an image filter generation system that generates a trained model using machine learning data that does not use learning data including image data of workpieces actually photographed during past manufacturing, the image filter generation system 1 of this embodiment can reduce the misrecognition of text when performing OCR on image data of workpieces actually manufactured.
[0123] Furthermore, according to the image filter generation system 1 of this embodiment, the object image data included in the learning data includes object image data of a workpiece captured in a simulated environment that envisions an actual environment.
[0124] By adopting the above-described manner, compared to an image filter generation system that generates a trained model using machine learning data that does not use learning data including workpiece image data captured in a simulated environment, the image filter generation system 1 of this embodiment can reduce text misrecognition when performing OCR on workpiece image data from actual manufacturing.
[0125] [Implementation Method 2]
[0126] In Embodiment 1, the inference device 300 generates and outputs inference result data of multiple types, but the inference device 300 may also not generate or output inference result data of multiple types. In the image filter generation system 1 according to Embodiment 2, the inference device 300 generates and outputs inference result data of only one type. Hereinafter, refer to... Figure 2 , Figure 5 , Figure 10 , Figure 11 The image filter generation system 1 according to Embodiment 2 will be described in detail. Furthermore, in Embodiment 2, structures different from those in Embodiment 1 will be described; structures identical to those in Embodiment 1 will be omitted due to their length.
[0127] (Details regarding the functional structure of the inference device 300 involved in Embodiment 2)
[0128] return Figure 2 In this embodiment 2, the article image data acquisition unit 320 acquires the article image data that has undergone OCR output from the vision sensor 100. In this embodiment, the article image data acquired by the article image data acquisition unit 320 that has undergone OCR is the article image data of the workpiece captured on the actual manufacturing production line when OCR is applied in the vision sensor 100.
[0129] In Implementation 2, the inference result data generation unit 330 inputs OCR-enhanced object image data into a trained model to generate inference result data.
[0130] The inference result data output unit 340 according to Embodiment 2 outputs the generated inference result data to the vision sensor 100.
[0131] Here, a summary of the processing of the inference result data output by the trained model, which is input from OCR-processed object image data, is provided. For example... Figure 5 As shown, the first item image data IMG-1 scored the highest (99 points) when processed and OCR was performed using an image filter based on the first image filter associated data F / P-001. The second item image data IMG-2 scored the highest (98 points) when processed and OCR was performed using an image filter based on the third image filter associated data F / P-003. The third item image data IMG-3 scored the highest (100 points) when processed and OCR was performed using an image filter based on the second image filter associated data F / P-002. The fourth item image data IMG-4 scored the highest (91 points) when processed and OCR was performed using either the second image filter associated data F / P-002 or the fifth image filter associated data F / P-005.
[0132] Therefore, for example, if the OCR image data IMG-0 is most similar to the first image data IMG-1, the trained model outputs the first image filter association data F / P-001 as the inference result data. Furthermore, if the OCR image data IMG-0 is most similar to the second image data IMG-2, the trained model outputs the third image filter association data F / P-003 as the inference result data. Similarly, if the OCR image data IMG-0 is most similar to the third image data IMG-3, the trained model outputs the second image filter association data F / P-002 as the inference result data. Finally, if the OCR image data IMG-0 is most similar to the second image data IMG-2, the trained model outputs either the second image filter association data F / P-002 or the fifth image filter association data F / P-005 as the inference result data. As a result, the inference result data generation unit 330 generates any one of the above-mentioned image filter association data F / P-001, F / P-002, F / P-003 and F / P-005 as inference result data.
[0133] (Regarding the inference result data generation and processing involved in Implementation Method 2)
[0134] Next, a flowchart will be used to explain the actions of the inference device 300 in generating and outputting inference result data. For example... Figure 10 As shown, after performing steps S201 and S202, the inference result data generation unit 330 inputs the newly OCRed object image data into the trained model and generates inference result data (step S213). Then, the inference result data output unit 340 outputs the generated inference result data (step S214), and the processing ends.
[0135] (Regarding the image filter generation process involved in Implementation Method 2)
[0136] Next, a flowchart will be used to illustrate the actions of the image filter generated and output by the vision sensor 100. For example... Figure 11 As shown, after the processing in step S301 is performed, the inference result data acquisition unit 190 acquires the inference result data output from the inference device 300 (step S312). Next, the image filter generation unit 130 generates an image filter based on the acquired inference result data (step S313). Then, the image filter output unit 140 outputs the generated image filter to the image processing unit 150 (step S314), and the processing ends.
[0137] As described above, according to the image filter generation system 1 of this embodiment, in the inference device 300, the inference result data generation unit 330 inputs the object image data to be OCRed into the trained model to generate inference result data. Furthermore, in the vision sensor 100, the image filter generation unit 130 generates an image filter based on the inference result data, and the image filter output unit 140 outputs the image filter. Additionally, the image processing unit 150 performs image processing on the object image data using the image filter, and the OCR unit 160 performs OCR on the object image data after image processing using the image filter.
[0138] By configuring it as described above, the vision sensor 100 can use the image filter with the highest score for character recognition during OCR, inferred by the trained model, to perform image processing on the object image data. For example, the vision sensor 100 generates an image filter based on the inference result data F / P-001, and uses this image filter to perform image processing and OCR on the object image data IMG-0 of the workpiece captured on the actual manufacturing production line during OCR application. Therefore, instead of using two types of image filters for image processing and OCR as in Embodiment 1, the result of character recognition with the highest score value can be used. As a result, each time the camera 110 captures an image of the workpiece, the vision sensor 100 can obtain the most suitable inference result data for the object image data to be OCR from the inference device 300, and perform image processing and OCR using the image filter based on this inference result data.
[0139] Furthermore, the image filter generation system 1 according to this embodiment achieves the same effect as the image filter generation system 1 according to Embodiment 1.
[0140] [Implementation Method 3]
[0141] In embodiments 1 and 2, the visual sensor 100, learning device 200, inference device 300, and storage device 400 are separate devices, but this is not a limitation; they can also be integrated. For example, the image filter generation device, i.e., the visual sensor 100, may also have the functions of the other devices 200, 300, and 400. The visual sensor 100 according to embodiment 3 may also have the functions of all these learning devices 200, inference devices 300, and storage devices 400. Hereinafter, refer to... Figure 12 The vision sensor 100 according to Embodiment 3 will be described in detail. Furthermore, in Embodiment 3, structures that are different from those in Embodiments 1 and 2 will be described, while structures that are the same as those in Embodiments 1 and 2 will be omitted from the description due to their length.
[0142] (Regarding the visual sensor 100 involved in Embodiment 3)
[0143] like Figure 12As shown, the visual sensor 100 omits the object image data output unit 170, the learning data output unit 180, and the inference result data acquisition unit 190. However, the visual sensor 100 also includes a learning data acquisition unit 210, a trained model generation unit 220, a reward calculation unit 221, a value function update unit 222, an object image data acquisition unit 320, an inference result data generation unit 330, and a trained model storage unit 410. Furthermore, the trained model acquisition unit 310 acquires the trained model stored in the trained model storage unit 410, and the object image data acquisition unit 320 acquires object image data for OCR from the camera 110.
[0144] As described above, the visual sensor 100 of this embodiment can perform the functions of the learning device 200, the inference device 300, and the storage device 400 of embodiments 1 and 2.
[0145] By configuring it in the manner described above, the visual sensor 100 according to this embodiment achieves the same effect as the image filter generation system 1 according to embodiments 1 and 2.
[0146] [Example of Change]
[0147] Furthermore, in Embodiment 3 described above, the devices 100, 200, 300, and 400 involved in Embodiments 1 and 2 are integrated into one device, but the combination of integrated devices is not limited to this. For example, the learning device 200 and storage device 400 involved in Embodiments 1 and 2 may be integrated into one device, with the remaining devices 100 and 300 being separate devices. Alternatively, the inference device 300 and storage device 400 involved in Embodiments 1 and 2 may be integrated into one device, with the remaining devices 100 and 200 being separate devices. Additionally, for example, the learning device 200, inference device 300, and storage device 400 involved in Embodiments 1 and 2 may be integrated into one device, with only the visual sensor 100 being a separate device.
[0148] Furthermore, in embodiments 1 and 2 described above, the visual sensor 100, learning device 200, inference device 300, and storage device 400 can transmit and receive data via a LAN, but the data transmission structure is not limited to this. For example, data can be transmitted and received via a communication cable connecting the visual sensor 100, learning device 200, inference device 300, and storage device 400, or via the Internet. In this case, for example, the learning device 200, inference device 300, and storage device 400 can also function as a so-called cloud server. In this case, the cloud server can also generate and store a trained model using machine learning based on learning data obtained from the visual sensor 100. Additionally, in this case, the cloud server can also input newly acquired OCR-enhanced object image data from the visual sensor 100 into the trained model to generate inference result data, and output it to the visual sensor 100.
[0149] Furthermore, in embodiments 1 to 3 described above, the trained model generation unit 220 generates a trained model using Q-learning, which is an example of a reinforcement learning algorithm. However, it is not limited to this; other reinforcement learning algorithms can also be used to generate the trained model. For example, the trained model generation unit 220 can also use TD-learning to generate the trained model.
[0150] Furthermore, in embodiments 1 to 3 described above, the trained model generation unit 220 generates a trained model using a reinforcement learning algorithm, but it is not limited to this. Other known learning algorithms, such as deep learning, neural networks, genetic programming, functional logic programming, and support vector machines, can also be used to generate the trained model. Additionally, the learning method is not limited to reinforcement learning; for example, known algorithms can be used to generate trained models for different learning methods such as supervised learning, unsupervised learning, and semi-unsupervised learning.
[0151] Here, when the model generation unit 220 generates a trained model through supervised learning, the learning data needs to include, for example, data representing the correct answer to the text that should be recognized in the OCR-recognized object image data. Furthermore, the correct answer data can be pre-input manually or automatically input based on the result of comparing strings recognized by OCR on multiple object image data.
[0152] Furthermore, when the model generation unit 220 generates a trained model through unsupervised learning, the learning data needs to include classification data that allows for the classification of various object image data requiring image processing, such as blurry object image data with text inscribed on the object, object image data of the object taken in bright indoor conditions, and object image data of the object taken in dark indoor conditions. Additionally, the image filter association data included in the learning data, such as image filter association data that requires suitable image filters for image processing of each category of object image data, needs to be pre-selected as described above.
[0153] Furthermore, when the trained model generation unit 220 generates a trained model through semi-unsupervised learning, the learning data needs to include, for example, the classification data and correct answer data mentioned above.
[0154] Furthermore, in embodiments 1 and 2 described above, the learning device 200 obtains learning data from the visual sensor 100 installed in the image filter generation system 1. In embodiment 3 described above, the visual sensor 100 obtains learning data generated by itself, but this is not a limitation. For example, the learning device 200 and the visual sensor 100 may also obtain learning data from other devices or systems performing OCR. The learning device 200 and the visual sensor 100 may, for example, obtain learning data from multiple image filter generation systems operating in the same area, or from image filter generation systems operating independently in different areas. In this case, the learning device 200 and the visual sensor 100 may also add or remove other image filter generation systems that have obtained learning data at arbitrary timing.
[0155] Furthermore, in embodiments 1 and 2 described above, the learning device 200 pre-installed in the image filter generation system 1 generates and outputs a trained model by performing machine learning on the learning data obtained from the vision sensor 100, but this is not a limitation. For example, a learning device installed in another image filter generation system that performs machine learning on the learning data obtained from the vision sensor can be set as the learning device 200 of the image filter generation system 1, and the learning data obtained from the vision sensor 100 can be used to relearn, update, and output the trained model.
[0156] Furthermore, in embodiments 1 and 2 described above, the inference device 300 acquires a trained model generated and output by the learning device 200 provided in the image filter generation system 1 and stored in the storage device 400, but is not limited to this. For example, the inference device 300 may also acquire a trained model generated and output by other image filter generation devices or other image filter generation systems.
[0157] Furthermore, the central processing unit, comprising a vision sensor 100, a learning device 200, and an inference device 300, including a control unit 51, a main storage unit 52, an external storage unit 53, an operation unit 54, a transceiver unit 56, and an internal bus 50, can be implemented using a conventional computer system without relying on a dedicated system. For example, the vision sensor 100, learning device 200, and inference device 300 performing the aforementioned processing can be configured by storing the computer program for performing the aforementioned actions on a computer-readable recording medium, such as a floppy disk or DVD-ROM (Read-Only Memory), and installing the computer program on a computer. Alternatively, the vision sensor 100, learning device 200, and inference device 300 can be configured by storing the computer program in a storage device on a server device on a communication network and downloading it to a conventional computer system.
[0158] Alternatively, the functions of the vision sensor 100, learning device 200, and inference device 300 may be implemented by sharing the responsibilities of the OS (Operating System) and the application, or by the coordinated operation of the OS and the application, and only the application portion may be stored in the recording medium or storage device.
[0159] Alternatively, the computer program can be overlaid on a carrier wave and provided via a communication network. For example, the aforementioned computer program can be published on a bulletin board system (BBS) on a communication network and provided via the network. Furthermore, by launching the computer program, it can be executed under the control of the operating system in the same way as other applications, thereby performing the aforementioned processing.
[0160] Various embodiments and modifications can be implemented without departing from the broad spirit and scope of the present invention. Furthermore, the above-described embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. That is, the scope of the invention is defined not by the embodiments, but by the claims. Moreover, various modifications implemented within the scope of the claims and their equivalent inventive meaning are considered to be within the scope of the present invention.
[0161] Explanation of the label
[0162] 1…Image filter generation system, 50…Internal bus, 51…Control unit, 52…Main storage unit, 53…External storage unit, 54…Operation unit, 55…Display unit, 56…Transmitter / receiver unit, 59…Control program, 100…Vision sensor, 110…Camera, 120…Image filter associated data generation unit, 130…Image filter generation unit, 140…Image filter output unit, 150…Image processing unit, 160…OCR unit, 170…Object image data output unit, 180…Learning data According to the output unit, 190... inference result data acquisition unit, 200... learning device, 210... learning data acquisition unit, 220... trained model generation unit, 221... reward calculation unit, 222... value function update unit, 230... trained model output unit, 300... inference device, 310... trained model acquisition unit, 320... item image data acquisition unit, 330... inference result data generation unit, 340... inference result data output unit, 400... storage device, 410... trained model storage unit.
Claims
1. An image filter generation system that generates image filters for image processing of object image data prior to OCR, wherein the object image data is image data of an object captured by a camera component. This image filter generation system has the following features: An image filter generating apparatus that generates the image filter; A learning device that learns the correlation between pre-acquired image data of the object and the image filters used in image processing of the image data of the object; and The inference device infers the appropriate image filter for image processing of the object image data for performing the OCR. The learning device includes: The learning data acquisition unit acquires learning data including the item image data, image filter association data, and OCR score data. The image filter association data represents the combination of image filters used in the image processing of the item image data and the values of the parameters of each image filter. The OCR score data represents the score of text recognition output by the OCR when the item image data is processed using the image filters based on the image filter association data. A trained model generation unit generates a trained model representing the correlation between the item image data, the image filter association data, and the OCR score data using machine learning on the learning data; and The trained model outputs the trained model. The inference device includes: The item image data acquisition unit acquires the item image data for which the OCR is performed; The inference result data generation unit inputs the object image data used for OCR into the trained model, and generates inference result data, which represents the appropriate combination of image filters and the values of the parameters of each image filter for image processing of the object image data used for OCR; and The inference result data output unit outputs the inference result data. The image filter generation device includes: An image filter generation unit generates an image filter based on the inference result data; and An image filter output unit outputs the image filter.
2. The image filter generation system according to claim 1, wherein, The machine learning described is reinforcement learning that uses a value function. The well-trained model generation unit generates the well-trained model by increasing the reward assigned to the value function when the score shown in the OCR score data increases as a result of a change in at least one of the data in the item image data and the image filter association data, and decreasing the reward when the score shown in the OCR score data decreases.
3. The image filter generation system according to claim 1 or 2, wherein, The image data of the items included in the learning data are image data of the workpieces that were actually photographed during past manufacturing processes.
4. The image filter generation system according to claim 1 or 2, wherein, The image data of the objects included in the learning data are image data of the workpieces taken in a simulated environment that imagines the actual environment.
5. An image filter generation apparatus, which generates an image filter for image processing before performing OCR on object image data, wherein the object image data is image data of an object captured by a camera component. The image filter generation device has the following features: The learning data acquisition unit acquires learning data including image filter association data, OCR score data, and pre-acquired object image data. The image filter association data represents the combination of image filters used in image processing of the object image data and the values of the parameters of each image filter. The OCR score data represents the score of text recognition output by the OCR when image processing of the object image data is performed using the image filters based on the image filter association data. The trained model generation unit generates a trained model that represents the correlation between the item image data, the image filter association data, and the OCR score data by using machine learning on the learning data; The item image data acquisition unit acquires the item image data for which the OCR is performed; The inference result data generation unit inputs the object image data for the OCR to the trained model and generates data representing the appropriate combination of image filters and the values of the parameters of each image filter for image processing of the object image data for the OCR, i.e., inference result data; An image filter generation unit generates an image filter based on the inference result data; as well as An image filter output unit outputs the image filter.
6. A learning device that learns the correlation between an image filter and object image data, wherein the image filter is an image filter used for image processing of object image data captured by a shooting component before OCR is performed on the object image data. This learning device has the following features: The learning data acquisition unit acquires learning data including image filter association data, OCR score data, and pre-acquired object image data. The image filter association data represents the combination of image filters used in image processing of the object image data and the values of the parameters of each image filter. The OCR score data represents the score of text recognition output by the OCR when image processing of the object image data is performed using the image filters based on the image filter association data. A trained model generation unit generates a trained model representing the correlation between the item image data, the image filter association data, and the OCR score data using machine learning on the learning data; and The trained model outputs the trained model. The trained model is used as input for the object image data to be OCRed, and generates data representing the appropriate combination of image filters and the values of the parameters of each image filter for image processing of the object image data to be OCRed.
7. A learning method for learning the correlation between an image filter and object image data, wherein the image filter is an image filter used for image processing of object image data captured by a camera component before OCR is performed on the object image data. This learning method includes the following steps: The learning data acquisition step involves a computer acquiring learning data that includes image filter association data, OCR score data, and pre-acquired object image data. The image filter association data represents the combination of image filters used in image processing of the object image data and the values of the parameters of each image filter. The OCR score data represents the score of text recognition output by the OCR when image processing of the object image data is performed using the image filters based on the image filter association data. The well-trained model generation step involves the computer generating a well-trained model representing the correlation between the object image data, the image filter association data, and the OCR score data using machine learning on the learning data; and The trained model output step involves the computer outputting the trained model. The trained model is used as input for the object image data to be OCRed, and generates data representing the appropriate combination of image filters and the values of the parameters of each image filter for image processing of the object image data to be OCRed.
Citation Information
Patent Citations
Image correction method and image correction apparatus
JP2009010853A
Image processing system and program
JP2020154600A
Method, device and mobile terminal for recognizing information
CN102957963A