Information processing device, information processing method, and program
The information processing device uses machine learning to accurately identify lenses in cross-sectional images, addressing variations in image quality and style, and enables efficient lens database creation and retrieval.
Patent Information
- Application Number
- JP2022579346
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-05
- Filing Date
- 2021-11-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-11-18
AI Technical Summary
Existing methods for identifying lenses from cross-sectional images are hindered by variations in image quality and drawing styles, leading to inefficiencies and difficulties in accurately identifying lenses, especially when multiple lenses are grouped together.
An information processing device utilizing machine learning to construct an identification model that detects lens presence areas and identifies lens types based on image features, enabling the creation of a database for lens information and calculating similarities between lens groups.
Effectively identifies lenses in cross-sectional images regardless of drawing style, allowing for efficient utilization of lens information and facilitating the creation of a database for lens verification and retrieval.
Smart Images

Figure 0007761600000001 
Figure 0007761600000002 
Figure 0007761600000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and a program, and more particularly to an information processing device, an information processing method, and a program for identifying a lens from an image showing a cross section of a lens provided in a device. [Background technology]
[0002] Information about lenses provided in devices such as cameras is used for a variety of purposes, and for example, it may be accumulated and used to build a database (see, for example, Patent Document 1).
[0003] In Patent Document 1, lens information is stored in a database. Also in Patent Document 1, a camera captures an image of the objective side of the lens unit, and the database is searched for lens information based on the captured image. The lens information found by the search is then sent to the camera. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-178704 Summary of the Invention [Problem to be solved by the invention]
[0005] As a prerequisite for utilizing information about lenses, such as building a lens information database in Patent Document 1, it is necessary to acquire lens information. One example of a method for acquiring lens information is to find an image showing a cross section of a lens installed in a device from documents such as papers and gazettes, identify the lens from the image, and acquire the identification result as lens information.
[0006] However, even if the images show the cross-section of the same type of lens, the appearance and characteristics of the images may differ depending on differences in image quality such as resolution. Also, if the image of the lens is a drawing (illustration), the image may change depending on the thickness or drawing style of the lines.
[0007] Here, if various cases of drawing styles, including image quality, are assumed and rules for identifying lenses from lens cross-sectional images are prepared for each style, it will be possible to handle differences in drawing styles. However, in this case, it will be time-consuming to prepare a large number of identification rules. Furthermore, for images drawn in drawing styles for which no identification rules are prepared, it will be difficult to identify the lenses that appear in the images.
[0008] On the other hand, if information about the lens identified from the image is stored, the information can be referenced later, and the stored information can be effectively utilized, for example, for lens verification.
[0009] The present invention has been made in consideration of the above circumstances, and aims to provide an information processing device, an information processing method, and a program that are capable of appropriately identifying a lens from an image showing a cross section of the lens. Another object of the present invention is to effectively utilize information about lenses identified from images. [Means for solving the problem]
[0010] In order to achieve the above object, the information processing device of the present invention is an information processing device equipped with a processor, which detects an area where a lens exists in a target image showing a cross-section of a part including a lens in a target device equipped with a lens, and identifies the lens of the target device that exists in the area where a lens exists based on the features of the area where a lens exists, using an identification model constructed by machine learning using multiple learning images showing the cross-sections of the lens.
[0011] In addition, the identification model may be constructed by machine learning using multiple learning images including two or more cross-sectional images of the same type of lens showing cross sections of the same type of lens with different drawing styles, and may be a model that identifies the lenses shown in each of the two or more cross-sectional images of the same type of lens as lenses of the same type.
[0012] In the information processing device of the present invention, the processor may also accumulate information relating to the lenses of the identified target devices to build an information database.
[0013] Furthermore, it is preferable that the processor acquires input information relating to a lens provided in the search device, and outputs information relating to the lens of the target device stored in the database in association with the search device based on the input information.
[0014] Furthermore, it is more preferable if the processor acquires input information regarding the lens from the search device, calculates the similarity between the lens of the search device and the lens of the target device based on the input information and information regarding the lens of the target device stored in the database, and outputs the information regarding the lens of the target device stored in the database in association with the similarity.
[0015] The processor may also detect a presence region for each lens in a target image showing a cross section of a portion including a lens group in a target device having a lens group arranged in a row. In this case, the processor may use an identification model to identify lenses in the lens group present in the presence region for each presence region, and accumulate information about the lenses in the lens group identified for each presence region in a database by aggregating the lens group as a unit.
[0016] In the above configuration, the processor may identify, for each presence area, the type of lens in the lens group identified for each presence area. Then, the processor may generate character string information representing the type of each lens in the order in which the lenses are arranged in the lens group based on the type of lens in the lens group identified for each presence area and the position of the presence area in the target image, and store the generated character string information in a database.
[0017] Furthermore, when multiple lens groups are shown in the target image, the processor may generate character string information for each lens group, and store the character string information generated for each lens group in the database for each lens group.
[0018] Furthermore, the processor may acquire input information about lenses from the search device, and if the search device has a lens group, acquire as input information character string information representing the type of each lens in the order in which the lenses are arranged in the lens group of the search device. In this case, the processor may calculate a first similarity between the lens group of the search device and the lens group of the target device based on the acquired character string information about the lens group of the search device and character string information about the lens group of the target device stored in the database. Then, the processor may output the character string information about the lens group of the target device stored in the database in association with the first similarity.
[0019] Furthermore, the processor may acquire input information about lenses from the search device, and if the search device has a lens group, acquire as input information character string information representing the type of each lens in the order in which the lenses are arranged in the lens group of the search device. In this case, the processor may change any characters in the acquired character string information about the lens group of the search device to blanks. The processor may then calculate a second similarity between the lens group of the search device and the lens group of the target device based on the changed character string information about the lens group of the search device and the character string information about the lens group of the target device stored in the database. The processor may then output the character string information about the lens group of the target device stored in the database in association with the second similarity.
[0020] The processor may also store, in a database, information about the lenses of the identified target devices in association with information about documents containing target images.
[0021] The processor may also use an object detection algorithm to detect the area where the lens is present from the target image.
[0022] Furthermore, the above-mentioned object can be achieved by an information processing method including the steps of: detecting, by a processor, an area where a lens is present in a target image showing a cross-section of a part including a lens in a target device equipped with a lens; and identifying, by the processor, the lens of the target device present in the area where a lens is present, based on the features of the area where a discrimination model is constructed by machine learning using a plurality of learning images showing cross-sections of the lens.
[0023] The information processing method may further include a step of accumulating, by the processor, information relating to the lenses of the identified target devices to build an information database.
[0024] Furthermore, according to the present invention, it is possible to realize a program that causes a processor to execute each step of the above-described information processing method. [Effects of the Invention]
[0025] According to the present invention, it is possible to appropriately identify a lens that appears in an image showing a cross section of the lens, regardless of the rendering style of the image, and to effectively utilize information about the lens identified from the image. [Brief explanation of the drawings]
[0026] [Figure 1] FIG. 10 is a diagram illustrating an example of a target image. [Figure 2] FIG. 4 is a diagram illustrating an example of the structure of a lens information database. [Figure 3] FIG. 1 is a conceptual diagram of a discrimination model. [Figure 4] FIG. 10 is a diagram showing an example of a result of detecting a region where a lens is present in a target image. [Figure 5] 1 is a diagram showing a configuration of an information processing apparatus according to an embodiment of the present invention; [Figure 6]FIG. 1 is a diagram showing a flow of information processing using an information processing device according to an embodiment of the present invention. [Figure 7] FIG. 10 is a diagram showing character string information about a lens group. [Figure 8] FIG. 10 is a diagram showing an example in which two lens groups are shown in an image. [Figure 9] FIG. 10 is a diagram illustrating an example of a screen showing the execution result of the information output process. DETAILED DESCRIPTION OF THE INVENTION
[0027] An information processing device, an information processing method, and a program according to one embodiment of the present invention (hereinafter referred to as "the present embodiment") will be described below with reference to the accompanying drawings.
[0028] It should be noted that the following embodiment is merely an example given for the purpose of explaining the present invention in an easy-to-understand manner, and is not intended to limit the present invention. That is, the present invention is not limited to the following embodiment, and various improvements or modifications may be made without departing from the spirit of the present invention. Furthermore, the present invention includes equivalents thereof.
[0029] In the following explanation, unless otherwise specified, "image" refers to computerized (digitized) image data that can be processed by a computer. Images also include images taken with a camera, drawings (illustrations) created with drawing software, and drawing data obtained by scanning hand-drawn or printed drawings.
[0030] <Outline of the information processing device according to this embodiment> The information processing device of this embodiment (hereinafter simply referred to as the information processing device) includes a processor and analyzes a target image of a target device. The target device is an optical device equipped with one or more lenses, and corresponds to, for example, a photographing device such as a camera, an observation device such as a camera viewfinder, and a terminal equipped with a photographing function such as a mobile phone or a smartphone.
[0031] As shown in FIG. 1, the target image is an image showing a cross section of a portion of the target device that includes at least a lens (hereinafter referred to as a lens cross-sectional image). For example, if the target device is a digital camera with an imaging lens, the target image corresponds to a cross-sectional image of the imaging lens. If the target device has a single lens, the lens cross-sectional image in which the lens appears is treated as the target image. On the other hand, if the target device has a lens group consisting of multiple lenses, the lens cross-sectional image in which the entire lens group appears is treated as the target image. The lens group refers to multiple lenses lined up in a straight line, as shown in FIG. 1.
[0032] The target image also includes images published or inserted in documents such as papers, patent publications, magazines, and websites. The following description will be given taking as an example a case where a lens cross-sectional image published in a patent publication is treated as the target image. Naturally, the following content can also be applied to documents other than patent publications.
[0033] The information processing device analyzes a target image in a patent publication and, based on the analysis results, identifies the lens of the target device that appears in the target image. Specifically, the information processing device identifies the lens of the target device that appears in the target image using an identification model described below, and more specifically, specifies the type and position in the target image of the identified lens. The position in the target image refers to coordinates (strictly speaking, two-dimensional coordinates) when a reference point set in the target image is the origin.
[0034] Here, if a single lens appears in the target image, the type of that lens and its position in the target image are identified. On the other hand, if a group of lenses appears in the target image, the type of each lens in the group of lenses and its position in the target image are identified, as well as the order of each lens, i.e., the type of lens located at each location in the group of lenses.
[0035] In the following, it is assumed that the lens of the target device is a spherical lens, and that there are four types of lenses, specifically, a convex lens, a concave lens, a convex meniscus lens, and a concave meniscus lens. However, the types of lenses are not limited to the above four types, and may include types other than those above (for example, aspherical lenses, etc.).
[0036] In this embodiment, character information (specifically, a symbol) is assigned to the lens of the target device whose type has been identified. Specifically, for example, a "T" is assigned to a convex lens, an "O" is assigned to a concave lens, a "P" is assigned to a convex meniscus lens, and an "N" is assigned to a concave meniscus lens. Note that the characters assigned to each type of lens are not particularly limited.
[0037] On the other hand, for each lens group, character string information is assigned in which the above codes are arranged in accordance with the order of each lens in the lens group. The character string information indicates the type of each lens in the order in which each lens is arranged in the lens group. For example, the character string information "TON" is assigned to a lens group consisting of three lenses arranged in the order of a convex lens, a concave lens, and a concave meniscus lens from the front.
[0038] The character string information for the lens group is generated based on the type and position (coordinates) of each lens in the lens group. For example, the character string information is generated by arranging character information (codes) indicating the type of each lens in order from the lens closest to the reference position.
[0039] The information processing device also builds a database by accumulating information about the lens or lens group of the target device identified for each target image, specifically, character information indicating the type of lens or string information about the lens group.
[0040] In the database, information about the lens or lens group of the target device (hereinafter also referred to as lens information) is associated with the target image in which the lens or lens group appears, as shown in Figure 2. To explain in more detail, the lens information is stored in the database in association with information about the document containing the target image, specifically, the identification number of the patent publication in which the target image is published and the drawing number assigned to the target image in the patent publication.
[0041] Furthermore, the information processing device reads out and outputs lens information that satisfies predetermined conditions from the database, thereby making it possible to obtain, for example, information about the lens of a target device that is the same as or similar to the search device from the database using information about the lens included in the search device as key information. In this embodiment, "search" refers to extracting information corresponding to a search device from lens information stored in a database, or identifying information based on a relationship or relevance with the search device. A search device is a device selected as a search target by a user of an information processing device, and is equipped with a lens or a group of lenses, just like the target device.
[0042] The format for outputting the lens information is not particularly limited, and the lens information may be displayed on a screen, or a sound indicating the content of the lens information may be generated (played back). The method for selecting the target device for which the lens information is to be output is also not particularly limited, and for example, information about the lenses of a target device whose similarity with the search device (similarity will be described later) is equal to or greater than a reference value may be output. Alternatively, only information about the lenses of the top N target devices (N is a natural number equal to or greater than 1) in order of highest similarity may be output. Alternatively, lens information may be output for all target devices in descending order of similarity.
[0043] Furthermore, by identifying patent publications associated with the searched lens information, it is possible to find patent publications that contain lens cross-sectional images showing lenses or lens groups of the same or similar type as the lenses or lens groups provided in the search device.
[0044] As described above, in this embodiment, it is possible to identify the lens that appears in a lens cross-sectional image included in a document such as a patent publication, and to create a database of information about the lens. By using the database, it is possible to identify a lens that meets predetermined conditions and find a document that includes a lens cross-sectional image showing that lens.
[0045] <About the discrimination model> The identification model used in this embodiment (hereinafter referred to as the identification model M1) will be described with reference to Fig. 3. The identification model M1 is a model for identifying a lens or a group of lenses appearing in a target image. As shown in Fig. 3, the identification model M1 of this embodiment is composed of a derived model Ma and a specific model Mb.
[0046] The derived model Ma is a model that derives the feature quantities of an existence region in a target image by inputting the target image. The existence region is a rectangular region in the target image where a lens exists, as shown in Fig. 4, and in this embodiment, one lens (more specifically, an image of a lens cross section) exists in one existence region. In a target image showing a lens group, as shown in Fig. 4, there will be the same number of existence regions as the number of lenses that make up the lens group.
[0047] The derived model Ma is configured, for example, by a convolutional neural network (CNN) having a convolutional layer and a pooling layer in the intermediate layer. Examples of CNN models include the 16-layer CNN (VGG16) from Oxford visual geometry group, Google's Inception model (GoogLeNet), Kaiming He's 152-layer CNN (Resnet), and Chollet's improved Iception model (Xception).
[0048] When deriving the feature quantities of the existence region using the derived model Ma, the existence region is identified in the target image. Specifically, an image of one lens cross section is detected for each lens in the target image, and a rectangular region surrounding the detected image is set for each lens in the target image. The function of identifying the existence region in the target image is implemented in the derived model Ma by machine learning, which will be described later.
[0049] The features output from the derived model Ma are learned features in the convolutional neural network CNN, and are features identified in the process of general image recognition (pattern recognition).The features derived by the derived model Ma are input to the identification model Mb for each region.
[0050] The specific model Mb is a model that, when the feature quantities of the existence region derived by the derived model Ma are input, identifies the lens type corresponding to the feature quantities and geometric information of the existence region. The geometric information of the existence region includes the position of the existence region in the target image, such as the coordinates (x and y coordinates) of one vertex in the existence region, the width of the existence region (length in the x direction), and the height of the existence region (length in the y direction).
[0051] In this embodiment, the identification model Mb is configured, for example, by a neural network (NN), and identifies multiple candidates (candidate lens types) when identifying the lens type corresponding to the feature amount of the existence area. A softmax function is applied to the multiple identified candidates, and a confidence level is calculated for each candidate. The confidence level is a numerical value indicating the probability (likelihood, prediction accuracy) that each of the multiple candidates corresponds to the lens type present in the existence area. The sum of n confidence levels (n is a natural number) to which the softmax function is applied is 1.0.
[0052] The identification model Mb identifies, among multiple candidates identified for one presence area, a candidate determined based on a certainty level, for example, the candidate with the highest certainty level, as the type of lens present in the presence area. As described above, according to the identification model Mb, the type of lens appearing in the target image is selected from multiple candidates identified based on the feature values of the lens presence area, based on the certainty level of each candidate. When outputting the candidate selected as the type of lens, the certainty level of that candidate may also be output.
[0053] According to the identification model M1 described above, it is possible to identify a lens present in an existence region based on the feature amount of the existence region in the target image.
[0054] The identification model M1 (more specifically, each of the two models Ma and Mb constituting the identification model M1) is constructed by machine learning using a training dataset consisting of multiple training images showing lens cross sections and ground truth labels indicating the type of lens. In this embodiment, the lens cross sections shown in the training images are the cross sections of any of four types of spherical lenses. On the other hand, if the target device's lenses include aspherical lenses, machine learning may be performed using training images showing the cross sections of the aspherical lenses to construct an identification model capable of identifying aspherical lenses. However, even if the target device's lenses include aspherical lenses, an identification model that does not detect aspherical lenses may be constructed. Alternatively, an identification model that detects aspherical lenses by replacing them with the closest shape of one of the four types of spherical lenses may be constructed. Regarding the number of learning data sets used in machine learning, the more the better from the viewpoint of improving the accuracy of learning, and preferably 50,000 or more.
[0055] For the lens cross-sectional images used as learning images, an image in the normal position and images rotated 90 degrees, 180 degrees, and 270 degrees from the normal position may be prepared, and these four types of images may be included in the learning images.
[0056] The learning images may be randomly selected lens cross-sectional images, or lens cross-sectional images selected under certain conditions may be used as learning images.
[0057] The machine learning performed to construct the discriminative model M1 is, for example, supervised learning, and its method is deep learning (i.e., a multi-layer neural network). However, the type of machine learning (algorithm) is not limited to this, and the type of machine learning may be unsupervised learning, semi-supervised learning, reinforcement learning, or transduction. The machine learning technique may be genetic programming, inductive logic programming, Boltzmann machine, matrix factorization (MF), factorization machine (FM), support vector machine, clustering, Bayesian network, extreme learning machine (ELM), or decision tree learning. In neural network machine learning, methods such as gradient descent or backpropagation may be used to minimize the objective function (loss function).
[0058] Furthermore, in the machine learning of this embodiment, multiple training images including two or more lens cross-sectional images showing cross sections of the same type of lens but with different drawing styles (hereinafter referred to as "same type lens cross-sectional images") may be used. When the lens cross-section is a line drawing, the drawing style may include the line thickness, color, line type, and degree of line inclination, the way curved portions are drawn (curvature, etc.), the lens orientation, the dimensions of each part of the lens, the scale ratio, the presence or absence of auxiliary lines such as a center line, the presence or absence of lines indicating the optical path of light passing through the lens, the background color, the hatching style, and the presence or absence of suffixes such as symbols. Furthermore, when the lens cross-sectional image is represented as image data in a bitmap format, for example, the resolution (the density of the grid representing the image) is also included in the drawing style.
[0059] In the above case, machine learning is performed to construct an identification model M1 (strictly speaking, a derived model Ma) that derives common features from two or more cross-sectional images of lenses of the same type. For example, supervised learning is performed by attaching a correct answer label indicating the type of lens to each of multiple learning images of the same type of lens that have different drawing styles. This constructs an identification model M1 that derives common features from two or more cross-sectional images of lenses of the same type and can identify lenses shown in each cross-sectional image of lenses of the same type as lenses of the same type.
[0060] <Configuration of the information processing device according to this embodiment> Next, a description will be given of an example of the configuration of the information processing device 10 shown in Fig. 5. In Fig. 5, the external interface is described as "external I / F".
[0061] 5, the information processing device 10 is a computer in which a processor 11, a memory 12, an external interface 13, an input device 14, an output device 15, and a storage 16 are electrically connected to one another. In this embodiment, the information processing device 10 is configured by one computer, but the information processing device 10 may also be configured by multiple computers.
[0062] The processor 11 is configured to execute a program 21, which will be described later, and to perform processing for realizing the functions of the information processing device 10. The processor 11 is configured, for example, by one or more CPUs (Central Processing Units) and the program 21, which will be described later.
[0063] The hardware processor constituting the processor 11 is not limited to a CPU, but may be a field programmable gate array (FPGA), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or other integrated circuits (ICs), or a combination thereof. Furthermore, the processor 11 may be a single integrated circuit (IC) chip that performs the overall functions of the information processing device 10, such as a system on chip (SoC). Furthermore, one processing unit possessed by the information processing device of the present invention may be configured by one of the various processors described above, or may be configured by a combination of two or more processors of the same or different types, for example, a combination of multiple FPGAs, or a combination of an FPGA and a CPU, etc. Furthermore, the multiple functions of the information processing device of the present invention may be configured by one of various processors, or two or more of the multiple functions may be configured together by one processor. Alternatively, one or more CPUs and software may be combined to form one processor, which may then implement multiple functions. The above-mentioned hardware processor may be an electric circuit that combines circuit elements such as semiconductor elements.
[0064] The memory 12 is configured by semiconductor memory such as a ROM (Read Only Memory) and a RAM (Random Access Memory). The memory 12 provides a working area for the processor 11 and temporarily stores various data generated by the processing executed by the processor 11.
[0065] The memory 12 stores a program 21 for causing a computer to function as the information processing device 10 of this embodiment. The program 21 includes the following programs p1 to p5. p1: A program for constructing a discriminative model M1 using machine learning p2: A program for detecting the presence of lenses in a target image p3: A program to identify lenses present in the detected area p4: A program for storing information about identified lenses and building a database p5: A program for outputting lens information stored in a database
[0066] The program 21 may be acquired by reading it from a computer-readable recording medium, or may be acquired by receiving (downloading) it via a network such as the Internet or an intranet.
[0067] The external interface 13 is an interface for connecting to an external device. The information processing device 10 communicates with the external device, such as a scanner or another computer on the Internet, via the external interface 13. Through such communication, the information processing device 10 can acquire data for machine learning or data on patent publications in which target images are published.
[0068] The input device 14 includes, for example, a mouse and a keyboard, and accepts input operations from the user. For example, the user can draw a lens cross-sectional image through the input device 14, thereby allowing the information processing device 10 to acquire a learning image to be used for machine learning. Furthermore, for example, when searching for information about lenses of a target device of the same or similar type as the search device, the user operates the input device 14 to input information about the lenses provided in the search device. This allows the information processing device 10 to acquire input information about the lenses provided in the search device.
[0069] The output device 15 is comprised of, for example, a display and a speaker, and is a device for displaying or playing back the search results based on the input information, for example, information about lenses of target devices of the same or similar type as the searched device. The output device 15 can also output lens information stored in the database 22.
[0070] The storage 16 is configured by, for example, a flash memory, a hard disc drive (HDD), a solid state drive (SSD), a flexible disc (FD), a magneto-optical disc (MO disc), a compact disc (CD), a digital versatile disc (DVD), a secure digital card (SD card), and a universal serial bus memory (USB memory). The storage 16 stores various data including data for machine learning. Furthermore, the storage 16 also stores various models constructed by machine learning, including the identification model M1.
[0071] Furthermore, information about the lens of the target device identified from the target image (lens information) is stored in storage 16 in association with information about the document containing the target image, more specifically, the identification number of the patent publication in which the target image is published, etc. In other words, a database 22 of lens information is constructed in storage 16.
[0072] In the case shown in Figure 2, the information associated with the lens information includes the identification information of the patent publication as well as the figure number assigned to the target image in the patent publication. If the document containing the target image is a paper, the paper title and the page on which the target image is published can be associated with the lens information. If the document containing the target image is a book, the book title and the page on which the target image is published can be associated with the lens information.
[0073] As shown in FIG. 2, the lens information stored in the database 22 includes character information indicating the types of lenses that appear in the target image. In particular, when a lens group appears in the target image, the lens information includes string information that represents the types of lenses in the lens group in the order in which the lenses are arranged. In the case shown in FIG. 2, the database 22 stores, as string information about the lens group, two types of string information about the lens group: one that represents the lens types in the order from one end of the lens group to the other end, and one that represents the lens types in the order from the other end to the first end (hereinafter also referred to as mirror-image string information). This is to enable searching of both string information in which the types of lenses in the lens group are arranged in the forward direction and string information in which the types are arranged in the reverse direction (for example, the strings NTOP and POTN). However, this is not limited to this, and it is also possible to store only string information in which the types of lenses in the lens group are arranged in one direction without storing mirror-image string information. In this case, the stored string information may be converted into string information in which the order is reversed during searching.
[0074] In the present embodiment, the storage 16 is a device built into the information processing device 10, but the present invention is not limited to this, and the storage 16 may include an external device connected to the information processing device 10. The storage 16 may also include an external computer (for example, a server computer for a cloud service) communicably connected via a network. In this case, part or all of the database 22 described above may be stored in the external computer constituting the storage 16.
[0075] The hardware configuration of the information processing device 10 is not limited to the above-described configuration, and components can be added, omitted, or replaced as appropriate depending on the specific embodiment.
[0076] <Information processing flow> Next, an information processing flow using the information processing device 10 will be described. The information processing flow described below employs the information processing method of the present invention. That is, each step in the information processing flow described below constitutes the information processing method of the present invention and is performed by processor 11 of the computer that constitutes information processing device 10. Specifically, each step in the information processing flow is performed by processor 11 executing program 21.
[0077] The information processing flow of this embodiment proceeds in the order of a learning phase S001, a database construction phase S002, and an information output phase S003, as shown in Fig. 6. Each phase will be described below.
[0078] [Learning Phase] The learning phase S001 is a phase in which machine learning is performed to build models required in subsequent phases. In the learning phase S001, a first machine learning S011 and a second machine learning S012 are performed, as shown in FIG.
[0079] The first machine learning S011 is machine learning for constructing the discrimination model M1, and as described above, is performed using a plurality of training images showing lens cross sections. In this embodiment, supervised learning is performed as the first machine learning S011. In supervised learning, a training image showing a single lens cross section and a ground truth label indicating the type of lens that appears in the training image are used as a training dataset.
[0080] Furthermore, in the first machine learning S011, as described above, there are cases where a plurality of learning images including two or more cross-sectional images of the same type of lens showing cross sections of the same type of lens with different drawing styles are used. In this case, an identification model M1 (strictly speaking, a derivation model Ma) is constructed so as to derive common features from the two or more cross-sectional images of the same type of lens.
[0081] The second machine learning S012 is machine learning for constructing a detection model for detecting a lens presence area in a target image. The detection model is a model for detecting a presence area from a target image using an object detection algorithm.
[0082] As object detection algorithms, R-CNN (Region-based CNN), Fast R-CNN, YOLO (You Only Look Once), and SDD (Single Shot Multibox Detector) can be used. In this embodiment, an image detection model using YOLO is constructed from the viewpoint of detection speed. As YOLO, Yolo-v (version) 3 and Yolo-v4 can be used.
[0083] The training data (teacher data) used in the second machine learning S012 is created by applying an annotation tool to the training images. The annotation tool is a tool that annotates the target data with relevant information such as a correct label (tag) and the coordinates of the target object. Examples of annotation tools that can be used include LabeImg by Tzutalin and VoTT by Microsoft.
[0084] To create the learning data used in the second machine learning S012, for example, an image showing the cross section of a lens or a group of lenses is prepared as a learning image. Specifically, a lens cross-sectional image is extracted from a patent publication in which the lens cross-sectional image is published. Then, an annotation tool is started to display the learning image, and the area where the lens is located is enclosed by a bounding box, and the area is annotated (labeled) to create the learning data.
[0085] Furthermore, the method for creating learning data may be a method other than the above method; for example, learning images may be prepared using a lens image generation program. The lens image generation program is a program that automatically draws lens cross-sectional images by specifying the type of spherical lens and setting parameters for each part of the lens (e.g., the radius of curvature of the curved portion and the thickness of the central portion, etc.). By randomly setting each parameter for each type of spherical lens using the lens image generation program, it is possible to acquire a large number of lens cross-sectional images of each type. The acquired lens cross-sectional images are used as learning data together with the lens type specified when the image was created.
[0086] By performing the second machine learning S012 using the above learning data, a detection model that is a YOLO-style object detection model is constructed.
[0087] [Database construction phase] The database construction phase S002 is a phase in which the lens or lens group appearing in the target image published in the patent publication is identified, and information (lens information) relating to the identified lens or lens group is accumulated to construct the database 22.
[0088] In the database construction phase S002, first, the processor 11 of the information processing device 10 extracts a target image from a patent publication, and applies the above-described detection model to the extracted target image to detect a presence area in the target image (S021). That is, in this step S021, the processor 11 detects a presence area of a lens in the target image using an object detection algorithm (specifically, YOLO).
[0089] In this case, if the target image contains images of multiple lens cross sections, such as a target image showing a cross section of a part including a lens group, the processor 11 detects the presence area for each lens in the target image (see Figure 4).
[0090] Next, the processor 11 identifies the lens present in the existence area based on the feature amount of the existence area using the identification model M1 (S022). Specifically, the processor 11 inputs the image fragment of the existence area detected in step S021 to the identification model M1. In the identification model M1, the feature amount of the existence area is derived in the previous-stage derived model Ma.
[0091] In the subsequent identification model Mb, the type of lens present in the presence area and geometric information of the presence area are identified based on the feature quantities of the presence area input from the derived model Ma. At this time, multiple lens type candidates are identified based on the feature quantities of the presence area, and a certainty factor is calculated for each candidate. In the identification model Mb, for example, the candidate with the highest certainty factor is identified as the type of lens present in the presence area. However, this is not limited thereto, and for example, all candidates whose certainty factors satisfy a predetermined condition may be identified as the type of lens present in the presence area.
[0092] If multiple presence regions are detected in step S021, the lens identification process using the identification model M1 (i.e., step S022) is repeatedly executed for each presence region. As a result, for a target image showing a cross section of a portion including a lens group, each lens in the lens group can be identified for each presence region. That is, for each of multiple presence regions included in the target image, candidate types of lenses present in the presence region, the certainty of the candidates, and geometric information of the presence region (position in the target image) are identified for each region.
[0093] Furthermore, when each lens in the lens group is identified for each existence area, the processor 11 aggregates information about each lens in the lens group identified for each existence area as a unit of the lens group. Specifically, the processor 11 generates information about the lens group, for example, character string information, based on the type of each lens in the lens group identified for each existence area and the existence area position (coordinates).
[0094] To explain an example of the procedure for generating character string information, processor 11 calculates the center position of the x coordinate or y coordinate for each of multiple presence areas in the target image. If there are two or more presence areas with calculated center positions close to each other, the lenses present in each of the two or more presence areas are considered to belong to the same lens group. Then, for two or more lenses belonging to the same lens group, character information (code) indicating the lens type is arranged in order starting from the lens closest to the reference position. By this procedure, character string information such as that shown in FIG. 7 is obtained. At this time, mirror image character string information in the reverse order may also be generated.
[0095] Note that, as shown in Fig. 8, there are cases where a target image shows multiple lens groups for reasons such as illustrating the operation or state transition of each lens in the lens group. In this case, according to the generation procedure described above, character string information can be generated for each lens group, and in the case shown in Fig. 8, character string information can be generated for each of the upper lens group and the lower lens group.
[0096] Next, processor 11 stores information about the lenses of the identified target device, specifically, character information indicating the type of lens, in storage 16 (S023). Furthermore, for target devices equipped with lens groups, the character information indicating the type of each lens in the lens group identified for each presence area is aggregated with the lens group as a unit to generate character string information, and the generated character string information is stored (accumulated) in storage 16.
[0097] 8, when multiple lens groups are shown in the target image, character string information is generated for each lens group as described above. In this case, the character string information generated for each lens group is stored (accumulated) in the storage 16 for each lens group.
[0098] Furthermore, the processor 11 stores (accumulates) the lens information consisting of character information or character string information in association with information about the document including the target image (for example, identification information of a patent publication, etc.).
[0099] The above-described series of steps S021 to S023 are repeatedly performed for each target image, with the target image being changed. As a result, information about the lens of the target device (lens information) is stored and accumulated in storage 16 for each target image, and as a result, a database 22 of lens information about the lenses of the target device is constructed. Then, according to database 22, the target device in which the lens information is stored, the target image showing the lens of the target device, and the patent publication in which the target image is published can be searched using the lens information as a key.
[0100] [Information output phase] The information output phase S003 is a phase in which lens information stored in the database 22 is output in accordance with predetermined conditions. Specifically, at the start of the information output phase S003, the user performs an input operation related to the lenses provided in the search device. Here, it is assumed that the lenses provided in the search device are a lens group consisting of multiple lenses.
[0101] The processor 11 of the information processing device 10 acquires input information indicating the content of the input operation (S031). In this step S031, the processor 11 acquires, as the input information, character information indicating the types of lenses included in the search device, more specifically, character string information in which the types of lenses in the lens group are input in order.
[0102] After acquiring the input information, processor 11 compares the lens group of the target device, whose lens information is stored in database 22, with the lens group provided in the search device (S032). Specifically, processor 11 calculates the similarity between the character string information about the lens group of the search device indicated by the input information acquired in S031 and the character string information stored for each target image in database 22. The calculated similarity corresponds to the similarity between the lens (lens group) of the search device and the lens (lens group) of the target device.
[0103] In this embodiment, the Levenshtein distance method is used to evaluate the similarity between character strings. However, the algorithm for calculating the similarity between character strings is not particularly limited, and may be, for example, a Gestalt pattern matching method, a Jaro-Winkler distance method, or other similarity calculation methods.
[0104] In this embodiment, two types of similarity can be calculated. When calculating one of the two types, the first similarity, character string information about the lens group of the search device acquired as input information and character string information about the lens group of the target device stored in database 22 are used. Based on this character string information, the first similarity between the lens group of the search device and the lens group of the target device is calculated.
[0105] An example of calculating the first similarity will be described using the following two pieces of character string information A and B as an example. String information A:NNNOTNTTOTNTOT String information B:NNNOTNOOOTNTOT The first similarity is calculated using the Levenshtein distance method, which evaluates the number of times characters are deleted and added until two pieces of string information being compared match, and the number of times (score) is used as the similarity. Since the seventh and eighth characters of the above string information A and B differ, two characters must be deleted and two characters added to match the two pieces of string information. Therefore, the similarity between the above two pieces of string information A and B, i.e., the first similarity, is 4 (= 2 + 2). Note that the smaller the first similarity, the more similar the two pieces of string information are.
[0106] According to the first similarity, the pieces of character string information are compared as they are, so that the similarity between the pieces of character string information can be evaluated simply and directly.
[0107] When calculating the second similarity, which corresponds to the other of the two types of similarity, any character in the string information indicated by the input information, for example, one character, is changed to blank (a blank space). The second similarity is calculated based on the changed string information and the string information indicated by the lens information accumulated in database 22. In other words, the second similarity between the lens group of the search device and the lens group of the target device is calculated based on the string information in which some characters have been changed to blank for the lens group of the search device acquired as input information, and the string information for the lens group of the target device accumulated in database 22.
[0108] Like the first similarity, the second similarity is calculated using the Levenshtein distance method. Specifically, the number of times (score) characters are deleted and added until the two pieces of string information being compared match is evaluated, and this number is used as the similarity. Furthermore, when comparing the string information, characters that have been changed to blanks are not judged as to whether they match. In other words, the second similarity is calculated by comparing the characters in the string information excluding the blanks. For example, taking the above-mentioned two pieces of string information A and B as an example, if the seventh character of string information A is changed to a blank, the score will be 2, and this score may be used as the second similarity. String information A:NNNOTN_TOTNTOT String information B:NNNOTNOOOTNTOT
[0109] As a method for calculating the second similarity, the parts of the string information indicated by the input information that are to be changed to blanks may be changed in order, a score may be calculated for each piece of string information after the change, and the average value may be calculated as the second similarity.
[0110] According to the second similarity described above, even if an error or missed detection occurs when identifying each lens in a group of lenses appearing in a target image using the identification model M1, the similarity between character string information can be appropriately calculated taking this into account.
[0111] In this embodiment, as described above, the character string information is compared between the lens group of the search device and the lens group of the target device, and the similarity is evaluated as the comparison result. At this time, the character string information stored in the database 22 may be clustered, and the similarity may be evaluated by identifying the cluster to which the character string information for the lens group of the input search device belongs.
[0112] After calculating the similarity, processor 11 outputs information about the lenses of the target device stored in database 22, more specifically, character string information about the group of lenses included in the target device, in association with the search device based on the above-mentioned input information (S033). "Outputting information about the lenses of the target device (e.g., character string information) in association with the search device" means outputting information about the group of lenses of the target device (e.g., character string information) stored in database 22 in a manner that allows the relationship between the lenses of the search device and the lenses of the target device (specifically, the degree of similarity between the lens groups) to be recognized.
[0113] In step S033, the processor 11 outputs information about the lenses of the target device stored in the database 22 in association with the calculated similarity, specifically in association with one or both of the first similarity and the second similarity. To explain in more detail, for example, from the character string information about the lens group of the target device stored in the database 22, only the character string information having the highest first similarity or the highest second similarity may be extracted and output.
[0114] However, the present invention is not limited to the above case, and for example, character string information in which the first similarity or the second similarity exceeds a preset reference value may be extracted and the extracted character string information may be output. Alternatively, as shown in FIG. 9, M items (M is a natural number of 2 or more) having the highest first similarity or second similarity may be extracted in descending order, and the extracted M items of character string information may be output. Alternatively, all of the character string information about the lens groups of the target devices stored in the database 22 may be output in order of the highest (or lowest) first similarity or second similarity.
[0115] Alternatively, character string information extracted based on both the first similarity and the second similarity may be output. For example, the average value of the first similarity and the second similarity may be calculated, and the character string information with the highest average value, or M pieces of character string information extracted in descending order of average value may be output. Alternatively, all character string information may be output in order of average value. In this case, the average value may be a simple average value of the first similarity and the second similarity. Alternatively, the average value may be calculated by multiplying each of the first similarity and the second similarity by a weight (coefficient). In this case, the weight multiplied by the first similarity may be greater than or smaller than the weight multiplied by the second similarity.
[0116] The manner in which lens information (specifically, character string information about a lens group) is output in association with the similarity is not particularly limited as long as it is possible to recognize a lens group that is more similar to the lens group included in the search device.
[0117] Furthermore, when the character string information is output, for example, a document in which a lens cross-sectional image of the lens group indicated by the character string information (i.e., the target image) is published, specifically, an identification number of a patent publication, etc., is also output (see FIG. 9). This makes it possible to find a patent publication in which a lens cross-sectional image of a lens group similar to the lens group of the search device into which the user has input the character string information is published.
[0118] <Effectiveness of this embodiment> The information processing device 10 of this embodiment can identify a lens based on features of a region where a lens is present in a target image showing a cross section of a portion of a target device that includes the lens, using an identification model M1 constructed by machine learning. Furthermore, the information processing device 10 of this embodiment stores information about the identified lens of the target device (lens information) in association with information about a document in which the target image is published, thereby constructing a database 22. The lens information, document information, and other information stored in the database 22 are used in a searchable state.
[0119] To elaborate on the above effect, conventional technology establishes rules for the correspondence between lens features appearing in an image and lens types, and identifies lenses in an image according to those rules. However, if the drawing style of a lens differs from the normal style, and no identification rules are prepared that can accommodate that drawing style, there is a risk that the lens drawn in the abnormal style cannot be identified. In such cases, the lens identification result cannot be obtained, and the lens information obtained from the identification result is difficult to use.
[0120] In contrast, in this embodiment, by utilizing the identification model M1, which is the result of machine learning, it is possible to effectively identify lenses present in an existence region based on the feature amounts of the existence region in a target image. In other words, in this embodiment, even if the drawing style has changed, it is possible to identify the feature amounts of the existence region of a lens drawn in that drawing style, and once the feature amounts have been identified, it is possible to identify the lens from the feature amounts.
[0121] Then, lens information about the identified lenses is accumulated for each target image and stored in a database, so that the lens information can be used in a searchable manner thereafter. Furthermore, in the database, the lens information is associated with information about the document in which the target image is published. This allows the lens information to be used as key information to find the desired document. For example, it is possible to find documents that contain cross-sectional images of lenses of a group of lenses that are the same as or similar to the group of lenses included in the search device.
[0122] <Other embodiments> Up to this point, the information processing device, information processing method, and program of the present invention have been described using specific examples, but the above-described embodiment is merely an example, and other embodiments are also possible. For example, the computer constituting the information processing device may be a server used for an ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), IaaS (Infrastructure as a Service), or the like. In this case, a user using a service such as the ASP operates a terminal (not shown) to send input information related to a search device to the server. Upon receiving the input information, the server outputs lens information stored in database 22 to the user's terminal based on the input information. The information sent from the server is displayed or played aloud on the user's terminal.
[0123] Furthermore, in the above embodiment, the machine learning (first and second machine learning) for constructing various models is performed by the information processing device 10, but this is not limited to this. Part or all of the machine learning may be performed by another device (computer) different from the information processing device 10. In this case, the information processing device 10 acquires a model constructed by machine learning performed by the other device. For example, when the first machine learning is performed by the other device, the information processing device 10 acquires an identification model M1 from the other device and identifies lenses that appear in the target image using the acquired identification model M1.
[0124] In the above embodiment, the input information acquired in the information output phase S003 is string information indicating the type of each lens in the lens group provided by the search device. In the above embodiment, the similarity between the search device and the target device is calculated by comparing the input string information with string information about the lens group of the target device stored in the database 22. However, this is not limited to this, and the input information may be, for example, a cross-sectional image of the lens group provided by the search device. In this case, each lens in the lens group provided by the search device is identified from the lens cross-sectional image, which is the input information, and string information about the lens group is generated based on the identification result. When identifying each lens in the lens group, the identification model M1 constructed by the first machine learning may be used, and in this case, transfer learning may also be performed.
[0125] Furthermore, when the input information is a lens cross-sectional image, a model for calculating the similarity between images may be used. That is, the similarity may be calculated between a lens cross-sectional image showing the lens group of the search device and a lens cross-sectional image showing the lens group of the target device (i.e., the target image). In this case, the similarity calculation model may be a model that highly evaluates the similarity between two lens cross-sectional images (same type lens cross-sectional images) in which the same type of lens is drawn in different drawing styles. Specifically, the same label (correct answer label) may be assigned to multiple lens cross-sectional images of the same type, and machine learning may be performed using the labeled image data to construct the similarity calculation model.
[0126] In the above-described embodiment, the area where the lens exists in the target image is automatically detected using a detection model constructed by machine learning, but the present invention is not limited to this. For example, the target image may be displayed on a screen, and the user may specify the area where the lens exists through the screen (for example, by surrounding it with a bounding box or by inputting the coordinates of each vertex of the area), and the area where the lens exists may be detected based on that operation.
[0127] In the above embodiment, one lens exists in one presence area in the target image, and the identification model M1 identifies the one lens existing in one presence area. However, this is not limited to this, and one presence area may contain multiple lenses. In this case, the identification model M1 determines whether multiple lenses exist in one presence area based on the feature amount of the presence area, and if multiple lenses exist, identifies the combination of lenses. [Explanation of symbols]
[0128] 10. Information processing equipment 11 processors 12 Memory 13 Communication Interface 14 Input Devices 15 Output Devices 16. Storage 21 Programs 22 Databases M1 discrimination model Ma derivation model Mb Output Model
Claims
1. An information processing device including a processor, The processor: Detecting a region where the lens is present in a target image showing a cross section of a portion including the lens in a target device equipped with the lens; identifying the lens of the target device present in the presence area based on the feature amount of the presence area using an identification model constructed by machine learning using a plurality of learning images showing cross sections of the lens; The identification model is constructed by machine learning using the plurality of learning images, which include two or more cross-sectional images of the same type of lens showing cross sections of the same type of lens with different drawing styles, and is a model that identifies the lenses shown in each of the two or more cross-sectional images of the same type of lens as lenses of the same type.
2. The information processing device according to claim 1 , wherein the processor accumulates information about the lenses of the identified target devices to build a database of the information.
3. The processor: Obtain input information regarding the lens provided in the search device; The information processing apparatus according to claim 2 , wherein information relating to the lens of the target device stored in the database is output in association with the search device based on the input information.
4. The processor: Obtain input information regarding the lens provided in the search device; Calculating a similarity between the lens of the search device and the lens of the target device based on the input information and information about the lens of the target device stored in the database; The information processing apparatus according to claim 2 , wherein information about the lens of the target device stored in the database is output in association with the degree of similarity.
5. The processor: detecting the presence region for each lens in the target image showing a cross section of a portion including a group of lenses arranged in a row in the target device; Identifying the lenses in the lens group present in each of the presence areas by the identification model, The information processing apparatus according to claim 2 , wherein information about the lenses in the lens group identified for each of the existence regions is collected as a unit of the lens group and stored in the database.
6. The processor: Identifying the type of lens in the lens group identified for each of the existence regions, generating character string information representing the type of each lens in the lens group in the order in which the lenses are arranged in the lens group based on the type of lens in the lens group identified for each of the existence areas and the position of the existence area in the target image; The information processing apparatus according to claim 5 , wherein the generated character string information is stored in the database.
7. When a plurality of lens groups are shown in the target image, the processor generates the character string information for each lens group; The information processing apparatus according to claim 6 , wherein the character string information generated for each lens group is stored in the database for each lens group.
8. The processor: Obtain input information regarding the lens provided in the search device; In the case where the search device includes a lens group, acquiring, as the input information, character string information representing the type of each lens in the order in which the lenses are arranged in the lens group of the search device; calculating a first similarity between the lens group of the search device and the lens group of the target device based on the acquired character string information about the lens group of the search device and the character string information about the lens group of the target device stored in the database; The information processing apparatus according to claim 6 , wherein the character string information about the lens group of the target device stored in the database is output in association with the first similarity.
9. The processor: Obtain input information regarding the lens provided in the search device; In the case where the search device includes a lens group, acquiring, as the input information, character string information representing the type of each lens in the order in which the lenses are arranged in the lens group of the search device; changing any character in the character string information about the lens group of the search device obtained to a blank; calculating a second similarity between the lens group of the search device and the lens group of the target device based on the changed character string information about the lens group of the search device and the character string information about the lens group of the target device stored in the database; The information processing apparatus according to claim 6 , wherein the character string information about the lens group of the target device stored in the database is output in association with the second similarity.
10. The information processing apparatus according to claim 2 , wherein the processor stores, in the database, information about the lens of the identified target device in association with information about a document including the target image.
11. The information processing device according to claim 1 , wherein the processor detects an area where the lens is present from the target image by using an object detection algorithm.
12. The processor: Obtain input information regarding the lens provided in the search device; In the case where the search device includes a lens group, acquiring, as the input information, character string information representing the type of each lens in the order in which the lenses are arranged in the lens group of the search device; In the character string information about the lens group of the search device that has been acquired, any character is excluded from comparison; calculating a second similarity between the lens group of the search device and the lens group of the target device based on the character string information for the lens group of the search device excluding the characters not to be compared and the character string information for the lens group of the target device stored in the database; The information processing apparatus according to claim 6 , wherein the character string information about the lens group of the target device stored in the database is output in association with the second similarity.
13. A step of detecting, by a processor, a region where the lens is present in a target image showing a cross section of a portion including the lens in a target device equipped with the lens; Identifying the lens of the target device present in the presence area based on the feature amount of the presence area using an identification model constructed by machine learning using a plurality of learning images showing cross sections of lenses by a processor; Including, The identification model is constructed by machine learning using the plurality of learning images, which include two or more cross-sectional images of the same type of lens showing cross sections of the same type of lens with different drawing styles, and is a model that identifies the lenses shown in each of the two or more cross-sectional images of the same type of lens as lenses of the same type.
14. The information processing method of claim 13 , further comprising the step of: accumulating, by a processor, information relating to the identified lenses of the target devices to build a database of said information.
15. A program that causes a processor to execute each step of the information processing method according to claim 13 or 14.
Citation Information
Patent Citations
Lens
JP1999337707A
Lens information registration system, lens information server used for the same and camera body
JP2014183565A
Lens information registration system, lens information server used for the same, and method for controlling operation thereof
JP2016178704A
Image processing apparatus, method and program
JP2018175217A