Unstructured data searching method, searching device, and program
Patent Information
- Application Number
- JP2024571462
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-01-16
- Filing Date
- 2023-01-16
- Publication Date
- 2025-09-30
Abstract
Description
Unstructured data search method, search device, and program
[0001] The present invention relates to a method, device, and program for searching unstructured data.
[0002] Conventionally, it has been known to "predict drug discovery target proteins for diseases to be treated" (see, for example, Patent Document 1). Problem to be solved
[0003] Provides a more efficient way to explore unstructured data. General disclosure
[0004] In a first aspect of the present invention, there is provided a method for searching unstructured data, comprising the steps of: acquiring original data, which is unstructured data; determining predetermined first constraint conditions according to the original data; generating a first group of candidate data by changing parameters included in the original data according to the first constraint conditions; determining at least one candidate data based on the first group of candidate data; determining additional constraint conditions different from the first constraint conditions; generating an additional group of candidate data by changing parameters included in the at least one candidate data according to the additional constraint conditions; and outputting the additional group of candidate data.
[0005] In the above-described method for searching unstructured data, the step of determining the additional constraints may include a step of determining the additional constraints based on the first group of candidate data.
[0006] In any of the above methods for searching unstructured data, the step of determining the at least one candidate data may include a step of determining as the candidate data at least one of unstructured data included in the first candidate data group or structured data that indicates the properties of the first candidate data group.
[0007] In any of the above methods for searching unstructured data, the step of determining the additional constraints may include a step of determining the additional constraints based on at least one of unstructured data included in the first candidate data group or structured data indicating the properties of the first candidate data group.
[0008] In any of the above methods for searching unstructured data, at least one of the first constraint or the additional constraint may include fixed parameters, which are parameters to which the constraint is applied, and variable parameters, which are parameters to which the constraint is applied, that are variable.
[0009] In any of the above methods for searching unstructured data, the variable parameters may include at least one of an increase parameter that increases a parameter of structured data indicating the nature of the object of application, or a decrease parameter that decreases a parameter of structured data indicating the nature of the object of application.
[0010] In any of the above methods for searching unstructured data, the fixed parameters may include a direction specification parameter that specifies a direction of change in the unstructured data to which the fixed parameters are applied.
[0011] In any of the above methods for searching unstructured data, the direction specification parameter may be set based on a difference between the unstructured data.
[0012] In any of the above methods for searching unstructured data, the step of generating the first group of candidate data may include a step of generating the first group of candidate data with different variable parameters of the first constraint condition under a condition in which the fixed parameters of the first constraint condition are fixed.
[0013] In any of the above methods for searching unstructured data, the step of generating the additional candidate data group may include a step of generating the additional candidate data group in which the variable parameters of the additional constraint conditions are different under conditions in which the fixed parameters of the first constraint conditions and the fixed parameters of the additional constraint conditions are fixed.
[0014] In a second aspect of the present invention, there is provided a program for causing a computer to execute any one of the above methods for searching unstructured data.
[0015] In a third aspect of the present invention, there is provided an unstructured data search device comprising: an original data acquisition unit that acquires original data, which is unstructured data; a first determination unit that determines a first constraint condition that is predetermined in accordance with the original data; a first generation unit that generates a first group of candidate data by changing parameters included in the original data according to the first constraint condition; a candidate data determination unit that determines at least one candidate data based on the first group of candidate data; an additional determination unit that determines an additional constraint condition that is different from the first constraint condition; an additional generation unit that generates an additional group of candidate data by changing parameters included in the at least one candidate data according to the additional constraint condition; and an output unit that outputs the additional candidate data group.
[0016] The unstructured data search device may include a parameter determination unit that determines parameters to be constrained by the first constraint or the additional constraint.
[0017] The above summary of the invention does not list all of the features of the present invention, and subcombinations of these features may also be inventions.
[0018] 2 shows an overview of the configuration of the search device 100. FIG. 2 shows an example of a search method using the search device 100. FIG. 3 shows an example of a flowchart for realizing the search method. FIG. 4 shows a search method of a comparative example. FIG. 5 shows an example of a method for determining a parameter P by the parameter determination unit 25. FIG. 6 shows an example of a method for determining a parameter P by the parameter determination unit 25. FIG. 7 shows an example of a method for determining a parameter P by the parameter determination unit 25. FIG. 8 shows an example of a method for applying a constraint using a directionality designation parameter 63. FIG. 9 shows an example of a method for applying a constraint using a directionality designation parameter 63. FIG. 10 shows an example of a method for applying a constraint using a directionality designation parameter 63. FIG. 11 shows an example of a method for applying a constraint using a directionality designation parameter 63. FIG. 12 shows an example of a computer 2200 in which multiple aspects of the present invention may be embodied in whole or in part.
[0019] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.
[0020] 1A shows an overview of the configuration of a search device 100. The search device 100 includes an original data acquisition unit 10, a constraint condition determination unit 20, a data generation unit 30, a candidate data determination unit 40, and an output unit 50.
[0021] The searching device 100 obtains a search result Rs obtained by searching unstructured data that meets predetermined constraints. For example, the searching device 100 searches for candidate compounds for drug discovery under predetermined constraints and outputs the candidate compounds as the search result Rs. Note that although the description may use compounds as unstructured data, the type of unstructured data is not limited to compounds.
[0022] Unstructured data may be chemical compounds, character strings, sentences, images, videos, music, 3D objects, graph structures, blueprints, object layouts, or human layouts. Structured data may be data structured, such as rows and columns.
[0023] The original data acquiring unit 10 acquires original data Do, which is unstructured data. The original data acquiring unit 10 may acquire, as the original data Do, unstructured data manually input by a user of the search device 100. The original data acquiring unit 10 may acquire, as the original data Do, any unstructured data generated using a computer or the like.
[0024] The original data Do is unstructured data that is input to obtain the desired search results Rs. The original data Do may include chemical compounds such as inorganic or organic compounds. The original data Do may include character strings, sentences, images, videos, music, 3D objects, graph structures, blueprints, object layouts, or human layouts.
[0025] The constraint condition determination unit 20 determines constraint conditions to be applied to the original data Do. The constraint condition determination unit 20 may acquire the original data Do from the original data acquisition unit 10. The constraint condition determination unit 20 may determine constraint conditions for obtaining a desired search result Rs according to the original data Do. The constraint condition determination unit 20 may determine multiple constraint conditions. The constraint condition determination unit 20 in this example has a first determination unit 21 and an additional determination unit 22 for determining multiple constraint conditions.
[0026] The first determination unit 21 determines a predetermined first constraint C1 according to the original data Do. The first constraint C1 may be a fixed parameter for fixing a parameter P of the original data Do, or a variable parameter for changing the parameter P of the original data Do under a predetermined condition. The fixed parameter and the variable parameter will be described later.
[0027] The data generation unit 30 generates a candidate data group from the original data Do using the constraint conditions determined by the constraint condition determination unit 20. The data generation unit 30 may acquire the original data Do from the original data acquisition unit 10. The data generation unit 30 in this example has a first generation unit 31 and an additional generation unit 32 for applying the constraint conditions multiple times to generate candidate data groups multiple times.
[0028] The first generator 31 generates a first candidate data group G1 by changing parameters included in the original data Do according to a first constraint C1. The first candidate data group G1 may include a large number of unstructured data items obtained by applying the first constraint C1 to the original data Do.
[0029] The parameter P is information contained in the unstructured data. The parameter P may be the structure of a compound, such as the number of atoms or the number of aromatic rings. If the unstructured data is an image containing a person's face, the parameter P may be the angle of the person's face, the attractiveness level, or the size of the face. If the unstructured data is a video, the parameter P may be the number of characters, the rating of the video, or the genre. If the unstructured data is a character string, the parameter P may be the classification of the character string (e.g., a noun or a verb) or the number of characters. If the unstructured data is a sentence, the parameter P may be whether the sentence is an official document or a chat, or if the unstructured data is a written answer to a test, the parameter P may be the score, etc. If the unstructured data is audio, the parameter P may be gender, voice quality, or voice pitch.
[0030] The parameter P may be explicitly determined from the unstructured data, or may be calculated based on the unstructured data or predicted from the unstructured data by machine learning. That is, the parameter P may be a property of the unstructured data, such as the activity or water solubility of a compound.
[0031] The candidate data determination unit 40 determines at least one candidate data Dc based on the first candidate data group G1. The candidate data determination unit 40 may determine, as the candidate data Dc, at least one of unstructured data included in the first candidate data group G1 or structured data indicating the properties of the first candidate data group G1. If the unstructured data is a compound, the properties of the first candidate data group G1 may be the characteristics of the compound. The candidate data determination unit 40 may determine the candidate data Dc based on information input by a user who references the first candidate data group G1. The candidate data determination unit 40 may determine the candidate data Dc based on predetermined conditions, or may randomly determine the candidate data Dc from the first candidate data group G1.
[0032] The additional determination unit 22 determines additional constraint conditions Ca that are different from the first constraint conditions C1. The additional determination unit 22 may acquire the first candidate data group G1 from the first generation unit 31. The additional determination unit 22 may determine the additional constraint conditions Ca based on the first candidate data group G1. Specifically, the additional determination unit 22 may determine the additional constraint conditions Ca based on at least one of the unstructured data included in the first candidate data group G1 and the structured data that indicates the properties of the first candidate data group G1. In this way, by determining the additional constraint conditions Ca based on the result of applying the first constraint conditions C1, it is possible to gradually restrict the degree of freedom of search for unstructured data. Furthermore, the additional determination unit 22 may determine the additional constraint conditions Ca based on the candidate data Dc acquired from the candidate data determination unit 40.
[0033] The additional generation unit 32 generates an additional candidate data group Ga by changing a parameter P included in at least one candidate data Dc according to an additional constraint condition Ca. By using an appropriately determined additional constraint condition Ca, the additional generation unit 32 can generate an additional candidate data group Ga that is closer to the desired search result Rs.
[0034] The output unit 50 outputs the search result Rs. The output unit 50 may output the additional candidate data group Ga as the search result Rs. The output unit 50 may store the search result Rs in an arbitrary storage device, may display it on a display unit such as a display, or may transmit the search result Rs to an external device. The output unit 50 may display, as the search result Rs, the properties of each of the additional candidate data group Ga together with the structures of the compounds in the additional candidate data group Ga.
[0035] The user may determine the compound to be synthesized from the additional candidate data group Ga of the search result Rs. If a desirable candidate compound is obtained as the search result Rs, the user may actually synthesize the selected compound and measure its physical properties experimentally. On the other hand, even if a desirable candidate compound is not obtained as the search result Rs, the user can gain new insights by learning about the search process. Based on suggestions obtained from the presented unstructured data and data indicating the properties of the unstructured data, the user can impose constraints on the conditions for generating the next unstructured data and generate multiple pieces of unstructured data under those constraints.
[0036] The parameter determination unit 25 determines a parameter P to be constrained in order to determine the constraint condition. The parameter determination unit 25 may determine fixed parameters and variable parameters of the constraint condition. The parameter determination unit 25 may determine whether to increase, decrease, or maintain the parameter P. The method of determining the parameter P by the parameter determination unit 25 will be described later.
[0037] The search device 100 of this example can support the search for better unstructured data candidates while interacting with the user step by step. The search device 100 can reduce the number of search steps to the desired unstructured data by repeating the process of determining constraints and generating candidate data multiple times.
[0038] 1B shows an example of a search method using the search device 100. The search device 100 of this example searches for compounds while interacting with a user step by step.
[0039] The first generation unit 31 generates a first candidate data group G1 based on the original data Do. The first candidate data group G1 may include candidate data 1 to candidate data n. In addition to the first candidate data group G1, the first generation unit 31 may predict the characteristics of each of the first candidate data group G1 and provide the predicted characteristics to the user.
[0040] The user may determine candidate data Dc to which the next constraint condition is applied from the first candidate data group G1, or may refer to the first candidate data group G1 and determine data different from candidate data 1 to candidate data n as candidate data Dc. The user may actually synthesize any of the candidate data in the first candidate data group G1 and measure the physical property values, and then determine the candidate data Dc to be input to the additional generation unit 32.
[0041] The additional generation unit 32 generates an additional candidate data group Ga based on the candidate data Dc. The additional candidate data group Ga may include additional candidate data 1 to additional candidate data n. In addition to the additional candidate data group Ga, the additional generation unit 32 may predict the properties of each of the additional candidate data group Ga and provide the predicted properties to the user. The prediction of the properties of the compound may be achieved by machine learning or the like.
[0042] The search method of this example can improve the efficiency of the search process up to obtaining the search result Rs by generating multiple candidate data groups such as the first candidate data group G1 and the additional candidate data group Ga. By knowing the process up to the generation of the additional candidate data group Ga, the user can select any additional candidate data from the additional candidate data group Ga with higher accuracy, actually combine the data, and measure the physical property values.
[0043] 1C is an example of a flowchart for implementing a search method. In step S100, original data Do is acquired. The original data Do may be input by a user, or may be generated by the search device 100 in response to a user request. In step S102, a first constraint C1 is determined according to the original data Do. A constraint with a high priority from among multiple constraints may be determined as the first constraint C1.
[0044] In step S104, a first candidate data group G1 is generated from the original data Do. By using a constraint condition with a high priority as the first constraint condition C1, it is possible to extract a first candidate data group G1 that meets the request. In step S106, at least one candidate data Dc is determined. In step S108, an additional constraint condition Ca that is different from the first constraint condition C1 is determined. In step S110, an additional candidate data group Ga is generated from the candidate data Dc. In step S112, the additional candidate data group Ga is output.
[0045] In step S112, it is determined whether or not to continue the search. If the search is to be continued, the process may return to step S106 and candidate data Dc may be determined. Here, the candidate data Dc may be determined based on the additional candidate data group Ga. Any candidate data in the additional candidate data group Ga may be determined as candidate data Dc, or data different from the candidate data included in the additional candidate data group Ga may be determined as candidate data Dc by referring to the additional candidate data group Ga. In this way, by repeating the search and adding constraints, the degree of freedom can be reduced and the search efficiency can be improved.
[0046] FIG. 2 shows a comparative search method. In this search method, original data is input to a data generation unit, and a candidate data set is directly generated. In other words, in the comparative search method, the user cannot know how the candidate data set was generated, and cannot obtain suggestions for improvement. Therefore, if the generated candidate data set does not produce desirable results, it is difficult to improve it. Furthermore, when multiple parameters P are changed simultaneously, it is difficult to understand the contribution of each parameter P, and it is not easy to determine the relative merits of the parameters P.
[0047] 3A shows an example of a method for determining the parameter P by the parameter determination unit 25. In this example, a method for determining the parameter P will be described assuming that the unstructured data is a compound, but the method for determining the parameter P in this example may be similarly applied to unstructured data other than compounds.
[0048] The parameter determination unit 25 may determine fixed parameters 61 and variable parameters 62. In this example, the fixed parameters 61 and variable parameters 62 are structures of compounds. At least one of the first constraint C1 or the additional constraint Ca includes fixed parameters 61 for fixing a parameter P to which the constraint is applied, and variable parameters 62 for changing a parameter P to which the constraint is applied.
[0049] The data generating unit 30 may generate a first candidate data group G1 in which the variable parameters 62 of the first constraint C1 are different under the condition that the fixed parameters 61 of the first constraint C1 are fixed. In this example, the data generating unit 30 generates compounds in the first candidate data group G1 by varying the structure specified in the variable parameters 62 under the condition that the structure specified in the fixed parameters 61 is fixed.
[0050] In addition to the first candidate data group G1, the searching device 100 may display the characteristics, such as the physical properties, of each of the first candidate data group G1. This makes it possible to visualize which structure affects the activity of the compound, and to obtain suggestions for determining the candidate data Dc and the additional constraint condition Ca. In this example, one of the compounds in the first candidate data group G1 is designated as the candidate data Dc.
[0051] The structure to be set as the fixed parameters 61 may be specified by the user. The searching device 100 may specify the properties of a compound as the fixed parameters 61. The searching device 100 may fix the activity of a compound by specifying the fixed parameters 61. The activity of a compound may be its binding ability with a target molecule that causes a disease. The searching device 100 may specify a portion of the compound that contributes to the activity as the fixed parameter 61. The searching device 100 may change the compound so that it exhibits water solubility within a predetermined range by specifying the variable parameters 62.
[0052] 3B shows an example of a method for determining the parameter P by the parameter determination unit 25. In this example, a case will be described in which a candidate data group Ga is generated by applying an additional constraint condition Ca to the candidate data Dc in FIG. 3A.
[0053] In this example, the additional constraint Ca applies a constraint to the candidate data Dc using fixed parameters 61 and variable parameters 62. The fixed parameters 61 in this example fix the structure and properties of the compound. The variable parameters 62 in this example are the structure of the compound other than the fixed parameters 61.
[0054] The data generation unit 30 of this example generates an additional candidate data group Ga in which the variable parameters 62 of the additional constraints Ca are different under conditions in which the fixed parameters 61 of the first constraints C1 and the fixed parameters 61 of the additional constraints Ca are fixed. That is, the fixed parameters 61 of this example specify the same compound structure as the fixed parameters 61 of FIG. 3A , as well as property 1 as a fixed parameter. This makes it possible to extract an additional candidate data group Ga that further satisfies the property 1 condition from the first candidate data group G1 generated in FIG. 3A . In this way, by obtaining only the results limited by the results obtained in the first synthesis in the second synthesis, it is possible to efficiently search for a desired compound.
[0055] The searching device 100 may specify at least one candidate data Dc' based on the additional candidate data group Ga. In this example, one of the compounds included in the additional candidate data group Ga is specified as the candidate data Dc'.
[0056] 3C shows an example of a method for determining the parameter P by the parameter determination unit 25. In this example, a case will be described in which a candidate data group Ga′ is generated by applying an additional constraint Ca′ to the candidate data Dc′ in FIG.
[0057] The additional constraint Ca in this example applies constraints to the candidate data Dc′ using fixed parameters 61 and variable parameters 62. The fixed parameters 61 in this example fix the structure and properties of the compound. The variable parameters 62 in this example are the structure of the compound other than the fixed parameters 61 and property 2, which is different from property 1. That is, the additional constraint Ca in this example specifies property 2 as variable parameter 62 in addition to the fixed parameters 61 and variable parameters 62 specified in FIG. 3B . This makes it possible to extract an additional candidate data group Ga′ by further changing property 2 by a predetermined value from the additional candidate data group Ga generated in FIG. 3B . In this way, the desired compound can be efficiently searched for by obtaining, in the third synthesis, only the results limited by the results obtained in the first and second synthesis.
[0058] [0] The search device 100 may be used for various purposes other than searching for compounds. The search device 100 may be used to search for defective product images for visual inspection. The search device 100 can generate candidate data in which the position or shape of the defective part is changed by setting parameters other than the defective part as fixed parameters 61 and the defective part as variable parameters 62. The search device 100 may generate candidate data for defective product images by specifying parameters such as the color of scratches, the amount of light, or the pattern, based on the original data Do, in order to create a more realistic defective product image. By learning the generated candidate data for defective product images as training data, the accuracy of detecting defective products can be improved.
[0059] 4 shows an example of a method for determining the parameter P by the parameter determination unit 25. The parameter P in this example is a property of a compound, which is unstructured data.
[0060] The variable parameters 62 may include at least one of an increase parameter that increases the parameter P of the structured data indicating the nature of the application object, or a decrease parameter that decreases the parameter P of the structured data indicating the nature of the application object.
[0061] In this example, for the parameters P of Characteristics 1 to 5, Characteristics 1 to 4 are designated as variable parameters 62, and Characteristic 5 is designated as a fixed parameter 61. Characteristics 1 and 2 are increasing parameters designated so that the characteristics of the compound after the change are improved compared to the compound before the change. Characteristics 3 and 4 are decreasing parameters designated so that the characteristics of the compound after the change are worse compared to the compound before the change. Characteristic 5 is a fixed parameter 61 designated so that the characteristics of the compound before and after the change are maintained.
[0062] The searching device 100 can acquire the search result Rs more efficiently by specifying the directionality of the variable parameters 62. For example, the searching device 100 can acquire a desired compound by specifying that the water solubility value of the compound be increased and that the toxicity value be decreased.
[0063] When the unstructured data is an image, the search device 100 may specify the gender of a person included in the image as the fixed parameter 61 and the height and face size as the variable parameters 62. The search device 100 may output a search result Rs in which the height of the person included in the image is reduced and the face is made smaller while maintaining the gender.
[0064] Similarly, when the unstructured data is music, the searching device 100 can change the time proportions of the chorus, verse, and bridge without changing the overall time. The searching device 100 may change the melody while fixing the musical range of the music. The searching device 100 may also change the musical range in stages.
[0065] When the unstructured data is a video, the search device 100 can change the appearance time of characters or the number of characters. The search device 100 may gradually change and optimize the background video, behavior patterns of characters, appearance order, or conversation content of characters that appear in the video.
[0066] In this way, the search device 100 can refine unstructured data by specifying the parameter P as a characteristic and repeatedly applying constraints, thereby enabling a search for unstructured data that meets the user's needs.
[0067] 5A shows an example of a method for applying constraints using a direction specification parameter 63. The fixed parameters 61 in this example include the direction specification parameter 63.
[0068] The directionality specification parameter 63 specifies the directionality of change in the unstructured data A to which it is applied. The search device 100 can specify a direction and search for the unstructured data group A' by using the directionality specification parameter 63 as unstructured data separate from the unstructured data A. The search device 100 in this example uses text A as the directionality specification parameter 63, but is not limited to this.
[0069] In this example, we will explain the case where an additional candidate data group Ga is generated from candidate data Dc, but the direction specification parameter 63 can also be used in the same way when generating a first candidate data group G1 from original data Do.
[0070] The direction specification parameter 63 may be a sentence such as "A compound with high water solubility and small molecular weight, which is likely to be made into a drug" for a compound. The direction specification parameter 63 may be a sentence such as "A girl standing alone in nature" for an image. Furthermore, when the candidate data Dc is an "image of nature," the direction specification parameter 63 may be a sentence such as "Add a seated man over 60 years old."
[0071] The direction specification parameter 63 may be a sentence such as "change the tuning midway through and raise the range" for music, or a sentence such as "featuring a Japanese person in a Hollywood-style movie" for video.
[0072] The search device 100 of this example can easily search for candidates that are in line with the direction requested by the user by using the direction specification parameter 63. The search device 100 of this example can support a more interactive search with the user.
[0073] 5B shows an example of a method for applying a constraint using a direction specification parameter 63. In this example, the direction specification parameter 63 is a single piece of unstructured data. The direction specification parameter 63 includes unstructured data B.
[0074] The data generation unit 30 uses the unstructured data B of the directionality specification parameter 63 as the additional constraint condition Ca to generate an additional candidate data group Ga based on the unstructured data A, which is the candidate data Dc. For example, the data generation unit 30 generates an unstructured data group A' in which unstructured data B is added to unstructured data A, as the additional candidate data group Ga. For example, the data generation unit 30 generates an unstructured data group A' in which an ornament in unstructured data B is added to a character image in unstructured data A.
[0075] When the unstructured data is a compound, the data generation unit 30 may specify a specific core (i.e., main structure) or a specific substituent (i.e., partial structure) as unstructured data B and generate an unstructured data group A' having a specific structure as an additional candidate data group Ga.
[0076] When the directionality specification parameter 63 indicates a defective portion, the data generation unit 30 can generate an unstructured data group A' in which the defective portion is added to the candidate data Dc as an additional candidate data group Ga.
[0077] 5C shows an example of a method for applying a constraint using a direction specification parameter 63. In this example, the direction specification parameter 63 is a set of unstructured data. The direction specification parameter 63 includes an unstructured data group B.
[0078] The data generation unit 30 generates an additional candidate data group Ga based on the unstructured data A, which is the candidate data Dc, using the unstructured data group B of the directionality specification parameter 63 as an additional constraint condition Ca. For example, the data generation unit 30 generates an unstructured data group A' by specifying the art style of the unstructured data group B for the character image of the unstructured data A.
[0079] If unstructured data A is music, the style of the music can be specified by using multiple songs as structured data group B. If unstructured data A is video, the atmosphere or genre of the video can be specified by using multiple videos as structured data group B.
[0080] The search device 100 of this example can specify a more abstract directionality as a constraint by using an unstructured data group as the directionality specification parameter 63 .
[0081] 5D shows an example of a method for applying constraint conditions using a direction specification parameter 63. The direction specification parameter 63 in this example is set based on the difference between the unstructured data.
[0082] The data generation unit 30 generates an additional candidate data group Ga based on the unstructured data A, using the difference between unstructured data B and unstructured data C of the directionality specification parameter 63 as an additional constraint condition Ca. For example, the data generation unit 30 generates an unstructured data group A' by specifying the painting style of the difference between unstructured data B and unstructured data C for the image of unstructured data A. For example, the search device 100 can input the difference between a Van Gogh-style painting and a Picasso-style painting and convert the Van Gogh-style painting into a Picasso-style painting.
[0083] The search device 100 of this example can add constraints from a new perspective by using the difference in unstructured data as the direction specification parameter 63 .
[0084] 6 illustrates an example of a computer 2200 in which aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 may cause the computer 2200 to function as or perform operations associated with an apparatus or one or more sections of the apparatus according to embodiments of the present invention, and / or to perform a process or steps of a process according to embodiments of the present invention. Such programs may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.
[0085] A computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphics controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes input / output units such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.
[0086] The CPU 2212 operates according to programs stored in the ROM 2230 and RAM 2214, thereby controlling each unit. The graphics controller 2216 acquires image data generated by the CPU 2212 into a frame buffer or the like provided in the RAM 2214 or into the graphics controller 2216 itself, and causes the image data to be displayed on the display device 2218.
[0087] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads programs or data from the DVD-ROM 2201 and provides the programs or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.
[0088] ROM 2230 stores therein a boot program or the like that is executed by computer 2200 upon activation, and / or programs that depend on the hardware of computer 2200. I / O chip 2240 may also connect various I / O units to I / O controller 2220 via a parallel port, a serial port, a keyboard port, a mouse port, etc.
[0089] The programs are provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The programs are read from the computer-readable medium, installed in the hard disk drive 2224, RAM 2214, or ROM 2230, which are also examples of computer-readable media, and executed by the CPU 2212. Information processing described in these programs is read by the computer 2200, and brings about cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by realizing information manipulation or processing in accordance with the use of the computer 2200.
[0090] For example, when communication is performed between computer 2200 and an external device, CPU 2212 may execute a communication program loaded into RAM 2214 and instruct communication interface 2222 to perform communication processing based on the processing described in the communication program. Under the control of CPU 2212, communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in RAM 2214, hard disk drive 2224, DVD-ROM 2201, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes received data received from the network to a reception buffer processing area or the like provided on the recording medium.
[0091] Furthermore, the CPU 2212 may cause all or a necessary portion of a file or database stored on an external recording medium such as the hard disk drive 2224, the DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. to be read into the RAM 2214, and may perform various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording medium.
[0092] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 2212 may perform various types of processing on data read from the RAM 2214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 2214. The CPU 2212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored on the recording medium, the CPU 2212 may search for an entry that matches a condition specified by the attribute value of the first attribute from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.
[0093] The above-described programs or software modules may be stored in a computer-readable medium on or near the computer 2200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable medium, thereby providing the programs to the computer 2200 via the network.
[0094] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.
[0095] It should be noted that the order of execution of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a subsequent process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order.
[0096] 10...Original data acquisition unit, 20...Constraint condition determination unit, 21...First determination unit, 22...Additional determination unit, 25...Parameter determination unit, 30...Data generation unit, 31...First generation unit, 32...Additional generation unit, 40...Candidate data determination unit, 50...Output unit, 100...Search device, 61...Fixed parameter, 62...Variable parameter, 63...Direction specification parameter, 2200...Computer, 2201...DVD-ROM, 2210...Host controller, 2212...CPU, 2214...RAM, 2216...Graphics controller, 2218...Display device, 2220...Input / output controller, 2222...Communication interface, 2224...Hard disk drive, 2226...DVD-ROM drive, 2230...ROM, 2240...Input / output chip, 2242...Keyboard
Claims
1. obtaining raw data, which is unstructured data; determining a first constraint condition in response to the original data; generating a first set of candidate data using the first constraint; determining additional constraints different from the first constraint; generating an additional candidate data set using the additional constraints; outputting the additional candidate data group; A method for searching unstructured data comprising:
2. The step of generating the first candidate data group includes a step of obtaining parameters corresponding to the original data using the first constraint condition. The method of claim 1 .
3. The unstructured data includes at least one of a chemical compound, a character string, a sentence, an image, a video, music, a 3D object, a graph structure, a blueprint, an arrangement of objects, or an arrangement of people. The method of claim 1 .
4. The parameter includes a sentence The method of searching unstructured data according to claim 2 .
5. Determining the additional constraints includes determining the additional constraints based on the first set of candidate data. The method of claim 1 .
6. A step of determining at least one candidate data based on the first candidate data group, The step of determining at least one candidate data includes a step of determining, as the candidate data, at least one of unstructured data included in the first candidate data group and structured data indicating a property of the first candidate data group. The method of claim 1 .
7. Determining the additional constraints includes determining the additional constraints based on at least one of unstructured data included in the first set of candidate data or structured data indicative of a characteristic of the first set of candidate data. The method of searching unstructured data according to claim 5.
8. At least one of the first constraint or the additional constraint includes a fixed parameter, which is a parameter to which the constraint is applied, and a variable parameter, which is a parameter to which the constraint is applied, that is variable. The method of claim 1 .
9. The variable parameters include at least one of an increase parameter that increases a parameter of the structured data that indicates the nature of the application target, or a decrease parameter that decreases a parameter of the structured data that indicates the nature of the application target. The method of searching unstructured data according to claim 8.
10. The fixed parameters include a direction specification parameter that specifies the direction of change of the unstructured data to which the fixed parameters are applied. The method of searching unstructured data according to claim 8.
11. The direction specification parameter is set based on the difference between the unstructured data. The method of searching unstructured data according to claim 10.
12. The step of generating the first candidate data group includes a step of generating the first candidate data group with different variable parameters of the first constraint condition under a condition in which the fixed parameters of the first constraint condition are fixed. The method of searching unstructured data according to claim 8.
13. The step of generating the additional candidate data group includes a step of generating the additional candidate data group with different variable parameters of the additional constraint conditions under a condition in which the fixed parameters of the first constraint conditions and the fixed parameters of the additional constraint conditions are fixed. The method of searching unstructured data according to claim 8.
14. A program for causing a computer to execute the method for searching unstructured data according to any one of claims 1 to 13.
15. an original data acquisition unit that acquires original data that is unstructured data; a first determination unit that determines a first constraint condition according to the original data; a first generation unit that generates a first candidate data group using the first constraint condition; an additional constraint determining unit that determines an additional constraint different from the first constraint; an additional generation unit that generates an additional candidate data group using the additional constraint condition; an output unit that outputs the additional candidate data group; An unstructured data search device comprising:
16. a parameter acquisition unit that acquires parameters using the first constraint or the additional constraint; The apparatus for searching unstructured data according to claim 15.