Information processing apparatus, inspection evaluating system, and inspection evaluating method
The information processing device effectively evaluates text content by segmenting and quantifying natural language descriptions, addressing the limitations of existing systems to enhance inspection accuracy and early detection of issues.
Patent Information
- Application Number
- JP2025197242
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-07-06
- Filing Date
- 2025-11-18
- Publication Date
- 2026-01-23
AI Technical Summary
Existing document ranking systems fail to properly evaluate the content of documents, particularly qualitative information such as natural language descriptions of inspection results, which are often omitted in quantitative evaluations.
An information processing device that includes a condition receiving unit, counting unit, blacklist and whitelist processing units, and an evaluation calculation unit to quantify and evaluate text data by dividing it into segments, applying N-grams or morphological analysis, and using dummy variables to reflect user inputs for accurate evaluation.
Enables appropriate evaluation of text content, enhancing the accuracy of inspections by incorporating qualitative information, thereby improving the effectiveness of evaluations and early detection of potential issues.
Smart Images

Figure 2026012579000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an inspection and evaluation system, and an inspection and evaluation method. [Background technology]
[0002] Patent Document 1 describes a document ranking device for ranking electronic documents (Di) in a file path of a file system taking into account the relevance of the documents to a search term (t), the device comprising: a semantic description generation module configured to generate a semantic description (SDi) of the document using the content of the document and store the semantic description in a semantic description repository; a similarity-based scoring module configured to calculate a similarity score based on the similarity between the semantic description of the document and the search term; a quality indicator-based scoring module configured to calculate a quality score of the document based on the completeness, accuracy, and freshness of the document; a combining module configured to receive user input for relative weighting of the similarity score and the quality score and to combine the resulting relatively weighted similarity score and quality score to provide a final score for the document; and a ranking module configured to rank the documents in the file path based on the final score. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-076208 Summary of the Invention [Problem to be solved by the invention]
[0004] Although the above techniques can formally rank documents, they are unable to properly evaluate the content of the documents.
[0005] An object of the present invention is to provide a technique that can appropriately evaluate the content of a text. [Means for solving the problem]
[0006] The present application includes multiple means for solving at least part of the above problems, examples of which are as follows: An information processing device according to one aspect of the present invention includes: a condition receiving unit that receives quantification conditions for text data, a counting processing unit that divides the text data into predetermined segments and counts the number of occurrences or the presence or absence of occurrences, an output unit that displays a result in which the quantification conditions are reflected in the counted number of occurrences or the presence or absence of occurrences, a dummy variable receiving unit that receives a user selection input for each segment on the display of the output unit as a designation of a dummy variable, and a dummy variable conversion unit that acquires a result of counting the number of occurrences or the presence or absence of occurrences of the segments accepted as the designation of the dummy variable, and the output unit outputs the result of counting the number of occurrences or the presence or absence of occurrences acquired by the dummy variable conversion unit.
[0007] Furthermore, the above-mentioned information processing device may be characterized in that the quantification conditions include a blacklist that specifies one or more words to be excluded from the category, and the output unit, in the process of reflecting the quantification conditions, is provided with a blacklist processing unit that uses the blacklist to exclude the words to be excluded from the category.
[0008] Furthermore, the above-mentioned information processing device may be characterized in that the quantification condition includes a whitelist that specifies one or more words to be added to the category, and in the process of the output unit reflecting the quantification condition, it is equipped with a whitelist processing unit that uses the whitelist to add the words to be added to the category and re-counts the number of occurrences or whether or not they occur.
[0009] In the information processing device, the aggregation processing unit may divide the text data into the predetermined categories using N-grams.
[0010] In addition, the above-mentioned information processing device may be characterized in that the quantification conditions include a whitelist that specifies one or more words to be added to the category, and in the process of the output unit reflecting the quantification conditions, it is provided with a whitelist processing unit that uses the whitelist to add the words to the category and re-aggregate the number of occurrences or whether or not they appear, and the aggregation processing unit uses N-grams to divide the text data into the specified categories, and combines the specified categories using the number of occurrences or whether or not they appear, and proposes words with a word length exceeding the value of N as candidates for the whitelist.
[0011] In addition, the above-mentioned information processing device may be characterized in that the quantification conditions include specification of parts of speech to be used as the categories, and the aggregation processing unit uses morphological analysis to divide the text data into the specified categories, and excludes from the aggregation any of the specified categories that do not correspond to the parts of speech.
[0012] The information processing device may also be characterized in that the text data is accompanied by one or more values of a predetermined measurement result, and the output unit adds the result of counting the number of occurrences or presence or absence of occurrence of the category obtained by the dummy variable conversion unit as the value of the measurement result for each category.
[0013] The information processing device may further include an evaluation calculation unit in which the text data includes a natural language description of the inspection results of the structure and is accompanied by one or more values of predetermined measurement results of the structure, and the output unit adds the results of counting the number of occurrences or presence or absence of occurrence of the categories obtained by the dummy variable conversion unit as the value of the measurement result for each category.
[0014] In addition, another aspect of the present invention provides an inspection and evaluation system that uses an information processing device, the information processing device comprising a control unit and a memory unit, wherein the memory unit stores one or more text data containing natural language descriptions related to inspection results of the structure, along with one or more values of predetermined measurement results of the structure, and the control unit performs the following steps: a condition receiving step for receiving quantification conditions for the text data; a counting step for dividing the text data into predetermined categories and tallying up the number of occurrences or presence or absence of occurrences; an output step for displaying a result in which the quantification conditions are reflected in the tallied number of occurrences or presence or absence of occurrences; a dummy variable receiving step for receiving a user's selection input for each category in the display of the output step as a designation of a dummy variable; a dummy variable conversion step for acquiring the result of counting the number of occurrences or presence or absence of occurrence of the category accepted as the designation of the dummy variable; a result output step for adding the result of counting the number of occurrences or presence or absence of occurrence acquired in the dummy variable conversion step as the value of the measurement result for each category; and an evaluation calculation step for calculating a predetermined evaluation index of the structure using the value of the measurement result.
[0015] In addition, another aspect of the present invention provides an inspection and evaluation method using an information processing device, the information processing device comprising a control unit and a memory unit, wherein the memory unit stores one or more text data containing a natural language description of the inspection results of the structure, along with one or more values of predetermined measurement results of the structure, and the control unit performs the following steps: a condition receiving step for receiving quantification conditions for the text data; a counting step for dividing the text data into predetermined categories and tallying up the number of occurrences or presence or absence of occurrence; an output step for displaying a result in which the quantification conditions are reflected in the tallied number of occurrences or presence or absence of occurrence; a dummy variable receiving step for receiving a user's selection input for each category in the display of the output step as a designation of a dummy variable; a dummy variable conversion step for acquiring the result of counting the number of occurrences or presence or absence of occurrence of the category accepted as the designation of the dummy variable; a result output step for adding the result of counting the number of occurrences or presence or absence of occurrence acquired in the dummy variable conversion step as the value of the measurement result for each category; and an evaluation calculation step for calculating a predetermined evaluation index of the structure using the value of the measurement result. [Effects of the Invention]
[0016] According to the present invention, it is possible to provide a technique that can appropriately evaluate the content of a text.
[0017] Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 1 is a block diagram illustrating an example of an inspection and evaluation system according to an embodiment. [Figure 2] FIG. 2 illustrates an example of a data structure of a conversion target data storage unit; [Figure 3] FIG. 10 illustrates an example of a data structure of a frequency storage unit; [Figure 4] FIG. 10 is a diagram illustrating an example of a data structure of a whitelist storage unit. [Figure 5] FIG. 2 is a diagram illustrating an example of the hardware configuration of a data quantification server device. [Figure 6] FIG. 10 is a diagram illustrating an example of a flow of data quantification processing. [Figure 7] FIG. 10 is a diagram showing an example of a conversion target specification screen. [Figure 8] FIG. 10 is a diagram showing an example of a frequently used word acquisition condition specification screen. [Figure 9] FIG. 10 is a diagram showing an example of the flow of word appearance count processing (using N-gram). [Figure 10] FIG. 10 is a diagram illustrating an example of the flow of a blacklist application process. [Figure 11] FIG. 10 is a diagram illustrating an example of a flow of a whitelist application process. [Figure 12] FIG. 10 is a diagram illustrating an example of a flow of a dummy variable selection process. [Figure 13] FIG. 10 is a diagram showing an example of a dummy variable designation screen. [Figure 14] FIG. 10 is a diagram illustrating an example of a flow of a dummy variable conversion process. [Figure 15] FIG. 10 is a diagram illustrating an example of a conversion result confirmation screen. [Figure 16] FIG. 10 is a diagram showing another example of the frequently occurring word acquisition condition specification screen. [Figure 17] FIG. 10 is a diagram showing an example of the flow of word appearance count processing (using morphological analysis). [Figure 18] FIG. 10 is a diagram showing another example of the dummy variable specification screen. [Figure 19] FIG. 10 is a diagram showing yet another example of the dummy variable specification screen. [Figure 20] FIG. 10 is a block diagram illustrating an example of a local information collection system according to a fourth embodiment. [Figure 21] FIG. 2 illustrates an example of a data structure of a conversion target data storage unit; [Figure 22] FIG. 10 is a diagram illustrating an example of a data structure of a time-address priority storage unit. [Figure 23] FIG. 10 is a diagram illustrating an example of the flow of information integration processing. [Figure 24] FIG. 10 is a diagram illustrating an example of a flow of a regional information aggregation process. [Figure 25] FIG. 10 is a diagram illustrating an example of a summary item setting screen. [Figure 26] FIG. 10 is a diagram showing an example of a single variable totalization screen. [Figure 27] FIG. 10 is a diagram illustrating an example of a multiple variable tabulation screen. [Figure 28] FIG. 10 is a diagram showing an example of a first-level cross table screen. [Figure 29] FIG. 10 is a diagram showing an example of a multi-layer cross table screen. [Figure 30] FIG. 10 is a diagram showing another example (time slice) of a one-layer cross table screen. [Figure 31] FIG. 10 is a diagram showing another example of a first-level cross table screen (display limited to ongoing items). [Figure 32] FIG. 10 is a diagram illustrating an example of a map display screen. [Figure 33] FIG. 10 is a diagram showing another example of a map display screen. [Figure 34] FIG. 10 is a block diagram of another example of the local information collection system according to the fourth embodiment. [Figure 35] FIG. 2 illustrates an example of a data structure of a conversion target data storage unit; [Figure 36] FIG. 10 is a diagram illustrating an example of a data structure of a dummy tag storage unit. [Figure 37] FIG. 10 is a diagram illustrating an example of a data structure of an inter-image tag similarity storage unit. [Figure 38] FIG. 10 is a diagram showing an example of an image fuzzy search screen. [Figure 39] FIG. 10 is a diagram showing an example of a tag similar image search screen. DETAILED DESCRIPTION OF THE INVENTION
[0019] An inspection and evaluation system 1 to which an embodiment according to one aspect of the present invention is applied will be described below with reference to the drawings. In the following embodiments, when necessary for convenience, the description will be divided into multiple sections or embodiments. However, unless otherwise specified, they are not unrelated to each other, and one is related to the other as a partial or complete modification, detail, supplementary explanation, etc.
[0020] Furthermore, in the following embodiments, when referring to the number of elements (including the number, numerical value, amount, range, etc.), unless otherwise specified or when it is clearly limited to a specific number in principle, it is not limited to that specific number and may be more or less than the specific number.
[0021] Furthermore, it goes without saying that in the following embodiments, the components (including element steps, etc.) are not necessarily essential unless otherwise specified or unless they are clearly considered essential in principle.
[0022] Similarly, in the following embodiments, when referring to the shapes, positional relationships, etc. of components, etc., it is intended to include those that are substantially similar or similar to those shapes, etc., unless otherwise specified or when it is considered that this is clearly not the case in principle. This also applies to the above numerical values and ranges.
[0023] In addition, in all the drawings for explaining the embodiments, the same components are generally designated by the same reference numerals, and repeated explanations thereof will be omitted.
[0024] 1 is a block diagram of an inspection and evaluation system 1 according to this embodiment. The inspection and evaluation system 1 is used by a user 10 using an information processing terminal (not shown) to connect to a data quantification server device 100 via a browser or the like, but is not limited to this and may be used by directly operating the data quantification server device 100.
[0025] Although not shown, when connecting from an information processing terminal to the data quantification server device 100, the connection is made via a network such as a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, a mobile phone network, or a combination of these. The network may be a VPN (Virtual Private Network) on a wireless communication network such as a mobile phone communication network.
[0026] Examples of applications of the inspection and evaluation system 1 include a business system that handles inspection results for the maintenance of specified structures (for example, public facilities such as bridges and roads), or a business system that handles inspection results for manufacturing deliverables.
[0027] In this case, the user 10 uses one or more text data containing a description of the test results in natural language in addition to one or more predetermined test items as test results for the test evaluation. These test results are updated as needed by an examiner or testing device (not shown).
[0028] The user 10 understands the actual state of the test subject by evaluating the test results, but since the test items are limited to quantitative ones, qualitative evaluations such as findings are often written in natural language. Furthermore, such qualitative evaluations often include the tester's experience and know-how, and are information that should be used for test evaluation.
[0029] However, in order to evaluate a large number of test results, it is more efficient to process large amounts of information using computers, so natural language descriptions of findings, etc. are often omitted when evaluating test results.
[0030] If natural language descriptions that reflect such experience and know-how could be reflected in the evaluation of inspection results, it would be possible to make highly effective use of information, further increasing the accuracy of evaluations and helping to detect and prevent serious incidents early.
[0031] In this embodiment, a business system that handles bridge inspection results will be described as an example. The data quantification server device 100 includes a storage unit 110, a control unit 120, an input unit 130, and an output unit 140, which are communicatively connected to one another via a bus or the like.
[0032] The storage unit 110 includes a conversion target data storage unit 111 , a frequency storage unit 112 , a blacklist storage unit 113 , and a whitelist storage unit 114 .
[0033] 2 is a diagram showing an example of the data structure of the conversion target data storage unit. The conversion target data storage unit 111 contains text data written in natural language. The text data also contains a description in natural language of the inspection results of the structure, along with one or more values of predetermined measurement results of the structure.
[0034] More specifically, the conversion target data storage unit 111 includes a bridge code 111A, which is an identifier that distinguishes a bridge from other bridges, an inspection date 111B, an inspector comment 111C, an X measurement value 111D, and a Y measurement value 111E.
[0035] The inspector comments 111C are text data written using the natural language described above. For example, they include comments such as "There is a crack in the main girder, so measures are needed." The X measurement value 111D and the Y measurement value 111E are the values of the inspection results of predetermined quantifiable inspection items.
[0036] 3 is a diagram showing an example of the data structure of the frequency storage unit 112. The frequency storage unit 112 includes a ranking 112A, a word 112B, an occurrence count 112C, a black designation 112D, a white designation 112E, a user display target 112F, and a quantification target 112G.
[0037] Rank 112A is a ranking assigned in descending order of occurrence count 112C, which is the frequency of occurrence (which may be the number of occurrences or whether or not an occurrence occurs) of a word (not limited to a word, but a predetermined segment into which text data is divided) identified by word 112B.
[0038] Black designation 112D is information indicating whether a word is included in the blacklist. Similarly, white designation 112E is information indicating whether a word is included in the whitelist. User display target 112F is information specifying whether a word is a target word (a predetermined segment into which text data is divided) to be displayed on a screen used by user 10. Quantification target 112G is information specifying whether a word is a target word (a predetermined segment into which text data is divided) to be quantified by user 10.
[0039] The blacklist storage unit 113 is a list that specifies one or more words (predetermined segments into which text data is divided) to be excluded from evaluation, and is a data structure, such as a list or array, that allows each word to be read and edited individually.
[0040] 4 is a diagram showing an example of the data structure of the whitelist storage unit 114. The whitelist storage unit 114 is a list that specifies one or more words to be added as words (predetermined segments into which text data is divided), but since there is a lot of accompanying information, it is illustrated as a table structure.
[0041] Word ID 114A is an identifier that identifies a word (a predetermined segment into which text data is divided). Word 114B is a predetermined segment into which text data is divided. Applicability 114C is information that identifies whether a word (a predetermined segment into which text data is divided) is to be applied as a whitelist.
[0042] The compound word 114D is information specifying whether or not it is a compounded word (a predetermined segment into which the text data is divided). The compound basis word 114E and compound adjacent word 114F are information indicating the base words (predetermined segments into which the text data is divided) when words (predetermined segments into which the text data is divided) are compounded. For example, if the compound word "ABC" is compounded with "A", which is often used adjacent to the beginning of "BC", the compound basis word 114E is "BC" and the compound adjacent word 114F is "A". In other words, if the segment of the word 114B is a compound word, "True" is stored in the compound word 114D, and the compound source words are stored in the compound basis word 114E and compound adjacent word 114F.
[0043] The control unit 120 includes a condition receiving unit 121, a counting unit 122, a blacklist processing unit 123, a whitelist processing unit 124, a dummy variable receiving unit 125, a dummy variable conversion unit 126, and an evaluation calculation unit 127.
[0044] The condition receiving unit 121 receives quantification conditions for text data. More specifically, the condition receiving unit 121 receives quantification conditions such as designation of text data to be quantified, one or more blacklisted words, one or more whitelisted words, a frequently occurring word acquisition condition (number of items to be displayed), or a part-of-speech filter.
[0045] The counting unit 122 divides the text data into predetermined segments and counts the frequency of occurrence thereof. More specifically, the counting unit 122 divides the text data into predetermined segments (words) using N-gram or morphological analysis. When the counting unit 122 divides the text data using morphological analysis, the counting unit 122 can also exclude from the counting of the frequency of occurrence any predetermined segment (word) that does not correspond to a specified part of speech (common noun, proper noun, person's name, place name, symbol, etc.).
[0046] Furthermore, when the text data is divided using N-grams, the counting unit 122 may combine predetermined categories (words) using the counted occurrence frequencies, and may propose words with a word length exceeding the value of "N" of the N-gram as candidates for the whitelist. For example, the counting unit 122 may propose co-occurrence expressions (expressions that are often used together when a certain word is used), or may propose synonyms, abbreviations (such as "crack" and "crackware"), or foreign words (such as "crack" and "crack").
[0047] For example, the aggregation processing unit 122 may also suggest compound words specific to the inspection object that are used by combining one word with another word (for example, for bridge inspections, "main girder" from "main" + "girder," "Gerber girder" from "Gerber" + "girder," "deck slab" from "deck" + "slab," etc.) as candidates for the whitelist.
[0048] In the process of reflecting the quantification conditions, the blacklist processing unit 123 uses the blacklist to exclude words that should be excluded from a predetermined category (word). That is, the blacklist processing unit 123 excludes words registered in the blacklist from the appearance frequencies tallied by the counting unit 122.
[0049] In the process of reflecting the quantification conditions, the whitelist processing unit 124 uses the whitelist to add words to a predetermined category (word) and re-counts the frequency of occurrence. The whitelist processing unit 124 also adds to the whitelist candidates specified by the user 10 from among the whitelist candidates proposed by the counting processing unit 122. The whitelist processing unit 124 also adds to the whitelist a predetermined category (word or free text) specified by the user 10. In the process of adding to the whitelist, the whitelist processing unit 124 changes the applicability 114C of the whitelist storage unit 114 from "False" to "True."
[0050] The dummy variable receiving unit 125 receives an input selected by the user for each predetermined category (word) on the display of the output unit 140, and designates a dummy variable. That is, when the dummy variable receiving unit 125 receives the selection of a word to be quantified, it changes the quantification target 112G in the frequency storage unit 112 from "False" to "True."
[0051] The dummy variable conversion unit 126 obtains the results of counting the frequency of occurrence of a specific category (word) accepted as a dummy variable specification from the aggregation processing unit 122, and stores the results of the dummy variable conversion for each word in the conversion target data storage unit 111 as one of the measurement values.
[0052] The evaluation calculation unit 127 uses the values of the measurement results to calculate a predetermined evaluation index for the structure being inspected. Various evaluation methods are possible, and any method may be used in this embodiment. For example, words that predict danger may be used as a factor in calculating a low evaluation for the object. In other words, the evaluation calculation unit 127 allows various processes as long as they use feature values of text data that are quantitatively calculated as dummy variables for evaluation.
[0053] The input unit 130 accepts input from the user 10 to the data quantification server device 100. For example, the input unit 130 accepts various types of input, such as various contact inputs such as typing, touching, and flick input, or various types of inputs such as voice input or eye gaze input.
[0054] The output unit 140 outputs information from the data quantification server device 100 to the user 10. The information that is output is various types of output information such as screens and forms.
[0055] 5 is a diagram showing an example of the hardware configuration of a data quantification server device. The data quantification server device 100 has a hardware configuration realized by the housing of a so-called server device, workstation, personal computer, smartphone, or tablet terminal. The data quantification server device 100 includes a calculation device 101, a main memory device 102, an auxiliary memory device 103, and a bus 107 connecting each device. In addition, the data quantification server device 100 includes a communication device for communicating with other devices via a network, and input / output devices such as a touch panel, keyboard, microphone, and display.
[0056] The arithmetic device 101 is, for example, a CPU (Central Processing Unit).
[0057] The main storage device 102 is a memory device such as a RAM (Random Access Memory).
[0058] The auxiliary storage device 103 is a non-volatile storage device capable of storing digital information, such as a so-called hard disk drive, a solid state drive (SSD), or a flash memory.
[0059] The input / output devices include various input / output devices such as a keyboard, a mouse, a touch panel, a display, a microphone, and a speaker.
[0060] The input / output device, the arithmetic unit 101, the main memory device 102, and the auxiliary memory device 103 are connected to one another by connecting wires such as a bus 107.
[0061] The condition receiving unit 121, the aggregation processing unit 122, the blacklist processing unit 123, the whitelist processing unit 124, the dummy variable receiving unit 125, the dummy variable conversion unit 126, and the evaluation calculation unit 127 of the data quantification server device 100 described above are realized by a program that causes the calculation device 101 to perform processing. This program is stored in the main memory device 102, the auxiliary memory device 103, or a ROM device (not shown), and is loaded onto the main memory device 102 for execution and executed by the calculation device 101.
[0062] The storage unit 110 of the data quantification server device 100 is realized by the main storage device 102 and the auxiliary storage device 103. The input unit 130 and the output unit 140 are realized by input / output devices. The above is an example of the hardware configuration of the data quantification server device 100.
[0063] The configuration of the data quantification server device 100 can be further divided into components according to the processing content, or one component can be divided into components that perform more processing.
[0064] Furthermore, each control unit (condition receiving unit 121, aggregation processing unit 122, blacklist processing unit 123, whitelist processing unit 124, dummy variable receiving unit 125, dummy variable conversion unit 126, and evaluation calculation unit 127) may be constructed using dedicated hardware (ASIC, GPU, etc.) that realizes each function. Furthermore, the processing of each control unit may be executed by one piece of hardware or by multiple pieces of hardware.
[0065] Next, the operation of the inspection and evaluation system 1 in this embodiment will be described.
[0066] 6 is a diagram showing an example of the flow of the data quantification process. The data quantification process is started in response to a start instruction from the user 10.
[0067] First, the condition receiving unit 121 performs a conversion target specification receiving process (step S001). Specifically, the condition receiving unit 121 receives information (file name, conversion target string) input on the conversion target specification screen 200.
[0068] 7 is a diagram showing an example of a conversion target specification screen 200. The conversion target specification screen 200 includes a file name input area 201, a file selection input area 202, a conversion target column selection input area 203, a file content display area 204, and a Next button 205.
[0069] The file name input area 201 accepts input of the data to be converted into a dummy variable, i.e., the path to the file. When the file selection input area 202 accepts input, it displays a directory tree and accepts the selection of the target file.
[0070] When the conversion target column selection input area 203 receives an input, it acquires information (column names) of columns in the file whose path is input in the file name input area 201, formats the information into a selectable list, and displays it.
[0071] The file content display area 204 acquires the contents of the file whose path was entered in the file name input area 201, formats it into display information that matches the file format (for example, if it is a CSV (Comma-Separated Values) file, formats it into a table format and displays it).
[0072] When the next button 205 receives an input, it transmits the input value of the file name input field 201 and the input value of the conversion target column selection input field 203 to the data quantification server device 100 .
[0073] Then, the condition receiving unit 121 performs a frequent word acquisition condition receiving process (step S002). Specifically, the condition receiving unit 121 receives information (top number of frequently used words to be acquired, blacklist, whitelist) input on the frequent word acquisition condition specification screen 300.
[0074] 8 is a diagram showing an example of a frequently used word acquisition condition specification screen 300. The frequently used word acquisition condition specification screen 300 includes an input area 301 for inputting the top number of frequently used words to be acquired, a blacklist input area 302, a whitelist input area 303, and a next button 304.
[0075] The input area 301 for inputting the number of frequently occurring words to be acquired accepts the specification of the maximum number of frequently occurring words (predetermined segments into which text data is divided) to be displayed or output as frequently occurring words.
[0076] The blacklist input area 302 accepts input of a blacklist, i.e., a list of words to be excluded based on the aggregated frequency of occurrence. This can include general words to be excluded, such as "mono" (thing), "koto" (thing), and "takara" (for), which have extremely little semantic information, as well as business-related words with little semantic information. These words to be excluded may also be displayed in advance. In this case, the user 10 can simply delete the words not to be excluded from the displayed list, which is more convenient for users 10 who are less familiar with the system.
[0077] The whitelist input area 303 accepts input of a whitelist, i.e., a list specifying one or more words (predetermined segments into which text data is divided) to be added to the aggregation of occurrence frequencies in addition to the aggregated occurrence frequencies. This may include neologisms, new words, foreign words, words of high business importance, or words with particularly large numbers of characters, which may be added as target words for addition. Furthermore, these target words for addition may be displayed in advance. In this case, the user 10 can simply delete the target words for addition that are not to be added from the pre-displayed target words for addition, which is more convenient for users 10 who are less familiar with the system.
[0078] When the next button 304 receives input, it sends the input value of the top frequently occurring word acquisition number input area 301, the input value of the blacklist input area 302, and the input value of the whitelist input area 303 to the data quantification server device 100.
[0079] Then, the counting unit 122 performs a word appearance counting process (step S003). Specifically, the counting unit 122 starts the word appearance counting process shown in FIG.
[0080] Then, the blacklist processing unit 123 performs the blacklist application process (step S004). Specifically, the blacklist processing unit 123 starts the blacklist application process shown in FIG.
[0081] Then, the whitelist processing unit 124 performs a whitelist application process (step S005). Specifically, the whitelist processing unit 124 starts the whitelist application process shown in FIG.
[0082] Then, the dummy variable receiving unit 125 receives the information input on the dummy variable designation screen 400 (the word for which the check input has been received and the input value of the word to be added to the whitelist) and starts the dummy variable selection process (step S006).
[0083] 13 is a diagram showing an example of a dummy variable specification screen. The dummy variable specification screen 400 includes a word appearance count display area 401, an applied whitelist display area 404, a whitelist addition word input area 405, a frequently occurring word acquisition condition change button 406, and a next button 407.
[0084] The word appearance count display area 401 is an area for displaying words in a table in descending order of appearance count. This display reflects the processing results of steps S003 to S005. The word appearance count display area 401 includes a check input acceptance area 402 for specifying dummy variables (subjects to quantification), and whitelist words 403 are displayed in an emphasized manner (highlighted, reversed, etc.).
[0085] Applied whitelist words are displayed in the applied whitelist display area 404. An additional whitelist word input area 405 accepts input of additional whitelist words that the user wishes to apply.
[0086] When an input is received, the frequent word acquisition condition change button 406 transitions to the frequent word acquisition condition specification screen 300 in order to change the frequently used word acquisition conditions.
[0087] When the next button 407 receives an input, it transmits the word for which the check input has been received in the check input receiving area 402 and the input value in the whitelist addition word input area 405 to the data quantification server device 100 .
[0088] Then, the dummy variable conversion unit 126 determines whether the dummy variable selection completion flag is ON or not (step S007). The dummy variable selection completion flag is a flag whose ON / OFF is controlled in the dummy variable selection process. If the flag is "OFF" ("No" in step S007), the dummy variable conversion unit 126 returns the control to step S002.
[0089] If the dummy variable selection completion flag is "ON" ("Yes" in step S007), the dummy variable conversion unit 126 performs a dummy variable conversion process (step S008). Specifically, the dummy variable conversion unit 126 performs a dummy variable conversion process to acquire the aggregation results of the aggregation processing unit 122 for the words designated as dummy variables (subjects to quantification). In addition, the dummy variable conversion unit 126 creates a conversion result confirmation screen 500 using the acquired aggregation results.
[0090] 15 is a diagram showing an example of a conversion result confirmation screen 500. The conversion result confirmation screen 500 includes a file name input area 501, a file selection input area 502, a dummy variable display area 503, and an output button 504.
[0091] The file name input area 501 accepts input of data converted into dummy variables, i.e., the path to the file. When the file selection input area 502 accepts input, it displays a directory tree and accepts the selection of the target file.
[0092] The dummy variable display area 503 displays the contents of the file input in the file name input area 501. When a file to which the results of conversion processing performed by the dummy variable conversion unit 126 have been added is specified in the file name input area 501, the file contents including information to which dummy variables (the number of occurrences of each word specified as a dummy variable) have been added are displayed.
[0093] When the output button 504 receives an input, it transmits the file name input in the file name input area 501 to the data quantification server device 100 .
[0094] Then, the dummy variable conversion unit 126 performs a conversion result saving process (step S009). Specifically, the dummy variable conversion unit 126 accepts the file name input in the file name input area 501 on the conversion result confirmation screen 500 and saves it as a file of the conversion result. The dummy variable conversion unit 126 also passes the file name to the evaluation calculation unit 127, and the evaluation calculation unit 127 uses the file to perform a predetermined evaluation on the inspection object.
[0095] This is an example of the flow of data quantification processing. Data quantification processing can quantify information contained only in free text from the frequency of occurrence of important words contained in the free text portion of a file that includes a free text area. This allows the content of the text to be appropriately evaluated.
[0096] 9 is a diagram showing an example of the flow of the word appearance count counting process (using N-gram). The word appearance count counting process starts in step S003 of the data quantification process.
[0097] First, the tallying unit 122 initializes the word appearance count list (step S0031). Then, the tallying unit 122 reads the conversion target data (step S0032). Specifically, the tallying unit 122 reads the file specified by the input value in the file name input field 201.
[0098] Then, the counting unit 122 determines whether or not there are any lines that have not been acquired (step S0033). If there are no lines that have not been acquired ("No" in step S0033), the counting unit 122 ends the word appearance counting process.
[0099] If there is an unacquired row ("Yes" in step S0033), the tallying unit 122 acquires one of the unacquired rows (step S0034).
[0100] Then, the aggregation processor 122 acquires the value of the target column from the acquired row (step S0035). Specifically, the aggregation processor 122 acquires unprocessed information for the column specified by the input value in the conversion target column selection input field 203, one by one, starting from the top.
[0101] Then, the counting unit 122 counts the number of times words appear in the value (free description) using N-grams (for all words) (step S0036). Specifically, the counting unit 122 reads out, as text, information on each row for the column specified by the input value in the conversion target column selection input area 203, divides the text using N-grams to create segments, and specifies the frequency of appearance of each segment.
[0102] The counting unit 122 then adds the counting result to the word appearance count list (step S0037). Specifically, the counting unit 122 compares the characters in each category with the characters in each category in the word appearance count list, and if they are within the range of variation, considers them to be the same and adds them to the appearance frequency. For characters in a category that are different from characters in any category, the identified appearance frequency is recorded in the frequency storage unit 112 as the appearance frequency of that category. The counting unit 122 then returns control to step S0033.
[0103] This completes the flow of the word occurrence count processing. According to the word occurrence count processing, the free text to be analyzed is sequentially read and classified using N-grams, and the frequency of occurrence for each classification is recorded in a word occurrence count list.
[0104] 10 is a diagram showing an example of the flow of the blacklist application process. The blacklist application process starts in step S004 of the data quantification process.
[0105] First, the blacklist processing unit 123 acquires a word appearance count list (step S0041). Specifically, the blacklist processing unit 123 receives the word appearance count list added in step S0037 of the word appearance count count processing (using N-gram).
[0106] Then, the blacklist processing unit 123 reads the blacklist (step S0042). Specifically, the blacklist processing unit 123 reads the blacklist received in step S002 of the data quantification process.
[0107] Then, the blacklist processing unit 123 determines whether there are any unchecked words (step S0043). If there are no unchecked words ("No" in step S0043), the blacklist processing unit 123 ends the blacklist application process.
[0108] If there are any unchecked words ("Yes" in step S0043), the blacklist processing unit 123 acquires the unchecked words from the word appearance frequency list (step S0044).
[0109] Then, the blacklist processing unit 123 determines whether the word is included in the blacklist (step S0045). If the word is not included in the blacklist ("No" in step S0045), the blacklist processing unit 123 returns control to step S0043.
[0110] If the word is included in the blacklist ("Yes" in step S0045), the blacklist processing unit 123 excludes the word from the display or output targets (step S0046). Specifically, the blacklist processing unit 123 stores "True" in the black designation 112D of the word in the frequency storage unit 112.
[0111] The above is the flow of the blacklist application process. According to the blacklist application process, even if a word on the blacklist is detected, the word can be excluded from being displayed or output.
[0112] 11 is a diagram showing an example of the flow of the whitelist application process. The whitelist application process starts in step S005 of the data quantification process.
[0113] First, the whitelist processing unit 124 reads the whitelist (step S0051). Specifically, the whitelist processing unit 124 reads the whitelist received in step S002 of the data quantification process and the whitelist added in step S0064 of the dummy variable selection process, which will be described later.
[0114] Then, the whitelist processing unit 124 determines whether there are any unchecked words (step S0052). If there are no unchecked words ("No" in step S0052), the whitelist processing unit 124 ends the whitelist application process.
[0115] If there are any unchecked words ("Yes" in step S0052), the whitelist processing unit 124 determines whether or not any words that do not have a count result in the word appearance count list are included in the whitelist (step S0053). If there are count results ("No" in step S0053), the whitelist processing unit 124 returns control to step S0052.
[0116] If there is no counting result ("Yes" in step S0053), the whitelist processing unit 124 counts the number of occurrences of words for which there is no counting result (step S0054). Then, the whitelist processing unit 124 returns the control to step S0052.
[0117] The above is the flow of the whitelist application process. According to the whitelist application process, if the frequency of occurrence of a word on the whitelist has not been tallied, the frequency of occurrence of the word can be tallied and added to the list of words to be displayed or output.
[0118] 12 is a diagram showing an example of the flow of the dummy variable selection process. The dummy variable selection process starts in step S006 of the data quantification process.
[0119] First, the dummy variable receiving unit 125 displays a word appearance count list (step S0061). Specifically, the dummy variable receiving unit 125 displays the dummy variable designation screen 400. At this time, as described above, the dummy variable receiving unit 125 uses the processing results of steps S003 to S005 to display the words in order of appearance count.
[0120] Then, the dummy variable receiving unit 125 displays the number of occurrences of the whitelisted words (step S0062). Specifically, the dummy variable receiving unit 125 highlights, for example, the words in the whitelist obtained as a result of the whitelist application process on the dummy variable specification screen 400.
[0121] Then, the dummy variable receiving unit 125 receives the selection of the dummy variable (step S0063). Specifically, the dummy variable receiving unit 125 receives the words for which check input has been received in the check input receiving area 402 of the dummy variable specification screen 400 as a dummy conversion list.
[0122] Then, the dummy variable receiving unit 125 receives the word to be added to the whitelist (step S0064). Specifically, the dummy variable receiving unit 125 receives the value input in the whitelist addition word input area 405 on the dummy variable specification screen 400 as the word to be added to the whitelist.
[0123] Then, the dummy variable receiving unit 125 determines whether the word to be added already has a count result in the word appearance count list (step S0065). If there are multiple words to be added, the dummy variable receiving unit 125 determines whether each word already has a count result in the word appearance count list.
[0124] If the word to be added already has a count result in the word appearance count list (Yes in step S0065), the dummy variable receiving unit 125 sets the dummy variable selection completion flag to ON (step S0066).
[0125] If the words to be added include words that do not have a count result in the word appearance count list (No in step S0065), the dummy variable receiving unit 125 sets the dummy variable selection completion flag to OFF (step S0067).
[0126] The above is the flow of the dummy variable selection process. According to the dummy variable selection process, it is possible to select words to be used as dummy variables, that is, words to be quantified, from among the words whose occurrence counts are indicated.
[0127] 14 is a diagram showing an example of the flow of the dummy variable transformation process. The dummy variable transformation process starts in step S008 of the data quantification process.
[0128] First, the dummy variable transformation unit 126 reads the dummy transformation list (step S0081). Specifically, the dummy variable transformation unit 126 reads the dummy transformation list received in step S0063 of the dummy variable selection process.
[0129] Then, the dummy variable conversion unit 126 determines whether or not there are any unprocessed dummy variable words (step S0082). If there are no unprocessed dummy variable words ("No" in step S0082), the dummy variable conversion unit 126 ends the dummy variable conversion process.
[0130] If there are any unprocessed dummy variable words ("Yes" in step S0082), the dummy variable conversion unit 126 acquires the unprocessed dummy variable words (step S0083).
[0131] Then, the dummy variable conversion unit 126 adds a string of words of the dummy variables to the input data (step S0084). Specifically, the dummy variable conversion unit 126 provides a column for each word of the dummy variables in the input data, i.e., the data to be converted.
[0132] Then, the dummy variable conversion unit 126 acquires the value of the processing target column of each row of the input data (step S0085). Specifically, the dummy variable conversion unit 126 reads the conversion target column of the conversion target data.
[0133] Then, the dummy variable conversion unit 126 determines whether or not the acquired value of the processing target column includes a word of a dummy variable that has not been processed (step S0086).
[0134] If the word of the dummy variable is included in the acquired value of the column to be processed ("Yes" in step S0086), the dummy variable conversion unit 126 inputs "1" as the value of the dummy variable string included in the value of the column to be processed (step S0087). Then, the dummy variable conversion unit 126 returns the control to step S0082.
[0135] If the word of the dummy variable is not included in the acquired value of the target column ("No" in step S0086), the dummy variable conversion unit 126 inputs "0" as the value of the dummy variable string not included in the value of the target column (step S0088). Then, the dummy variable conversion unit 126 returns the control to step S0082.
[0136] This concludes the flow of the dummy variable conversion process, which allows free text to be converted into selected dummy variables.
[0137] The above is the inspection and evaluation system 1 according to the embodiment of the present invention. According to the inspection and evaluation system 1, dummy variables can be obtained from the contents of qualitative documents, and therefore appropriate evaluation can be performed.
[0138] The present invention is not limited to the above-described embodiment. Various modifications of the above-described embodiment are possible within the scope of the technical concept of the present invention. For example, in the above-described embodiment, N-grams are used to obtain categories (words) in the word occurrence counting process, but this is not limiting. For example, categories (words) may be obtained by other means, such as morphological analysis. In this case, it is possible to determine even the parts of speech, so that a part-of-speech filter can be used in the process of counting the occurrence frequency to improve accuracy. Such a second embodiment will be described with reference to FIGS. 16 to 18.
[0139] The second embodiment is basically the same as the first embodiment, but there are some differences, which will be mainly described below.
[0140] FIG. 16 is a diagram showing another example of a frequent word acquisition condition specification screen. A part-of-speech filter specification input area 310 has been added to the frequent word acquisition condition specification screen 300'. The part-of-speech filter specification input area 310 accepts input for narrowing down the words to be quantified as dummy variables by part of speech. On the frequent word acquisition condition specification screen 300', it is possible to specify and input parts of speech such as "common noun," "proper noun," "person's name," "place name," and "symbol." Parts of speech that are not checked here will not be words to be quantified as dummy variables. Therefore, if there are categories (words) that you want to quantify, you can specify them individually in the whitelist.
[0141] 17 is a diagram showing an example of the flow of word appearance count processing (using morphological analysis). This flow is basically the same as the flow of word appearance count processing (using N-gram), but the processing flow after step S0035 is different.
[0142] First, the counting unit 122 performs morphological analysis on the target value (free description) (step S0136).
[0143] Then, the counting unit 122 determines whether there are any words that have not been evaluated in the morphological analysis result (step S0137). If there are no words that have not been evaluated in the morphological analysis result ("No" in step S0137), the counting unit 122 returns the control to step S0033.
[0144] If there is a word that has not been evaluated in the morphological analysis result ("Yes" in step S0137), the counting unit 122 determines whether the part of speech of the word is a noun (specified part of speech) (step S0138). Specifically, the counting unit 122 determines whether the word is a "common noun," "proper noun," "person's name," "place name," or "symbol" that has been specified and input. If the part of speech of the word is not a noun (specified part of speech) ("No" in step S0138), the counting unit 122 returns control to step S0137.
[0145] If the part of speech of the word is a noun (specified part of speech) ("Yes" in step S0138), the counting unit 122 counts the number of times the word appears and adds it to the word appearance count list (step S0139). Then, the counting unit 122 returns the control to step S0137.
[0146] This is the flow of the word occurrence count processing (using morphological analysis). With the word occurrence count processing (using morphological analysis), categories (words) other than the specified part of speech are ignored as noise when counting the occurrence frequency, making it possible to perform highly accurate evaluations.
[0147] 18 is a diagram showing another example of a dummy variable specification screen. The dummy variable specification screen 400' basically has the same display content as the dummy variable specification screen 400, but includes a word occurrence count (noun) display area 410 and a whitelist word occurrence count display area 411.
[0148] The word occurrence count (noun) display area 410 is an area that displays words in a table in descending order of occurrence count. This display reflects the processing results of steps S003 to S005. The word occurrence count (noun) display area 410 includes a check input acceptance area for specifying dummy variables (subject to quantification). However, it does not include words on the whitelist. The counting results for words on the whitelist are displayed as a separate table in the whitelist word occurrence count display area 411.
[0149] The above is the inspection and evaluation system according to the second embodiment. According to the inspection and evaluation system according to the second embodiment, dummy variables can be obtained with higher accuracy, and therefore, appropriate evaluation can be performed.
[0150] Furthermore, in the second embodiment, the whitelist addition candidates may be proposed on the dummy variable specification screen. Such a modification will be described with reference to FIG.
[0151] The inspection and evaluation system according to the third embodiment basically has a configuration substantially similar to that of the second embodiment. However, there are some differences. The following description will focus on these differences.
[0152] FIG. 19 is a diagram showing yet another example of a dummy variable specification screen. The dummy variable specification screen 400'' includes a whitelist addition candidate input area 420. The aggregation processing unit 122 divides text data into predetermined segments using N-grams, and combines the predetermined segments using frequency of occurrence to propose words with a word length exceeding the value of N as candidates for the whitelist. Furthermore, when dividing text data into predetermined segments using morphological analysis, the aggregation processing unit 122 can also propose other expressions such as co-occurrence expressions, synonyms, and foreign words.
[0153] The whitelist candidates proposed by the aggregation processing unit 122 are then displayed in a list in the whitelist addition candidate input area 420, with a check box corresponding to each word (category). Words (categories) with an entry in the check box are treated as candidates to be added to the whitelist.
[0154] The above is the inspection and evaluation system according to the third embodiment.
[0155] Furthermore, the technology according to the present invention is not limited to the inspection and evaluation system described above, but may also be applied to a regional information collection system that receives reports and references at any time and collects and analyzes data at any time. Such an example will be described with reference to Figures 20 to 39.
[0156] FIG. 20 is a block diagram of an example of a local information collection system according to a fourth embodiment. The local information collection system 1000 is a system for freely sharing information about abnormal situations and public safety in a certain area among surrounding residents, local governments, and administrative officials. For example, if a resident notices a bump on the sidewalk, they can use the system to report the bump. This can be accepted by a local government organization or administrative official using the system, leading to arrangements for repairs. Alternatively, the system can be used to detect and contain food poisoning, epidemics, and disasters, and to share information about suspicious individuals.
[0157] The regional information collection system 1000 basically has almost the same configuration as the inspection and evaluation system 1, but there are some differences. The following mainly describes these differences.
[0158] The regional information collection system 1000 includes a data quantification server device 100'. An external user 20 who is a user of the system uses the data quantification server device 100' from a terminal such as a smartphone or a personal computer via a network 50 such as a public network such as the Internet, a mobile phone data communication network, a WAN (Wide Area Network), or a LAN (Local Area Network).
[0159] The storage unit 110' of the data quantification server device 100' includes a conversion target data storage unit 111' and a time-address priority storage unit 115.
[0160] 21 is a diagram showing an example of the data structure of the conversion target data storage unit. The conversion target data storage unit 111′ includes an event ID 111A′, an auxiliary ID 111B′, a comment (freely written content) 111C′, a commenter 111D′, a comment time 111E′, a site address 111F′, a site latitude and longitude 111G′, a commenter position 111H′, an image position 111J′, a text extraction position 111K′, and a status 111L′.
[0161] The event ID 111A' is an identifier that distinguishes a series of events, including the report and other reports related to that report, from other events. The auxiliary ID 111B' is an identifier that distinguishes the report, each report, contact, etc. within the event from others. The comment (freely written content) 111C' is free text that expresses the content of the report or contact in natural language. For example, it may include local disaster prevention information, disaster information, or information about malfunctions in the living environment.
[0162] The commenter 111D' and comment time 111E' are information that respectively identify the person who made the comment and the time when the commented event was observed. The site address 111F' and site latitude and longitude 111G' are information that respectively identify the area including the location where the abnormal situation or security problem occurred and the location itself.
[0163] Commenter location 111H' is information that identifies the location where the commenter was at the time they posted the comment. Image location 111J' is information that identifies the shooting location associated with an image if the commenter attached an image. Text extraction location 111K' is location information geocoded by extracting keywords corresponding to location information from text information in the comment. Status 111L' is information that identifies whether each comment is completed or ongoing.
[0164] 22 is a diagram showing an example of the data structure of the time-address priority storage unit. The time-address priority storage unit 115 includes a specific item 115A, a ranking 115B, and original information 115C. The specific item 115A is information specifying whether the criterion to be specified is time or address. The ranking 115B is information specifying the priority order for the item specified by the specific item 115A. The original information 115C is information specifying the original information for the item specified by the specific item 115A. For example, if the specific item 115A is "site address," the ranking 115B is "1," and the original information 115C is "commenter location," the rule indicates that the commenter location is given the highest priority when specifying the site address. Similarly, if the specified item 115A is "site address," the ranking 115B is "2," and the original information 115C is "image location," the rule indicates that when identifying the site address, the image location is given priority after the commenter location and is identified as the site address.
[0165] The control unit 120 also includes an information integration unit 128. The information integration unit 128 performs information integration processing, which will be described later.
[0166] 23 is a diagram showing an example of the flow of information integration processing. The information integration processing is started when a predetermined number of comments (for example, one or three) are added, or at predetermined time intervals (for example, every 10 minutes).
[0167] The information integration unit 128 performs the data quantification process of FIG. 6 (step S101), and then identifies the time and address of each comment (step S102).
[0168] To identify the time and address of each comment, the information integration unit 128 refers to the rules in the time-address priority storage unit 115 and identifies the comment time 111E' and the site address 111F' for each comment. Specifically, for the "text extraction location," which is the source information for the "site address," the information integration unit 128 prioritizes words that are listed as "place names" in a dictionary and extracts keywords corresponding to location information from the content of the comment, or extracts word parts that resemble place names using a well-known technique called named entity extraction and extracts latitude and longitude using a well-known technique called geocoding. The information integration unit 128 then refers to the time-address priority storage unit 115 and references the location information in the "priority order" of the "site address." If there is a missing value, the order is skipped and the lower-ranked location information is used as the "site address."
[0169] Similarly, for the time of the comment, the information integration unit 128 extracts the time of the report, the time of the image, and the time of text extraction, and refers to the time-address priority memory unit 115 to refer to the time information in the "priority order" of the "comment time." If there is a missing value, that order is skipped and the time information of the lower order is adopted as the "comment time."
[0170] This is the flow of the information integration process. The information integration process uses data quantification processing to extract frequently occurring words as dummy variables from the free text in the comments, and also identifies the time and address for each comment.
[0171] 24 is a diagram showing an example of the flow of the regional information compilation process. The regional information compilation process is started when a request is made by the external user 20.
[0172] First, the information integration unit 128 acquires dummy variables (step S201). Specifically, the information integration unit 128 reads the dummy variables created in step S101 of the information integration process. Then, the information integration unit 128 creates a summary item setting screen 600 and transmits it to the terminal used for access by the external user 20 for display.
[0173] 25 is a diagram showing an example of a summary item setting screen. The summary item setting screen 600 includes areas for inputting comment extraction conditions, classification axes to be used for display, and summary values that specify the target of summary display.
[0174] The area for entering comment extraction conditions includes an extraction condition (dummy variable) reception area 610, which is set to limit dummy variables, and an extraction condition (other than dummy variables) reception area 620, which is set to limit variables other than dummy variables (i.e., standard items).
[0175] More specifically, the extraction condition (dummy variable) reception area 610 includes check boxes for the dummy variables to be narrowed down, and check boxes 611 that control whether the condition is the presence or absence of each dummy variable.
[0176] The extraction condition (other than dummy variables) reception area 620 includes check boxes for variables other than the dummy variables to be narrowed down, and a detailed condition reception area for receiving input of detailed conditions for each dummy variable. For example, for comment time, it includes a time range specification reception area 621 for receiving input specifying either the start time or the end time, or both, that determine the comment time extraction range. For status, it includes a check box 622 for receiving whether the status is ongoing or completed.
[0177] The area for entering the classification axes to be used for display includes a classification axis X (dummy variable) reception area 630, which is used to set up the restriction of dummy variables, and a classification axis Y (other than dummy variables) reception area 640, which is used to set up the restriction of variables other than dummy variables (i.e., standard items).
[0178] The classification axis X (dummy variable) reception area 630 includes an area for receiving the number k to be selected and the original number n as parameters for determining the dummy variable combination nCk. The number k to be selected is specified in the number of dummy variables to be selected reception area 631, and the original number n is specified in the checkbox of the dummy variable that aggregates the value = 1.
[0179] The classification axis Y (other than dummy variables) receiving area 640 includes an area that receives the classification axis and its hierarchy.
[0180] The aggregated value receiving area 650 is provided with a receiving area for specifying whether or not to display the number of comments, and whether or not to display the comment contents in a linked manner.
[0181] Then, the information integration unit 128 receives the tally items set on the tally item setting screen 600 (step S202).
[0182] The information integration unit 128 determines whether or not only dummy variables have been selected as classification axes (step S203). Specifically, the information integration unit 128 determines whether or not the setting of classification axis X has been accepted in the classification axis X (dummy variable) acceptance area 630, and whether or not a check of a standard item has been accepted in the classification axis Y (other than dummy variables) acceptance area 640.
[0183] If only dummy variables are selected for the classification axis ("Yes" in step S203), the information integration unit 128 determines whether one dummy variable is selected for the classification axis (step S204). For example, the information integration unit 128 determines whether the value input in the number of dummy variables to be selected reception area 631 is 1.
[0184] If one dummy variable is selected as the classification axis (if "Yes" in step S204), the information integration unit 128 classifies the number of comments and the comments for each selected dummy variable and outputs them according to the specification received in the total value receiving area 650 (step S205). An example of this output is the single variable total screen 700 described below.
[0185] 26 is a diagram showing an example of a single variable tabulation screen. The single variable tabulation screen 700 includes a table 701 in which dummy variables are arranged as rows (vertical axis) and the number of comments or comment content is arranged as horizontal axis. For example, if both the number of comments and comment content are specified in the tabulation value reception area 650, the number of comments containing "sidewalk" and the content of the comments containing "sidewalk" are displayed in the row in which the dummy variable is "sidewalk."
[0186] If one dummy variable has not been selected for the classification axis (if "No" in step S204), the information integration unit 128 classifies the number of comments and the comments for each combination of the selected dummy variables, and outputs them according to the specification received in the total value receiving area 650 (step S206). An example of this output is the multi-variable total screen 710 described below.
[0187] 27 is a diagram showing an example of a multiple variable aggregation screen. The multiple variable aggregation screen 710 includes a table 711 that organizes combinations of dummy variables as rows (vertical axis) and the number of comments and comment content as horizontal axes. For example, if both the number of comments and the comment content are specified in the aggregate value reception area 650, the row for which the dummy variable is "sidewalk x repair" will display the number of comments that include both "sidewalk" and "repair" and the content of the comments that include both "sidewalk" and "repair."
[0188] If only dummy variables are not selected for the classification axis ("No" in step S203), the information integration unit 128 determines whether or not one variable other than dummy variables has been selected for the classification axis (step S207). For example, the information integration unit 128 determines whether or not only one level of classification axis has been selected in the classification axis Y (other than dummy variables) reception area 640.
[0189] If one variable other than a dummy variable is selected as a classification axis (if "Yes" in step S207), the number of comments and the comments are classified for each combination of the selected dummy variables and for each classification axis other than the dummy variables, and output according to the specification received in the aggregate value reception area 650 (step S208). An example of this output is the one-level cross table screen 750, which will be described later.
[0190] 28 is a diagram showing an example of a first-level cross table screen. In the first-level cross table screen 750, combinations of dummy variables are arranged as rows (vertical axis) 751, and the horizontal axis shows the address of the work site 752, which is the item specified for the classification axis Y. In other words, the content of the comment is displayed in the area where the combination of dummy variables included intersects with the address of the work site. For example, in the row where the dummy variable is "sidewalk x repair," if only the comment content is specified in the aggregate value reception area 650, the content of the comment that includes both "sidewalk" and "repair" will be displayed organized by the address of the work site.
[0191] If no variable other than the dummy variables has been selected for the classification axis ("No" in step S207), the number of comments and the comments are classified for each combination of the selected dummy variables and for each classification axis other than the dummy variables, and output according to the specification received in the aggregate value receiving area 650 (step S209). An example of this output is the multi-layer cross table screen 760 described below.
[0192] 29 is a diagram showing an example of a multi-hierarchical cross table screen. In the multi-hierarchical cross table screen 760, combinations of dummy variables are arranged as rows (vertical axis) 761, and combinations of site addresses 762 and commenters 763, which are items specified on classification axis Y, are arranged on the horizontal axis in the order of the numbers specified on classification axis Y. In other words, the content of a comment is displayed in the area where the included combination of dummy variables intersects with the combination of site addresses and commenters. For example, in a row where the dummy variable is "sidewalk x repair," if only the comment content is specified in the aggregate value reception area 650, the content of comments that include both "sidewalk" and "repair" are displayed organized by site address and commenter.
[0193] The above is an example of the flow of the regional information compilation process. According to the regional information compilation process, the report information for the region is organized and displayed in a classified manner according to the specified category axis.
[0194] The screens output in the flow of the regional information aggregation process are not limited to the above screens, and may be displayed as different screens depending on the type of standard item.
[0195] 30 is a diagram showing another example of the first-level cross table screen (time slice). The first-level cross table screen (time slice) 770 displays a first-level cross table by time zone. This is an example of arranging outputs for which comment times have been accepted in the extraction condition (other than dummy variables) acceptance area 620.
[0196] 31 is a diagram showing another example of a one-level cross-tabulation screen (displaying only ongoing cases). The one-level cross-tabulation screen (displaying only ongoing cases) 780 displays a one-level cross-tabulation of ongoing cases by time zone. This is an example in which the output of comment time and status received in the reception area 620 (excluding dummy variables) is arranged.
[0197] 32 is a diagram showing an example of a map display screen. The map display screen 800 includes a map image 801 of a region, a comment field 802 superimposed on the map image, display settings (dummy variables to be displayed) 805, and display settings (location information to be displayed) 806. The comment field 802 includes one of the dummy variables, the number of comments thereon 803, and comment content 804.
[0198] Display settings (dummy variables to be displayed) 805 accepts the specification of dummy variables to be selectively displayed in the comment field 802 or the criteria for determining the dummy variables. Display settings (location information to be displayed) 806 accepts the input of how to divide the boundaries of the areas of the map image 801.
[0199] 33 is a diagram showing another example of a map display screen. The map display screen 800 includes a map image 801 of a region, a comment field 802' superimposed on the map image, display settings (dummy variables to be displayed) 805, and display settings (location information to be displayed) 806. The comment field 802' includes one of the dummy variables, the number of comments thereon 803', and comment content 804'.
[0200] In the example of Fig. 33, the display setting (dummy variables to be displayed) 805 is in a state where it accepts the specification of the criteria for determining the dummy variables (maximum number of items at the relevant location). Therefore, for each area (T) of the map image 801, the dummy variable with the largest number of items is extracted and displayed.
[0201] The above is an example of the local information collection system according to the fourth embodiment. With the local information collection system according to the fourth embodiment, information on abnormal situations and public safety in a certain area can be freely shared among surrounding residents, local governments, and administrative officials.
[0202] 34 is a block diagram of another example of a local information collection system according to the fourth embodiment. The local information collection system 1000' basically has almost the same configuration as the local information collection system 1000, but there are some differences. The following description will focus on these differences.
[0203] The storage unit 110 ″ of the data quantification server device 100 ″ includes a conversion target data storage unit 111 ″, a dummy tag storage unit 116 , and an inter-image tag similarity storage unit 117 .
[0204] 35 is a diagram showing an example of the data structure of the conversion target data storage unit. The conversion target data storage unit 111'' further includes an image 111M'. This image is an image that the commenter attaches when commenting.
[0205] FIG. 36 is a diagram illustrating an example of the data structure of the dummy tag storage unit. The dummy tag storage unit 116 includes an image 116A arranged in rows and a first dummy tag 116B and a second dummy tag 116C arranged in columns. The image 116A is information that identifies an image. The first dummy tag 116B and the second dummy tag 116C are columns that are set according to dummy variables. The first dummy tag 116B and the second dummy tag 116C are columns that are set to prevent overlapping of dummy tags associated with any of the images 116A. Therefore, the first dummy tag 116B and the second dummy tag 116C also vary depending on the image included in the image 116A. At the intersection of the rows and columns, a "1" is stored if the dummy variable is included in the comment for the image, and a "0" is stored otherwise. The data structure of the dummy tag storage unit 116 is not limited to this, and may be such that, for example, only tags related to images are associated. In other words, the data structure may be such that tags of dummy variables not included in comments are not associated. Using such a data structure, images matching the search keywords can be searched for.
[0206] 37 is a diagram showing an example of the data structure of the inter-image tag similarity storage unit. The inter-image tag similarity storage unit 117 includes a brute force table between images, and the similarity of an image 117A in the column direction with respect to an image 117B in the row direction is calculated according to a predetermined standard and stored. In this example, the number of tags common to the images is calculated as the similarity.
[0207] In the local information collection system 1000′, following the process of identifying the time and address of each comment, which is carried out in step S102 of the information integration process, images 111M′ are extracted from the conversion target data storage unit 111″, and dummy variables extracted from comments (free-form content) 111C′ related to each image are associated with each image as tags by the similarity search unit 129. Then, the association is stored in the dummy tag storage unit 116 by the similarity search unit 129.
[0208] Furthermore, the similarity search unit 129 determines the similarity between the images and stores it in the inter-image tag similarity storage unit 117. In this process, the similarity search unit 129 compares tags based on the associated dummy variables for each image and counts the number of matching tags to determine the similarity. In other words, images attached to comments that have three common dummy variables are given a similarity of "3" and stored in the inter-image tag similarity storage unit 117.
[0209] Furthermore, the similarity search unit 129 can use these dummy tag storage units 116 to receive search words, search for images, and output them. This is called an image fuzzy search.
[0210] 38 is a diagram showing an example of an image fuzzy search screen. The image fuzzy search screen 900 includes a search word input area 901 and a search result display area 902. The search result display area 902 also includes a similarity display area 903 and an image information display area 904.
[0211] A search word input area 901 accepts keywords (dummy variables) for searching for images. A search result display area 902 displays an image information display area 904 that lists images and tags similar to the keyword input in the search word input area 901, and a similarity display area 903.
[0212] Here, for each image 116A in the dummy tag storage unit 116, the similarity search unit 129 treats a vector whose components are the values of the dummy tags as a feature vector that indicates the characteristics of the image, and searches for images that have a high similarity to the feature vector consisting of the search keyword. In this search, the similarity search unit 129 can determine the similarity by calculating the Euclidean distance between the vectors. However, this is not limited to this, and the similarity may also be determined by the number of matching tags.
[0213] Furthermore, the similarity search unit 129 can receive an image, search for other similar images, and output them using the dummy tag storage unit 116. This is called tag similarity image search.
[0214] 39 is a diagram showing an example of a tag similar image search screen. The tag similar image search screen 910 includes a search image area 911, a similar search execution instruction reception area 912, and a search result display area 920. The search result display area 920 also includes a similarity display area 921 and an image information display area 922.
[0215] The search image area 911 includes an image for which similar images are to be searched. For example, when a comment with an attached image is displayed and an image similar to the comment is searched for, the image attached to the comment corresponds to the image for which similar images are to be searched. When receiving an input, the similar search instruction receiving area 912 receives an instruction to search for images similar to the image specified in the corresponding search image area 911. The search result display area 920 displays an image information display area 922 that lists images and tags similar to the image included in the search image area 911, and a similarity display area 921.
[0216] Here, the similarity search unit 129 searches for image 117B in the inter-image tag similarity storage unit 117 to identify other images with high similarity. However, without being limited to this, the similarity search unit 129 may treat a vector whose components are the values of dummy tags at the time of execution as a feature vector indicating the features of the image, and search for images with high similarity to the feature vector of the search image. In this search, the similarity search unit 129 can determine the similarity by calculating the Euclidean distance between vectors. However, without being limited to this, the similarity may also be determined by the number of matching tags.
[0217] The above is another example of a local information collection system according to the fourth embodiment. According to this example of a local information collection system according to the fourth embodiment, for an image associated with a free text comment, a related dummy variable can be associated as tag information for the image. Therefore, when performing an image search, rather than comparing the images themselves, tag information can be treated as vector information, and similar images can be identified according to the similarity of the vectors. This improves the speed of image search. In particular, when there is a large number of images, it is possible to increase the search speed for those images while reducing search noise.
[0218] In addition, in another example of the local information collection system according to the fourth embodiment, an example of searching for images was given, but this is not limited to this, and unstructured data such as video, audio, etc., or a combination thereof, may be posted along with comments and searched.
[0219] Furthermore, the technical elements of the above-described embodiments may be applied independently, or may be divided into multiple parts such as program parts and hardware parts and applied.
[0220] The present invention has been described above mainly with reference to the embodiments. [Explanation of symbols]
[0221] 1···Inspection evaluation system, 10···User, 100···Data quantification server device, 110···Memory unit, 111···Conversion target data memory unit, 112···Frequency memory unit, 113···Blacklist memory unit, 114···Whitelist memory unit, 120···Control unit, 121···Condition reception unit, 122···Aggregation processing unit, 123···Blacklist processing unit, 124···Whitelist processing unit, 125···Dummy variable reception unit, 126···Dummy variable transformation unit, 127···Evaluation calculation unit, 130···Input unit, 140···Output unit, 20···External user, 50···Network, 115···Time and address priority memory unit, 116···Dummy tag memory unit, 117···Inter-image tag similarity memory unit, 128···Information integration unit, 129···Similarity search unit, 1000···Regional information collection system.
Claims
1. a condition receiving unit that receives a quantification condition for the text data; a counting unit that divides the text data into predetermined segments and counts the number of occurrences or the presence or absence of occurrences; an output unit that displays the counted number of occurrences or the presence or absence of occurrences, reflecting the quantification conditions; a dummy variable receiving unit that receives a user's selection input for each of the categories in the display of the output unit as a designation of a dummy variable; a dummy variable conversion unit that acquires a result of counting the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable, the output unit outputs the result of counting the number of occurrences or the presence or absence of occurrences acquired by the dummy variable transformation unit.
1. An information processing device comprising:
2. 2. The information processing device according to claim 1, the quantification condition includes a blacklist specifying one or more words to be excluded from the classification; a blacklist processing unit that uses the blacklist to exclude the words to be excluded from the classification in the process of reflecting the quantification conditions by the output unit; An information processing device comprising:
3. 2. The information processing device according to claim 1, The quantification condition includes a whitelist that specifies one or more words to be added as the category; a whitelist processing unit that adds the word to be added to the category using the whitelist and recounts the number of occurrences or the presence or absence of occurrences in the process of the output unit reflecting the quantification conditions; An information processing device comprising:
4. 2. The information processing device according to claim 1, The aggregation processing unit divides the text data into the predetermined categories using N-grams.
1. An information processing device comprising:
5. 2. The information processing device according to claim 1, The quantification condition includes a whitelist that specifies one or more words to be added as the category; a whitelist processing unit that adds the word to be added to the category using the whitelist and recounts the number of occurrences or the presence or absence of occurrences in the process of the output unit reflecting the quantification conditions; The aggregation processing unit divides the text data into the predetermined categories using N-grams, and combines the predetermined categories using the number of occurrences or the presence or absence of occurrences, and proposes words with a word length exceeding the value of N as candidates for the whitelist.
1. An information processing device comprising:
6. 2. The information processing device according to claim 1, the quantification condition includes a designation of a part of speech to be used as the classification, the tabulation processing unit divides the text data into the predetermined categories using morphological analysis, and excludes the predetermined categories that do not correspond to the parts of speech from the tabulation.
1. An information processing device comprising:
7. 2. The information processing device according to claim 1, The text data is accompanied by one or more values of predetermined measurement results, The output unit the result of counting the number of occurrences or the presence or absence of occurrence of the category acquired by the dummy variable transformation unit is added as the value of the measurement result for each category; 1. An information processing device comprising:
8. 2. The information processing device according to claim 1, the text data includes a natural language description of the inspection results of the structure, and is accompanied by one or more values of predetermined measurement results of the structure; an evaluation calculation unit that calculates a predetermined evaluation index of the structure using the values of the measurement results; The output unit the result of counting the number of occurrences or the presence or absence of occurrence of the category acquired by the dummy variable transformation unit is added as the value of the measurement result for each category; 1. An information processing device comprising:
9. An inspection and evaluation system using an information processing device, the information processing device includes a control unit and a storage unit; the storage unit stores one or more pieces of text data including a natural language description of an inspection result of the structure, together with one or more values of a predetermined measurement result of the structure; The control unit a condition receiving step of receiving a quantification condition for the text data; a counting step of dividing the text data into predetermined segments and counting the number of occurrences or the presence or absence of occurrences; an output step of displaying the counted number of occurrences or the presence or absence of occurrences, reflecting the quantification conditions; a dummy variable receiving step of receiving a selection input from the user for each of the categories in the display of the output step as a designation of a dummy variable; a dummy variable transformation step of acquiring a result of counting the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable; a result output step of adding the result of counting the number of occurrences or the presence or absence of occurrences acquired in the dummy variable transformation step as the value of the measurement result for each of the categories; an evaluation calculation step of calculating a predetermined evaluation index of the structure using the values of the measurement results; An inspection and evaluation system characterized by carrying out the above.
10. An inspection and evaluation method using an information processing device, the information processing device includes a control unit and a storage unit; the storage unit stores one or more pieces of text data including a natural language description of an inspection result of the structure, together with one or more values of a predetermined measurement result of the structure; The control unit a condition receiving step of receiving a quantification condition for the text data; a counting step of dividing the text data into predetermined segments and counting the number of occurrences or the presence or absence of occurrences; an output step of displaying the counted number of occurrences or the presence or absence of occurrences, reflecting the quantification conditions; a dummy variable receiving step of receiving a selection input from the user for each of the categories in the display of the output step as a designation of a dummy variable; a dummy variable transformation step of acquiring a result of counting the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable; a result output step of adding the result of counting the number of occurrences or the presence or absence of occurrences acquired in the dummy variable transformation step as the value of the measurement result for each of the categories; an evaluation calculation step of calculating a predetermined evaluation index of the structure using the values of the measurement results; A testing and evaluation method characterized by carrying out the steps of:
11. 8. The information processing device according to claim 1, The text data includes information on local disaster prevention, disaster information, and information on problems in the living environment. a natural language description of any one or any combination thereof, The output unit outputting the number of cases extracted from the text data as a result of aggregation based on the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable; 1. An information processing device comprising:
12. 8. The information processing device according to claim 7, The text data includes a natural language description of one of local disaster prevention information, disaster information, and information on problems in the living environment, or a combination thereof; the predetermined measurement result value includes at least position information, the output unit outputs, as a count result, the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable and the number of cases extracted from the text data using the value of the location information as conditions.
1. An information processing device comprising:
13. 8. The information processing device according to claim 7, The text data includes a natural language description of one of local disaster prevention information, disaster information, and information on problems in the living environment, or a combination thereof; the predetermined measurement result value includes a plurality of pieces of position information; an information integration unit that determines which of the plurality of pieces of location information to adopt in accordance with a predetermined priority; the output unit outputs, as a count result, the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable and the number of cases extracted from the text data using the value of the location information as conditions.
1. An information processing device comprising:
14. 14. The information processing device according to claim 12, the output unit displays the number of extracted items superimposed on a map according to the adopted location information.
1. An information processing device comprising:
15. 8. The information processing device according to claim 7, The text data includes a natural language description of one of local disaster prevention information, disaster information, and information on problems in the living environment, or a combination thereof; the predetermined measurement result value includes at least date and time information, the output unit outputs, as a counting result, the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable and the number of items extracted from the text data using the value of the date and time information as conditions.
1. An information processing device comprising:
16. 16. The information processing device according to claim 11, the output unit outputs the text data corresponding to the aggregation result.
1. An information processing device comprising:
17. 8. The information processing device according to claim 7, The predetermined measurement result value includes at least unstructured data of any one of image, video, and audio, or a combination thereof; a similarity search unit that associates the name of the dummy variable as a tag of the unstructured data and uses the name for the search based on the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable; An information processing device comprising:
18. 18. The information processing device according to claim 17, the similarity search unit associates the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable as a feature vector representing the characteristics of the unstructured data, calculates similarity between the feature vectors, and uses the similarity between the feature vectors for similarity search of the unstructured data; 1. An information processing device comprising:
19. 18. The information processing device according to claim 17, The similarity search unit The search keywords are obtained as feature vectors, the number of occurrences or the presence or absence of occurrences of the category accepted as the designation of the dummy variable is associated with a feature vector representing the characteristics of the unstructured data, and the similarity with the feature vector acquired as the search keyword is calculated and used for similarity search of the unstructured data; 1. An information processing device comprising:
20. A data quantification method using an information processing device, comprising: the information processing device includes a control unit, The control unit a condition receiving step of receiving a quantification condition for the text data; a counting step of dividing the text data into predetermined segments and counting the number of occurrences or the presence or absence of occurrences; an output step of displaying the counted number of occurrences or the presence or absence of occurrences, reflecting the quantification conditions; a dummy variable receiving step of receiving a selection input from the user for each of the categories in the display of the output step as a designation of a dummy variable; a dummy variable transformation step of acquiring a result of counting the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable; a second output step of outputting the result of counting the number of occurrences or the presence or absence of occurrences obtained in the dummy variable transformation step; A data quantification method comprising:
21. 21. A data quantification method according to claim 20, comprising: The text data includes information on local disaster prevention, disaster information, and information on problems in the living environment. a natural language description of any one or any combination thereof, a third output step of outputting the number of cases extracted from the text data as a count result, based on the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable; A data quantification method comprising:
22. 21. A data quantification method according to claim 20, comprising: The text data includes a natural language description of one or a combination of local disaster prevention information, disaster information, and information on problems in the living environment, and is accompanied by one or more values of predetermined measurement results; The predetermined measurement result value includes at least position information, a third output step of outputting, as a count result, the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable and the number of cases extracted from the text data using the value of the location information as conditions; A data quantification method comprising:
23. 21. A data quantification method according to claim 20, comprising: The text data includes a natural language description of one or a combination of local disaster prevention information, disaster information, and information on problems in the living environment, and is accompanied by one or more values of predetermined measurement results; the predetermined measurement result value includes a plurality of pieces of position information; performing an information integration step of determining adoption of any one of the plurality of pieces of location information in accordance with a predetermined priority; a third output step of outputting, as a count result, the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable and the number of cases extracted from the text data using the value of the location information as conditions; A data quantification method comprising:
24. 24. A data quantification method according to claim 22 or 23, comprising the steps of: In the third output step, the number of extracted cases is superimposed on a map in accordance with the adopted location information. A data quantification method comprising:
25. 21. A data quantification method according to claim 20, comprising: The text data includes a natural language description of one or a combination of local disaster prevention information, disaster information, and information on problems in the living environment, and is accompanied by one or more values of predetermined measurement results; the predetermined measurement result value includes at least date and time information, a third output step of outputting, as a count result, the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable and the number of items extracted from the text data using the value of the date and time information as conditions; A data quantification method comprising:
26. 26. A data quantification method according to any one of claims 21 to 25, comprising the steps of: In the third output step, the text data corresponding to the counting result is output. A data quantification method comprising:
27. 21. A data quantification method according to claim 20, comprising: The text data is accompanied by one or more values of predetermined measurement results, and the values of the predetermined measurement results include unstructured data such as at least one of an image, a video, and an audio, or a combination thereof; a similarity search step in which the name of the dummy variable is associated as a tag of the unstructured data and used for search, based on the number of occurrences or the presence or absence of occurrence of the category accepted as the designation of the dummy variable; A data quantification method comprising:
28. 28. A data quantification method according to claim 27, comprising: In the similarity search step, the number of occurrences or the presence or absence of occurrences of the categories accepted as the designation of the dummy variables are associated as feature vectors representing the characteristics of the unstructured data, and similarity between the feature vectors is calculated and used for similarity search of the unstructured data. A data quantification method comprising:
29. 28. A data quantification method according to claim 27, comprising: In the similarity search step, The search keywords are obtained as feature vectors, the number of occurrences or the presence or absence of occurrences of the category accepted as the designation of the dummy variable is associated with a feature vector representing the characteristics of the unstructured data, and the similarity with the feature vector acquired as the search keyword is calculated and used for similarity search of the unstructured data; A data quantification method comprising:
Citation Information
Patent Citations
Document ranking device, method, and computer program
JP2016076208A