Information processing system

The information processing system addresses inefficiencies in managing and searching construction documents by chunking, attributing, and keywording documents, ensuring rapid retrieval and context-aware search results.

JP2025186093APending Publication Date: 2025-12-23BASIS CONSULTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024094680
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing construction management systems face challenges in managing and searching documents related to structures, buildings, and equipment due to non-digitized paper and microfilm formats, lack of uniform terminology, and large document volumes, leading to inefficient information retrieval.

Method used

An information processing system that divides documents into chunks, assigns attributes and keywords, and utilizes a ledger and vocabulary system to facilitate efficient searching, including organization and managed object information, document classification, and location data, with score-based result display.

Benefits of technology

Enables rapid retrieval of desired documents by leveraging attributes, keywords, and vocabulary relationships, accommodating organizational and document changes, and providing context-aware search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186093000001_ABST
    Figure 2025186093000001_ABST
Patent Text Reader

Abstract

To provide an information processing system for retrieving documents, materials, etc., in a crossing manner.SOLUTION: An information processing system manages documents related to a management object and includes: a chunk generation processing unit that generates a chunk from data such as a document; an attribute assignment processing unit that assigns an attribute to the generated chunk based on information recorded in a ledger recording unit; an important word assignment processing unit that assigns an important word to the generated chunk based on an important word dictionary recorded in a vocabulary and mapping definition recording unit; a chunk recording unit that associates the attribute and / or important word assigned to the generated chunk with the chunk and records it; and a search and display processing unit that receives a search query including the attribute and / or important word from a searcher terminal used by a searcher, refers to the chunk recording unit, and causes the searcher terminal to display a matching chunk.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system for cross-sectionally searching documents, materials, and the like. [Background technology]

[0002] In the construction and civil engineering fields, it is necessary to manage structures such as roads, tunnels, and bridges, buildings such as facilities, and equipment such as electricity, communications, and accessories for operation and management. In addition, in order to carry out the survey, design, construction, and management of structures, buildings, and equipment, various documents and materials such as numerous design documents, drawings, and measurement data (hereinafter referred to as "documents, etc.") must be managed appropriately. Note that tangible objects such as structures, buildings, and equipment that are managed by organizations such as companies, groups, and government agencies in the construction and civil engineering fields are called "managed items."

[0003] To address this issue, there is a device such as that described in Patent Document 1 below. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2022-187749 Summary of the Invention [Problem to be solved by the invention]

[0005] The construction management support device of Patent Document 1 requires that documents and other information be digitized in advance and organized in a specified format. However, for managed objects such as structures, buildings, and equipment, work continues for decades, from survey and design to construction and management. While it would be possible to digitize and manage all surveys and designs from now on, the reality is that many of the documents and other information related to the managed objects that already exist are not digitized and are stored in various formats, such as paper and microfilm. Therefore, if you want to know information about a managed object, you must search through the paper or microfilm documents and other information related to that managed object.

[0006] Furthermore, each of the tasks of surveying, designing, construction, and management requires different know-how, departments in charge, and contractors. Furthermore, each task is passed on to multiple people over the long term, and the people in charge are often transferred over time. As a result, there is no uniform definition of terms used in the systems used for each task, nor in the documents generated. The same terms are sometimes written in different ways. Therefore, even if documents are digitized and a cross-sectional search is performed, it may be difficult to find the information you need.

[0007] Furthermore, because documents related to the managed items have been accumulated over decades, the volume of documents is enormous. Even if cross-sectional searches are possible, the number of documents returned may be so large that it may be necessary to repeat the search process multiple times to find the document you are actually looking for.

[0008] As described above, the management of documents relating to objects to be managed, such as structures, buildings, and equipment, involves the above-mentioned problems. [Means for solving the problem]

[0009] In view of the above problems, the inventors have invented an information processing system that properly manages documents related to objects to be managed, such as structures, buildings, and equipment, and that enables users to reach desired documents more quickly than before when searching.

[0010] The first invention is an information processing system for managing documents, etc. related to managed objects, the information processing system having: a chunk generation processing unit that generates chunks from data of documents, etc.; an attribute assignment processing unit that assigns attributes to the generated chunks based on information recorded in a ledger recording unit; a keyword assignment processing unit that assigns keyword words to the generated chunks based on a keyword dictionary recorded in a vocabulary and mapping definition recording unit; a chunk recording unit that records the attributes and / or keyword words assigned to the generated chunks in association with the chunks; and a search and display processing unit that receives search queries including attributes and / or keyword words from a searcher terminal used by a searcher, refers to the chunk recording unit, and displays matching chunks on the searcher terminal.

[0011] By configuring the present invention, a document or the like is divided into chunks, and attributes, keywords, etc. are assigned to the chunks, so that a desired document or the like can be reached more quickly than before.

[0012] In the above-mentioned invention, the ledger recording unit has an organization / managed object ledger recording unit that records information about the organization that manages the managed objects and / or information about the managed objects, and the organization / managed object ledger recording unit stores information indicating changes in the organization as information about the organization, and stores information indicating changes in the managed objects as information about the managed objects, and can be configured as an information processing system.

[0013] By configuring the present invention as described above, the organization that manages the managed object and changes in the information about the managed object are stored as attributes assigned to the chunk, so even if there are changes in the information about the organization or managed object, information can be searched for without being affected by those changes.

[0014] In the above-mentioned invention, the information processing system has an organization / managed object ledger generation processing unit that generates and / or updates the information to be recorded in the organization / managed object ledger recording unit, and the organization / managed object ledger generation processing unit can be configured as an information processing system that accepts input of information about the organization or information about the managed object, obtains information about the organization or the managed object from a computer system that records information about the organization or a computer system that records information about the managed object, and records the information in the organization / managed object ledger recording unit.

[0015] By executing the processing of the present invention, it is possible to reflect the organization that manages the managed object and the transitions that are appropriate for the information on the managed object as attributes that are assigned to chunks.

[0016] In the above-mentioned invention, the ledger recording unit has a document ledger recording unit that records one or more of the names and creation dates of documents, etc. related to the managed object, information on the organization that created the documents, etc. related to the managed object, information on the classification of the documents, etc., and information on the classification of the business, and the document ledger recording unit stores the higher concept of the document, etc., and its subordinate concept classification and notation in correspondence with each other as information on the classification of the documents, etc., and stores the classification of the business, etc., and its expression in correspondence with each other as information on the classification of the business, and can be configured as an information processing system.

[0017] By configuring the present invention as described above, information on documents, etc., particularly information on the classification of documents, etc. and the classification of business, is stored as attributes assigned to chunks, making it possible to search for information based on the classification of documents, etc. and the classification of business, etc.

[0018] In the above-mentioned invention, the information processing system has a document ledger generation processing unit that generates and / or updates information to be recorded in the document ledger recording unit, and the document ledger generation processing unit can be configured as an information processing system that extracts the name and creation date of the document from a specified location in the document, extracts the document classification and business classification from the content of the document, and records them in the document ledger recording unit.

[0019] By executing the process of the present invention, information such as the name of a document, the classification of a document, and the classification of a business can be extracted from data such as a document, and reflected as attributes to be assigned to chunks.

[0020] In the above-mentioned invention, the ledger recording unit has a location ledger recording unit that records the location information of the managed object, and the location ledger recording unit can be configured as an information processing system that stores one or more of the location name, location type, coordinates, distance, and time as the location information of the managed object.

[0021] With the configuration of the present invention, the location information of the managed object is stored as an attribute assigned to the chunk, so that information can be searched for based on the location information.

[0022] In the above-mentioned invention, the information processing system has a location register generation processing unit that generates and / or updates the information to be recorded in the location register recording unit, and the location register generation processing unit can be configured as an information processing system that accepts input of location information of the managed object, obtains the location information of the managed object from a computer system that records the location information of the managed object, and records it in the location register recording unit.

[0023] By executing the processing of the present invention, the location information of the managed object can be reflected as an attribute given to the chunk.

[0024] In the above invention, the vocabulary and mapping definition recording unit records an ontology definition representing a conceptual system of basic vocabulary having a graph structure, definitions of related terms related to the basic vocabulary, and a key word dictionary, the ontology definition defines the basic vocabulary included in the document etc. in a graph structure, the related word definitions define one or more of synonyms, which are vocabulary that indicate the same concept as the concept of the basic vocabulary expressed in the ontology definition, synonyms, which are vocabulary that indicate concepts similar to the concept of the basic vocabulary expressed in the ontology definition, and related words, which are vocabulary that indicate concepts that appear in relation to the concept of the basic vocabulary expressed in the ontology definition, by extending the graph structure of the ontology definition, and the key word dictionary is a list of important vocabulary selected from vocabulary that appear in the ontology definition and the related word definition.

[0025] By configuring the present invention, information on basic vocabulary, related words, etc., and a dictionary of important words is stored, so information can be searched for based on vocabulary and mapping definition information.

[0026] In the above-mentioned invention, the information processing system has a vocabulary and mapping definition generation processing unit that generates and / or updates the information to be recorded in the vocabulary and mapping definition recording unit, and the vocabulary and mapping definition generation processing unit can be configured as an information processing system having: a vocabulary extraction processing unit that extracts some or all of technical terms, spelling variations, synonyms, similar words, related words, proper nouns, and word classifications from a training corpus constructed based on data of documents, etc. recorded in the document, etc. recording unit; a mapping definition generation processing unit that generates mapping definitions of the spelling variations and superordinate and subordinate concepts of words; and a key word dictionary generation processing unit that generates a key word dictionary, which is a dictionary of the key words, based on data of documents, etc. recorded in the document, etc. recording unit.

[0027] By configuring the present invention, it is possible to automatically generate and update vocabulary relationships and key word dictionaries.

[0028] In the above-mentioned invention, the search and display processing unit can be configured as an information processing system that receives a search query including a search keyword and attributes and / or key words from the searcher terminal, searches document data etc. in a document recording unit based on the search keyword in the search query, extracts chunks of documents etc. that match the search keywords to create an initial result set, and narrows down the search results using the attributes and / or key words in the search query and the attributes and / or key words assigned to chunks in the created result set.

[0029] In the above-mentioned invention, the search and display processing unit can be configured as an information processing system that receives a search query including a search keyword and attributes and / or key words from the searcher terminal, identifies synonyms, similar words and / or related words by referring to the vocabulary and mapping definition recording unit using the search keyword in the search query, searches document data in the document recording unit using the search keyword and the identified synonyms, similar words and / or related words, extracts chunks of matching documents, etc. to create an initial result set, and narrows down the search results using the attributes and / or key words in the search query and the attributes and / or key words assigned to chunks in the created result set.

[0030] Conventional searches using only search keywords can result in a large number of matching documents, making it difficult to find the desired document. However, by narrowing down the search results based on attributes and key words, as in these inventions, and taking into account vocabulary relationships, it is possible to find documents more easily than before.

[0031] In the above-mentioned invention, the search and display processing unit can be configured as an information processing system that calculates a score using search conditions for attributes and / or key words assigned to chunks in the initial result set, and displays a predetermined number of results from the top scores as search results.

[0032] By configuring it as in the present invention, chunks that match the search criteria and are likely to contain attributes and key words closely related to the search keywords are displayed at the top while taking into account the user's context, making it easier to find documents, etc. than with conventional search methods that use only search keywords.In addition, it is possible to provide a function that records score calculation patterns and manages optimal patterns that take into account the user's context, such as the target time period, target area, and target business.

[0033] In the above-mentioned invention, the search and display process can be configured as an information processing system that displays chunks in the initial result set on the searcher terminal, accepts input of attributes and / or key words from the searcher terminal, and displays chunks in the initial result set that match the attributes and / or key words as search results on the searcher terminal.

[0034] As in the present invention, after an initial result set is displayed, the search may be narrowed down by accepting input of search conditions such as attributes and keywords, and search keywords.

[0035] The first invention can be realized by loading and executing an information processing program of the present invention into a computer. That is, the information processing program causes a computer that manages documents, etc. related to managed objects to function as a chunk generation processing unit that generates chunks from data of documents, etc., an attribute assignment processing unit that assigns attributes to the generated chunks based on information recorded in a ledger recording unit, a keyword assignment processing unit that assigns keyword words to the generated chunks based on a keyword dictionary recorded in a vocabulary and mapping definition recording unit, and a search and display processing unit that receives a search query including attributes and / or keyword words from a searcher terminal used by a searcher, refers to the chunk recording unit that records the attributes and / or keyword words assigned to the generated chunks in association with the chunks, and displays matching chunks on the searcher terminal. [Effects of the Invention]

[0036] By using the information processing system of the present invention, documents relating to objects to be managed such as structures, buildings, and equipment can be managed appropriately, and when searching, desired documents can be found more quickly than before. [Brief explanation of the drawings]

[0037] [Figure 1] 1 is a diagram schematically illustrating an example of the overall configuration of an information processing system according to the present invention. [Figure 2] FIG. 1 is a diagram schematically illustrating an example of a hardware configuration of a computer that executes an information processing system according to the present invention. [Figure 3] 3 is a flowchart showing an example of the overall processing in the information processing system of the present invention. [Figure 4] FIG. 10 is a diagram schematically illustrating chunk generation processing in a chunk generation processing unit. [Figure 5] FIG. 10 is a diagram illustrating a process for converting a table into semi-structured text. [Figure 6] 10 is a diagram schematically illustrating processing in an attribute assignment processing unit and a keyword assignment processing unit. FIG. [Figure 7] 10 is a diagram showing an example of a table in the organization / management object ledger recording unit. FIG. [Figure 8] FIG. 2 is a diagram illustrating an example of a table in a document ledger recording unit. [Figure 9] FIG. 2 is a diagram illustrating an example of a location ledger recording unit. [Figure 10] FIG. 10 is a diagram illustrating an example of an ontology definition. [Figure 11] FIG. 10 is a diagram illustrating an example of a graph structure expanded with synonyms and related words. [Figure 12] FIG. 2 is a diagram illustrating an example of a key word dictionary. [Figure 13] FIG. 10 is a diagram illustrating an example of a search screen. [Figure 14] FIG. 10 is a diagram showing another example of a search screen. [Figure 15] FIG. 10 is a diagram showing another example of a search screen. [Figure 16]FIG. 10 is a diagram illustrating an example of the overall configuration of an information processing system according to a second embodiment. [Figure 17] FIG. 10 is a diagram illustrating an example of the overall configuration of an information processing system according to a third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0038] FIG. 1 shows an example of the overall system configuration of an information processing system 1 of the present invention, and FIG. 2 shows an example of the hardware configuration of a computer used in the information processing system 1.

[0039] The information processing system 1 of the present invention uses a management terminal 2. The management terminal 2 in the information processing system 1 is realized using a computer. Fig. 2 schematically shows an example of the hardware configuration of the computer. The computer has a calculation device 70 such as a CPU that executes the calculation processing of a program, a storage device 71 such as a RAM, HDD, or SSD that stores information, a display device 72 such as a display that displays information, an input device 73 such as a keyboard or mouse that can input information, and a communication device 74 that sends and receives the processing results of the calculation device 70 and the information stored in the storage device 71 via a network such as the Internet or a LAN.

[0040] If the computer is equipped with a touch panel display, the display device 72 may be integrated with the input device 73. Touch panel displays are often used in portable communication terminals such as tablet computers and smartphones, but are not limited to these.

[0041] The touch panel display is a device that integrates the functions of the display device 72 and the input device 73 in that input can be made directly on the display using a predetermined input device (such as a touch panel pen) or a finger.

[0042] The management terminal 2 may be configured from one or more computers, and some of its functions may be in the cloud.

[0043] The management terminal 2 accepts input of search conditions from a searcher terminal 3 used by a searcher, and returns search results.

[0044] The management terminal 2 has a document recording unit 20, a chunk generation processing unit 21, an attribute assignment processing unit 22, a keyword assignment processing unit 23, a chunk recording unit 24, a ledger recording unit 25, a vocabulary and mapping definition recording unit 26, and a search and display processing unit 27.

[0045] The document recording unit 20 records electronic data of documents related to the managed objects. Documents related to the managed objects exist in various formats such as paper and microfilm, but these are converted into electronic data and recorded.

[0046] The chunk generation processing unit 21 performs character recognition processing such as OCR on electronic data such as documents related to managed objects to convert the data into text and divide it into sentence groups (chunks) of a predetermined length. Optical character recognition is one method of character recognition processing, but other methods such as deep learning and LLM (large-scale language model) may also be used to convert the data into text. Furthermore, if the electronic data such as documents to be recorded in the document recording unit 20 has already been converted into text, the character recognition processing can be omitted.

[0047] The predetermined length is preferably about the length of a paragraph, and is preferably long enough to assign attributes and keywords, as described below, to the sentence group. For example, one sentence group may be divided into chunks of about 300 to 500 characters. To divide a text document or the like into chunks, the text is divided at delimiters such as sentences and bullet points, and then a sentence group of about 300 to 500 characters in length is extracted along the delimiters, and this is generated as a chunk. Figure 4 shows a schematic diagram of the chunk generation process in the chunk generation processing unit 21.

[0048] If a text document contains a table, the table title and the first row and / or column of the table are recognized as data items, and the text is converted into semi-structured text using a known method, generating the table title and table as a single chunk. Known methods include, for example, "Interactive web-wrapper construction for extracting relational information from web documents" (https: / / doi.org / 10.1145 / 1062745.1062822), but other methods are also acceptable. Figure 5 shows a schematic diagram of the process of converting a table into semi-structured text.

[0049] Furthermore, if the text document contains figures or images, the title of the figure or image and the figure or image are generated as one chunk.

[0050] The attribute assignment processing unit 22 analyzes the chunks generated by the chunk generation processing unit 21, extracts information that matches the information in the ledger recording unit 25 (described later), the information in the organization / managed object ledger recording unit 251, the information in the document etc. ledger recording unit 252, and the information in the location ledger recording unit 253, and assigns the extracted information as an attribute of the chunk by a method such as character string pattern matching. Also, in order to record where the chunk is written in the document etc., the attribute assignment processing unit 22 assigns to each chunk identification information such as the file name of the document etc. and the appearance position (page number, line number, coordinates on the page, etc.).

[0051] The keyword assignment processing unit 23 analyzes the chunks generated by the chunk generation processing unit 21, extracts those that match the keyword dictionary in the vocabulary and mapping definition recording unit 26 described later, and assigns the extracted keyword to the chunk as a keyword using techniques such as string pattern matching.

[0052] 6 shows a schematic diagram of the processing in the attribute assignment processing unit 22 and the keyword assignment processing unit 23. For example, by using the Aho-Corasick algorithm for exact matches and the Wu-Manber algorithm for fuzzy matches as a string pattern matching method, attributes can be detected and assigned quickly in a single text scan.

[0053] The chunk recording unit 24 associates the chunks generated by the chunk generation processing unit 21 with the attributes assigned by the attribute assignment processing unit 22 and the keyword assigned by the keyword assignment processing unit 23, and records them as chunks. The chunk recording unit 24 may also associate and record identification information such as the file name of a document or the like, the position of appearance, and the like with the chunks, in addition to the attributes assigned by the attribute assignment processing unit 22.

[0054] The ledger recording unit 25 records information such as organizations, managed objects, documents, etc., locations, etc. The ledger recording unit 25 has an organization / managed object ledger recording unit 251, a document etc. ledger recording unit 252, and a location ledger recording unit 253.

[0055] The organization / managed object ledger recording unit 251 records information on the organization that manages the managed object and / or information on the managed object.

[0056] Organizational information includes the organization's identification name, such as the organization name, and location information. Organizational identification information includes a general organizational structure, such as a head office, branch office, office, and affiliated companies. Organizational information is recorded periodically or irregularly (such as after each change). This allows changes to the organization's identification name and location to be recorded. The organization's new name, old name, period, etc. are recorded in association with each other.

[0057] Information on managed objects includes the types of managed objects managed by an organization, the names of the managed objects, etc. The names of managed objects are recorded periodically or irregularly (after changes, etc.). This allows the names of managed objects to be recorded when they are changed, such as at the time of construction, completion, or transfer. The new and old names of the managed objects, the period, etc. are recorded in association with each other.

[0058] Fig. 7 shows an example of a table in the organization / management object ledger recording unit 251. Fig. 7(a) is an example of a table showing information about organizations, and Fig. 7(b) is an example of a table showing information about management objects.

[0059] The document ledger recording unit 252 records some or all of the following information: the name and creation date of documents related to the managed objects, information on the organization that created the documents related to the managed objects, information on the classification (type) of documents related to the managed objects, and information indicating the classification (type) of work performed by the organization.

[0060] The name and creation date of documents, etc. related to the managed items include identification information such as the unique name of the document, etc., and information on the date of creation. Usually, the name and date recorded on the cover of the document, etc. are recorded. In addition, if there are documents with the same name, the version number is recorded based on the date of creation.

[0061] The organizational information may be the same as the organizational information (FIG. 7(a)) in the organization / management target object ledger recording unit 251. The same table may be used, or the same information may be recorded in a different table.

[0062] Information on document classifications includes identification information such as names frequently used in documents to indicate their type. Higher-level classifications such as "report" and "design document" are associated with lower-level classifications such as "subdivision details" that refine the higher-level classifications, such as "construction report" and "inspection report." Furthermore, "notation," which is the expression of documents that appears in actual documents, is also associated.

[0063] The information on the business classification includes identification information such as names that indicate the types of business performed by an organization. The business classification (survey, design, construction, etc.) is associated with the "notation" that expresses the business classification as it appears in actual documents, etc.

[0064] Fig. 8 shows an example of a table in the document ledger recording unit 252. Fig. 8(a) is an example of a table of document names and creation dates, Fig. 8(b) is an example of a table of document classification information, and Fig. 8(c) is an example of a table of business classification information.

[0065] The position ledger recording unit 253 records the position information of the managed object. The position ledger recording unit 253 records one or more of the position name, position type, coordinates, distance, time, etc. of the managed object in association with each other.

[0066] The location name of the managed object is an item that indicates the type of managed object and evidence, and examples of location names include route names, structure names, and facility names. Location names may change during construction and completion, and are recorded periodically or irregularly (after changes, etc.). The location type is information that indicates the type of managed object. The coordinate and distance information records the coordinates (latitude, longitude, altitude) of the managed object and distance information indicating the distance from a specified point. Distance can be converted to latitude and longitude using the LRS (Linear Referencing System) conversion formula, and latitude and longitude can be converted to distance using map matching. The period records the validity period of the managed object. Figure 9 shows an example of the location ledger recording unit 253.

[0067] The vocabulary and mapping definition recording unit 26 records ontology definitions that represent the conceptual system of basic vocabulary having a graph structure, definitions of synonyms, similar words, and related words related to the basic vocabulary (hereinafter referred to as "related word definitions"), and a key word dictionary. The ontology definitions, related word definitions, and key word dictionary may be recorded as data sets in advance, or may be generated from documents, etc. recorded in the document, etc. recording unit 20.

[0068] The ontology definition defines basic vocabulary contained in documents such as technical documents and business manuals (vocabulary related to the structure of the managed object, vocabulary related to the state of the managed object, etc.) in a graph structure as superordinate and subordinate concepts, and synonyms of the basic vocabulary as basic terms. Figure 10 shows an example of an ontology definition.

[0069] Synonyms in the definition of related terms are vocabulary that indicates concepts that have the same meaning as concepts of the basic vocabulary expressed in the ontology definition (for example, managed objects or phenomena with the same meaning). For example, if "Hokuriku Expressway" is defined in the ontology definition, then "Hokuriku Road" is a synonym. Synonyms in the definition of related terms are vocabulary that indicates concepts similar to concepts of the basic vocabulary expressed in the ontology definition (for example, managed objects or phenomena with similar meaning). For example, if "crack" is defined as a type of deformation in the ontology definition, similar concepts such as "fissure," "crack," and "crack" are synonyms. Related terms in the definition of related terms are vocabulary that indicates concepts that frequently appear in relation to concepts of the basic vocabulary expressed in the ontology definition (for example, cause and effect, phenomenon and action, composition and material, etc.). Synonyms, similar words, and related terms are extracted by expanding the graph structure of the basic vocabulary expressed in the ontology definition as a framework to reveal their relationships. A graph is constructed using the basic vocabulary of the ontology definition, and then expanded by associating it with words in the target document, etc. Figure 11 shows an example of a graph structure expanded with synonyms and related words. The dotted lines in Figure 11 represent synonyms and related words expanded from the ontology definition in Figure 10.

[0070] The key word dictionary is a list of contextually and semantically important vocabulary extracted from several dozen documents recorded in the document recording unit 20, from among the vocabulary appearing in ontology definitions and related term definitions; for example, approximately 1,000 words are extracted. Figure 12 shows an example of a key word dictionary. Note that the number of words extracted as key words is not limited to 1,000, but can be any number. For example, it can be approximately 100 words or 10,000 words, and can be selected appropriately depending on the purpose. Furthermore, documents selected from the document recording unit 20 are selected to contain high-quality text, taking into consideration the comprehensiveness and consistency of the content.

[0071] The search and display processing unit 27 receives a search query input from the searcher terminal 3 used by a person searching for documents, etc., extracts documents, etc. and / or chunks that match the search query, and displays them as search results on the searcher terminal 3. The search and display processing unit 27 has an index management unit 271, a search processing unit 272, and a display processing unit 273.

[0072] The index management unit 271 creates and manages custom dictionaries for each region, organization, and domain to enable flexible management of search attributes according to the characteristics of each region and organization. Custom dictionaries are created that include technical terms, abbreviations, place names, etc. specific to each different region or organization, and are managed as indexes. This enables searches that reflect the characteristics of each region and organization. The index management unit 271 also manages the creation and update history of custom dictionaries. This makes it possible to analyze the reproducibility of past search results and the impact of updating custom dictionaries. Furthermore, the index management unit 271 automatically selects and applies the dictionary that is optimal for the document, etc. to be searched, and performs the search.

[0073] The search processing unit 272 receives a search query for documents, etc. from the searcher terminal 3, and searches for documents, etc. that match the search query by referring to the custom dictionary recorded in the index management unit 271. The search processing unit 272 receives, as a search query, search keywords, attributes, keyword words, etc. that have been assigned to each chunk by the attribute assignment processing unit 22 and keyword assignment processing unit 23 from the searcher terminal 3, and extracts matching documents, etc. from the document etc. recording unit 20 and chunks from the chunk recording unit 24 by referring to the custom dictionary in the index management unit 271 and synonyms, similar words, and related words in the vocabulary and mapping definition recording unit 26.

[0074] The search processing unit 272 accepts input of search conditions such as search keywords, attributes, and key words as a search query. Search keywords are keywords freely entered by the person conducting the search. Search conditions are attributes and key words associated with chunks and are the attributes and key words for which a search is desired. While the search conditions are illustrated as attributes and / or key words, other conditions may also be used. Note that it is not necessary to accept input of either the search keywords or the search conditions. Based on the search keywords and / or search conditions in the search query, the search processing unit 272 references the custom dictionary in the index management unit 271 and the synonyms, similar words, and related words in the vocabulary and mapping definition recording unit 26, and extracts data and chunks of documents, etc., containing the search keywords and / or their synonyms, similar words, and related words from among the documents, etc., recorded in the document, etc. recording unit 20 and the chunks recorded in the chunk recording unit 24, and creates an initial result set. If the search query contains attributes, key words, etc. as search conditions, a score calculation is performed to narrow down the results, such as increasing the score, and a predetermined number of the top scores from the score calculation are displayed as search results by the display processing unit 273, which will be described later.

[0075] There are various methods for calculating scores using search conditions from the initial result set, but for example, the number of occurrences of attributes and key words for each chunk can be counted and the counted number can be calculated as the score. Also, weighting can be set according to attributes and key words, with attributes counted as 1 and key words counted as 2.

[0076] The display processing unit 273 displays the results of the search processing performed by the search processing unit 272 on the searcher terminal 3. At this time, for the search results, based on the identification information such as the file name of the document etc. assigned by the attribute assignment processing unit 22 and the appearance position, the display processing unit 273 extracts the data of the document etc. from the document etc. recording unit 20 and / or extracts the chunks from the chunk recording unit 24 and displays them.

[0077] An example of a search screen is shown in Fig. 13. In Fig. 13, an input field for search keywords and input fields for search conditions such as attributes and keywords are provided, which can be input as a search query. Any attribute, keyword, etc. can be selected as a search condition. Then, the display processing unit 273 displays search results that match the search keywords and search conditions input as the search query.

[0078] 13, a search keyword is first entered, and chunks that match the search keyword are searched for based on that search keyword. The display processing unit 273 displays those chunks as search results, and may also display the top attributes, keywords, etc. assigned to the matching chunks in the search condition input field, allowing the searcher to select those search conditions, and the display processing unit 273 may display chunks that match the search conditions selected by the searcher as narrowed-down results.

[0079] Alternatively, a search screen such as that shown in Fig. 14 may be used. Search conditions such as search keywords, attributes, and keywords are entered as a search query in the search pane of Fig. 14, and the search results are received from the display processing unit 273 and displayed in the display pane.

[0080] Furthermore, a search screen such as that shown in Fig. 15 may also be used. Fig. 15 shows a case where, when a photograph, map, image, or the like of a managed object is selected from the list of chunks displayed in the display pane of Fig. 14, the photograph is displayed in the display pane, and information about the managed object located near the photograph is displayed in the detailed display pane.

[0081] The search screen can be configured arbitrarily, and examples are shown in Figures 13 to 15. Display using other screen configurations may also be used. [Example]

[0082] Next, an example of processing performed by the information processing system 1 of the present invention will be described with reference to the flowchart of FIG.

[0083] In the information processing system 1, digitized documents and the like are recorded in the document and the like recording unit 20. Therefore, the chunk generation processing unit 21 reads out the electronic data of the documents and the like recorded in the document and the like recording unit 20 at a predetermined timing, for example, at a predetermined time every day, a predetermined time every week, or at the timing when the documents and the like are recorded in the document and the like recording unit 20, and converts the data into text by performing character recognition processing such as OCR (S100).

[0084] Then, the chunk generation processing unit 21 divides the read document or the like into sentence groups (chunks) of a predetermined length (S110).

[0085] Furthermore, the attribute assignment processing unit 22 assigns matching information as an attribute of the chunk by a method such as character string pattern matching based on the information in the ledger recording unit 25 (S120). That is, the chunk generation processing unit 21 determines whether the chunk contains information that matches the organization information in the organization / managed object ledger recording unit 251, the information on the managed object, the names and creation dates of documents etc. related to the managed object in the document etc. ledger recording unit 252, the information on the organization that created the documents etc. related to the managed object, the classification of documents etc. related to the managed object, the information on the business classification, and the information on the location name, location type, coordinates, distance, time etc. of the managed object in the location ledger recording unit 253, and if matching information is found, assigns it to the chunk as an attribute.

[0086] Furthermore, the attribute assignment processing unit 22 assigns to each chunk identification information such as the file name of the document, etc., and the position of appearance (page number, line number, coordinates within the page, etc.) in order to record where the chunk is written in the document, etc. When the display processing unit 273 of the search and display processing unit 27 displays chunks that are search results, the data of the document, etc. to be recorded in the document, etc. recording unit 20 and / or the chunk to be recorded in the chunk recording unit 24 are displayed based on this position of appearance.

[0087] The keyword assignment processing unit 23 assigns matching information as an attribute of the chunk using a method such as character string pattern matching based on the keyword dictionary in the vocabulary and mapping definition recording unit 26 (S130).

[0088] By performing the above processing, attributes and key words are assigned to each chunk, so that the information of the target document, etc. can be converted into text and recorded in the form of chunks to which attributes and key words have been assigned, and the information can be recorded in the chunk recording unit 24.

[0089] When a searcher wishes to extract documents containing a predetermined keyword from documents, the searcher inputs a search query including search keywords, attributes, keywords, and other search conditions on the search screen.

[0090] When the search processing unit 272 of the search and display processing unit 27 receives a search query from the searcher terminal 3, it executes search processing (S140). That is, based on the search keywords in the search query, the search processing unit 272 references the custom dictionary of the index management unit 271 and synonyms, similar words, and related words recorded in the vocabulary and mapping definition recording unit 26, and extracts chunks of documents, etc. containing the search keywords from among the data of documents, etc. recorded in the document, etc. recording unit 20 and the chunks recorded in the chunk recording unit 24, to create an initial result set. Then, for each chunk included in the initial result set, if the search query includes search conditions such as attributes or keywords, it performs score calculations to narrow down the results, such as increasing the score, and causes the display processing unit 273 to display a predetermined number of chunks from the top of the score calculations, for example, the top 10 chunks, as search results on the searcher terminal 3.

[0091] By performing the above-mentioned process, a searcher can easily find documents that meet the search criteria. In addition, in the ledger recording unit 25, past organizations, names of managed objects, classifications of documents, classifications of work, etc. are managed as unified attributes and are assigned to chunks as attributes, so even if these names change, the changes can be accommodated. [Example]

[0092] Next, a modified example of the processing of the information processing system 1 of the present invention will be described. In this embodiment, the processing when each piece of information recorded in the organization / managed object ledger recording unit 251, the document ledger recording unit 252, and the location ledger recording unit 253 in the ledger recording unit 25 is newly recorded will be described. An example of the configuration of the information processing system 1 in this embodiment is shown in Fig. 16.

[0093] The information processing system 1 in this embodiment includes a ledger generation processing unit 28 in addition to the configuration of the first embodiment.

[0094] The ledger generation processing unit 28 updates the information recorded in the ledger recording unit 25. The ledger generation processing unit 28 has processing units corresponding to the information recorded in the ledger recording unit 25, and has an organization / managed object ledger generation processing unit 281, a document etc. ledger generation processing unit 282, and a location ledger generation processing unit 283.

[0095] The organization and managed object ledger generation processing unit 281 accepts input of changes to the organization information and managed object information, which are data sets to be recorded in the organization and managed object ledger recording unit 251, and records them in the organization and managed object ledger recording unit 251.

[0096] Information on the identification name and location of the organization that manages the managed object is received by the organization and managed object ledger generation processing unit 281 whenever an organization is split or merged, or when the name or address is changed, and is reflected in the organization and managed object ledger recording unit 251. The organization information may be received by manual input, or information on the change may be acquired periodically or irregularly by referring to a data table in another computer system that records information on the organizational structure, and recorded in the organization information in the organization and managed object ledger recording unit 251.

[0097] The names of the information on the managed objects may change during the construction period and the management period. Therefore, whenever there is a change in the information on the managed objects, the organization and managed object ledger generation processing unit 281 accepts the input of that information and reflects it in the organization and managed object ledger recording unit 251. Note that the information on the managed objects may be manually entered, or the information on the changes may be acquired periodically or irregularly by referring to a data table in another computer system that records the information on the managed objects, and recorded in the information on the managed objects in the organization and managed object ledger recording unit 251. The names in the facility information may represent a certain range, such as an AA road, or a point, such as a BB interchange, so it is advisable to organize the relationship and associate the names of the managed objects with coordinate expressions.

[0098] The document etc. ledger generation processing unit 282 updates (updates) the names and creation dates of documents etc. related to the managed objects, which are data sets recorded in the document etc. ledger recording unit 252, information on the organization that created the documents etc. related to the managed objects, information such as the classification (type) of documents etc. related to the managed objects, and information indicating the classification (type) of work performed by the organization, by referring to the data of documents etc. recorded in the document etc. recording unit 20 (data of documents etc. after being converted to text).

[0099] The document ledger generation processing unit 282 extracts the name and creation date of the document etc. according to predetermined rules for extracting the name and creation date of the document etc. from the data of the document etc., and records them in the document ledger recording unit 252.

[0100] For example, the name of a document is extracted from the first block of text on the first page of the document, written in the largest font size. The creation date is extracted from the date written on the first and last pages of the document. There are multiple ways to write the date, so you can set the pattern in advance and extract the date that matches one of the patterns as the creation date.

[0101] The document ledger generation processing unit 282 may execute the same processing as that in the organization / management object ledger generation processing unit 281 as organizational information, or may acquire the organizational information generated by the organization / management object ledger generation processing unit 281 and record it in the document ledger recording unit 252.

[0102] The document ledger generation processing unit 282 automatically extracts the classification of a document, etc. from the contents of the document, etc. For example, it determines the classification of a document, etc. based on names frequently used as the type of material, such as "drawing," "design document," or "research report," that appear in the text on the first page of the document, etc., related to the managed object, and records the classification in the document ledger recording unit 252. The names of material types are set in advance as initial values ​​in the ontology definition, related term definition, key word dictionary, etc., of the vocabulary and mapping definition recording unit 26, and can be referenced. The same material classification is then assigned to chunks generated from the document, etc.

[0103] The document ledger generation processing unit 282 automatically extracts task classifications from the contents of a document, etc. For example, it determines the task classification based on the name indicating the task classification written in the text on the first page of the document, such as a commonly used name for task classification, such as "inspection and design," "construction," "management," or "repair work," and records the determined task classification in the document ledger recording unit 252. The task classification names are set in advance as initial values ​​in the ontology definition, related term definition, and key word dictionary of the vocabulary and mapping definition recording unit 26, and can be referenced. The same task classification is then assigned to chunks generated from the document, etc.

[0104] The location register generation processing unit 283 accepts input of identification information, distance, coordinates, place name, etc. of managed objects to be recorded in the location register recording unit 253, and reflects the information in the location register recording unit 253. Information about managed objects may, for example, change in name during construction and management periods. Therefore, whenever there is a change in the information about the managed object, the organization and managed object register generation processing unit 281 accepts the input of that information and reflects it in the organization and managed object register recording unit 251. Note that the information about the managed object may be manually entered, or information about the change may be acquired periodically or irregularly by referencing a data table in another computer system that records the information about the managed object, and recorded in the information about the managed object in the organization and managed object register recording unit 251. The names in the information about the managed object may represent a certain range, such as an AA road, or a point, such as a BB interchange. Therefore, it is advisable to organize the relationship and associate the names of the managed objects with coordinate expressions.

[0105] Coordinate representation may be defined as a point, a line string, or a polygon.

[0106] By the ledger generation processing unit 28 executing the above-described processing, each piece of information to be recorded in the ledger recording unit 25 can be input and updated. [Example]

[0107] Next, another modified example of the processing of the information processing system 1 of the present invention will be described. In this embodiment, a case will be described in which the vocabulary and mapping definitions in the vocabulary and mapping definition recording unit 26 are automatically expanded. An example of the configuration of the information processing system 1 in this embodiment is shown in FIG.

[0108] The information processing system 1 in this embodiment has a vocabulary and mapping definition generation processing unit 29 in addition to the configuration of embodiment 2. In the following description, a case where the vocabulary and mapping definition generation processing unit 29 is added to the configuration of embodiment 2 will be described, but the vocabulary and mapping definition generation processing unit 29 may also be added to the configuration of embodiment 1.

[0109] The vocabulary and mapping definition generation processing unit 29 has a vocabulary extraction function (vocabulary extraction processing unit), a mapping definition generation function (mapping definition generation processing unit), and a key word dictionary generation function (key word dictionary generation processing unit).

[0110] The vocabulary extraction function selects data to be used as training data from the digitized document data recorded in the document recording unit 20, and constructs a training corpus. Then, some or all of technical terms, spelling variations, synonyms, related words, proper nouns, and phrase classifications are extracted from the training corpus. It is preferable that the documents and other data used as the training corpus be selected by knowledgeable experts. Technical terms, spelling variations, synonyms, related words, proper nouns, and phrase classifications may be extracted automatically using known techniques, or may be manually input.

[0111] The mapping definition generation function generates mapping definitions for orthographic variations and for multiple subordinate concepts of a broader concept term in word classification. For example, as shown in Figure 10, concepts such as "crack," "peeling," and "efflorescence" are acquired from the training corpus as subordinate concepts of "damage," and concepts such as "bidirectional crack" are acquired as subordinate concepts of "crack." The accuracy of the mapping definition can be improved by using data such as document classifications and business classifications as initial vocabulary for the broader concepts. For example, concepts such as "damage," "crack," "peeling," and "efflorescence" can be acquired using a bootstrap method or similar method. These processes can be performed using known methods, such as those described in Mikolov, Tomas et al., "Efficient Estimation of Word Representations in Vector Space," International Conference on Learning Representations (2013).

[0112] The key word dictionary generation function may accept manual input, or may use frequently occurring words as key words in data such as documents recorded in the document recording unit 20. Also, known methods such as those described in LEGAL-BERT: The Muppets straight out of Law School, Chalkidis et al., Findings (2020) may be used.

[0113] By performing the above-described processing, the vocabulary and mapping definition recording unit 26 can be automatically generated and updated. [Industrial Applicability]

[0114] By using the information processing system 1 of the present invention, documents and the like relating to objects to be managed such as structures, buildings, and equipment can be managed appropriately, and when performing a search, desired documents and the like can be reached more quickly than before. [Explanation of symbols]

[0115] 1: Information processing system 2: Management terminal 3: Searcher terminal 20: Document Records Department 21: Chunk generation processing unit 22: Attribute assignment processing unit 23: Keyword assignment processing section 24: Chunk recording section 25: Ledger Recording Department 26: Vocabulary and mapping definition record section 27: Search and display processing section 28: Ledger generation processing unit 29: Vocabulary and mapping definition generation processing unit 70: Arithmetic device 71:Storage device 72:Display device 73: Input device 74:Communication equipment 251: Organization and Management Object Ledger Recording Department 252: Document Ledger Record Division 253: Location register recording section 271: Index Management Department 272: Search processing unit 273: Display processing unit 281: Organization and managed object ledger generation processing unit 282: Document ledger generation processing unit 283: Location register generation processing unit

Claims

1. An information processing system for managing documents and the like related to managed objects, The information processing system includes: a chunk generation processing unit that generates chunks from data such as documents; an attribute assignment processing unit that assigns attributes to the generated chunks based on information to be recorded in a ledger recording unit; a keyword assignment processing unit that assigns a keyword to the generated chunk based on a keyword dictionary recorded in a vocabulary and mapping definition recording unit; a chunk recording unit that records attributes and / or keywords assigned to the generated chunks in association with the chunks; a search and display processing unit that receives a search query including attributes and / or keywords from a searcher terminal used by a searcher, refers to the chunk recording unit, and displays matching chunks on the searcher terminal; An information processing system comprising:

2. The ledger recording unit The system has an organization / managed object ledger recording unit that records information on the organization that manages the managed object and / or information on the managed object, The organization / management object ledger recording unit The information on the organization stores information indicating changes in the organization, information indicating a transition of the managed object is stored as the information of the managed object; 2. The information processing system according to claim 1, wherein:

3. The information processing system includes: an organization / management object ledger generation processing unit that generates and / or updates information to be recorded in the organization / management object ledger recording unit, The organization / management object ledger generation processing unit accepting input of information about the organization or information about the managed object, acquiring information about the organization or the managed object from a computer system that records information about the organization or a computer system that records information about the managed object, and recording the information in the organization / managed object ledger recording unit; 3. The information processing system according to claim 2.

4. The ledger recording unit a document ledger recording unit for recording one or more of the names and creation dates of documents, etc. related to the managed objects, information on the organization that created the documents, etc. related to the managed objects, information on the classification of the documents, etc., and information on the classification of business, The document ledger recording unit includes: As information on the classification of the document, etc., a superordinate concept of the document, etc., and a classification and notation of its subordinate concept are stored in association with each other; The information on the business classification is stored in association with the business classification and its expression.

2. The information processing system according to claim 1, wherein:

5. The information processing system includes: a document ledger generation processing unit that generates and / or updates information to be recorded in the document ledger recording unit; The document ledger generation processing unit extracting the name and creation date of the document from a predetermined portion of the document, extracting the classification of the document and the classification of the business from the content of the document, and recording them in the document ledger recording unit; 5. The information processing system according to claim 4.

6. The ledger recording unit a location register recording unit for recording location information of the managed object; The position register recording unit As the location information of the managed object, one or more of a location name, a location type, coordinates, a distance, and a time are stored.

2. The information processing system according to claim 1, wherein:

7. The information processing system includes: a location register generation processing unit that generates and / or updates information to be recorded in the location register recording unit; The location register generation processing unit receiving an input of location information of the managed object; acquiring the location information of the managed object from a computer system that records the location information of the managed object, and recording the location information in the location book recording unit; 7. The information processing system according to claim 6.

8. The vocabulary and mapping definition recording unit an ontology definition that represents a conceptual system of basic vocabulary having a graph structure, definitions of related terms related to the basic vocabulary, and a dictionary of important words are recorded; The ontology definition: A basic vocabulary contained in the document etc. is defined in a graph structure, The definitions of related terms are as follows: one or more of synonyms, which are vocabulary terms indicating the same concepts as the concepts of the basic vocabulary expressed in the ontology definition; synonyms, which are vocabulary terms indicating concepts similar to the concepts of the basic vocabulary expressed in the ontology definition; and related words, which are vocabulary terms indicating concepts that appear in relation to the concepts of the basic vocabulary expressed in the ontology definition, are defined by extending the graph structure of the ontology definition; The key word dictionary is a list of important vocabulary selected from vocabulary appearing in the ontology definition, related term definition, etc.

2. The information processing system according to claim 1, wherein:

9. The information processing system includes: a vocabulary and mapping definition generation processing unit that generates and / or updates information to be recorded in the vocabulary and mapping definition recording unit, The vocabulary and mapping definition generation processing unit a vocabulary extraction processing unit that extracts some or all of technical terms, spelling variations, synonyms, similar words, related words, proper nouns, and word classifications from a training corpus that is constructed based on data such as documents recorded in the document recording unit; a mapping definition generation processing unit that generates mapping definitions of the spelling variations and the superordinate and subordinate concepts of the words and phrases; a key word dictionary generation processing unit that generates a key word dictionary, which is a dictionary of key words, based on data of documents, etc. recorded in the document, etc. recording unit; 9. The information processing system according to claim 8, further comprising:

10. The search and display processing unit receiving a search query including a search keyword or an attribute and / or a keyword from the searcher terminal; Searching for document data in a document storage unit based on the search keywords in the search query, extracting chunks of documents that match the search keywords, and creating an initial result set; narrowing down the search results using the attributes and / or keywords in the search query and the attributes and / or keywords assigned to the chunks in the created result set; 2. The information processing system according to claim 1, wherein:

11. The search and display processing unit receiving a search query including a search keyword and an attribute and / or a keyword from the searcher terminal; identifying synonyms, similar words, and / or related words by referring to the vocabulary and mapping definition recorder using search keywords in the search query; searching for document data in a document storage unit using the search keyword and the identified synonyms, similar words, and / or related words, and extracting chunks of matching documents to create an initial result set; narrowing down the search results using the attributes and / or keywords in the search query and the attributes and / or keywords assigned to the chunks in the created result set; 2. The information processing system according to claim 1, wherein:

12. The search and display processing unit Calculating a score using search conditions of attributes and / or keywords assigned to chunks in the initial result set, and displaying a predetermined number of results from the top scores as search results.

12. The information processing system according to claim 10 or 11.

13. The search and display process includes: displaying chunks in the initial result set on the searcher terminal; Accepting input of search conditions for attributes and / or key words from the searcher terminal, and displaying chunks of the initial result set that match the search conditions for attributes and / or key words as search results on the searcher terminal; 12. The information processing system according to claim 10 or 11.

14. A computer that manages documents related to managed items, a chunk generation processing unit that generates chunks from data such as documents; an attribute assignment processing unit that assigns attributes to the generated chunks based on information to be recorded in a ledger recording unit; a keyword assignment processing unit that assigns a keyword to the generated chunk based on a keyword dictionary recorded in a vocabulary and mapping definition recording unit; a search and display processing unit that receives a search query including attributes and / or keywords from a searcher terminal used by a searcher, refers to a chunk recording unit that records the attributes and / or keywords assigned to the generated chunks in association with the chunks, and displays matching chunks on the searcher terminal; An information processing program that functions as a

Citation Information

Patent Citations

  • Construction management support apparatus

    JP2022187749A