Structured ontology construction and application method for information search demand
By constructing a structured ontology model for intelligence, the problem of lack of purpose and comprehensiveness in intelligence search is solved, and more efficient search and analysis results are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CSSC SYST ENG RES INST
- Filing Date
- 2021-10-30
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies lack purposefulness and comprehensiveness in intelligence gathering, resulting in one-dimensional search behavior and incomplete results.
We construct a structured intelligence ontology by mining representative sets of words related to intelligence search needs to form a structured ontology model, and use entity recognition technology to guide searches and aggregate results.
It enhances the purposefulness and comprehensiveness of intelligence gathering, and improves analysis efficiency and the accuracy of results.
Smart Images

Figure CN114238547B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of natural language processing technology, specifically relating to a structured ontology construction and application method for intelligence search needs. Background Technology
[0002] With the rapid development of internet technology, the amount of open-source intelligence data is increasing, gradually exhibiting characteristics of massive volume, disorganization, low information density, and information overload. This significantly hinders the efficiency and performance of intelligence search and analysis. When intelligence analysts are assigned a specific target analysis task, they often lack effective search guidance, wasting a lot of time and resulting in low overall efficiency.
[0003] Currently, most traditional solutions for intelligence retrieval needs are based on keywords or keyword expansion methods. They then use fuzzy keyword search and document relevance calculations to return the documents with the highest relevance to the keyword as search results. This traditional approach has two typical shortcomings:
[0004] (1) Lack of purpose in search behavior: The method of using simple keywords or keyword synonyms as search conditions can only perform searches with a single-dimensional condition. In order to obtain more detailed search results, users need to list and combine keywords themselves, and this listing and combination work requires a lot of background and prior knowledge.
[0005] (2) Lack of comprehensive search results: Current keyword-based intelligence search results are simply aggregated results obtained by fuzzy matching of a certain intelligence search keyword, and cannot provide more detailed information about that keyword. Users need to further summarize and generalize based on this. Summary of the Invention
[0006] To address the aforementioned problems, this invention provides a structured ontology construction and application method for intelligence retrieval needs, comprising the following steps:
[0007] The construction of structured intelligence terminology involves mining a set of representative words from intelligence search needs, which serves as the data source for the construction of structured ontology.
[0008] The construction of a structured intelligence ontology reveals a structured ontology model for intelligence search needs. The intelligence ontology model includes information such as entity type, entity attribute type, and entity relationship type.
[0009] The application of structured intelligence ontology involves using a pre-constructed structured intelligence ontology to return the attribute dimensions of entities based on the user's search query through entity recognition.
[0010] Furthermore, the construction of the aforementioned structured intelligence terminology specifically includes the following steps:
[0011] To address intelligence search needs, a domain terminology dictionary was created through a combination of expert development, manual collection, and analysis from vertical websites.
[0012] To address intelligence search needs, an unstructured text corpus for intelligence search needs was created through web crawling and manual collection.
[0013] For an unstructured text corpus of intelligence search needs, a terminology extraction model is used to mine an intelligence search needs terminology dictionary.
[0014] The collected unstructured text corpus of intelligence search requests is combined with open encyclopedic data, and a model is used to train the data to form a word vector file for intelligence search requests.
[0015] Based on the word vector file of intelligence search requirements, the similarity between each word and its extended words is calculated, and words with similarity higher than a threshold are retained. All words with similarity higher than the threshold constitute the term dictionary of intelligence search requirements.
[0016] Furthermore, using the intelligence gathering requirements terminology dictionary as input, each word in the dictionary is expanded using a thesaurus to obtain expanded words for each word.
[0017] Furthermore, the construction of the structured intelligence ontology includes the following steps:
[0018] Based on the terminology dictionary database for intelligence search needs, we will construct a set of encyclopedic attributes and a set of encyclopedic superordinate concepts using encyclopedia websites.
[0019] In the unstructured text of the acquired public intelligence, we mine the set of entity attributes and the set of superordinate concepts of the entities;
[0020] The encyclopedia-type attribute set, the encyclopedia-type superordinate concept set, the entity attribute set, and the entity superordinate concept set are merged, deduplicated, and then normalized.
[0021] The normalized result is transformed according to the ontology storage structure to finally obtain the structured ontology of intelligence.
[0022] Furthermore, constructing an encyclopedia-type attribute set specifically includes the following steps:
[0023] Using an intelligence search demand terminology dictionary as input, iterate through each entry in the intelligence search demand terminology dictionary;
[0024] Request an encyclopedia website to obtain the encyclopedia page for the term;
[0025] Parse the attribute box information on the encyclopedia page and use the attribute names in the attribute box as the first attribute set of the entry;
[0026] Parse the directory information of the encyclopedia page and use the hierarchical names in the directory information as the second attribute set of the entry;
[0027] The first attribute set and the second attribute set are merged to obtain the attribute set corresponding to the term;
[0028] The attribute sets of all the aforementioned entries are merged to obtain the attribute set of the encyclopedia category.
[0029] Furthermore, constructing a set of overarching concepts for encyclopedic purposes specifically includes the following steps:
[0030] Using an intelligence search demand terminology dictionary as input, iterate through each entry in the intelligence search demand terminology dictionary;
[0031] Request an encyclopedia website to obtain the encyclopedia page for the term;
[0032] The tags and related field tags in the encyclopedia page of the term are analyzed to form the first set of concepts for the term.
[0033] The term is segmented; if the term ends with a noun, the last noun in the term is taken as the second set of concepts for the term.
[0034] The first concept set and the second concept set are merged to obtain the superordinate concept set corresponding to the term;
[0035] The sets of superordinate concepts of all the aforementioned terms are merged to obtain the set of superordinate concepts for the encyclopedia category.
[0036] Furthermore, the application of the structured intelligence ontology includes the following steps:
[0037] The intelligence search terminology dictionary and the intelligence structured ontology are combined and then stored in a database.
[0038] The terminology names and ontology concept names in the structured intelligence terminology database are stored in the full-text search database as an index database for user searches.
[0039] Entity recognition is performed based on the constructed intelligence structured terminology database. For entities input by the user, the type and attribute information of the entity are obtained from the intelligence structured ontology database.
[0040] Return the entity type and attribute information, and use visualization methods to display the index directory;
[0041] After clicking on the index directory, the system uses the entity name and selected attributes as search keywords to search the collected unstructured intelligence text, obtain relevant search results, and then aggregate and display them.
[0042] Furthermore, the step of obtaining the type and attribute information of the entity from the structured intelligence ontology database specifically includes the following steps:
[0043] If the entity in the terminology database can be accurately found in the user input, the type and attribute information of the entity can be obtained directly from the intelligence structured ontology database.
[0044] If the entity in the terminology database cannot be accurately found from the user input, an approximate entity is obtained using the user input, and the type and attribute information of the approximate entity are obtained from the structured ontology database.
[0045] Furthermore, obtaining an approximate entity using user input specifically includes the following steps:
[0046] The user-input question is matched against the full-text search database.
[0047] Record entity names with a similarity greater than the threshold;
[0048] Linking the entity name to the terminology database yields an approximate entity.
[0049] The beneficial effects of this invention are as follows: This invention provides an ontology construction and application method oriented towards intelligence search needs, which can aggregate data from multiple sources and quickly mine terms and ontology data suitable for intelligence search needs; based on the constructed terms and ontology data, the purposefulness and comprehensiveness of current intelligence search and analysis can be effectively enhanced, and the analytical efficiency of user intelligence search can be effectively improved. Attached Figure Description
[0050] Figure 1 A schematic diagram of the process of this invention. Detailed Implementation
[0051] The present invention will be further described below with reference to the accompanying drawings and embodiments. The following embodiments are only used to explain the invention and are not intended to limit the scope of protection of the present invention.
[0052] This invention provides a fast and efficient ontology modeling method for intelligence search scenarios, forming an ontology knowledge base (prior knowledge) suitable for intelligence search needs. Based on this, practical search applications are validated, automatically guiding search and aggregating results based on the information dimensions involved in the user's search keywords, thereby improving the intelligence search, information aggregation, and analysis capabilities of intelligence analysts. Figure 1 As shown, the specific steps are as follows:
[0053] Step 1: Constructing an intelligence structured terminology database.
[0054] The aim is to uncover a representative set of terms within the domain of intelligence search needs. These terms clearly represent keywords or question formats that intelligence analysts might search for. This construction method includes the following five sub-steps:
[0055] 1.1 To meet intelligence search needs, a domain terminology dictionary is created through methods such as expert development, manual collection, and parsing from vertical websites.
[0056] 1.2 To address intelligence search needs, an unstructured text corpus for non-intelligence search needs is created through web crawling and manual collection.
[0057] 1.3. Using CRF terminology extraction models, an intelligence search terminology dictionary is mined from the unstructured text corpus of intelligence search requirements.
[0058] 1.4 The collected unstructured text corpus of intelligence search requirements and open encyclopedia data are used to train the Word2vec model to form a word vector file of intelligence search requirements.
[0059] 1.5. Using the intelligence search requirement terminology dictionary database formed in steps 1.1 and 1.3 as input, expand the terminology for each word in the dictionary database using the constructed thesaurus, and then use the steps...
[0060] The word vector files obtained in 1.4 are used to calculate the similarity between words, and words with similarity values above the threshold are retained, ultimately resulting in a dictionary of terms for intelligence search needs.
[0061] Step 2: Constructing the structured ontology of intelligence.
[0062] The aim is to mine a structured ontology model for the intelligence search demand domain. This ontology model includes information such as entity types, entity attribute types, and entity relationship types, representing the three-dimensional search dimensions within intelligence search demands. The structured intelligence ontology is represented as {O, C, R, A}, where: O represents the set of intelligence search demand concept ontology, C represents the set of intelligence search demand concepts, R represents the set of all relations related to the intelligence search demand, and A represents the set of all attributes related to the intelligence search demand. The specific implementation includes the following 15 sub-steps:
[0063] 2.1 In response to intelligence search needs, experts develop and manually collect existing ontologies, such as from cnschema.org and Freebase, as an existing set of structured ontologies;
[0064] 2.2 Using the intelligence search demand terminology dictionary database obtained in the construction of the intelligence structured terminology database as input, traverse each entry in it and perform the operations of steps 2.3-2.9 on encyclopedia websites such as Baidu Encyclopedia respectively;
[0065] 2.3. Request the encyclopedia website to obtain the encyclopedia page for the term;
[0066] 2.4. Parse the infobox information on the encyclopedia page and use the attribute names in the infobox as the attribute set for the entry. For example, for "Boeing 737 passenger plane" in Baidu Encyclopedia, the attribute set is {"Chinese name", "Foreign name", "Country", "Aircraft name"};
[0067] 2.5. Parse the directory information of the encyclopedia page and use the hierarchical names in the directory information as the attribute set of the entry. For example, for "Boeing 737 passenger plane" in Baidu Encyclopedia, the attribute set obtained is {"development history", "technical characteristics", "overall evaluation"};
[0068] 2.6. Merge the attribute set results obtained in steps 2.4 and 2.5 to form the attribute set corresponding to the term;
[0069] 2.7. Analyze the tags and domain tags in the encyclopedia page of this entry to form the set of parent concepts for this entry;
[0070] 2.8. Perform word segmentation on the entry. If the entry ends with a noun, then the last noun in the entry is taken as the concept set of the entry.
[0071] 2.9. Merge the concept sets obtained in steps 2.7 and 2.8 to obtain the set of superordinate concepts to which the term belongs.
[0072] 2.10. Merge the sets of superordinate concepts and attribute sets from multiple encyclopedia sources to obtain the set of superordinate concepts and attribute sets for the encyclopedia category.
[0073] 2.11. In the unstructured text of the obtained public intelligence, the set of entity attributes is mined by setting heuristic mining rules for entity attributes (such as "A's B").
[0074] 2.12. In the unstructured text of the obtained public intelligence, the set of superordinate concepts of entities is mined by setting heuristic mining rules for superordinate entities (such as "A is a kind of B" or "A is a kind of B").
[0075] 2.13. Merge the sets obtained in steps 2.10, 2.11, and 2.12, remove duplicates, and normalize them. The normalization method is achieved by using name similarity calculation and name thesaurus mapping.
[0076] 2.14. The results obtained in step 2.13 are transformed according to the ontology storage structure to finally form an intelligence structured ontology.
[0077] Step 3: Application of the structured intelligence ontology.
[0078] The aim is to leverage a pre-constructed structured intelligence ontology to intervene in question search and answer return based on user search needs, making user searches more targeted and specific, and providing a more hierarchical aggregation and display of search results. The specific implementation includes the following eight sub-steps:
[0079] 3.1 Store the intelligence search requirement terminology dictionary obtained from the intelligence structured terminology database construction step and the intelligence structured ontology collection obtained from the intelligence structured ontology construction step using databases such as MongoDB.
[0080] 3.2 Store the terminology names and ontology concept names from the structured intelligence terminology database into the Elasticsearch full-text search database as an index database for user searches.
[0081] 3.3. Based on the user's input, perform entity recognition using the constructed structured intelligence terminology database, and link the entities in the user's input to the terminology database. There are two scenarios, in which steps 3.4 and 3.5 are executed respectively.
[0082] 3.4 If the entity in the terminology database can be accurately found in the user's question, proceed directly to step 3.6;
[0083] 3.5 If the entity in the terminology database cannot be accurately found in the user's question, the user's question is matched with the elastic-search full-text search database established in step 2, a similarity threshold is set, and entities with a similarity greater than a certain threshold are used as the accurate entity names and linked to the terminology database.
[0084] 3.6 Obtain the entity's ontology information from the structured ontology database, including the entity's type and attribute information.
[0085] 3.7 Return the entity type and attribute information, and display the index directory using visualization methods.
[0086] 3.8 After clicking the index directory, use the entity name and selected attribute as search keywords to search the collected unstructured intelligence text, obtain relevant search results, and aggregate and display them.
[0087] In summary, these are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. All equivalent changes and modifications made in accordance with the scope of the present invention and the contents of the specification are within the scope of the present invention.
Claims
1. A structured ontology construction and application method for intelligence search needs, characterized in that, Includes the following steps: The construction of structured intelligence terminology involves mining a set of representative words from intelligence search needs, which serves as the data source for the construction of structured ontology. The construction of a structured intelligence ontology reveals a structured ontology model for intelligence search needs. The intelligence ontology model includes entity types, entity attribute types, and entity relationship types. The application of structured intelligence ontology, based on the constructed structured intelligence ontology, returns the attribute dimensions of entities through entity recognition in response to user search queries; The construction of the aforementioned structured intelligence terminology specifically includes the following steps: To address intelligence search needs, a domain terminology dictionary was created through a combination of expert development, manual collection, and analysis from vertical websites. To address intelligence search needs, an unstructured text corpus for intelligence search needs was created through web crawling and manual collection. For an unstructured text corpus of intelligence search needs, a terminology extraction model is used to mine an intelligence search needs terminology dictionary. The collected unstructured text corpus of intelligence search requests is combined with open encyclopedic data, and a model is used to train the data to form a word vector file for intelligence search requests. Based on the word vector file of intelligence search requirements, the similarity between each word and its extended words is calculated, and words with similarity higher than the threshold are retained. All words with similarity higher than the threshold constitute the term dictionary of intelligence search requirements. The construction of the structured intelligence ontology includes the following steps: Based on the terminology dictionary database for intelligence search needs, we will construct a set of encyclopedic attributes and a set of encyclopedic superordinate concepts using encyclopedia websites. In the unstructured text of the acquired public intelligence, we mine the set of entity attributes and the set of superordinate concepts of the entities; The encyclopedia-type attribute set, the encyclopedia-type superordinate concept set, the entity attribute set, and the entity superordinate concept set are merged, deduplicated, and then normalized. The normalized result is transformed according to the ontology storage structure to finally obtain the structured ontology of intelligence.
2. The method for constructing and applying a structured ontology oriented towards intelligence search needs as described in claim 1, characterized in that: Using an intelligence gathering requirements terminology dictionary as input, each word in the dictionary is expanded using a thesaurus to obtain expanded terms for each word.
3. The method for constructing and applying a structured ontology oriented towards intelligence search needs as described in claim 1, characterized in that, Constructing an encyclopedia-type attribute set includes the following steps: Using an intelligence search demand terminology dictionary as input, iterate through each entry in the intelligence search demand terminology dictionary; Request an encyclopedia website to obtain the encyclopedia page for the term; Parse the attribute box information on the encyclopedia page and use the attribute names in the attribute box as the first attribute set of the entry; Parse the directory information of the encyclopedia page and use the hierarchical names in the directory information as the second attribute set of the entry; The first attribute set and the second attribute set are merged to obtain the attribute set corresponding to the term; The attribute sets of all the aforementioned entries are merged to obtain the attribute set of the encyclopedia category.
4. The method for constructing and applying a structured ontology oriented towards intelligence search needs as described in claim 1, characterized in that, Constructing a set of overarching concepts for encyclopedias specifically includes the following steps: Using an intelligence search demand terminology dictionary as input, iterate through each entry in the intelligence search demand terminology dictionary; Request an encyclopedia website to obtain the encyclopedia page for the term; The tags and related field tags in the encyclopedia page of the term are analyzed to form the first set of concepts for the term. The term is segmented; if the term ends with a noun, the last noun in the term is taken as the second set of concepts for the term. The first concept set and the second concept set are merged to obtain the superordinate concept set corresponding to the term; The sets of superordinate concepts of all the aforementioned terms are merged to obtain the set of superordinate concepts for the encyclopedia category.
5. The method for constructing and applying a structured ontology oriented towards intelligence search needs as described in claim 1, characterized in that, The application of the structured intelligence ontology includes the following steps: The intelligence search terminology dictionary and the intelligence structured ontology are combined and then stored in a database. The terminology names and ontology concept names in the structured intelligence terminology database are stored in the full-text search database as an index database for user searches. Entity recognition is performed based on the constructed intelligence structured terminology database. For entities input by the user, the type and attribute information of the entity are obtained from the intelligence structured ontology database. Return the entity type and attribute information, and use visualization methods to display the index directory; After clicking on the index directory, the system uses the entity name and selected attributes as search keywords to search the collected unstructured intelligence text, obtain relevant search results, and then aggregate and display them.
6. The method for constructing and applying a structured ontology oriented towards intelligence search needs as described in claim 5, characterized in that, The specific steps for obtaining the type and attribute information of the entity from the structured intelligence ontology database include: If the entity in the terminology database can be accurately found in the user input, the type and attribute information of the entity can be obtained directly from the intelligence structured ontology database. If the entity in the terminology database cannot be accurately found from the user input, an approximate entity is obtained using the user input, and the type and attribute information of the approximate entity are obtained from the structured ontology database.
7. The method for constructing and applying a structured ontology oriented towards intelligence search needs as described in claim 6, characterized in that, Obtaining an approximate entity using user input specifically includes the following steps: The user-input question is matched against the full-text search database. Record entity names with a similarity greater than the threshold; Linking the entity name to the terminology database yields an approximate entity.
Citation Information
Patent Citations
System and method for constructing information-analysis-oriented knowledge maps
CN106815293A
Segmented semantic annotation method in weak annotation environment
CN110888991A