Address data storage method and device based on knowledge graph

By applying knowledge graph technology in address data storage to process and parse address data, the problem that traditional technology is difficult to deal with complex address data is solved, efficient storage and query of address data is achieved, and user experience is improved.

CN120067399APending Publication Date: 2025-05-30GUANGDONG NORMAL UNIV WEIZHI INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510081522.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional address data analysis technology is difficult to effectively process complex, irregular and incomplete address data, resulting in reduced parsing accuracy, and existing knowledge graphs have challenges in dynamic updates and efficient query of address data.

Method used

The address data storage method based on the knowledge graph is adopted, and the original address data is obtained, and the standardized and hierarchical address data is determined using a pre-established data processing model and hierarchical analysis model, and the preset matching algorithm is used to store it in the graph database. Finally, the data is synchronized to the search engine through the data synchronization model.

Benefits of technology

It realizes synchronous changes of hierarchical address data in the graph database and search engine, ensures the accuracy and timeliness of address information, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067399A_ABST
    Figure CN120067399A_ABST
Patent Text Reader

Abstract

The invention discloses an address data storage method and device based on a knowledge graph. The method comprises the following steps: acquiring original address data; determining standardized address data according to the original address data based on a pre-established data processing model; based on a pre-established hierarchical analysis model, according to the standardized address data, determining multiple pieces of hierarchical address data; based on a preset matching algorithm, storing the plurality of hierarchical address data in a matched graph database; and synchronously marking the plurality of hierarchical address data in the graph database to a search engine based on a pre-established data synchronization model. Therefore, the hierarchical address data can be synchronously changed in the graph database and the search engine, the accuracy and timeliness of the address information are ensured, and the use experience of a user is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of address data, and particularly relates to a method and device for storing address data based on a knowledge graph. Background Art

[0002] Address data has a complex structure and diverse expression forms. Especially in natural language, users may input the same address data in different ways. Traditional address data parsing technologies mainly rely on regular expressions or manually defined parsing rules. These methods are suitable for processing addresses with relatively standardized structures, but when dealing with complex, irregular, and incomplete address data, the parsing accuracy is greatly reduced. The method of parsing rule matching requires a large amount of manual rule writing and is difficult to expand.

[0003] With the increase in data scale and the improvement of data structure complexity, the knowledge graph technology has become an important way to structurally store and manage complex data. The knowledge graph can effectively describe complex domain knowledge through the representation of entities (nodes) and relationships (edges), and perform intelligent queries and inferences. In the prior art, the application of the knowledge graph in address data storage and management helps to improve the structural level of data, but still faces significant challenges in dynamic update and efficient query. Summary of the Invention

[0004] The present invention provides a method for storing address data based on a knowledge graph, which can synchronously change hierarchical address data in a graph database and a search engine, ensure the accuracy and timeliness of address information, and effectively improve the user experience.

[0005] To solve the above technical problems, in a first aspect, the present invention discloses a method for storing address data based on a knowledge graph, the method comprising: Obtaining original address data; Based on a pre-established data processing model, determining standardized address data according to the original address data; Based on a pre-established hierarchical parsing model, determining multiple hierarchical address data according to the standardized address data; Based on a preset matching algorithm, storing the multiple hierarchical address data in a matching graph database; Based on a pre-established data synchronization model, synchronously marking the multiple hierarchical address data in the graph database in a search engine.

[0006] As a preferred embodiment, the step of determining standardized address data according to the original address data based on the pre-established data processing model includes: Determine target address data based on a preset ratio threshold of the original address data and the address data in the standard address library; Complete the target address data according to the address data in the standard address library to determine standardized address data.

[0007] As a preferred implementation manner, the step of determining standardized address data based on a pre-established data processing model according to the original address data includes: Based on a pre-trained natural language processing model, obtain several similar address data related to the semantics according to the original address data, and determine standardized address data.

[0008] As a preferred implementation manner, the step of storing the multiple-level address data in a matching graph database based on a preset matching algorithm includes: Create nodes in the graph database; Store each of the level address data in the graph database in the form of nodes; Establish an edge relationship between any two adjacent level address data; Generate an index of the nodes and the edge relationships.

[0009] As a preferred implementation manner, in the step of synchronously marking the multiple-level address data in the graph database on a search engine based on a pre-established data synchronization model, it includes: Identify information changes of the level address data; Mark the information changes of the level address data; Trigger the synchronization mechanism preset in the data synchronization model; Synchronously mark the information changes of the level address data on the search engine.

[0010] As a preferred implementation manner, the step of determining standardized address data based on a pre-established data processing model according to the original address data includes: Perform natural language parsing on the original address data based on the data processing model, and determine whether the original address data includes clear address information. If the determination is yes, match it with the address data in the standard address library to determine standardized address data; if the determination is no, terminate the natural language parsing and report an error.

[0011] The second aspect of the present invention discloses an address data storage device based on a knowledge graph, and the device includes: An acquisition module for acquiring original address data; An output module, configured to determine standardized address data based on a pre-established data processing model according to the original address data; A splitting module, configured to determine a plurality of hierarchical address data based on a pre-established hierarchical parsing model according to the standardized address data; A storage module, configured to store the plurality of hierarchical address data in a matching graph database based on a preset matching algorithm; A synchronization module, configured to synchronously mark the plurality of hierarchical address data in the graph database on a search engine based on a pre-established data synchronization model.

[0012] A third aspect of the present invention discloses another address data storage device based on a knowledge graph, the device comprising: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the steps of the address data storage method based on a knowledge graph disclosed in the first aspect of the present invention.

[0013] A fourth aspect of the present invention discloses a computer storage medium, the computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the address data storage device based on a knowledge graph disclosed in the first aspect of the present invention.

[0014] Compared with the prior art, the present invention has the following beneficial effects: In an embodiment of the present invention, original address data is obtained; based on a pre-established data processing model, standardized address data is determined according to the original address data; based on a pre-established hierarchical parsing model, a plurality of hierarchical address data is determined according to the standardized address data; based on a preset matching algorithm, the plurality of hierarchical address data is stored in a matching graph database; based on a pre-established data synchronization model, the plurality of hierarchical address data in the graph database is synchronously marked on a search engine. It can be seen that implementing the present invention can synchronously change the hierarchical address data in the graph database and the search engine, ensure the accuracy and timeliness of the address information, and effectively improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a schematic flowchart of an address data storage method based on a knowledge graph disclosed in an embodiment of the present invention; Figure 2 is a schematic diagram of nodes and variable relationships of an address data storage method based on a knowledge graph disclosed in an embodiment of the present invention; Figure 3Schematic diagram of the structure of an address data storage device based on a knowledge graph disclosed in an embodiment of the present invention; Figure 4 Schematic diagram of the structure of another address data storage device based on a knowledge graph disclosed in an embodiment of the present invention. Detailed implementation manners

[0016] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.

[0018] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in conjunction with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments.

[0019] Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0020] The present invention discloses an address data storage method and device based on a knowledge graph, which can synchronously change hierarchical address data in a graph database and a search engine, ensure the accuracy and timeliness of address information, and effectively improve the user experience.

[0021] Embodiment 1 Please refer to Figure 1 , Figure 1 which is a flowchart of an address data storage method based on a knowledge graph disclosed in an embodiment of the present invention. Among them, Figure 1The described method for storing address data based on a knowledge graph can be applied to a device for storing address data based on a knowledge graph. Among them, the storage of address data based on a knowledge graph can be integrated in a local server or a cloud server, which is not limited in the embodiments of the present invention. As Figure 1 shown, the method for storing address data based on a knowledge graph may include the following operations: 101. Obtain the original address data.

[0022] In the embodiments of the present invention, the original address data may be in Chinese text, or in English text, Japanese text, etc. The language type is not limited here. For the convenience of understanding, Chinese text is taken as an example below.

[0023] 102. Based on a pre-established data processing model, determine the standardized address data according to the original address data.

[0024] Specifically, it includes: determining the target address data based on a preset ratio threshold between the original address data and the address data in the standard address library; Due to differences in databases, data formats, and naming methods, the accuracy of address data parsing will be severely affected. By processing the original address data through the data processing model and screening according to the preset ratio threshold, redundant information can be removed, the noise in the original address data can be filtered, and the target address data can be determined.

[0025] Complete the target address data according to the address data in the standard address library to determine the standardized address data.

[0026] Some data in the target address data may be missing. In some necessary cases, the data processing model in this embodiment can complete it to determine the standardized address data.

[0027] As a preferred implementation, the embodiment of the present invention is based on a data processing model and can perform natural language analysis on the original address data to determine whether the original address data includes clear address information. If it is determined to be yes, it is matched with the address data in the standard address library to determine the standardized address data; if it is determined to be no, the natural language analysis is terminated and an error is reported. The natural language analysis model can convert the original address data into query instructions that the system can understand. The standard input format of the query command mainly includes four parts: center point, location word constraint, distance word constraint, and POI function constraint. Specifically, the center point extracted by the natural language parsing model supports most standard address inputs, such as "No. 22 Qiaotou Street, Garden Square, Inner Ring East Road" and other words; the location constraint word supports "northeast, northwest, southeast, southwest, east, south, west, north, nearby" and other words, and other words such as "around, around, surrounding" will be uniformly converted to "nearby"; the distance constraint word requires the input number to be an integer, such as "5 kilometers / km / KM, 50 meters / m" and other words; the POI function constraint extracted by the natural language parsing model supports most function constraint words, such as "hotel, restaurant, shopping center" and other words, and does not support the single input of multiple function constraint words. For example, when the original address data obtained is "restaurant 500 meters south of Hongtai Smart Valley", the natural language parsing model splits it, and the standard output of the center point, location constraint word, distance constraint word and POI function constraint after the split is "Hongtai Smart Valley", "South", "500 meters", "restaurant" in order. It can be understood that among the above four components, except for the center point which must have content, the remaining components are allowed to be empty.

[0028] It can be understood that this step also includes: based on the pre-trained natural language processing model, according to the original address data, several semantically related similar address data are obtained to determine the standardized address data; since the original address data in this embodiment are all Chinese texts as an example, the semantically related similar address data can be understood as typos, or homophones or similar characters caused by spelling errors. The pre-trained natural language processing model can enhance the compatibility with diversified address expressions during data storage and parsing, improve fault tolerance by analyzing semantic similarity, and ensure efficient parsing and consistent hierarchical structure under different input formats.

[0029] 103. Based on a pre-established hierarchical parsing model, multiple hierarchical address data are determined according to the standardized address data.

[0030] As an optional implementation, the hierarchical parsing model in this embodiment is the MGeo model. The MGeo model can parse the standardized address data, decompose the geographical elements in the standardized address data into hierarchical fields such as province, city, district, street, etc., and identify different hierarchies. The parsing of the MGeo model can be expressed as:

[0031] Among them, F is the parsing function of the MGeo model, which is used to extract address hierarchy information.

[0032] For example, when the standardized address data is "No. 200, Science and Technology Park, Nanshan District, Shenzhen City, Guangdong Province", the hierarchical address data after parsing by the MGeo model are respectively "Province: Guangdong Province", "City: Shenzhen City", "District: Nanshan District", "Street: Science and Technology Park", "House Number: No. 200".

[0033] 104. Based on a preset matching algorithm, store multiple hierarchical address data in a matching graph database.

[0034] Specifically, it includes: creating nodes in the graph database; Store each hierarchical address data in the graph database in the form of nodes; Establish the edge relationship between any two adjacent hierarchical address data; Generate an index for the node and edge relationship.

[0035] In the embodiment of the present invention, the graph database is a Neo4j graph database. The Neo4j graph database stores each hierarchical address data in the form of nodes and establishes the relationship between address elements through edges. For example, the nested hierarchical structure between province - city and district - street is realized through the hierarchical edge relationship, as shown in the schematic diagram of nodes and variable relationships. In addition, after the hierarchical address data is stored in the graph database, necessary indexes will also be generated to accelerate the query of hierarchical address data and ensure that the graph can respond quickly when querying hierarchical address data. Figure 2 As shown in the schematic diagram of nodes and variable relationships. In addition, after the hierarchical address data is stored in the graph database, necessary indexes will also be generated to accelerate the query of hierarchical address data and ensure that the graph can respond quickly when querying hierarchical address data.

[0036] As a preferred implementation manner, the Neo4j graph database has an index strategy built - in for optimizing query performance. In order to achieve address matching, considering various situations such as address text prefix, pinyin prefix, full text, full pinyin, first - letter prefix of address pinyin, full spelling of first - letter pinyin, synonym conversion, and custom word segmentation, a custom word library and a synonym dictionary are constructed, and indexes such as address slices and pinyin are stored.

[0037] 105. Based on a pre - established data synchronization model, synchronously mark multiple hierarchical address data in the graph database on the search engine.

[0038] Specifically, it includes: identifying the information change of the hierarchical address data; Marking the information change of the hierarchical address data; Triggering the synchronization mechanism preset in the data synchronization model; Synchronously marking the information change of the hierarchical address data on the search engine.

[0039] In the embodiment of the present invention, the search engine is ElasticSearch, and the data synchronization model can establish a data synchronization mechanism between the Neo4j graph database and ElasticSearch. When new, modified, or deleted hierarchical address data occurs in the Neo4j graph database, the data synchronization model automatically synchronizes these changes to ElasticSearch. The data synchronization model combines batch synchronization and real-time updates to ensure that the hierarchical address data remains consistent between the graph database and the search engine, meeting the requirements of efficient query and hierarchical retrieval. The address data storage method of this embodiment combines the distributed architecture design of the Neo4j graph database and ElasticSearch, ensuring both the efficiency of complex address queries and the ability to support rapid retrieval of large-scale data. In addition, the high integration and flexibility of the MGeo model enable the system to have good scalability, and new functional modules can be added according to actual needs to adapt to future development requirements.

[0040] Embodiment 2 Please refer to Figure 3 , Figure 3 , which is a schematic structural diagram of an address data storage device based on a knowledge graph disclosed in the embodiment of the present invention. As Figure 3 shown, the address data storage device based on the knowledge graph may include: An acquisition module 201, configured to acquire original address data; An output module 202, configured to determine standardized address data based on the original address data according to a pre-established data processing model; A splitting module 203, configured to determine multiple hierarchical address data based on the standardized address data according to a pre-established hierarchical parsing model; A storage module 204, configured to store multiple hierarchical address data in a matching graph database based on a preset matching algorithm; A synchronization module 205, configured to synchronize and mark multiple hierarchical address data in the graph database to a search engine based on a pre-established data synchronization model.

[0041] It can be seen that the device described in the embodiment Figure 3 can synchronize changes to hierarchical address data in the graph database and the search engine, ensuring the accuracy and timeliness of address information and effectively improving the user experience.

[0042] Specifically, for the details and technical advantages of the module function steps in the above device, reference can be made to the corresponding descriptions in Embodiment 1, and details will not be repeated in this embodiment.

[0043] Embodiment 3 Please refer to Figure 4 , Figure 4The following is a schematic structural diagram of another address data storage device based on a knowledge graph disclosed in an embodiment of the present invention. As Figure 4 shown, the address data storage device based on a knowledge graph may include: A memory 301 storing executable program code; A processor 302 coupled to the memory 301; The processor calls the executable program code stored in the memory and executes the steps of the address data storage method based on a knowledge graph described in Embodiment 1 of the present invention.

[0044] Embodiment 4 An embodiment of the present invention discloses a computer storage medium. The computer storage medium stores computer instructions, which are used to execute the steps of the address data storage method based on a knowledge graph described in Embodiment 1 of the present invention when called.

[0045] Embodiment 5 An embodiment of the present invention discloses a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the address data storage method based on a knowledge graph described in Embodiment 1.

[0046] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0047] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each implementation manner can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium that can be used to carry or store data. Finally, it should be noted that: What is disclosed in an address data storage method and device based on a knowledge graph according to an embodiment of the present invention is only a preferred embodiment of the present invention, and is only used to illustrate the technical solution of the present invention, rather than to limit it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. As described above, it is only a preferred embodiment of the present invention, and there is no restriction on the present invention in any form. Therefore, any modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A method for storing address data based on knowledge graph, characterized in that: The method comprises: Get the original address data; Based on a pre-established data processing model, determining standardized address data according to the original address data; Based on a pre-established hierarchical parsing model, a plurality of hierarchical address data are determined according to the standardized address data; Based on a preset matching algorithm, storing the plurality of hierarchical address data in a matching graph database; Based on a pre-established data synchronization model, the multiple levels of address data in the graph database are synchronously marked on the search engine.

2. The address data storage method according to claim 1, characterized in that: The step of determining the standardized address data based on the original address data based on the pre-established data processing model includes: Determining target address data based on a preset ratio threshold between the original address data and address data in a standard address library; The target address data is completed according to the address data in the standard address library to determine the standardized address data.

3. The address data storage method according to claim 1, characterized in that: The step of determining the standardized address data based on the original address data based on the pre-established data processing model includes: Based on the pre-trained natural language processing model, several semantically related similar address data are obtained according to the original address data, and the standardized address data is determined.

4. The address data storage method according to claim 1, characterized in that: The step of storing the plurality of hierarchical address data in a matching graph database based on a preset matching algorithm includes: Create nodes in the graph database; Storing each of the hierarchical address data in the form of a node in the graph database; Establishing an edge relationship between any two adjacent hierarchical address data; Generate an index of the relationship between the node and the edge.

5. The address data storage method according to claim 4, characterized in that: The step of synchronously marking the multiple levels of address data in the graph database on the search engine based on the pre-established data synchronization model includes: Identifying information changes to the hierarchical address data; Marking information changes of the hierarchical address data; Triggering the synchronization mechanism preset by the data synchronization model; The information changes of the hierarchical address data are synchronously marked on the search engine.

6. The address data storage method according to claim 1, characterized in that: The step of determining the standardized address data based on the original address data based on the pre-established data processing model includes: Based on the data processing model, the original address data is parsed in natural language to determine whether the original address data includes clear address information. If it is, it is matched with the address data in the standard address library to determine the standardized address data; if it is not, the natural language parsing is terminated and an error is reported.

7. An address data storage device based on knowledge graph, characterized in that: The device comprises: An acquisition module, used to obtain original address data; An output module, used to determine standardized address data according to the original address data based on a pre-established data processing model; A splitting module, for determining a plurality of hierarchical address data according to the standardized address data based on a pre-established hierarchical parsing model; A storage module, used for storing the plurality of hierarchical address data in a matching graph database based on a preset matching algorithm; The synchronization module is used to synchronize and mark the multiple hierarchical address data in the graph database with the search engine based on a pre-established data synchronization model.

8. An address data storage device based on knowledge graph, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the steps of the address data storage method based on the knowledge graph as described in any one of claims 1 to 5.

9. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, which, when called, are used to execute the address data storage method based on the knowledge graph as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Map model-based place name address matching query method and system

    CN110147420A

  • Search system and method and storage medium

    CN111897836A

  • Search engine method and system based on knowledge graph, storage medium and electronic equipment

    CN116108194A

  • Method for constructing knowledge graph based on digital retina system

    CN116701663A

  • Knowledge graph construction method, electronic equipment and computer readable storage medium

    CN118504668A