Data import method, system and equipment of Neo4j database and medium
By combining graphical user interface configuration and LOAD commands with APOC components, the problem of inefficient import of large amounts of data into the Neo4j database is solved, support for multiple data source formats and efficient import are achieved, service interruptions are avoided, and visual data display is provided.
Patent Information
- Application Number
- CN202510925894.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
AI Technical Summary
The data import method of the Neo4j database has problems such as low efficiency and frequent service interruptions that affect service quality. The efficiency is extremely low when importing large amounts of data. In addition, the existing method has a single data source and cannot meet the import requirements of multiple data formats.
Use a graphical user interface to configure the data source and target database, combine the LOAD command and APOC components, import data into the Neo4j database through preset import strategies, support multiple data source formats, and merge entity relationships through batch inserts and APOC components to achieve data import.
It improves the data import speed, supports multiple data source formats, avoids Neo4j service downtime, and achieves efficient data import and visualization.
Smart Images

Figure CN120763236A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of database import, and in particular relates to a method, system, device and medium for importing data into a Neo4j database. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] The graph database Neo4j has received increasing attention due to its advantages such as embeddedness, high performance, and lightweightness. However, the speed and applicability of its own data import method are poor. There are many ways to import data into the Neo4j graph database. Commonly used methods include CREATE statement, LOAD CSV statement, Batch Inserter, Batch Import, and Neo4j-import. The defects of the above methods are: for CREATE statement, it is generally only applicable to importing less than 10,000 data items, and the import speed is extremely slow; LOAD CSV statement is generally applicable to importing less than 100,000 data items, and the import speed is at an intermediate level; Batch Inserter has file format requirements and can only import CSV (Comma-Separated Values) files, and the Neo4j service must be stopped during import; Batch Import method also requires stopping the Neo4j service; Neo4j-improt method must generate a new database and can only be used when initializing the database, and cannot be operated on an existing database.
[0004] At the same time, Neo4j has high requirements for the quantity and quality of data, which is directly related to the effect of data map display. After the Neo4j database is initialized at the beginning of the project, it becomes more difficult to import large amounts of data. The import methods of CREATE statements and LOAD CSV statements are not suitable for importing large amounts of data. When the data volume is in the millions, these two import methods take several days or even dozens of days, which is extremely inefficient. The Batch Inserter and BatchImport methods require frequent suspension of the Neo4j service, which will seriously affect the service quality. Summary of the Invention
[0005] The embodiments of the present invention provide a method, system, device, and medium for importing data into a Neo4j database to solve the problems of a single data source in traditional solutions, the low efficiency of importing large amounts of data, and the problem that frequent Neo4j service shutdowns seriously affect service quality.
[0006] According to a first aspect of an embodiment of the present invention, a data import system for a Neo4j database is provided, comprising: A data import task configuration unit, which is used to configure the data source and the Neo4j target database based on a graphical user interface; and upload the data in the configured data source to the server of the Neo4j target database; wherein the data source includes a database and a file; A data import task execution unit is used to display the entities and relationships of the data in the Neo4j target database in an interface; based on a mouse action event on the interface, an event handling menu corresponding to the mouse action event pops up; based on the event handling menu, a storage path of the current entity or relationship in the server is configured; based on the data in the storage path, the data is imported into the current entity or relationship through a preset import strategy to realize data import into the Neo4j database; wherein, the preset import strategy loads data based on the LOAD command and combines the APOC component to merge entities and relationships.
[0007] Furthermore, the data in the configured data source is uploaded to the server of the Neo4j target database. Specifically, for relational database type data sources, entity and relationship data are collected through preset SQL statements, and different categories of entities and relationships are divided into different data sets; for file type data sources, unified storage is performed.
[0008] Furthermore, the data is imported into the current entity or relationship through a preset import strategy, specifically: data loading based on LAOD commands; importing data into the Neo4j database through batch inserts; merging entities and relationships of data based on APOC components.
[0009] Furthermore, the interface display of entities and relationships of data in the Neo4j target database specifically includes list mode and canvas mode, wherein the list mode displays entities and relationships separately, and each entity and relationship is editable; the canvas mode displays entities and relationships in an integrated manner, with entities as nodes and relationships as edges connecting two entities, and each entity and node is editable; wherein the editable settings include adding, deleting, modifying and importing entities or relationships.
[0010] Furthermore, based on the mouse action event on the interface, an event handling menu corresponding to the mouse action event pops up, wherein, when importing an entity or relationship, the event handling menu includes the entity or relationship name, target database, import method, data set to be imported, and configuration of entity or relationship attribute information; when adding a new entity or relationship, the event handling menu includes the entity or relationship name, entity identifier, entity or relationship data set storage path, and configuration of entity or relationship attribute information.
[0011] Furthermore, it supports simultaneous configuration of multiple entity and relationship data sources and synchronous import of data.
[0012] Furthermore, the data source includes but is not limited to a relational database, a CSV file, and an Excel file.
[0013] According to a second aspect of an embodiment of the present invention, a method for importing data into a Neo4j database is provided. The method is based on the aforementioned Neo4j database data import system and includes: Configure the data source and Neo4j target database based on a graphical user interface; wherein the data source includes databases and files; Upload the data in the configured data source to the server where the Neo4j target database belongs; Display entities and relationships of data in the Neo4j target database; Based on the mouse action event on the interface, an event handling menu corresponding to the mouse action event is popped up; Based on the event handling menu, the storage path of the current entity or relationship in the server is configured; based on the data under the storage path, the data is imported into the current entity or relationship through a preset import strategy to realize data import into the Neo4j database; wherein, the preset import strategy imports the Neo4j database based on the LOAD command, and merges entities and relationships in combination with the APOC component.
[0014] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored and running on the memory, wherein when the processor executes the program, the method for importing data into a Neo4j database is implemented.
[0015] According to a fourth aspect of an embodiment of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for importing data into a Neo4j database is implemented.
[0016] One or more of the above technical solutions have the following beneficial effects: (1) The present invention transfers the data source data to the server of the Neo4j database in advance and imports entities by combining the LOAD command with the APOC component, thereby effectively improving the data import speed without stopping the Neo4j service.
[0017] (2) Different from the traditional method and the method that supports a single CSV data source, the solution described in the present invention can support multiple data sources such as databases, CSV and Excel files.
[0018] (3) The solution described in the present invention can visualize the data in the Neo4j database through interface design, and realize the selective import of entities and relationships through editable settings of entities and relationships in the interface, thereby meeting the targeted needs of users for data import.
[0019] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0021] Figure 1 This is a schematic diagram of the data source configuration interface according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the list mode according to an embodiment of the present invention; Figure 3 A schematic diagram of the canvas mode according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a newly added entity according to an embodiment of the present invention; Figure 5 A schematic diagram for creating a target database according to an embodiment of the present invention; Figure 6 This is a schematic diagram of an import entity according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the import relationship according to an embodiment of the present invention; Figure 8 This is a flow chart of a method for importing data into a Neo4j database according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0023] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.
[0024] Explanation of terms: Neo4j database: is a popular graph database. In this embodiment, the community edition is used. The community edition is open source and free.
[0025] MySQL database: an open source relational database management system.
[0026] APOC (Awesome Procedures on Cypher) is an official Neo4j component that provides a series of query, statistics, and modification methods. This embodiment utilizes the method used in this component to merge data based on specified attributes (e.g., business ID), ensuring data security and integrity. Furthermore, the APOC component is added when the Neo4j database is started, allowing the component library to be called at any time without restarting the Neo4j service.
[0027] SQL statement: Structured Query Language (SQL statement for short) is a database query and programming statement used to access data and query, update and manage relational databases.
[0028] CSV file: The full name is Comma Separated Values, which is a tabular data stored in plain text format, using commas as delimiters between fields.
[0029] LAOD command: A command used to load data from a file.
[0030] In one or more implementations, the solution described in this embodiment provides a data import system for a Neo4j database, including: A data import task configuration unit, which is used to configure the data source and Neo4j target database based on a graphical user interface; and upload the data in the configured data source to the server of the Neo4j target database; wherein the data source includes databases and files, specifically relational databases, CSV files, and Excel files; In specific implementation, data import configuration specifically includes: Data source configuration includes data source configuration of databases such as MySQL and file import. After processing, the two methods obtain the data set, save it, and choose to use it when importing data.
[0031] like Figure 1 As shown, the data source configuration method is: you need to configure the address and account password of a database such as MySQL. After the test link is successful, you need to write SQL statements to collect the data set to be imported into Neo4j.
[0032] Wherein, the data source is a database, and a data collection needs to be created. One database can correspond to numerous data collections. Each data collection needs to write a corresponding SQL statement, that is, one data collection corresponds to one SQL statement.
[0033] Specifically, the SQL statement corresponding to one data collection is fixed, but multiple data collections can be created according to different requirements. One data source can have multiple data collections. Regardless of the data source, the data set obtained ultimately needs to be stored in the physical disk of the server where the Neo4j target database is located. The data set file is stored in a CSV file. When importing data, the link of the Neo4j database is obtained, the data set file is read, and the data is imported into the Neo4j database.
[0034] As shown in Figure 2 , the file import method: supports the import of CSV and Excel files. Upload the specified format document to the server. Like the SQL written by the data source configuration, the data set to be imported can be viewed after successful import. In specific implementation, any data source ultimately obtains a normalized CSV file, and regardless of the data source, normalization processing is required. The CSV file also needs to be normalized.
[0035] The normalization processing specifically includes: reading data, creating a CSV file, and then writing data into the CSV file. When writing, the data needs to be operated accordingly, including processing of special characters not accepted by Neo4j, processing of English commas, and the like. The result set after processing needs to ensure that there are no special characters not supported by Neo4j, so as to ensure that the LOAD command can be successfully loaded.
[0036] In specific implementation, data import target maintenance supports configuring multiple Neo4j, which can be selected according to requirements. Specifically, the data import target database can maintain multiple, such as a first target database neo4j1 and a second target database neo4j2. As shown in Figure 5 , a neo4j target database is created. When importing data, the data is imported into which neo4j target database is selected in the target database column (such as neo4j1 or neo4j2 or test neo4j), as shown in the target database in Figure 6 and Figure 7 .
[0037] A data import task execution unit is used to display the entities and relationships of the data in the Neo4j target database on the interface; based on the mouse action event on the interface, an event handling menu corresponding to the mouse action event is popped up; based on the event handling menu, the storage path of the current entity or relationship in the server is configured; based on the data under the storage path, the data is imported into the current entity or relationship through a preset import strategy to realize data import into the Neo4j database; wherein, the preset import strategy imports the Neo4j database based on the LOAD command and combines the APOC component to merge the entities and relationships.
[0038] In the specific implementation, the data is imported into the current entity or relationship through the preset import strategy, specifically: data loading based on the LAOD command; importing the data into the Neo4j database through batch insertion; merging entities and relationships of the data based on the APOC component.
[0039] It should be noted that the LOAD command mentioned in this example refers to the LOAD CSV command, which is a built-in import command for Neo4j. Its import process is as follows: read a CSV file; load the data in the file; and, by default, verify it against the existing database data every 300 records. If there are duplicates, the data is modified; if there are no duplicates, a new one is added. The main reason for the slow import process using the LOAD CSV command is data verification. The larger the existing database data volume, the slower the import process (for example, if the data volume exceeds 1 million, importing 100 records can take over 15 minutes). However, this method of data loading is very fast.
[0040] To address the above issues, the solution described in this embodiment uses only the LOAD CSV command to load data during data import. Data verification and merging are handled using the APOC component, which has a built-in optimized data merging method and extremely fast processing speed. Specifically, the solution described in this embodiment divides data import into two parts: data source maintenance and data import. If the data source is a database, SQL statements must be written. The solution described in this embodiment obtains data results based on the SQL statements and processes the data results into a format acceptable to the LOAD command, namely, the CSV file format. If the data source is a file, the file is read, the data is loaded, and then the data is processed into a format acceptable to the LOAD command, namely, a CSV file. Therefore, the data set selected during data import is in a format that the LOAD command can normally operate on.
[0041] As described above, the specific import process of the solution described in this embodiment is as follows: the first step is to maintain the data source, process data in other formats into a format acceptable to the solution, and save the data set; the second step is to read the file in LOAD mode and load the data; the third step is to fill the data into the database using the batch insert method accepted by Neo4j, which is extremely fast; the fourth step is to merge the data using the merge method in the APOC component.
[0042] In the specific implementation, the data in the configured data source is uploaded to the server of the Neo4j target database. Specifically, for relational database type data sources, entity and relationship data are collected through preset SQL statements, and different categories of entities and relationships are divided into different data sets; for file type data sources, unified storage is performed.
[0043] In a specific implementation, the interface display of entities and relationships of data in the Neo4j target database specifically includes list mode and canvas mode, wherein the list mode displays entities and relationships separately, and each entity and relationship is editable; the canvas mode displays entities and relationships in an integrated manner, with entities as nodes and relationships as edges connecting two entities, and each entity and node is editable; wherein the editable settings include adding, deleting, modifying and importing entities or relationships.
[0044] In a specific implementation, based on the mouse action event on the interface, an event handling menu corresponding to the mouse action event pops up, wherein, when importing an entity or relationship, the event handling menu includes the entity or relationship name, target database, import method, data set to be imported, and configuration of entity or relationship attribute information; when adding a new entity or relationship, the event handling menu includes the entity or relationship name, entity identifier, entity or relationship data set storage path, and configuration of entity or relationship attribute information.
[0045] In a specific implementation, the system supports simultaneous configuration of multiple entity and relationship data sources and then synchronous import of data.
[0046] Specifically, for ease of understanding, the solution described in this embodiment is described in detail below with reference to specific examples: During data import, first select an existing Neo4j target database to display the database's NodeLabels (i.e., entities) and Relationship Types (i.e., relationships). You can choose between list or canvas modes to intuitively display the graph database's existing entities and relationships, supporting maintenance of entities and relationships in both display modes.
[0047] Specifically, such asFigure 2 and Figure 3 As shown, the canvas mode is the same as the list mode, and the function menu is bound to the right mouse button. Right-clicking a blank area of the graph will pop up a menu for adding relationships or entities. Right-clicking on nodes and lines allows you to edit, delete, and import. The following describes how to add entities, import entities, and import relationships: like Figure 4 As shown in the figure, the newly added entities are as follows: In canvas mode, right-click on a blank area of the graph to pop up a menu for adding a new relationship or entity. Configure the entity on the menu to create a new entity. In list mode, click the Add Entity interface button to pop up the Add Relationship or Entity menu. Configure the entity on the menu to create a new entity. like Figure 6 As shown, the entity import is as follows: Select the entity to import. All of the entity's attributes will be displayed by default. Select the data source. If you select a database, you'll need to select the collected dataset. If you select the file import method, you'll need to select the corresponding file. Then, map the entity's attributes to the imported dataset fields and select the key attributes for the merged data.
[0048] Taking canvas mode as an example, right-click the pharmaceutical factory, select Import, and fill in the import configuration. Select the target database in the data source configuration maintenance mentioned above. There are two import methods: Database and File. Both methods require selecting a dataset.
[0049] In the entity attribute configuration, the attribute name is the attribute of the pharmaceutical factory entity in the graph database, and the corresponding field is the table header in the dataset. After the one-to-one correspondence is completed, the migration is executed to migrate the data in the dataset to the neo4j database.
[0050] like Figure 7 As shown, the relationship import is as follows: Select the relationship to import. By default, all the relationship data and the relationship's start and end entities will be displayed. Select the data source. If you select Database, you need to select the collected dataset; if you select File as the import method, you need to select the corresponding file. Then, map the relationship attributes to the imported database fields, and match the key fields of the start and end entities. You also need to select the key attributes required to integrate and merge the relationship data.
[0051] Import a relationship: Right-click on the canvas to create a connection. The relationship name, target database, start entity, and end entity are not editable. Select the import method and the dataset to import.
[0052] The relationship attribute configuration is consistent with the above entity attribute configuration.
[0053] Furthermore, the solution described in this embodiment supports one-click data import after configuring multiple entities and relationships at the same time. It also supports one-click sorting of imported data, that is, merging identical data based on key attributes.
[0054] Specifically, such as Figure 2 As shown, multiple entities (such as diseases, representations) can be configured. Click the Migrate button to migrate all configured entities (diseases, representations) at the same time. Among them, the sorting function is a step that must be performed after importing data, which is used to merge the same data.
[0055] In specific implementation, the solution described in this embodiment connects to Neo4j via the LOAD CSV method. By adjusting the import logic and statements, the import logic is aligned with the official large-scale import logic, enabling rapid data import into the graph database. The official APOC component is introduced, and its merging of Node Labels (entities) and Relationships (relationships) method is used to ensure data integrity. This allows the solution described in this embodiment to import data without stopping the Neo4j service, achieving import speeds far exceeding conventional import methods.
[0056] In one or more embodiments, Figure 8 As shown, this embodiment corresponds to the above-mentioned Neo4j database data import system and provides a Neo4j database data import method, including: Configure data sources and Neo4j target databases based on a graphical user interface; Upload the data in the configured data source to the server where the Neo4j target database belongs; Display entities and relationships of data in the Neo4j target database; Based on the mouse action event on the interface, an event handling menu corresponding to the mouse action event is popped up; Based on the event handling menu, the storage path of the current entity or relationship in the server is configured; based on the data under the storage path, the data is imported into the current entity or relationship through a preset import strategy to realize data import into the Neo4j database; wherein, the preset import strategy imports the Neo4j database based on the LOAD command, and merges entities and relationships in combination with the APOC component.
[0057] In further embodiments, there is also provided: An electronic device includes a memory, a processor, and a computer program stored and running on the memory, wherein the processor implements the open vocabulary object detection method when executing the program.
[0058] A non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements the open vocabulary object detection method.
[0059] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. Neo4j database data import system, characterized by: include: Data import task configuration unit, which is used to configure the data source and Neo4j target database based on the graphical user interface; And, uploading the data in the configured data source to the server of the Neo4j target database; wherein the data source includes databases and files; A data import task execution unit is used to display the entities and relationships of the data in the Neo4j target database in an interface; based on a mouse action event on the interface, an event handling menu corresponding to the mouse action event pops up; based on the event handling menu, a storage path of the current entity or relationship in the server is configured; based on the data in the storage path, the data is imported into the current entity or relationship through a preset import strategy to realize data import into the Neo4j database; wherein, the preset import strategy loads data based on the LOAD command and combines the APOC component to merge entities and relationships.
2. The data import system for Neo4j database according to claim 1, characterized in that: The data in the configured data source is uploaded to the server of the Neo4j target database. Specifically, for relational database type data sources, entity and relationship data are collected through preset SQL statements, and different types of entities and relationships are divided into different data sets; for file type data sources, unified storage is performed.
3. The data import system for Neo4j database according to claim 1, characterized in that: The method of importing data into the current entity or relationship through the preset import strategy is as follows: the method of importing data into the current entity or relationship through the preset import strategy is as follows: loading data based on the LAOD command; importing data into the Neo4j database through batch insertion; merging entities and relationships of data based on the APOC component.
4. The data import system for Neo4j database according to claim 1, characterized in that: The interface display of entities and relationships of data in the Neo4j target database specifically includes list mode and canvas mode, wherein the list mode displays entities and relationships separately, and each entity and relationship is editable; the canvas mode displays entities and relationships in an integrated manner, using entities as nodes and relationships as edges connecting two entities, and each entity and node is editable; wherein the editable settings include adding, deleting, modifying and importing entities or relationships.
5. The data import system for Neo4j database according to claim 1, characterized in that: Based on the mouse action event on the interface, an event handling menu corresponding to the mouse action event pops up, wherein, when importing an entity or relationship, the event handling menu includes the entity or relationship name, target database, import method, data set to be imported, and configuration of entity or relationship attribute information; when adding a new entity or relationship, the event handling menu includes the entity or relationship name, entity identifier, entity or relationship data set storage path, and configuration of entity or relationship attribute information.
6. The data import system for Neo4j database according to claim 1, characterized in that: Supports simultaneous configuration of multiple entity and relationship data sources and synchronous import of data.
7. The data import system for Neo4j database according to claim 1, characterized in that: The data sources include but are not limited to relational databases, CSV files, and Excel files.
8. A method for importing data into a Neo4j database, based on the Neo4j database data import system according to any one of claims 1 to 7, characterized in that: include: Configure the data source and Neo4j target database based on a graphical user interface; wherein the data source includes databases and files; Upload the data in the configured data source to the server where the Neo4j target database belongs; Display entities and relationships of data in the Neo4j target database; Based on the mouse action event on the interface, an event handling menu corresponding to the mouse action event is popped up; Based on the event handling menu, the storage path of the current entity or relationship in the server is configured; based on the data under the storage path, the data is imported into the current entity or relationship through a preset import strategy to realize data import into the Neo4j database; wherein, the preset import strategy loads data based on the LOAD command and combines the APOC component to merge entities and relationships.
9. An electronic device comprising a memory, a processor, and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the method for importing data into the Neo4j database as claimed in claim 8 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for importing data into a Neo4j database as claimed in claim 8 is implemented.
Citation Information
Patent Citations
Data migration method, system and device and computer readable storage medium
CN111324595A
Graph database Neo4J interactive visualization operation method and system
CN111813926A
Knowledge graph service processing method and device
CN112287114A
Whole-process visual configuration system and method for knowledge graph
CN114417018A
Power equipment fault tracing method and device
CN116993043A