Data importing method based on similarity calculation

By using similarity calculation technology in the data import method, combined with editing distance and word semantic similarity algorithm, automatic mapping of file table header fields and database fields is realized, solving the problems of inefficiency and high workload in the existing technology, and improving the accuracy and automation of imports.

CN119938754APending Publication Date: 2025-05-06AEROSPACE SCI & IND INTELLIGENT OPERATION RES & INFORMATION SECURITY RES INST (WUHAN) CO LTD

Patent Information

Application Number
CN202411956524.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-29
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing data import methods are inefficient and prone to errors. They need to maintain a large number of templates when processing multiple header information, which increases the workload.

Method used

The data import method based on similarity calculation is adopted, and the improved edit distance calculation similarity algorithm is combined with the word semantic similarity algorithm based on "Knownology Network" to realize the automatic mapping of file table header fields and database fields.

Benefits of technology

It effectively reduces the workload of manual configuration, improves the success rate and accuracy of matching, can automatically record user configuration content, and automatically load it the next time you import it, avoiding duplicate configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938754A_ABST
    Figure CN119938754A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of informatization management, and particularly relates to a data import method based on similarity calculation, automatic matching of file fields and database fields is realized by calculating the similarity of the file fields and the fields in a database, and the workload of manual configuration can be effectively reduced; a common editing distance calculation similarity algorithm is improved, the influence of the longest common subsequence on similarity calculation is considered, a similarity calculation formula is improved, and the matching success rate is increased; when the similarity is calculated, literal matching and semantic matching can be carried out by adopting a method of combining an improved editing distance similarity calculation algorithm and word semantic similarity calculation based on the Zhinetting, so that the matching success rate is effectively improved; besides automatic matching, a user configuration interface is provided, a user can manually adjust the mapping relation, and the flexibility and usability of the method are improved; user configuration content can be automatically recorded, and repeated configuration work of a user is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of information management, and in particular relates to a data import method based on similarity calculation. Background Art

[0002] With the rapid development of Internet technology, the degree of social informatization is constantly improving. Strengthening information management and promoting digital transformation have become an important part of improving the competitiveness of enterprises. The construction of an enterprise information system requires the batch import of massive excel tables, log files, database files, etc. stored into the database of the enterprise information system for unified storage, management and use. Currently, the commonly used file data import methods include manual entry and template import. Manual entry means manually arranging the data in the file and entering it into the system. This method is not only inefficient, wastes a lot of manpower, but also prone to errors; template import means first defining the header information template, and then importing the data according to the template. The disadvantage of this method is that when the header information of the file is diverse, a large number of templates need to be maintained, resulting in an increase in the workload of the user. In order to solve these problems, the present invention proposes a data import method based on similarity calculation, which can realize the automatic mapping of file header fields and database fields, and conveniently and quickly import file data into the information system, effectively reducing the workload of personnel configuration and maintenance. Summary of the invention

[0003] 1. Technical issues to be resolved

[0004] The technical problem to be solved by the present invention is: how to provide a data import method based on similarity calculation.

[0005] (II) Technical solution

[0006] In order to solve the above technical problems, the present invention provides a data import method based on similarity calculation, characterized in that the method comprises the following steps:

[0007] Step 1: Upload files to the system. Multiple file types are supported, including common table file types such as excel and txt, and common database file types such as dbf and json. Multiple files can be uploaded, but the types and headers of all files must be consistent.

[0008] Step 2: Automatically read the header field column of the file. For Excel type files, the system reads the header information of the first sheet page by default. If you need to read other sheet page information, you need to configure the sheet page;

[0009] Step 3: The user selects the file fields to be imported into the database;

[0010] Step 4: Determine whether the user has configured a matching relationship before based on the field selected by the user. If the user has configured it before, the previous configuration is automatically loaded; if not, proceed to the next similarity calculation;

[0011] Step 5: Before calculating the similarity, it is necessary to confirm that a database field information table has been established in the database. This table counts the field information of all tables in the database. The field information includes the field name (the database field name is generally in English), the field Chinese name, the table to which it belongs, and the field alias (other names that are synonymous with the English name of the field); after confirming that the field information table has been established, calculate the similarity between the file field and the field name, field Chinese name, and field alias in the field information table. The similarity calculation adopts a method that combines the improved edit distance algorithm for calculating similarity with the word semantic similarity algorithm based on HowNet;

[0012] Step 6: Automatically match the data tables and fields in the database based on the similarity calculation results. Multiple data tables can be matched, that is, the data in the file is imported into multiple data tables respectively. To avoid matching too many data tables, the ratio of the number of matched fields in the data table to the number of data columns in the file to be imported will be calculated here. Only data tables with a ratio greater than the set threshold will be matched;

[0013] Step 7: The user modifies the matching results. For fields that are not matched or that the user believes are not matched accurately, the user can manually select fields in the data table for matching;

[0014] Step 8: The system will automatically record the user configuration content. The next time you import the same field of the same file, the configuration content will be automatically loaded to avoid repeated configuration by the user.

[0015] Step 9: Import the data in the file into the database according to user configuration.

[0016] The similarity calculation method in step 5 is specifically as follows:

[0017] Step 51: First, the similarity between the file field name and the field name, field Chinese name, and field alias in the field information table is obtained by using an improved edit distance similarity calculation algorithm;

[0018] Edit distance refers to the minimum number of operations from one string to another. The smaller the edit distance, the more similar the two strings are, that is, the more similar the two words are. The operations include deletion, insertion, and replacement.

[0019] The calculation process of this method is as follows:

[0020] Assume that two strings consisting of multiple characters S = S1S2S3…S mand T = T1T2T3…T n , construct the matching relationship matrix LD[m+1,n+1] between S and T, and fill the matrix LD as follows:

[0021]

[0022] in:

[0023]

[0024] i, j represent the serial numbers of characters in the strings S and T respectively; m and n represent the number of characters in the strings S and T respectively;

[0025] The element d at the lower right corner of the matrix LD mn The value of is the edit distance between the two strings. The smaller the edit distance between the two strings, the higher the similarity between the two strings. The calculation formula for calculating the similarity Sim(S,T) using the edit distance is as follows:

[0026]

[0027] However, the above method only considers the number of edits, and does not consider the impact of the longest common subsequence on similarity; for the two strings "ID number" and "resident identity document number", the similarity calculated by this formula is sim(S,T)=1-4 / 8=0.5, which is not high, but in fact the two words have a high similarity; therefore, the impact of the longest common subsequence on similarity is considered here, and the algorithm for calculating similarity by edit distance is improved as follows:

[0028]

[0029] Wherein, LCS(S, T) is the longest common subsequence of two strings, and α is an adjustable parameter, which reflects the influence of the longest common subsequence on the similarity. The similarity of the above strings is calculated using the improved similarity calculation formula, where α is taken as 1, and the calculated similarity is sim(S, T) = 1-4 / 11 = 0.636. The improved similarity calculation algorithm can effectively improve the similarity of similar words.

[0030] When the similarity between the calculated file field name and the field name in the field information table or the field Chinese name or field alias is greater than the set threshold, the file field is considered to match the field in the database; it is worth noting that the field alias is another name that is synonymous with the English name of the field. Users can add multiple field aliases based on experience, which can effectively improve the success rate of matching;

[0031] Step 52: The algorithm for calculating similarity using the improved edit distance solves the problem of literal matching, that is, the two strings are literally the same or similar, including the text field column name is "mobile phone number", and the Chinese name of the field in the database is "mobile phone number"; but if the two strings are semantically similar words, including the text field column name is "mobile phone number", and the Chinese name of the field in the database is "telephone number", the algorithm for calculating similarity using the improved edit distance cannot achieve good results, and it is very likely that they will not match; therefore, when the improved edit distance similarity calculation algorithm fails to match the database field, the semantic similarity algorithm based on the HowNet word is used to calculate the similarity and further match; and when the database field has been matched by the improved edit distance similarity calculation algorithm, this step is no longer necessary; the semantic similarity calculation process based on the HowNet word is as follows:

[0032] In HowNet, two concepts, "sense" and "semene", are proposed. A sense is a description of the semantics of a word. Each word can be expressed as several senses, and a sense is composed of multiple basic semes. For two words ω1 and ω2, if ω1 has N senses s 11 ,s 12 ,…,s 1N , ω2 has M meanings s 21 ,s 22 ,…,s 2M , then the similarity between two words ω1 and ω2 is:

[0033]

[0034] Where sim(s 1I ,s 2J ) is the similarity of the meaning. Since the meaning is composed of four types of semantic primitives: the first basic semantic primitive, other basic semantic primitives, relational semantic primitives, and relational symbol description semantic primitives, the similarity of the two meanings s1 and s2 is:

[0035]

[0036] Among them, β i1 is an adjustable parameter, where 1≤i1≤4, and satisfies β1+β2+β3+β4=1, β1≥β2≥β3≥β4, sim j1 (p1, p2) represents the similarity between the sememe sets of the j1th type;

[0037] The sememe similarity is usually calculated by the semantic distance of the sememes, which is calculated as follows:

[0038]

[0039] Where dis(p1,p2) is the semantic distance between the semantic primitives p1 and p2 in the semantic primitive hierarchy tree. When the semantic primitives are in different trees, this value takes a larger constant, and α is an adjustable parameter.

[0040] The above formula can be used to calculate the semantic similarity between the file field name and the database field name or the field Chinese name or field alias. When the calculated semantic similarity is greater than the set threshold, the field column is considered to match the field in the database. This method can be used to solve the matching of semantically similar words.

[0041] (III) Beneficial effects

[0042] Compared with the prior art, the key innovations of the present invention are as follows:

[0043] (1) By calculating the similarity between the file fields and the fields in the database, the file fields and the database fields can be automatically matched, which can effectively reduce the workload of manual configuration;

[0044] (2) Improve the commonly used edit distance similarity calculation algorithm, consider the impact of the longest common subsequence on similarity calculation, improve the similarity calculation formula, and improve the matching success rate;

[0045] (3) When calculating similarity, the method of combining edit distance calculation with word semantic similarity calculation based on HowNet is adopted, which can perform literal matching and semantic matching, effectively improving the matching success rate;

[0046] (4) When calculating similarity, not only the similarity with the field name and the Chinese name of the field in the database is calculated, but also the similarity with the field alias. Users can configure multiple field aliases based on experience, which can further improve the matching success rate and matching accuracy;

[0047] (5) It can automatically record user configuration content, and the next time the same file field is imported, the previous configuration can be automatically loaded, avoiding users from repeating configuration work;

[0048] (6) It can import various types of files, including common table file types such as excel, txt, etc., and common database file types such as dbf, json, etc.;

[0049] (7) Multiple files can be uploaded and imported at one time;

[0050] (8) The file fields to be imported can be manually configured, which provides greater flexibility.

[0051] The advantages of the present invention are as follows:

[0052] (1) By calculating the similarity between the file fields and the fields in the database, the file fields and the database fields can be automatically matched, which can effectively reduce the workload of manual configuration;

[0053] (2) Improve the commonly used edit distance similarity calculation algorithm, consider the impact of the longest common subsequence on similarity calculation, improve the similarity calculation formula, and improve the matching success rate;

[0054] (3) When calculating similarity, the improved edit distance similarity calculation algorithm is combined with the word semantic similarity calculation based on HowNet, which can perform literal matching and semantic matching, effectively improving the matching success rate;

[0055] (4) In addition to automatic matching, a user configuration interface is provided so that users can manually adjust the mapping relationship, which improves the flexibility and usability of the method;

[0056] (5) It can automatically record user configuration content to avoid users from repeating configuration work. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is the overall flow chart of the present invention.

[0058] Figure 2 This is a flow chart of the similarity calculation method. DETAILED DESCRIPTION

[0059] In order to make the purpose, content, and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the accompanying drawings and examples.

[0060] The present invention proposes a data import method and system based on similarity calculation, which improves the commonly used edit distance similarity calculation algorithm and combines it with the word semantic calculation similarity algorithm based on HowNet, calculates the similarity between the file header field and the database field name, and then realizes automatic mapping of the file header field to the database field according to the similarity, which can effectively eliminate the manual configuration work of the user; at the same time, a user configuration interface is provided, and the user can modify the automatic matching result. After the user configuration is completed, the user configuration will be automatically recorded. When the same file field is imported next time, the previous configuration will be automatically loaded to avoid repeated configuration by the user. The method and system proposed by the present invention can conveniently and quickly import data from a large number of files into the database of the information system.

[0061] In order to solve the above technical problems, the present invention provides a data import method based on similarity calculation, characterized in that the method comprises the following steps:

[0062] Step 1: Upload files to the system. Multiple file types are supported, including common table file types such as excel and txt, and common database file types such as dbf and json. Multiple files can be uploaded, but the types and headers of all files must be consistent.

[0063] Step 2: Automatically read the header field column of the file. For Excel type files, the system reads the header information of the first sheet page by default. If you need to read other sheet page information, you need to configure the sheet page;

[0064] Step 3: The user selects the file fields to be imported into the database;

[0065] Step 4: Determine whether the user has configured a matching relationship before based on the field selected by the user. If the user has configured it before, the previous configuration is automatically loaded; if not, proceed to the next similarity calculation;

[0066] Step 5: Before calculating the similarity, it is necessary to confirm that a database field information table has been established in the database. This table counts the field information of all tables in the database. The field information includes the field name (the database field name is generally in English), the field Chinese name, the table to which it belongs, and the field alias (other names that are synonymous with the English name of the field); after confirming that the field information table has been established, calculate the similarity between the file field and the field name, field Chinese name, and field alias in the field information table. The similarity calculation adopts a method that combines the improved edit distance algorithm for calculating similarity with the word semantic similarity algorithm based on HowNet;

[0067] Step 6: Automatically match the data tables and fields in the database based on the similarity calculation results. Multiple data tables can be matched, that is, the data in the file is imported into multiple data tables respectively. To avoid matching too many data tables, the ratio of the number of matched fields in the data table to the number of data columns in the file to be imported will be calculated here. Only data tables with a ratio greater than the set threshold will be matched;

[0068] Step 7: The user modifies the matching results. For fields that are not matched or that the user believes are not matched accurately, the user can manually select fields in the data table for matching;

[0069] Step 8: The system will automatically record the user configuration content. The next time you import the same field of the same file, the configuration content will be automatically loaded to avoid repeated configuration by the user.

[0070] Step 9: Import the data in the file into the database according to user configuration.

[0071] The similarity calculation method in step 5 is specifically as follows:

[0072] Step 51: First, the similarity between the file field name and the field name, field Chinese name, and field alias in the field information table is obtained by using an improved edit distance similarity calculation algorithm;

[0073] Edit distance refers to the minimum number of operations from one string to another. The smaller the edit distance, the more similar the two strings are, that is, the more similar the two words are. The operations include deletion, insertion, and replacement.

[0074] The calculation process of this method is as follows:

[0075] Assume that two strings consisting of multiple characters S = S1S2S3…S m and T = T1T2T3…T n , construct the matching relationship matrix LD[m+1,n+1] between S and T, and fill the matrix LD as follows:

[0076]

[0077] in:

[0078]

[0079] i, j represent the serial numbers of characters in the strings S and T respectively; m and n represent the number of characters in the strings S and T respectively;

[0080] The element d at the lower right corner of the matrix LD mn The value of is the edit distance between the two strings. The smaller the edit distance between the two strings, the higher the similarity between the two strings. The calculation formula for calculating the similarity Sim(S,T) using the edit distance is as follows:

[0081]

[0082] However, the above method only considers the number of edits, and does not consider the impact of the longest common subsequence on similarity; for the two strings "ID number" and "resident identity document number", the similarity calculated by this formula is sim(S,T)=1-4 / 8=0.5, which is not high, but in fact the two words have a high similarity; therefore, the impact of the longest common subsequence on similarity is considered here, and the algorithm for calculating similarity by edit distance is improved as follows:

[0083]

[0084] Wherein, LCS(S, T) is the longest common subsequence of two strings, and α is an adjustable parameter, which reflects the influence of the longest common subsequence on the similarity. The similarity of the above strings is calculated using the improved similarity calculation formula, where α is taken as 1, and the calculated similarity is sim(S, T) = 1-4 / 11 = 0.636. The improved similarity calculation algorithm can effectively improve the similarity of similar words.

[0085] When the similarity between the calculated file field name and the field name in the field information table or the field Chinese name or field alias is greater than the set threshold, the file field is considered to match the field in the database; it is worth noting that the field alias is another name that is synonymous with the English name of the field. Users can add multiple field aliases based on experience, which can effectively improve the success rate of matching;

[0086] Step 52: The algorithm for calculating similarity using the improved edit distance solves the problem of literal matching, that is, the two strings are literally the same or similar, including the text field column name is "mobile phone number", and the Chinese name of the field in the database is "mobile phone number"; but if the two strings are semantically similar words, including the text field column name is "mobile phone number", and the Chinese name of the field in the database is "telephone number", the algorithm for calculating similarity using the improved edit distance cannot achieve good results, and it is very likely that they will not match; therefore, when the improved edit distance similarity calculation algorithm fails to match the database field, the semantic similarity algorithm based on the HowNet word is used to calculate the similarity and further match; and when the database field has been matched by the improved edit distance similarity calculation algorithm, this step is no longer necessary; the semantic similarity calculation process based on the HowNet word is as follows:

[0087] In HowNet, two concepts, "sense" and "semene", are proposed. A sense is a description of the semantics of a word. Each word can be expressed as several senses, and a sense is composed of multiple basic semes. For two words ω1 and ω2, if ω1 has N senses s 11 ,s 12 ,…,s 1N , ω2 has M meanings s 21 ,s 22 ,…,s 2M , then the similarity between two words ω1 and ω2 is:

[0088]

[0089] Where sim(s 1I ,s 2J ) is the similarity of the meaning. Since the meaning is composed of four types of semantic primitives: the first basic semantic primitive, other basic semantic primitives, relational semantic primitives, and relational symbol description semantic primitives, the similarity of the two meanings s1 and s2 is:

[0090]

[0091] Among them, β i1 is an adjustable parameter, where 1≤i1≤4, and satisfies β1+β2+β3+β4=1, β1≥β2≥β3≥β4, sim j1 (p1, p2) represents the similarity between the sememe sets of the j1th type;

[0092] The sememe similarity is usually calculated by the semantic distance of the sememes, which is calculated as follows:

[0093]

[0094] Where dis(p1,p2) is the semantic distance between the semantic primitives p1 and p2 in the semantic primitive hierarchy tree. When the semantic primitives are in different trees, this value takes a larger constant, and α is an adjustable parameter.

[0095] The above formula can be used to calculate the semantic similarity between the file field name and the database field name or the field Chinese name or field alias. When the calculated semantic similarity is greater than the set threshold, the field column is considered to match the field in the database. This method can be used to solve the matching of semantically similar words.

[0096] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A data import method based on similarity calculation, characterized in that: The method comprises the following steps: Step 1: Upload files to the system. Multiple file types are supported, including common table file types and common database file types. Multiple files can be uploaded, but the types and headers of all files must be consistent. Step 2: Automatically read the header field column of the file. For Excel type files, the system reads the header information of the first sheet page by default. If you need to read other sheet page information, you need to configure the sheet page; Step 3: The user selects the file fields to be imported into the database; Step 4: Determine whether the user has configured a matching relationship before based on the field selected by the user. If the user has configured it before, the previous configuration is automatically loaded; if not, proceed to the next similarity calculation; Step 5: Before calculating the similarity, it is necessary to confirm that the field information table of the database has been established in the database. The table counts the field information of all tables in the database. The field information includes the field name, the Chinese name of the field, the table to which it belongs, and the field alias. After confirming that the field information table has been established, calculate the similarity between the file field and the field name, the Chinese name of the field, and the field alias in the field information table. The similarity calculation adopts a method combining the improved edit distance algorithm for calculating similarity and the word semantic similarity algorithm based on HowNet. Step 6: Automatically match the data tables and fields in the database based on the similarity calculation results. Multiple data tables can be matched, that is, the data in the file is imported into multiple data tables respectively. To avoid matching too many data tables, the ratio of the number of matched fields in the data table to the number of data columns in the file to be imported will be calculated here. Only data tables with a ratio greater than the set threshold will be matched; Step 7: The user modifies the matching results. For fields that are not matched or that the user believes are not matched accurately, the user can manually select fields in the data table for matching; Step 8: The system will automatically record the user configuration content. The next time you import the same field of the same file, the configuration content will be automatically loaded to avoid repeated configuration by the user. Step 9: Import the data in the file into the database according to user configuration.

2. The data import method based on similarity calculation according to claim 1, characterized in that: The similarity calculation method in step 5 is specifically as follows: Step 51: First, the similarity between the file field name and the field name, field Chinese name, and field alias in the field information table is obtained by using an improved edit distance similarity calculation algorithm; Edit distance refers to the minimum number of operations from one string to another. The smaller the edit distance, the more similar the two strings are, that is, the more similar the two words are. The operations include deletion, insertion, and replacement. The calculation process of this method is as follows: Assume that two strings consisting of multiple characters S = S1S2S3…S m and T = T1T2T3…T n , construct the matching relationship matrix LD[m+1,n+1] between S and T, and fill the matrix LD as follows: in: i, j represent the serial numbers of characters in the strings S and T respectively; m and n represent the number of characters in the strings S and T respectively; The element d at the lower right corner of the matrix LD mn The value of is the edit distance between the two strings. The smaller the edit distance between the two strings, the higher the similarity between the two strings. The calculation formula for calculating the similarity Sim(S,T) using the edit distance is as follows: However, the above method only considers the number of edits and does not take into account the impact of the longest common subsequence on similarity; for the two strings "ID number" and "Resident Identity Certificate Number", the similarity calculated using this formula is sim(S,T) = 1 - 4 / 8 = 0.5, and the calculated similarity is not high, but in fact these two words have a high degree of similarity; therefore, here we consider the impact of the longest common subsequence on similarity and improve the algorithm for calculating similarity based on the edit distance as follows: Among them, LCS(S,T) is the longest common subsequence of the two strings, and α is an adjustable parameter, which reflects the degree of influence of the longest common subsequence on similarity; use the improved similarity calculation formula to calculate the similarity of the above strings. Here, α is taken as 1, and the calculated similarity is sim(S,T) = 1 - 4 / 11 = 0.

636. The improved similarity calculation algorithm can effectively improve the similarity of similar words; When the similarity between the file field name and the field name or Chinese field name or field alias in the field information table is greater than the set threshold, it is considered that the file field matches the field in the database; it should be noted that the field alias is another name synonymous with the Chinese and English names of the field. Users can add multiple field aliases according to experience, which can effectively improve the success rate of matching; Step 52: The algorithm using the improved edit distance to calculate similarity solves the problem of literal matching, that is, the two strings are literally the same or similar, including the text field column name "Mobile Phone Number" and the Chinese field name in the database "Mobile Phone Number"; but if the two strings are semantically similar words, including the text field column name "Mobile Phone Number" and the Chinese field name in the database "Telephone Number", the algorithm using the improved edit distance to calculate similarity cannot achieve good results and is very likely to fail to match; therefore, when the database field is not matched according to the algorithm using the improved edit distance to calculate similarity, then use the algorithm for calculating similarity based on the WordNet to calculate similarity for further matching; and when the database field has been matched through the improved edit distance similarity calculation algorithm, this step is not required; The specific process of calculating the semantic similarity based on the WordNet is as follows: In HowNet, two concepts, "sense" and "semene", are proposed. A sense is a description of the semantics of a word. Each word can be expressed as several senses, and a sense is composed of multiple basic semes. For two words ω1 and ω2, if ω1 has N senses s 11 ,s 12 ,…,s 1N , ω2 has M meanings s 21 ,s 22 ,…,s 2M , then the similarity between two words ω1 and ω2 is: Where sim(s 1I ,s 2J ) is the similarity of the meaning. Since the meaning is composed of four types of semantic primitives: the first basic semantic primitive, other basic semantic primitives, relational semantic primitives, and relational symbol description semantic primitives, the similarity of the two meanings s1 and s2 is: Among them, β i1 is an adjustable parameter, where 1≤i1≤4, and satisfies β1+β2+β3+β4=1, β1≥β2≥β3≥β4, Sim j1 (p1, p2) represents the similarity between the sememe sets of the j1th type; And the similarity of sememes is usually calculated from the semantic distance of sememes, and the calculation is as follows: Among them, dis(p1, p2) is the semantic distance of sememes p1 and p2 in the sememe hierarchy tree. When the sememes are in different trees, this value takes a relatively large constant, and α is an adjustable parameter; Through the above formula, the semantic similarity between the file field name and the field name or Chinese field name or field alias in the database can be calculated. When the calculated semantic similarity is greater than the set threshold, it is considered that the field column matches the field in the database; using this method can solve the matching of semantically similar words.

3. The data import method based on similarity calculation according to claim 1, characterized in that: The above method realizes the automatic matching of file fields and database fields by calculating the similarity between file fields and fields in the database, and can effectively reduce the workload of manual configuration.

4. The data import method based on similarity calculation according to claim 1, characterized in that: The method improves the commonly used edit distance similarity calculation algorithm, considers the influence of the longest common subsequence on the similarity calculation, improves the similarity calculation formula, and improves the matching success rate.

5. The data import method based on similarity calculation according to claim 1, characterized in that: The method adopts a method combining edit distance calculation with word semantic similarity calculation based on HowNet when calculating similarity, which can perform literal matching and semantic matching, and effectively improves the matching success rate.

6. The data import method based on similarity calculation according to claim 1, characterized in that: When calculating similarity, the method not only calculates the similarity with the field name and the Chinese name of the field in the database, but also calculates the similarity with the field alias. The user can configure multiple field aliases based on experience, which can further improve the matching success rate and matching accuracy.

7. The data import method based on similarity calculation according to claim 1, characterized in that: The method can automatically record the user configuration content, and the previous configuration can be automatically loaded when the same file field is imported next time, thereby avoiding the user from repeating the configuration work.

8. The data import method based on similarity calculation according to claim 1, characterized in that: The method can realize the import of various types of files, including commonly used table file types such as excel and txt, and commonly used database file types such as dbf and json.

9. The data import method based on similarity calculation according to claim 1, characterized in that: The method can upload multiple files for import at one time.

10. The data import method based on similarity calculation according to claim 1, characterized in that: The method can manually configure the file fields to be imported, which has greater flexibility.

Citation Information

Patent Citations

  • Data lead-in method and device

    CN101866364A

  • Chinese statement similarity calculation method and apparatus, and computer storage medium

    CN106970912A

  • Network hot-point topic discovery method and system

    CN108509490A

  • Data import method and terminal equipment

    CN110471901A

Cited By

  • Method and device for processing multi-version spreadsheet import

    CN120235128A

  • Data importing method and device, electronic equipment and storage medium

    CN121144292A