Multi-dimensional data matching training processing method based on multi-type database files
By cleaning the text data of multi-type databases, unifying the format and matching the keywords, a data matching level framework is built, which solves the problem of inaccurate identification of newly uploaded text data category variables in multi-type databases, and realizes the flexibility and accuracy of data matching.
Patent Information
- Application Number
- CN202510433037.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Existing multi-type databases cannot automatically identify categorical variables for newly uploaded text data, resulting in inaccurate data matching and inflexible addition.
By cleaning and unifying the format of newly uploaded text data, multi-field similarity analysis and deduplication are performed, keyword fields are extracted for matching weight coefficient analysis, data matching level framework is constructed, category variables are dynamically identified and matched and added.
It realizes automatic category identification and dynamic matching of newly uploaded text data in multiple types of databases, ensuring the flexibility and accuracy of data upload, eliminating invalid data, and improving data matching efficiency.
Smart Images

Figure CN120354147A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and specifically to a multi-dimensional data matching training and processing method based on multi-type database files. Background Art
[0002] Multi-type data classification is the first step in data asset management. Whether it is cataloging, standardizing data assets, or data rights confirmation, management, or providing data asset services, effective data classification is the primary task. Data classification is more considered from the perspective of business or data management, including industry dimension, business domain dimension, data source dimension, sharing dimension, data opening dimension, etc. At the same time, according to these dimensions, data with the same attributes or characteristics are classified according to certain principles and methods.
[0003] Data category matching refers to ensuring that the categories in different data sets are consistent in definition, label, coding, or structure during data processing, integration, or analysis, so as to enable effective comparison, integration, or analysis. It is an important link in data preprocessing, aiming to eliminate data conflicts, analysis errors, or result deviations caused by inconsistent categories. Among them, identifying categorical variables is a key step in data analysis and preprocessing, which refers to finding out those variables in the data set that represent classification or grouping. The characteristic of categorical variables is that their values are a finite number of discrete categories or labels, rather than continuous numerical values. The purpose of identifying categorical variables is to better understand the data structure and lay a foundation for subsequent data cleaning, transformation, and analysis.
[0004] Categorical variables refer to variables used to represent data classification or grouping, and their values are usually discrete and finite categories or labels. At present, multi-type databases can only add data after manually selecting categories, and cannot accurately and automatically identify complex categorical variables in multi-type databases for data matching. The present application aims to automatically identify the categories of newly updated text data in multi-type databases, perform dynamic multi-category matching on newly updated text data, and analyze the matching weight coefficients of different added updated texts with different categories of samples in the multi-type database through the updated text data, so as to effectively identify categorical variables for data matching and addition, and ensure the flexibility of data upload and addition in the database. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-dimensional data matching training and processing method based on multi-type database files to solve the problems in the prior art.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A multi - dimensional data matching training and processing method based on multi - type database files:
[0008] S1: Obtain the newly uploaded and updated text data inside the database file for cleaning, standardize and unify the format of the updated text data. At the same time, perform multi - field similarity analysis on different updated text data, perform normalization processing on the data with similarity higher than the threshold, perform segmented optimization and duplicate removal according to the similarity of different updated text data, and eliminate invalid data;
[0009] S2: Extract the sample data inside the existing categories in the multi - type database, extract and mark the data features of the optimized updated text data and the sample data respectively. Traverse the keyword field information in the optimized updated text data to perform matching analysis with the corresponding sample data content, and analyze the matching weight coefficients with the existing categories in the multi - type database;
[0010] S3: Obtain the matching weight coefficients of the updated text data and all sample data, construct a data matching level framework, including a perfect match level, a similar match level, and a fuzzy match level. Obtain the content of the updated text data corresponding to the perfect match level, and perform matching processing on the updated text data and its corresponding sample data category;
[0011] S4: Obtain the content of the updated text data corresponding to the similar match level and the fuzzy match level, calculate the exact match ratio of different updated text data and the corresponding sample data. When the exact match ratio is higher than the set threshold, perform full - range associated matching training on the updated text data and the complete data content inside its corresponding sample data category, and upload the matching training data to the data matching level framework for secondary matching operations;
[0012] S5: When the exact match ratio is lower than the set threshold, index the semantic information of the uploaded and updated data, extract keywords according to the indexed semantic information, and mark it as a new category inside the database file.
[0013] Further settings: In S1, obtain the newly uploaded and updated text data inside the database file for cleaning, standardize and unify the format of the updated text data. At the same time, perform multi - field similarity analysis on different updated text data, perform normalization processing on the data with similarity higher than the threshold, perform segmented optimization and duplicate removal according to the similarity of different updated text data, and eliminate invalid data, including the following steps:
[0014] S11: Obtain the updated text data to be uploaded inside the multi - type database, and perform normalization processing on the updated text data according to the standardized format, including unifying the text format, delimiters, removing blank characters, removing common meaningless words, and clearing invalid text data;
[0015] S12: Divide the updated text data in the unified format into different fields, perform similarity comparison on different fields of the updated text data, and set the number of fields divided for the current two strings of updated text data to be n i 、n j , where the number of similar fields for the two strings of updated text data is k, and the set similarity threshold is L k , when it is determined that the similarity of the current two strings of updated text data is high. Identify synonyms for the data fields in the two strings of updated text data whose similarity has not been determined. After identifying synonyms, perform similarity analysis again. When the similarity threshold of the data fields in the two strings of updated text data is equal to 100%, it is determined that the current two strings of updated text data are extremely similar. Remove duplicates from one of the updated text data. When the similarity threshold of the data fields in the two strings of updated text data is greater than or equal to the set threshold, normalize and merge the similar fields in the two strings of updated text data, and remove the duplicate fields.
[0016] Further settings: In S2, extract the sample data within the existing categories in the multi-type database, extract and mark the data features of the optimized updated text data and the sample data respectively, traverse the keyword field information in the optimized updated text data and match and analyze the corresponding sample data content, and analyze the matching weight coefficients with the existing categories in the multi-type database. It also includes the following steps:
[0017] S21: Obtain the optimized and deduplicated updated text data and the sample data within the existing categories in the multi-type database respectively, capture the word order and semantics within the updated text data and the sample data, collect the high-frequency words within the updated text data for marking, correspond the high-frequency words within the updated text data with the text semantics, and mark the data fields that match the semantic information and contain high-frequency words as keyword fields;
[0018] S22: Calculate the semantic matching overlap degree between the keyword fields marked within the updated text data and different sample data, initially screen and mark the sample data with a higher overlap degree. At the same time, perform correlation word analysis on the keyword fields according to the content of the initially screened sample data, obtain the correlation words analyzed for each keyword field in the sample data, for the keyword fields with correlation words, pre-replace the original data text in the keyword fields with the correlation words, and mark the keyword fields after replacing the correlation words as preset keyword fields. Respectively perform overlap matching between the keyword fields and the preset keyword fields and the initially screened sample data, and analyze the matching weight coefficients of different keyword fields with different sample data categories.
[0019] Further settings: In S22, it includes the following steps:
[0020] S22-1: Perform a coincidence analysis between the keyword fields marked in the updated text data and different sample data. When the keyword fields marked in the updated text data are exactly the same as the text data inside the sample data, define the matching weight coefficient of this keyword field as N1;
[0021] S22-2: Obtain the preset keyword fields analyzed based on the text content of the sample data, and perform a coincidence analysis on the preset keyword fields. When the preset keyword fields replaced according to the keyword fields in the updated text data are exactly the same as the text data inside the sample data, define the matching weight coefficient of this keyword field as 0.8N1;
[0022] S22-3: When the keyword fields marked in the updated text data are semantically similar and coincide with the text data inside the sample data, define the matching weight coefficient of this keyword field as 0.6N1;
[0023] S22-4: When the replaced preset keyword fields are semantically similar and coincide with the text data inside the sample data, define the matching weight coefficient of this keyword field as 0.5N1;
[0024] 22-5: When the keyword fields marked in the updated text data and the preset keyword fields are vaguely coincident with different sample data, define the matching weight coefficient of this keyword field as 0;
[0025] S22-6: Analyze the matching weight coefficients between each updated text data and the sample data according to the number of keyword fields marked in the different updated text data. Among them, when the matching weight coefficients of all keyword fields inside a certain updated text data and the sample data are all N1, screen whether this updated text data and the sample data are duplicates, and delete them if this updated text data and the sample data are duplicate data.
[0026] Further setting: In S3, obtain the matching weight coefficients between the updated text data and all sample data, construct a data matching level framework, including a perfect match level, a similar match level, and a fuzzy match level, obtain the content of the updated text data corresponding to the perfect match level, and perform a matching process between the updated text data and its corresponding sample data category, which also includes the following steps:
[0027] S31: Analyze the matching weight coefficients of different updated text data according to the keyword fields marked inside the updated text data and the matching weight coefficients with different sample data. Construct a data matching level framework based on different matching weight coefficients. Define the matching weight coefficient of the updated text data and the sample data as 0.8N1 - N1 as the perfect match level, the matching weight coefficient of the updated text data and the sample data as 0.5N1 - 0.8N1 as the approximate match level, and the matching weight coefficient of the updated text data and the sample data as 0 - 0.5N1 as the fuzzy match level.
[0028] S32: Obtain the updated sample data at the perfect match level and the sample data it matches, and perform a matching mark on the category of the updated sample data at the perfect match level according to the category of the matched sample data.
[0029] Further set: In S4, obtain the content of the updated text data corresponding to the approximate match level and the fuzzy match level, calculate the exact match ratio of different updated text data and the corresponding sample data. When the exact match ratio is higher than the set threshold, perform a full-range associated matching training on the content of the complete data inside the category of the updated text data and its corresponding sample data, and upload the matching training data to the data matching level framework for a secondary matching operation, including the following steps:
[0030] Obtain the updated text data with the matching weight levels of the approximate match level and the fuzzy match level, and set the number of keyword fields inside a certain updated text data as M, M = {m1, m2, …, m k}, and the matching weight coefficients of each keyword field and the preset keyword field with a certain matched sample data are respectively Calculate the exact match ratio of this updated text data. Set the exact ratio threshold for the current different updated text data and the sample data match as 0.6N1, and set the exact match ratio f of the current updated text data M , according to the formula:
[0031]
[0032] When the exact ratio of the updated text data with the matching weight levels of approximate matching level and fuzzy matching level to the matched sample data is greater than the set threshold, open the complete data within the category corresponding to the current sample data. Conduct a keyword field coincidence analysis between the updated text data with the exact ratio greater than the set threshold and the full-range complete text data under the current sample data category. Re-analyze the associated words based on the content of the complete text data to construct the preset keyword fields. At the same time, conduct a coincidence match between the keyword fields within the updated text data and the preset keyword fields with the full-range complete text data, analyze the matching weight coefficients of different keyword fields with the full-range complete text data, and send the analyzed matching weight coefficients of the updated text data to the data matching level framework for re-determination;
[0033] When the matching result is the exact matching level, mark the category of the updated text data according to the determination result. When the matching result is the approximate matching level, send the exact ratio of the updated text data to the matched sample data and the matching weight coefficient with the full-range complete text data under the corresponding sample data category to the remote PC for manual discrimination.
[0034] Further set: In S5, when the exact matching ratio is lower than the set threshold, index the semantic information of the uploaded and updated data, extract keywords according to the indexed semantic information, and mark it as a new category within the database file, including the following steps:
[0035] S51: Obtain the updated text data with the exact ratio of matching with the sample data lower than the set threshold, extract the keyword fields within the updated text data, mark the updated text data as new category data within the multi-type database, and extract the keyword fields with a matching weight coefficient of 0 for the currently different keyword fields and the sample data to be matched;
[0036] S52: Index the filtered keyword fields according to the semantic information of the updated text data, mark the keyword fields with prominent semantic information expression, and define them as the new category index of the updated text data.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows: It aims to automatically identify the categories of newly updated text data in a multi-type database, conduct dynamic multi-category matching on the newly updated text data, preprocess the updated text data, and perform segmented optimization and deduplication according to the similarity of different updated text data to eliminate invalid data;
[0038] Secondly, by updating the text data and analyzing the sample data under different categories, the matching weight coefficients between the updated text added and different categories existing in the multi-type database are analyzed, a data matching level framework is constructed, and the updated text data is matched with the corresponding sample data categories. Meanwhile, for the updated text data with a low matching weight coefficient, keyword extraction is performed according to the semantic information of the index and added as a new category inside the database file, so as to effectively identify the category variables for data matching and addition, and ensure the flexibility of data upload and addition in the database. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to make the content of the present invention easier to be clearly understood, the present invention will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0040] Figure 1 It is a schematic diagram of the general steps of a multi-dimensional data matching training and processing method based on a multi-type database file according to the present invention;
[0041] Figure 2 It is a schematic diagram of the specific steps of S1 in a multi-dimensional data matching training and processing method based on a multi-type database file according to the present invention;
[0042] Figure 3 It is a schematic diagram of the specific steps of S2 in a multi-dimensional data matching training and processing method based on a multi-type database file according to the present invention;
[0043] Figure 4 It is a schematic diagram of the specific steps of S22 in a multi-dimensional data matching training and processing method based on a multi-type database file according to the present invention;
[0044] Figure 5 It is a schematic diagram of the specific steps of S3 in a multi-dimensional data matching training and processing method based on a multi-type database file according to the present invention;
[0045] Figure 6 It is a schematic diagram of the specific steps of S5 in a multi-dimensional data matching training and processing method based on a multi-type database file according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Please refer to Figures 1 to 6 , in the embodiments of the present invention, a multi-dimensional data matching training and processing method based on a multi-type database file:
[0048] S1: Obtain the newly uploaded and updated text data inside the database file for cleaning, standardize and unify the format of the updated text data. At the same time, perform multi-field similarity analysis on different updated text data, perform normalization processing on the data with similarity higher than the threshold, perform segmented optimization and duplicate removal according to the similarity of different updated text data, and eliminate invalid data;
[0049] As Figure 2 shown, step S1 needs further explanation;
[0050] S11: Obtain the updated text data to be uploaded inside multi-type databases, and perform normalization processing on the updated text data according to the standardized format, including unifying the text format, delimiter, removing blank characters, removing common meaningless words, and clearing invalid text data;
[0051] S12: Divide the updated text data with unified format into different fields, perform similarity comparison on different fields of the updated text data. Assume the number of fields divided for the current two strings of updated text data are n i 、n j respectively, where the number of similar fields for the two strings of updated text data is k, and the set similarity threshold is L k . When it is determined that the similarity of the current two strings of updated text data is high, perform synonym identification on the data fields in the two strings of updated text data whose similarity has not been determined. After identifying synonyms, perform similarity analysis again. When the similarity threshold of the fields of the two strings of updated text data is equal to 100%, it is determined that the current two strings of updated text data are extremely similar, and one of the updated text data is removed and de-duplicated. When the similarity threshold of the fields of the two strings of updated text data is greater than or equal to the set threshold, normalize and merge the similar fields in the two strings of updated text data, and remove the duplicate fields.
[0052] S2: Extract the sample data inside the existing categories in multi-type databases, extract and mark the data characteristics of the optimized updated text data and the sample data respectively, traverse the keyword field information in the optimized updated text data and perform matching analysis with the sample data content, and analyze the matching weight coefficient with the existing categories in multi-type databases;
[0053] As Figure 3 shown, step S2 needs further explanation;
[0054] S21: Obtain the updated text data after optimization and deduplication and the sample data within existing categories in the multi-type database respectively, capture the word order and semantics within the updated text data and the sample data, collect the high-frequency words within the updated text data for marking, correspond the high-frequency words within the updated text data with the text semantics, and mark the data fields that match the semantic information and have high-frequency words as keyword fields;
[0055] S22: Calculate the semantic matching overlap degree between the keyword fields marked within the updated text data and different sample data, preliminarily screen and mark the sample data with a relatively high overlap degree. At the same time, conduct a correlation word analysis on the keyword fields according to the content of the preliminarily screened sample data, obtain the correlation words analyzed for each keyword field in the sample data, for the keyword fields with existing correlation words, pre-replace the original data text in the keyword fields with the correlation words, mark the keyword fields after replacing the correlation words as preset keyword fields, respectively conduct an overlap match between the keyword fields and the preset keyword fields and the preliminarily screened sample data, and analyze the matching weight coefficients of different keyword fields with different sample data categories.
[0056] Among them, as Figure 4 , it should also be specifically noted that in step S22, it includes the following steps:
[0057] S22-1: Conduct an overlap analysis between the keyword fields marked within the updated text data and different sample data. When the keyword fields marked within the updated text data are completely overlapped with the text data within the sample data, define the matching weight coefficient of this keyword field as N1;
[0058] S22-2: Obtain the preset keyword fields analyzed according to the text content of the sample data, conduct an overlap analysis on the preset keyword fields. When the preset keyword fields replaced according to the keyword fields within the updated text data are completely overlapped with the text data within the sample data, define the matching weight coefficient of this keyword field as 0.8N1;
[0059] S22-3: When the keyword fields marked within the updated text data are semantically similar and overlapped with the text data within the sample data, define the matching weight coefficient of this keyword field as 0.6N1;
[0060] S22-4: When the replaced preset keyword fields are semantically similar and overlapped with the text data within the sample data, define the matching weight coefficient of this keyword field as 0.5N1;
[0061] S22-5: When the keyword fields marked within the updated text data and the preset keyword fields are vaguely overlapped with different sample data, define the matching weight coefficient of this keyword field as 0;
[0062] S22-6: Analyze the matching weight coefficients between each updated text data and the sample data according to the number of keyword fields with internal tags of different updated text data. Among them, when the matching weight coefficients of all keyword fields inside a certain updated text data with the sample data are all N1, screen whether this updated text data is duplicate with the sample data, and delete it if this updated text data and the sample data are duplicate data.
[0063] S3: Obtain the matching weight coefficients between the updated text data and all sample data, construct a data matching level framework, including a perfect matching level, a similar matching level, and a fuzzy matching level, obtain the content of the updated text data corresponding to the perfect matching level, and perform matching processing on the updated text data and its corresponding sample data category;
[0064] As Figure 5 shown, step S3 needs further explanation;
[0065] S31: Analyze the matching weight coefficients of different updated text data according to the matching weight coefficients between the keyword fields with internal tags of the updated text data and different sample data, construct a data matching level framework according to different matching weight coefficients, define the matching weight coefficient between the updated text data and the sample data as 0.8N1 - N1 as the perfect matching level, define the matching weight coefficient between the updated text data and the sample data as 0.5N1 - 0.8N1 as the similar matching level, and define the matching weight coefficient between the updated text data and the sample data as 0 - 0.5N1 as the fuzzy matching level;
[0066] S32: Obtain the updated sample data at the perfect matching level and the sample data it matches, and perform matching marking on the category of the updated sample data at its perfect matching level according to the category of the matching sample data.
[0067] S4: Obtain the content of the updated text data corresponding to the similar matching level and the fuzzy matching level, calculate the exact matching ratio between different updated text data and the corresponding sample data. When the exact matching ratio is higher than the set threshold, perform full-range associated matching training on the content of the complete data inside the category of this updated text data and its corresponding sample data, upload the matching training data to the data matching level framework, and perform a secondary matching operation;
[0068] Step S4 needs further explanation;
[0069] Obtain the updated text data with the matching weight levels of the similar matching level and the fuzzy matching level, set the number of keyword fields inside a certain updated text data as M, M = {m1, m2,..., m k}, and the matching weight coefficients of each keyword field and the preset keyword field with a certain matching sample data are respectively Calculate the exact matching ratio of the updated text data. Set the exact ratio threshold for the current different updated text data matching the sample data to 0.6N1, and set the exact matching ratio f of the current updated text data M , according to the formula:
[0070]
[0071] When the exact ratio of the updated text data with the matching weight levels of approximate matching level and fuzzy matching level to the matched sample data is greater than the set threshold, open the complete data within the category corresponding to the current sample data. Conduct a keyword field coincidence analysis on the updated text data with the exact ratio greater than the set threshold and the full-range complete text data under the current sample data category. Re-analyze the associated words according to the content of the complete text data to construct the preset keyword fields. At the same time, conduct a coincidence match between the keyword fields inside the updated text data and the preset keyword fields and the full-range complete text data, analyze the matching weight coefficients of different keyword fields with the full-range complete text data, and send the analyzed matching weight coefficients of the updated text data to the data matching level framework for re-judgment;
[0072] When the matching result is the exact matching level, mark the category of the updated text data according to the judgment result. When the matching result is the approximate matching level, send the exact ratio of the updated text data to the matched sample data and the matching weight coefficient with the full-range complete text data under the corresponding sample data category to the remote PC for manual discrimination.
[0073] S5: When the exact matching ratio is lower than the set threshold, index the semantic information of the uploaded and updated data, extract keywords according to the indexed semantic information, and mark it as a new category inside the database file.
[0074] As Figure 6 shown, step S5 needs further explanation;
[0075] S51: Obtain the updated text data whose exact ratio of matching with the sample data is lower than the set threshold, extract the keyword fields inside the updated text data, mark the updated text data as new category data inside the multi-type database, and extract the keyword fields with a matching weight coefficient of 0 for the current different keyword fields and the sample data to be matched;
[0076] S52: Index the filtered keyword fields according to the semantic information of the updated text data, mark the keyword fields with prominent semantic information expression, and define them as the new category index of the updated text data.
[0077] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A multi-dimensional data matching training and processing method based on multi-type database files, characterized in that: S1: Obtain the newly uploaded and updated text data inside the database file for cleaning, standardize and unify the format of the updated text data. At the same time, perform multi-field similarity analysis on different updated text data, perform normalization processing on the data with similarity higher than the threshold, perform segmented optimization and deduplication according to the similarity of different updated text data, and eliminate invalid data; S2: Extract the sample data inside the existing categories in the multi-type database, respectively extract and mark the data features of the optimized updated text data and the sample data, traverse the keyword field information in the optimized updated text data and perform matching analysis on the corresponding sample data content, and analyze the matching weight coefficients with the existing categories in the multi-type database; S3: Obtain the matching weight coefficients of the updated text data and all sample data, construct a data matching level framework, including a perfect match level, a similar match level, and a fuzzy match level. Obtain the content of the updated text data corresponding to the perfect match level, and perform matching processing on the updated text data and its corresponding sample data category; S4: Obtain the content of the updated text data corresponding to the similar match level and the fuzzy match level, calculate the exact match ratio of different updated text data and the corresponding sample data. When the exact match ratio is higher than the set threshold, perform full-range associated matching training on the complete data content inside the updated text data and its corresponding sample data category, upload the matching training data to the data matching level framework, and perform a secondary matching operation; S5: When the exact match ratio is lower than the set threshold, index the semantic information of the uploaded and updated data, extract keywords according to the indexed semantic information, and mark it as a new category inside the database file.
2. The multi-dimensional data matching training and processing method based on multi-type database files according to claim 1, characterized in that : In the above S1, obtaining the newly uploaded and updated text data inside the database file for cleaning, standardizing and unifying the format of the updated text data, at the same time performing multi-field similarity analysis on different updated text data, performing normalization processing on the data with similarity higher than the threshold, performing segmented optimization and deduplication according to the similarity of different updated text data, and eliminating invalid data, includes the following steps: S11: Obtain the updated text data to be uploaded inside the multi-type database, and perform normalization processing on the updated text data according to the standard format, including unifying the text format, delimiters, removing blank characters, removing common meaningless words, and clearing invalid text data; S12: Divide the updated text data in the unified format into different fields, compare the similarity of different updated text data fields, and set the number of fields divided by the current two strings of updated text data to n i and n j , where the number of similar fields of the two strings of updated text data is k, and the set similarity threshold is L k When it is determined that the similarity of the current two strings of updated text data is high, identify synonyms for the data fields whose similarity has not been determined in the two strings of updated text data, and perform similarity analysis again after identifying synonyms. When the similarity threshold of the two strings of updated text data fields is equal to 100%, it is determined that the current two strings of updated text data are extremely similar, remove duplicates from one of the updated text data, and when the similarity threshold of the two strings of updated text data fields is greater than or equal to the set threshold, normalize and merge the similar fields in the two strings of updated text data, and remove the duplicate fields.
3. A multi-dimensional data matching training and processing method based on multi-type database files according to claim 1, characterized in that : In the above S2, extracting the sample data inside the existing categories in the multi-type database, respectively extracting and marking the data features of the optimized updated text data and the sample data, traversing the keyword field information in the optimized updated text data and performing matching analysis on the corresponding sample data content, and analyzing the matching weight coefficients with the existing categories in the multi-type database, further includes the following steps: S21: Obtain the updated text data after optimization and deduplication and the sample data within existing categories in the multi-type database respectively, capture the word order and semantics within the updated text data and the sample data, collect and mark the high-frequency words within the updated text data, correspond the high-frequency words within the updated text data to the text semantics, and mark the data fields that match the semantic information and have high-frequency words as keyword fields; S22: Analyze the semantic matching coincidence degree between the marked keyword fields within the updated text data and different sample data, preliminarily screen and mark the sample data with a relatively high coincidence degree. At the same time, conduct a correlation word analysis on the keyword fields according to the content of the preliminarily screened sample data, obtain the correlation words analyzed for each keyword field in the sample data, for the keyword fields with correlation words, pre-replace the original data text in the keyword fields with the correlation words, mark the keyword fields after replacing the correlation words as preset keyword fields, respectively conduct coincidence matching between the keyword fields and the preset keyword fields and the preliminarily screened sample data, and analyze the matching weight coefficients of different keyword fields and different sample data categories.
4. A multi-dimensional data matching training and processing method based on multi-type database files according to claim 3, characterized in that In S22, the following steps are included: S22-1: Conduct a coincidence analysis between the marked keyword fields within the updated text data and different sample data. When the marked keyword fields within the updated text data are completely coincident with the text data within the sample data, define the matching weight coefficient of this keyword field as N1; S22-2: Obtain the preset keyword fields analyzed according to the text content of the sample data, conduct a coincidence analysis on the preset keyword fields. When the preset keyword fields replaced according to the keyword fields within the updated text data are completely coincident with the text data within the sample data, define the matching weight coefficient of this keyword field as 0.8N1; S22-3: When the marked keyword fields within the updated text data are semantically similar and coincident with the text data within the sample data, define the matching weight coefficient of this keyword field as 0.6N1; S22-4: When the replaced preset keyword fields are semantically similar and coincident with the text data within the sample data, define the matching weight coefficient of this keyword field as 0.5N1; S22-5: When the marked keyword fields within the updated text data and the preset keyword fields are vaguely coincident with different sample data, define the matching weight coefficient of this keyword field as 0; S22-6: Analyze the matching weight coefficients of each updated text data and the sample data according to the number of marked keyword fields within different updated text data. Among them, when the matching weight coefficients of all keyword fields within a certain updated text data and the sample data are all N1, screen whether this updated text data and the sample data are repeated, and delete them if this updated text data and the sample data are duplicate data.
5. A multi-dimensional data matching training and processing method based on multi-type database files according to claim 1, characterized in that : In S3, obtain the matching weight coefficients of the updated text data and all sample data, and construct a data matching level framework, including a perfect match level, a similar match level, and a fuzzy match level. Obtain the content of the updated text data corresponding to the perfect match level, and perform matching processing on the updated text data and its corresponding sample data category. The following steps are also included: S31: Analyze the matching weight coefficients of different updated text data according to the keyword fields marked inside the updated text data and the matching weight coefficients of different sample data. Construct a data matching level framework according to different matching weight coefficients. Define the matching weight coefficient of the updated text data and the sample data as 0.8N1 to N1 as the perfect match level, define the matching weight coefficient of the updated text data and the sample data as 0.5N1 to 0.8N1 as the similar match level, and define the matching weight coefficient of the updated text data and the sample data as 0 to 0.5N1 as the fuzzy match level; S32: Obtain the updated sample data at the perfect match level and the sample data it matches, and perform matching marking on the category of the updated sample data at its perfect match level according to the category of the matched sample data.
6. A multi-dimensional data matching training and processing method based on multi-type database files according to claim 1, characterized in that : In S4, obtain the content of the updated text data corresponding to the similar match level and the fuzzy match level, calculate the exact match ratio of different updated text data and the corresponding sample data. When the exact match ratio is higher than the set threshold, perform full-range associated matching training on the updated text data and the complete data content inside its corresponding sample data category, and upload the matching training data to the data matching level framework for secondary matching operations. The following steps are included: Obtain updated text data with matching weight levels of approximate matching level and fuzzy matching level, and set the number of keyword fields within a certain updated text data to M, M = {m1, m2, …, m k}, and the matching weight coefficients of each keyword field and the preset keyword field with a certain matched sample data are respectively Calculate the exact matching ratio of this updated text data. Set the exact ratio threshold for the current different updated text data to match the sample data to 0.6N1, and set the exact matching ratio f of the current updated text data M , according to the formula: When the exact ratio of the updated text data with the matching weight levels of the similar match level and the fuzzy match level and the matched sample data is greater than the set threshold, open the complete data inside the category corresponding to the current sample data. Analyze the coincidence of keyword fields between the updated text data with the current exact ratio greater than the set threshold and the full-range complete text data under the current sample data category. Re-analyze the related words according to the complete text data content to construct a preset keyword field. At the same time, perform coincidence matching on the keyword field inside the updated text data and the preset keyword field with the full-range complete text data, analyze the matching weight coefficients of different keyword fields and the full-range complete text data, and send the analyzed matching weight coefficients of the updated text data to the data matching level framework for re-determination; When the matching result is at the perfect match level, mark the category of the updated text data according to the determination result. When the matching result is at the similar match level, send the exact ratio of the updated text data and the matched sample data and the matching weight coefficient with the full-range complete text data under the corresponding sample data category to the remote PC for manual discrimination.
7. A multi-dimensional data matching training and processing method based on multi-type database files according to claim 1, characterized in that : In S5, when the exact match ratio is lower than the set threshold, index the semantic information of the uploaded updated data, extract keywords according to the indexed semantic information, and mark it as a new category inside the database file. The following steps are included: S51: Obtain updated text data whose exact ratio matching the sample data is lower than the set threshold, extract the keyword fields inside the updated text data, mark the updated text data as brand-new category data inside the multi-type database, and extract the keyword fields with a matching weight coefficient of 0 between the current different keyword fields and the sample data to be matched; S52: Index the screened keyword fields according to the semantic information of the updated text data, mark the keyword fields with prominent semantic information expression, and define them as the brand-new category index of the updated text data.
Citation Information
Patent Citations
Method and device for matching texts
CN102411583A
Rapid cleaning method for adverse reaction public database data
CN117632940A
Data preparation method, system and device for AIGC interaction analysis and medium
CN118093795A
Semantic similarity-based text extraction data similarity matching method
CN119720991A
Semantic parsing method and apparatus
WO2017198031A1