Data acquisition and processing method and device based on data item matching, equipment and medium

By identifying the data source type and metadata information and using preset matching databases for automatic matching, the problem of inefficient manual intervention in traditional data acquisition is solved, efficient and accurate data acquisition and processing is achieved, manual workload is reduced, and the needs of large-scale data processing is met.

CN120336593APending Publication Date: 2025-07-18CHENGDU WEISHITONG INFORMATION SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510477613.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

A lot of manual intervention is required in the traditional data acquisition process, especially when determining the matching relationship between the data items of different data sources and the target data storage structure or analysis model, it is inefficient and prone to errors, making it difficult to meet the needs of large-scale and real-time data acquisition and processing.

Method used

By identifying the data source type and obtaining metadata information, the preset matching database is used for automatic matching, including similarity checks for data item names, types and formats, the cosine similarity is calculated using Word2Vec, preliminary matching is performed in combination with preset matching rules, and optimization and adjustments are performed for items that do not meet the conditions, and data collection and storage are finally realized.

Benefits of technology

It realizes automatic matching system data items during data acquisition, reduces the repetitive work of manual matching, improves the efficiency and accuracy of data item matching, meets the needs of large-scale data acquisition and processing, and provides a high-quality data foundation for subsequent data mining and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336593A_ABST
    Figure CN120336593A_ABST
Patent Text Reader

Abstract

The invention discloses a data acquisition and processing method and device based on data item matching, equipment and a medium, and relates to the technical field of data processing. Comprising the following steps: identifying a data source type of to-be-collected data, and obtaining metadata information of a data source; storing a target data item obtained by analyzing the storage structure of the target data and the data feature requirement in a preset matching database; matching the metadata information with a target data item in a preset matching database to obtain a matching result, and judging whether the matching result meets a preset matching condition or not; if the matching result does not meet the preset matching condition, optimizing the matching result through a preset optimization mode; and if the matching result meets a preset matching condition, collecting corresponding data from the data source according to the matching result. Therefore, the system data items can be automatically matched in the data acquisition process, and repeated work of manual matching is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data set based on data item matching, as well as a processing method, device, equipment and medium therefor. Background Art

[0002] In today's digital age, data collection plays a crucial role in many fields such as business intelligence, big data analysis, and the Internet of Things. With the increasing diversification of data sources and the explosive growth of data volume, how to efficiently and accurately collect data from different data sources and integrate them has become a key issue. Traditional data collection often requires a large amount of manual intervention. Especially when determining the matching relationship between data items in different data sources and data items in the target data storage structure or analysis model, manual operations are not only inefficient but also error-prone, making it difficult to meet the requirements of large-scale and real-time data collection and processing.

[0003] As can be seen from the above, how to implement an automatic matching of system data items during the data collection process and reduce the repetitive work of manual matching is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a data collection and processing method, device, equipment and medium based on data item matching, which can realize the automatic matching of system data items during the data collection process and reduce the repetitive work of manual matching. The specific solutions are as follows:

[0005] In a first aspect, the present application provides a data collection and processing method based on data item matching, including:

[0006] Identifying the data source type of the data to be collected and obtaining the metadata information of the data source; the metadata information includes data item name, data type, data format, and data length;

[0007] Performing a parsing operation on the storage structure of the target data to obtain the target data item and the data feature requirements corresponding to the target data item, and storing the target data item and the data feature requirements in a preset matching database;

[0008] Matching the metadata information with the target data items in the preset matching database to obtain a matching result, and determining whether the matching result meets the preset matching conditions;

[0009] If the matching result does not meet the preset matching conditions, optimizing the matching result through a preset optimization method to obtain a new matching result, and jumping to the step of determining whether the matching result meets the preset matching conditions;

[0010] If the matching result meets the preset matching condition, collect the corresponding data from the data source according to the matching result, and perform processing and storage operations on the collected data based on the data feature requirements.

[0011] Optionally, identifying the data source type of the data to be collected and obtaining the metadata information of the data source includes:

[0012] Identify the data source type of the data to be collected through the connection information of the data source, and select the corresponding information acquisition method according to the data source type to obtain the metadata information of the data source;

[0013] Among them, the data source types include database data sources, file data sources, and network data sources.

[0014] Optionally, matching the metadata information with the target data items in the preset matching database to obtain a matching result includes:

[0015] Use the pre-trained Word2Vec to convert the data item names in the metadata information and the data item names of the target data items into word vector arrays to obtain the corresponding word vector arrays to be matched and target word vector arrays;

[0016] Calculate the cosine similarity between each word vector array to be matched and the target word vector array respectively, and determine the data item corresponding to the word vector array to be matched with the highest cosine similarity as the matching data item corresponding to the target data item;

[0017] Perform data type conversion and data format check on the target data item and its corresponding matching data item, and calculate the matching degree score of the current matching data item and the target data item according to the preset matching rules;

[0018] If the matching degree score exceeds the preset matching degree score threshold, obtain a matching result indicating successful matching. If the matching degree score does not exceed the preset matching degree score threshold, obtain a matching result indicating failed matching.

[0019] Optionally, the preset matching rules are constructed based on the name similarity of data items, data type compatibility, and data format consistency.

[0020] Optionally, determining whether the matching result meets the preset matching condition includes:

[0021] Determine whether the matching result meets the preset matching accuracy condition and the preset matching integrity condition;

[0022] Among them, the preset matching accuracy condition is the condition of the proportion of the number of correctly matched data items to the total number of matched data items; the preset matching integrity condition is the condition of the proportion of successfully matched data items in the target data items.

[0023] Optionally, optimizing the matching result through a preset optimization method to obtain a new matching result includes:

[0024] Correspondingly adjusting the preset matching rule based on the matching result, and rematching the metadata information with the target data items in the preset matching database according to the adjusted preset matching rule to obtain a new matching result.

[0025] Optionally, collecting corresponding data from the data source according to the matching result, and performing processing and storage operations on the collected data based on the data feature requirements, including:

[0026] Collecting corresponding data from the data source according to the matching result by using a data collection method corresponding to the data source type of the data source to obtain initial collected data;

[0027] Performing data type conversion, data cleaning, and data integration operations on the initial collected data based on the data feature requirements to obtain target collected data, and storing the target collected data in a preset storage space.

[0028] In a second aspect, the present application provides a data collection and processing device based on data item matching, including:

[0029] A data source identification module, configured to identify the data source type of the data to be collected and obtain the metadata information of the data source; the metadata information includes data item name, data type, data format, and data length;

[0030] A target data parsing module, configured to perform a parsing operation on the storage structure of the target data to obtain target data items and data feature requirements corresponding to the target data items, and store the target data items and the data feature requirements in a preset matching database;

[0031] A data item matching module, configured to match the metadata information with the target data items in the preset matching database to obtain a matching result, and determine whether the matching result meets the preset matching conditions;

[0032] An optimization and adjustment module, configured to, if the matching result does not meet the preset matching conditions, optimize the matching result through a preset optimization method to obtain a new matching result, and jump to the step of determining whether the matching result meets the preset matching conditions;

[0033] A data acquisition execution module, configured to, if the matching result meets the preset matching condition, collect corresponding data from the data source according to the matching result, and perform processing and storage operations on the collected data based on the data feature requirements.

[0034] In a third aspect, the present application provides an electronic device, including:

[0035] A memory, configured to store a computer program;

[0036] A processor, configured to execute the computer program to implement the foregoing data acquisition and processing method based on data item matching.

[0037] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, wherein when the computer program is executed by a processor, the foregoing data acquisition and processing method based on data item matching is implemented.

[0038] The present application provides a data acquisition and processing method based on data item matching. First, identify the data source type of the data to be collected, and obtain the metadata information of the data source; the metadata information includes data item name, data type, data format, and data length; then perform a parsing operation on the storage structure of the target data to obtain the target data item and the data feature requirements corresponding to the target data item, and store the target data item and the data feature requirements in a preset matching database; finally, match the metadata information with the target data item in the preset matching database to obtain a matching result, and determine whether the matching result meets the preset matching condition; if the matching result does not meet the preset matching condition, optimize the matching result through a preset optimization method to obtain a new matching result, and jump to the step of determining whether the matching result meets the preset matching condition; if the matching result meets the preset matching condition, collect corresponding data from the data source according to the matching result, and perform processing and storage operations on the collected data based on the data feature requirements.

[0039] As can be seen from the above, the present application performs a parsing operation on the storage structure of the target data to obtain the target data item and the data feature requirements corresponding to the target data item, and completes the automatic matching of system data items through a preset matching database. For data items that cannot be fully matched, recommended matching items can be given. Thus, automatic matching of system data items is realized during the data acquisition process, reducing the repetitive work of manual matching. Description of the Drawings

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.

[0041] Figure 1 Flowchart of a data collection and processing method based on data item matching disclosed in the present application;

[0042] Figure 2 Schematic structural diagram of a data collection and processing system based on data item matching disclosed in the present application;

[0043] Figure 3 Schematic diagram of a data item matching process disclosed in the present application;

[0044] Figure 4 Schematic diagram of a data collection and processing device based on data item matching disclosed in the present application;

[0045] Figure 5 Structural diagram of an electronic device disclosed in the present application. Detailed implementation manners

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0047] In today's digital age, data collection plays a crucial role in many fields such as business intelligence, big data analysis, and the Internet of Things. With the increasing diversification of data sources and the explosive growth of data volume, how to efficiently and accurately collect data from different data sources and integrate it has become a key issue. Traditional data collection often requires a large amount of manual intervention. Especially when determining the matching relationship between the data items of different data sources and the data items of the target data storage structure or analysis model, manual operations are not only inefficient but also error-prone, making it difficult to meet the requirements of large-scale and real-time data collection and processing. For this reason, the present application provides a data collection and processing solution based on data item matching, which can automatically match system data items during the data collection process and reduce the repetitive work of manual matching.

[0048] See Figure 1 As shown, the embodiments of the present application disclose a data collection and processing method based on data item matching, including:

[0049] Step S11: Identify the data source type of the data to be collected, and obtain the metadata information of the data source.

[0050] In this embodiment, by analyzing the input data source connection information, the data source type of the data to be collected is determined. Among them, the data source type includes database data source, file data source, and network data source. Specifically, identifying the data source type of the data to be collected and obtaining the metadata information of the data source may include: identifying the data source type of the data to be collected through the connection information of the data source, and for the database data source, using the metadata query statement of the database system to obtain table structure, field information, etc.; for the file data source, the metadata can be obtained by parsing the file header or predefined file format specifications; for the network data source, try to use known delimiters to split the received data to obtain the corresponding key-value pairs. That is, selecting the corresponding information acquisition method according to the data source type can ensure the integrity of the metadata information and improve the metadata acquisition efficiency.

[0051] Step S12: Parse the storage structure of the target data to obtain the target data items and the data feature requirements corresponding to the target data items, and store the target data items and the data feature requirements in a preset matching database.

[0052] In this embodiment, the storage structure of the target data is parsed to determine the required data items and their feature requirements, and the required data items and their feature requirements are stored in a preset matching database. Among them, the feature requirements include but are not limited to data type requirements, value range requirements; the content stored in the preset matching database includes but is not limited to data item names and their aliases, data types, and data formats.

[0053] Step S13: Match the metadata information with the target data items in the preset matching database to obtain a matching result, and determine whether the matching result meets the preset matching conditions.

[0054] In this embodiment, metadata information and data item requirements stored in a preset matching database are obtained, and matching calculations are performed on data items of a data source and target data items based on various factors such as name similarity of data items, data type compatibility, and data format consistency to determine corresponding relationships of data items with a relatively high matching degree. Specifically, the step of matching the metadata information with target data items in the preset matching database to obtain a matching result may include: using pre-trained Word2Vec to convert data item names in the metadata information and data item names of the target data items into word vector arrays to obtain corresponding word vector arrays to be matched and target word vector arrays; calculating the cosine similarity between each word vector array to be matched and the target word vector array respectively, and determining the data item corresponding to the word vector array to be matched with the highest cosine similarity as the matching data item corresponding to the target data item; performing data type conversion and data format check on the target data item and its corresponding matching data item, and calculating the matching degree score of the current matching data item and the target data item according to a preset matching rule; if the matching degree score exceeds a preset matching degree score threshold, a matching result indicating successful matching is obtained, and if the matching degree score does not exceed the preset matching degree score threshold, a matching result indicating failed matching is obtained. Among them, the preset matching rule is constructed based on name similarity of data items, data type compatibility, and data format consistency. That is, the data item names are converted into vectors through Word2Vec, and the name similarity score is calculated using cosine similarity, and matching data items with name similarity are initially screened based on the similarity score; then it is checked whether the data types are compatible. For example, numeric data and string data generally do not match, and matching data items with data type compatibility are further screened; finally, it is checked whether the data formats are consistent. For example, whether the date formats are the same, and the data items of the data source that are initially matched are finally confirmed; the matching degree score of each data item of the data source and the target data item is calculated according to a preset weighted algorithm, and the data item pair with a score exceeding the set threshold is regarded as a pair of data items with initially successful matching. Through the preset matching rule, data item matching can be automatically performed from the preset matching database, reducing the repetitive work of manual matching.

[0055] Further, the preliminary matching result is evaluated according to preset evaluation metrics. For example, calculate the matching accuracy rate, that is, the ratio of the number of correctly matched data items to the total number of matched data items; check the matching integrity, that is, whether the ratio of the successfully matched data items in the target data items meets the requirements. Specifically, determining whether the matching result meets the preset matching conditions may include: determining whether the matching result meets the preset matching accuracy rate condition and the preset matching integrity condition; wherein, the preset matching accuracy rate condition is the condition of the ratio of the number of correctly matched data items to the total number of matched data items; the preset matching integrity condition is the condition of the ratio of the successfully matched data items in the target data items. Evaluating the preliminary matching result can improve the accuracy rate and integrity of the matching.

[0056] Step S14: If the matching result does not meet the preset matching conditions, optimize the matching result through a preset optimization method to obtain a new matching result, and jump to the step of determining whether the matching result meets the preset matching conditions.

[0057] In this embodiment, if the matching result does not meet the preset matching conditions, the matching result is optimized according to the reason for the failure. For example, in a specific implementation, if a large number of data items fail to match due to a low name similarity score, a name synonym library is added to improve the matching effect; if there is a problem with the data type conversion matching, the conversion conditions or coefficients in the data type compatibility rules can be modified. Specifically, optimizing the matching result through a preset optimization method to obtain a new matching result may include: correspondingly adjusting the preset matching rules based on the matching result, and re-matching the metadata information with the target data items in the preset matching database according to the adjusted preset matching rules to obtain a new matching result. That is, adjust the name similarity, data type compatibility, and data format consistency construction methods of the data items in the preset matching rules according to the reason for the matching failure, and after the adjustment is completed, restart the automatic matching module for matching operations until the matching result meets the requirements. By adjusting the preset matching rules, when there are data items that cannot be fully matched, recommended matching items can be selected to meet the data collection requirements.

[0058] Step S15: If the matching result meets the preset matching conditions, collect the corresponding data from the data source, and perform processing and storage operations on the collected data based on the data feature requirements.

[0059] In this embodiment, according to the matching result, data is extracted from the data source by using a data collection method corresponding to the data source type of the data source, and after operations such as data type conversion, data cleaning, and data integration are performed according to the requirements of the target data storage structure, the data is stored at the target location. Specifically, the step of collecting corresponding data from the data source according to the matching result and performing processing and storage operations on the collected data based on the data feature requirements may include: collecting corresponding data from the data source by using a data collection method corresponding to the data source type of the data source according to the matching result to obtain initial collected data; performing data type conversion, data cleaning, and data integration operations on the initial collected data based on the data feature requirements to obtain target collected data, and storing the target collected data in a preset storage space. Among them, for a database data source, a data collection query statement is constructed for data collection; for a file data source, a data reading logic is used to collect data from the data source. That is, different data collection methods are adopted for different types of data sources, which can improve the efficiency of data collection, and increase the accuracy and integrity of data collection.

[0060] As can be seen from the above, in the embodiment of the present application, by parsing the storage structure of the target data to obtain the target data item and the data feature requirements corresponding to the target data item, and automatically matching the system data item through a preset matching database, for data items that cannot be fully matched, recommended matching items can be given. It can significantly reduce the manual workload in the data collection process, improve the efficiency and accuracy of data item matching, so as to better meet the requirements of large-scale data collection and processing, and provide a high-quality data basis for subsequent data mining, data analysis, etc. Thus, the system data item is automatically matched during the data collection process, reducing the repetitive work of manual matching.

[0061] See Figure 2 As shown, the embodiment of the present application provides a specific data collection and processing method based on data item matching, including:

[0062] In this embodiment, the data source identification module first analyzes the input data source connection information to determine the data source type. Among them, the data source type includes database data sources such as MySQL and Oracle, file data sources such as CSV, XML, and JSON files, and network data sources such as API interface data. Then, the metadata information is obtained through the corresponding data source access technology. For database data sources, the metadata query statements of the database system can be used to obtain table structures, field information, etc.; for file data sources, the metadata can be obtained by parsing the file header or predefined file format specifications. For network data sources, an attempt is made to split the received data using known delimiters to obtain the corresponding key-value pairs. Selecting the corresponding information acquisition method according to the data source type can ensure the integrity of the metadata information and improve the efficiency of metadata acquisition. For example, the reporting format of a Syslog protocol data source is:

[0063] <10>Oct 10 15:00:00 myauthserver.example.com authapp

[4711] : user=johndoe;ip=192.168.1.100;result=success

[0064] Among them, after removing the part related to the Syslog protocol, the remaining message body part is: user=johndoe;ip=192.168.1.100;result=success

[0065] Using ";" for attribute splitting and "=" for key-value pair splitting, the value of the user attribute is extracted as johndoe, the value of the ip attribute is 192.168.1.100, and the value of the result attribute is success.

[0066] Furthermore, in this embodiment, the target model parsing module reads the definition file or configuration information of the target data storage structure, extracts the required data items and their detailed feature requirements therefrom, and stores these features in the matching library. For example, for the login log data source, the target model includes UserName attribute, SrcIP attribute, and LoginResult attribute.

[0067] In this embodiment, after receiving the source data and the matching library data item information, the automatic matching module first performs fuzzy matching based on the data item name. It converts the data item name into a vector through Word2Vec and calculates the name similarity score using cosine similarity. For example, UserName is closest to user, ip is closest to SrcIP, and result is closest to LoginResult. Then it checks whether the data types are compatible. For example, numeric data and string data generally do not match, but in some specific scenarios, conversion matching can be performed and a certain matching penalty coefficient can be given. It then checks whether the data formats are consistent, such as whether the date formats are the same. Considering these factors, the matching degree score between each data source data item and the target data item is calculated according to the set weighted algorithm. Data item pairs with scores exceeding the set threshold are regarded as initially successfully matched. Through the preset matching rules, data item matching can be automatically performed from the preset matching database, reducing the repetitive work of manual matching.

[0068] In this embodiment, the matching result evaluation module evaluates the initial matching result according to the preset evaluation indicators. For example, it calculates the matching accuracy rate, that is, the proportion of the number of correctly matched data items to the total number of matched data items; it checks the matching integrity, that is, whether the proportion of successfully matched data items in the target data items meets the requirements. If the matching accuracy rate is lower than the set accuracy threshold or the matching integrity does not meet the standard, the optimization and adjustment module is triggered. By evaluating the initial matching result, the matching accuracy and integrity can be improved.

[0069] In this embodiment, the optimization and adjustment module analyzes the data item pairs that failed to match and finds out the possible reasons for the matching failure. For example, if a large number of data items fail to match due to low name similarity scores, manual adjustment is made to increase the name synonym library to improve the matching effect. For example, the data item name user is added as an alias for the data item UserName. If there are problems with data type conversion matching, the conversion conditions or coefficients in the data type compatibility rules can be modified. After the adjustment is completed, the automatic matching module is restarted for matching operations until the matching result meets the requirements. By optimizing and adjusting the preset matching rules, when there are data items that cannot be fully matched, recommended matching items can be screened out to meet the needs of data collection.

[0070] In this embodiment, the data collection execution module performs data collection for the database data source by constructing a data collection query statement, and performs data collection from the data source using the data reading logic for the file data source. After performing operations such as data type conversion, data cleaning, and data integration according to the requirements of the target data storage structure or analysis model, the data is stored in the target location.

[0071] As can be seen from the above, in the embodiments of the present application, by parsing the storage structure of the target data to obtain the target data items and the data feature requirements corresponding to the target data items, and completing the automatic matching of system data items through a preset matching database, for data items that cannot be fully matched, recommended matching items can be given. It can significantly reduce the manual workload in the data collection process, improve the efficiency and accuracy of data item matching, so as to better meet the requirements of large-scale data collection and processing, and provide a high-quality data basis for subsequent data mining, data analysis, etc. Thus, the automatic matching of system data items is realized in the data collection process, reducing the repetitive work of manual matching.

[0072] Correspondingly, as shown in Figure 4 the embodiments of the present application disclose a data collection and processing device based on data item matching, including:

[0073] A data source identification module 11, configured to identify the data source type of the data to be collected and obtain the metadata information of the data source; the metadata information includes data item name, data type, data format, and data length;

[0074] A target data parsing module 12, configured to perform a parsing operation on the storage structure of the target data to obtain the target data items and the data feature requirements corresponding to the target data items, and store the target data items and the data feature requirements in a preset matching database;

[0075] A data item matching module 13, configured to match the metadata information with the target data items in the preset matching database to obtain a matching result, and determine whether the matching result meets a preset matching condition;

[0076] An optimization and adjustment module 14, configured to, if the matching result does not meet the preset matching condition, optimize the matching result through a preset optimization method to obtain a new matching result, and jump to the step of determining whether the matching result meets the preset matching condition;

[0077] A data collection execution module 15, configured to, if the matching result meets the preset matching condition, collect corresponding data from the data source according to the matching result, and perform processing and storage operations on the collected data based on the data feature requirements.

[0078] As can be seen from the above, in the embodiments of the present application, by parsing the storage structure of the target data to obtain the target data items and the data feature requirements corresponding to the target data items, and completing the automatic matching of system data items through a preset matching database, for data items that cannot be fully matched, recommended matching items can be given. Thus, the automatic matching of system data items is realized in the data collection process, reducing the repetitive work of manual matching.

[0079] In some specific embodiments, the data source identification module 11 may specifically include:

[0080] A data source identification unit, configured to identify the data source type of the data to be collected through the connection information of the data source, and select a corresponding information acquisition method according to the data source type to acquire the metadata information of the data source; wherein, the data source type includes a database data source, a file data source, and a network data source.

[0081] In some specific embodiments, the data item matching module 13 may specifically include:

[0082] A word vector array conversion unit, configured to use the pre-trained Word2Vec to convert the data item names in the metadata information and the data item names of the target data items into word vector arrays, so as to obtain corresponding word vector arrays to be matched and target word vector arrays;

[0083] A matching data item determination unit, configured to calculate the cosine similarity between each word vector array to be matched and the target word vector array respectively, and determine the data item corresponding to the word vector array to be matched with the highest cosine similarity as the matching data item corresponding to the target data item;

[0084] A matching degree score calculation unit, configured to perform data type conversion and data format check on the target data item and the matching data item corresponding thereto, and calculate the matching degree score of the current matching data item and the target data item according to a preset matching rule; if the matching degree score exceeds a preset matching degree score threshold, a matching result indicating successful matching is obtained, and if the matching degree score does not exceed the preset matching degree score threshold, a matching result indicating failed matching is obtained.

[0085] A preset matching rule construction unit, configured to construct the preset matching rule based on the name similarity, data type compatibility, and data format consistency of the data items;

[0086] A matching result judgment unit, configured to judge whether the matching result meets a preset matching accuracy condition and a preset matching integrity condition; wherein, the preset matching accuracy condition is a condition for the proportion of the number of correctly matched data items in the total number of matched data items; the preset matching integrity condition is a condition for the proportion of the data items successfully matched in the target data items.

[0087] In some specific embodiments, the optimization and adjustment module 14 may specifically include:

[0088] A preset matching rule adjustment unit, configured to perform corresponding adjustment on the preset matching rule based on the matching result, and re-match the metadata information with the target data items in the preset matching database according to the adjusted preset matching rule to obtain a new matching result.

[0089] In some specific embodiments, the data acquisition execution module 15 may specifically include:

[0090] A data acquisition unit, configured to acquire corresponding data from the data source according to the matching result by using a data acquisition method corresponding to the data source type of the data source to obtain initial acquisition data;

[0091] A data storage unit, configured to perform data type conversion, data cleaning, and data integration operations on the initial acquisition data based on the data feature requirements to obtain target acquisition data, and store the target acquisition data in a preset storage space.

[0092] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 5 which is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure cannot be considered as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the data acquisition and processing method based on data item matching disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0093] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.

[0094] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.

[0095] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the data acquisition and processing method based on data item matching executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.

[0096] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the data acquisition and processing method based on data item matching disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0097] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0098] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0099] The steps of the method or algorithm described in combination with the embodiments disclosed in this document can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0100] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0101] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A data acquisition and processing method based on data item matching, characterized in that Including: Identifying the data source type of the data to be collected and obtaining the metadata information of the data source; The metadata information includes data item name, data type, data format, and data length; Performing a parsing operation on the storage structure of the target data to obtain the target data item and the data feature requirements corresponding to the target data item, and storing the target data item and the data feature requirements in a preset matching database; Matching the metadata information with the target data item in the preset matching database to obtain a matching result, and determining whether the matching result meets the preset matching condition; If the matching result does not meet the preset matching condition, optimizing the matching result through a preset optimization method to obtain a new matching result, and jumping to the step of determining whether the matching result meets the preset matching condition; If the matching result meets the preset matching condition, collecting the corresponding data from the data source according to the matching result, and performing processing and storage operations on the collected data based on the data feature requirements.

2. The data acquisition and processing method based on data item matching according to claim 1, characterized in that The identifying the data source type of the data to be collected and obtaining the metadata information of the data source includes: Identifying the data source type of the data to be collected through the connection information of the data source, and selecting a corresponding information acquisition method according to the data source type to obtain the metadata information of the data source; Wherein, the data source type includes database data source, file data source, and network data source.

3. The data acquisition and processing method based on data item matching according to claim 1, characterized in that The matching the metadata information with the target data item in the preset matching database to obtain a matching result includes: Using the pre-trained Word2Vec to convert the data item names in the metadata information and the data item names of the target data item into word vector arrays to obtain corresponding word vector arrays to be matched and target word vector arrays; Calculating the cosine similarity between each word vector array to be matched and the target word vector array respectively, and determining the data item corresponding to the word vector array to be matched with the highest cosine similarity as the matching data item corresponding to the target data item; Performing data type conversion and data format check on the target data item and the matching data item corresponding to it, and calculating the matching degree score of the current matching data item and the target data item according to the preset matching rule; If the matching degree score exceeds the preset matching degree score threshold, obtaining a matching result indicating successful matching, and if the matching degree score does not exceed the preset matching degree score threshold, obtaining a matching result indicating failed matching.

4. The data acquisition and processing method based on data item matching according to claim 3, characterized in that The preset matching rule is constructed based on the name similarity of data items, data type compatibility, and data format consistency.

5. The data acquisition and processing method based on data item matching according to claim 1, wherein The determining whether the matching result meets the preset matching condition includes: Determining whether the matching result meets the preset matching accuracy condition and the preset matching integrity condition; Wherein, the preset matching accuracy condition is the condition of the proportion of the number of correctly matched data items to the total number of matched data items; the preset matching integrity condition is the condition of the proportion of the data items successfully matched in the target data items.

6. The data acquisition and processing method based on data item matching according to claim 4, wherein Optimizing the matching result through a preset optimization method to obtain a new matching result, including: Correspondingly adjusting the preset matching rule based on the matching result, and rematching the metadata information with the target data items in the preset matching database according to the adjusted preset matching rule to obtain a new matching result.

7. The data acquisition and processing method based on data item matching according to any one of claims 1 to 6, characterized in that Collecting corresponding data from the data source according to the matching result, and performing processing and storage operations on the collected data based on the data feature requirements, including: Collecting corresponding data from the data source by using a data collection method corresponding to the data source type of the data source according to the matching result to obtain initial collected data; Performing data type conversion, data cleaning, and data integration operations on the initial collected data based on the data feature requirements to obtain target collected data, and storing the target collected data in a preset storage space.

8. A data acquisition and processing device based on data item matching, characterized in that, Including: A data source identification module, configured to identify the data source type of the data to be collected and obtain the metadata information of the data source; The metadata information includes data item name, data type, data format, and data length; A target data parsing module, configured to perform a parsing operation on the storage structure of the target data to obtain target data items and data feature requirements corresponding to the target data items, and store the target data items and the data feature requirements in a preset matching database; A data item matching module, configured to match the metadata information with the target data items in the preset matching database to obtain a matching result, and determine whether the matching result meets a preset matching condition; An optimization and adjustment module, configured to, if the matching result does not meet the preset matching condition, optimize the matching result through a preset optimization method to obtain a new matching result, and jump to the step of determining whether the matching result meets the preset matching condition; A data collection execution module, configured to, if the matching result meets the preset matching condition, collect corresponding data from the data source according to the matching result, and perform processing and storage operations on the collected data based on the data feature requirements.

9. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the data collection and processing method based on data item matching according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program, wherein the computer program, when executed by the processor, implements the data collection and processing method based on data item matching according to any one of claims 1 to 7.