A data management system based on data fusion
By filtering and supplementing data at the acquisition layer, integrating data at the processing layer, and adjusting data at the analysis layer within the data management system, the problem of low efficiency in discrete data processing is solved, achieving efficient data acquisition and integration.
Patent Information
- Application Number
- CN202410118672.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2044-01-29
AI Technical Summary
Existing technologies fail to effectively process the discrete data output from the processing layer, affecting the adjustment of operating parameters in the acquisition layer and resulting in low data acquisition efficiency.
The data acquisition layer obtains data and filters out redundancy and fills in missing data. The processing layer integrates data packets. The analysis layer performs secondary cleaning and merging of discrete data or adjusts the parameters of the acquisition layer. The application layer implements business requirements.
It improved data acquisition efficiency, reduced reliance on IT personnel, shortened response time, and enhanced data trust levels and integration quality.
Smart Images

Figure CN118013455B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and in particular to a data management system based on data fusion. Background Technology
[0002] Existing technologies utilize data storage and processing platforms such as big data platforms, data warehouses, and massive near real-time platforms. However, current data is still stored in a decentralized, thematic manner, and upper-layer data applications are still conducted in a siloed manner. There is a lack of a middleware data mart between the underlying basic data and upper-layer data applications, resulting in a lack of flexibility in application services, long response cycles, and difficulty in deeply mining the value of data.
[0003] Currently, there is no unified data resource management mechanism for enterprise-level data resources, the technical support is weak, and the data resources are relatively scattered. Effective management and comprehensive utilization of enterprise-level data resources are urgently needed.
[0004] Chinese Patent Publication No. CN111897644A discloses a multi-dimensional network data fusion and matching method, which includes decomposing and transforming fusion matching conditions to form a multi-dimensional matching pattern vector; converting each fusion matching rule into one piece of information in multiple matching pattern vectors, and caching the ID of the fusion matching rule in an array for fast access via array index; parsing and processing network data to extract quintuples, obtain application layer load, and parse and extract private protocol headers packaged according to custom specifications to obtain attribute information of different dimensions; and performing matching processing on the attribute information of different dimensions according to the corresponding matching conditions to form corresponding result vectors. It is evident that the prior art has the following problems: it does not consider processing the discrete data output by the processing layer, nor does it consider adjusting the operating parameters of the acquisition layer according to the specific situation of the discrete data, thus affecting the matching efficiency of the processing layer and consequently affecting the data acquisition efficiency of the application layer. Summary of the Invention
[0005] To address this, the present invention provides a data management system based on data fusion, which overcomes the problem in the prior art that does not consider processing the discrete data output by the processing layer, nor does it consider adjusting the operating parameters of the acquisition layer according to the specific situation of the discrete data, thus affecting the matching efficiency of the processing layer and consequently affecting the data acquisition efficiency of the application layer.
[0006] To achieve the above objectives, the present invention provides a data management system based on data fusion, comprising:
[0007] The acquisition layer is used to acquire data, filter out redundant data, and supplement missing data.
[0008] A processing layer, which is connected to the acquisition layer, is used to integrate the data output by the acquisition layer to output several data packets and discrete data that does not match each data packet.
[0009] An analysis layer, which is connected to the acquisition layer and the processing layer respectively, is used to determine whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data output by the processing layer, and when it is determined that the operating parameters of the acquisition layer do not meet the preset standard, the processing layer determines whether to perform secondary cleaning and merging of the data based on the discrete data output by the processing layer, or to adjust the filtering parameters of the acquisition layer.
[0010] Application layer: It is connected to the analysis layer and is used to realize specific business requirements and application scenarios.
[0011] Furthermore, the analysis layer determines whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of the discrete data output by the processing layer. When it is determined that the operating parameters of the acquisition layer do not meet the preset standard, the discrete data is re-merged according to the number of each data packet, or the processing method for the acquisition layer is determined according to the specific distribution of blank values and abnormal data in the discrete data.
[0012] Furthermore, the analysis layer obtains the original data through lineage analysis for the discrete data, and performs text rule filling based on the missing values of the original data. The analysis layer lists the filling results for a preset number of domains, and lists the filling results for a preset number of domains. The analysis layer matches each filling result with each data packet to add the discrete data to the data packet with the highest correlation.
[0013] The analysis layer has several matching adjustment methods for a preset enumeration number based on the number of data packets, and the adjustment range of each matching adjustment method for the preset enumeration number is different.
[0014] Furthermore, the analysis layer compares the number of blank values in the discrete data with the number of abnormal data to determine the acquisition processing method for the acquisition layer based on the comparison result. The acquisition processing method includes adjusting the preset area quantity to a corresponding value based on the number of abnormal data, or lowering the screening threshold for redundant data of the acquisition layer to a corresponding value based on the abnormal proportion of the number of blank values to the number of abnormal data. Here, the number of abnormal data is the difference between the total number of discrete data and the number of blank values.
[0015] Furthermore, the analysis layer has several redundancy adjustment methods for filtering redundancy data based on the abnormal proportion of the number of blank values to the number of abnormal data, and the adjustment range of each redundancy adjustment method for the filtering threshold is different.
[0016] After adjusting the screening threshold based on the proportion of abnormal data, the acquisition layer reprocesses the data based on the adjusted parameters. The analysis layer determines whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data re-output by the processing layer. Furthermore, under the condition of adjusting the screening threshold for redundant data in a secondary judgment, the preset number of fields is adjusted to the corresponding value based on the number of abnormal data.
[0017] Furthermore, the analysis layer determines the processing method for the processing layer based on the evaluation parameters of the discrete data. The processing methods include: re-determining new data packets based on the character values of each discrete data to repackage the data; or, determining the processing method for the acquisition layer a second time based on the number of abnormal data; or, lowering the screening threshold to the corresponding value based on the parameter difference between the evaluation parameter and the first preset evaluation parameter.
[0018] The analysis layer calculates the proportion of the length of a blank value in a single remaining data in the discrete data to the total length of the remaining data. The evaluation parameter is the average of the proportions of each length in the discrete data.
[0019] The discrete data includes several remaining data points.
[0020] Furthermore, the analysis layer maps the character values of each discrete data into nodes on a coordinate grid, with the horizontal axis representing the character value. For the same character value, it marks the same horizontal axis coordinate upwards along the vertical axis. The analysis layer delineates several domains based on the clustering of each node, and then packages the data corresponding to each node within a single domain to generate a new data packet.
[0021] Furthermore, the analysis layer has several data processing methods for filtering thresholds of the acquisition layer based on the number of abnormal data, and the adjustment range of each data processing method for the filtering threshold is different.
[0022] Furthermore, the analysis layer has several adjustment methods for the screening threshold based on the parameter difference between the evaluation parameter and the first preset evaluation parameter, and the adjustment range of each adjustment method for the screening threshold is different.
[0023] Furthermore, the analysis layer has several domain adjustment methods for the preset domain number based on the number of abnormal data, and the adjustment range of each domain adjustment method for the preset domain number is different;
[0024] After the analysis layer adjusts the number of abnormal data to the preset number of regions, the acquisition layer reprocesses the data based on the adjusted parameters. The analysis layer determines whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data re-output by the processing layer. Under the condition of adjusting the number of preset regions in the second determination, the redundancy data screening threshold is adjusted to the corresponding value according to the abnormal proportion of blank values to abnormal data.
[0025] Compared with existing technologies,
[0026] The acquisition layer obtains data and performs preliminary processing, namely filtering out redundant data and supplementing missing data, to ensure that the data formats are consistent. This allows the processing layer to classify and package the pre-processed data into several data packets. The analysis layer processes the packaged data packets and discrete data that do not match the data packets to further repackage the discrete data. By further optimizing the acquisition layer's algorithm based on the actual operation of the processing layer, the efficiency of data acquisition by the application layer is greatly improved.
[0027] Furthermore, the system determines whether to adjust the operating parameters of the acquisition layer based on the total number of bytes of discrete data. When the total number of bytes is small, only a small amount of discrete data needs to be re-matched with the data packets. A lineage analysis is performed on the discrete data to trace its source and the processing it has undergone, ensuring a higher level of data trust. The original data is then filled into multiple domains to select the filling results, which are then repackaged. The analysis layer determines the number of matching results generated for each domain based on the number of data packets. The number of data packets is proportional to the number of matching results generated for each domain, allowing for more targeted matching with each data packet. This effectively reduces the reliance on IT personnel for data processing maintenance and significantly improves the efficiency of data acquisition at the application layer.
[0028] Furthermore, when the total number of bytes is large, the actual situation of abnormal data and blank values in the discrete data is obtained. When there are too many abnormal values, it is determined that the number of fields generated when generating filling results for missing data is too low, resulting in a large number of data failing to match successfully. Therefore, the data filling of the analysis layer is corrected. The number of fields generated is determined based on the number of abnormal data, i.e., the preset number of fields. The number of abnormal data is proportional to the preset number of fields. In order to generate multi-field matching results according to the actual matching situation, the dependence of IT technicians on data processing maintenance is effectively reduced, and the efficiency of data acquisition by the application layer is greatly improved.
[0029] Furthermore, when the total number of bytes is large and there are too many blank values, it is determined that the acquisition layer has filtered out too much data during the process of filtering redundant data and replaced the filtered data with blank values, resulting in too many blank data failing to match the data packets. Therefore, the threshold for data filtering is lowered to allow more data to pass through the filtering process, and the data that passes through is filled in to reduce the number of blank values. This allows the analysis layer to make specific adjustments to the data processing based on the actual data, enabling the processing layer to better integrate the data and further greatly improving the efficiency of data acquisition by the application layer.
[0030] Furthermore, when the total number of bytes is too large, the processing method for the processing layer is determined based on the proportion of blank values in each data entry, i.e., the evaluation parameter of the discrete data. When the evaluation parameter is small, it is determined that there is a large amount of data with mismatches, so it is determined to generate several new data packets to repackage the discrete data, and the packaging of each new data packet is based on the character value of each data. When the evaluation parameter is large, it is determined that there are a large number of blank values. In this case, due to the excessive number of blank values, the threshold for data filtering is lowered, allowing more data to pass through the filtering, and the data that passes through is filled to reduce the number of blank values. This enables the processing layer to better integrate the data, further greatly improving the efficiency of data acquisition by the application layer and significantly reducing the response time for data acquisition.
[0031] Furthermore, structured data extracted from information resources to describe their characteristics and content, such as topic names, versions, related descriptions, and search points, are used to organize, describe, retrieve, store, and manage information and knowledge resources. Data is a crucial component of the application layer, while the acquisition layer is the core part for building, managing, maintaining, and using the application layer. The acquisition layer mainly includes: dimension & metric extraction, creating new grouped columns, value mapping, missing value imputation, column splitting, whitespace removal, type conversion, date expressions, data range definition, and data feature values. This involves filtering out redundant data by masking unnecessary fields and supplementing missing data; data governance simplifies modeling and automatically identifies problems in the data, helping users adjust and customize business datasets as needed, ultimately obtaining the required data, driving powerful visualization capabilities, and enabling faster discovery of new insights. Attached Figure Description
[0032] Figure 1 This is a structural block diagram of a data management system based on data fusion, as described in an embodiment of the present invention.
[0033] Figure 2 This is a flowchart illustrating the data determination method of the analysis layer in this embodiment of the invention, which determines whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes.
[0034] Figure 3 This is a flowchart illustrating the process by which the analysis layer determines the matching adjustment method for a preset enumeration number based on the number of data packets, according to an embodiment of the present invention.
[0035] Figure 4 This is a flowchart illustrating the process by which the analysis layer determines the acquisition and processing method for the acquisition layer based on the comparison between the number of blank values and the number of abnormal data in the discrete data, as described in this embodiment of the invention. Detailed Implementation
[0036] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0037] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0038] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0039] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0040] Please see Figure 1 , Figure 2 , Figure 3 as well as Figure 4 The diagrams shown are: a structural block diagram of a data management system based on data fusion according to an embodiment of the present invention; a flowchart of a data judgment method in which the analysis layer determines whether the operating parameters of the acquisition layer meet preset standards based on the total number of bytes; a flowchart of a matching adjustment method for a preset enumeration number in which the analysis layer determines the number of data packets; and a flowchart of a acquisition processing method for the acquisition layer in which the analysis layer determines the acquisition processing method for the acquisition layer based on the comparison result of the number of blank values and the number of abnormal data in the discrete data. An embodiment of the present invention provides a data management system based on data fusion, comprising:
[0041] The acquisition layer is used to acquire data, filter out redundant data, and supplement missing data.
[0042] A processing layer, which is connected to the acquisition layer, is used to integrate the data output by the acquisition layer to output several data packets and discrete data that does not match each data packet.
[0043] The analysis layer processes and transforms data through various operations to enable subsequent analysis and applications. The goal of the processing layer is to provide high-quality, consistent, and usable data to meet the needs of the subsequent application layer.
[0044] An analysis layer, which is connected to the acquisition layer and the processing layer respectively, is used to determine whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data output by the processing layer, and when it is determined that the operating parameters of the acquisition layer do not meet the preset standard, the processing layer determines whether to perform secondary cleaning and merging of the data based on the discrete data output by the processing layer, or to adjust the filtering parameters of the acquisition layer.
[0045] Application layer: It is connected to the analysis layer and is used to realize specific business requirements and application scenarios.
[0046] Specifically, the analysis layer determines whether the operating parameters of the acquisition layer meet a preset standard data judgment method based on the total number of bytes of the discrete data output by the processing layer, wherein:
[0047] The first data determination method is that the analysis layer determines that the operating parameters of the acquisition layer meet the preset standard, and determines that the integration of each data is completed. The analysis layer then outputs the integrated data packet of each data to the application layer. The first data determination method satisfies that the total number of bytes is less than or equal to the first preset total number of bytes.
[0048] The second data determination method is that the analysis layer determines that the operating parameters of the acquisition layer do not meet the preset standard, and re-merges the discrete data according to the number of each data packet; the second data determination method satisfies that the total number of bytes is less than or equal to the second preset total number of bytes and greater than the first preset total number of bytes.
[0049] The third data processing method is that the analysis layer determines that the operating parameters of the acquisition layer do not meet the preset standard, and determines the processing method for the acquisition layer based on the comparison result of the number of blank values and the number of abnormal data in the discrete data; the third data processing method satisfies that the total number of bytes is less than or equal to the third preset total number of bytes and greater than the second preset total number of bytes, and the second preset total number of bytes is less than the third preset total number of bytes.
[0050] The fourth data processing method involves the analysis layer determining that the operating parameters of the acquisition layer do not meet the preset standard, and determining the processing method for the acquisition layer based on the proportion of the length of blank values in each remaining data of the discrete data to the total length of a single remaining data character; the fourth data processing method satisfies that the total number of bytes is greater than the third preset total number of bytes;
[0051] The discrete data includes several remaining data points.
[0052] The decision to adjust the acquisition layer's operating parameters is based on the total number of bytes of discrete data. When the total number of bytes is small, only a small amount of discrete data needs to be re-matched with data packets. A lineage analysis is performed on the discrete data to trace its origins and the processing it has undergone, ensuring a higher level of data trust. The original data is then filled into multiple domains for secondary selection of the filling results, which are then repackaged. The analysis layer determines the number of matching results generated for each domain based on the number of data packets. The number of data packets is proportional to the number of matching results generated for each domain, allowing for more targeted matching with each data packet. This effectively reduces the reliance on IT personnel for data processing maintenance and significantly improves the application layer's efficiency in acquiring data.
[0053] Specifically, the analysis layer determines the data source of discrete data through lineage analysis and obtains the original data. Based on the missing values of the original data, it performs text rule filling and lists the filling results for a preset number of domains, with each domain listing a preset number of filling results. The analysis layer matches each filling result with each data packet to add the discrete data to the data packet with the highest correlation. The analysis layer marks the corresponding domains that have been matched as the priority domains for the processing layer for the corresponding server, so that when the processing layer processes data for that server, it will prioritize using the priority domains to fill the missing values.
[0054] The analysis layer determines the matching adjustment method for the preset enumeration number based on the number of data packets, where:
[0055] The first matching adjustment method involves the analysis layer using a first preset matching adjustment coefficient to adjust the preset enumeration number to a corresponding value; the first matching adjustment method satisfies the condition that the number of data packets is less than or equal to the first preset number.
[0056] The second matching adjustment method is that the analysis layer uses a second preset matching adjustment coefficient to adjust the preset enumeration number to the corresponding value; the second matching adjustment method satisfies that the number of data packets is less than or equal to the second preset number and greater than the first preset number, and the first preset number is less than the second preset number;
[0057] The third matching adjustment method is that the analysis layer uses a third preset matching adjustment coefficient to adjust the preset enumeration number to the corresponding value; the third matching adjustment method satisfies that the number of data packets is greater than the second preset number.
[0058] Specifically, the analysis layer compares the number of blank values with the number of outliers in the discrete data to determine the acquisition processing method for the acquisition layer based on the comparison result, wherein:
[0059] The first data acquisition and processing method involves the analysis layer adjusting the number of preset regions to a corresponding value based on the number of abnormal data; the first data acquisition and processing method satisfies the condition that the number of blank values is less than the number of abnormal data.
[0060] The second data acquisition and processing method involves the analysis layer lowering the filtering threshold for redundant data to a corresponding value based on the proportion of blank values to the abnormal data; the second data acquisition and processing method satisfies the condition that the number of blank values is greater than the number of abnormal data.
[0061] The number of outliers is the difference between the total number of discrete data and the number of blank values.
[0062] When the total number of bytes is large, the actual situation of abnormal data and blank values in the discrete data is obtained. When there are too many abnormal values, it is determined that the number of fields generated when generating filling results for missing data is too low, resulting in a large number of data failing to match successfully. Therefore, the data filling of the analysis layer is corrected. The number of fields generated is determined based on the number of abnormal data, i.e., the preset number of fields. The number of abnormal data is proportional to the preset number of fields. In order to generate multi-field matching results according to the actual matching situation, the dependence of IT technicians on data processing maintenance is effectively reduced, and the efficiency of data acquisition by the application layer is greatly improved.
[0063] Specifically, the analysis layer determines the redundancy adjustment method for the filtering threshold of redundant data based on the anomaly ratio of the number of blank values to the number of abnormal data, wherein:
[0064] The first redundancy adjustment method involves the analysis layer using a first preset redundancy adjustment coefficient to adjust the screening threshold to a corresponding value; the first redundancy adjustment method satisfies the condition that the abnormal proportion is less than or equal to the first preset abnormal proportion.
[0065] The second redundancy adjustment method is that the analysis layer uses a second preset redundancy adjustment coefficient to adjust the screening threshold to the corresponding value; the second redundancy adjustment method satisfies that the abnormal proportion is less than or equal to the second preset abnormal proportion and greater than the first preset abnormal proportion, and the first preset abnormal proportion is less than the second preset abnormal proportion.
[0066] The third redundancy adjustment method involves the analysis layer using a third preset redundancy adjustment coefficient to adjust the screening threshold to a corresponding value; the third redundancy adjustment method satisfies the condition that the abnormal proportion is greater than the second preset abnormal proportion.
[0067] After adjusting the screening threshold based on the proportion of abnormal data, the acquisition layer reprocesses the data based on the adjusted parameters. The analysis layer determines whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data re-output by the processing layer. Furthermore, under the condition of adjusting the screening threshold for redundant data in a secondary judgment, the preset number of fields is adjusted to the corresponding value based on the number of abnormal data.
[0068] Furthermore, when the total number of bytes is large and there are too many blank values, it is determined that the acquisition layer has filtered out too much data during the process of filtering redundant data and replaced the filtered data with blank values, resulting in too many blank data failing to match the data packets. Therefore, the threshold for data filtering is lowered to allow more data to pass through the filtering process, and the data that passes through is filled in to reduce the number of blank values. This allows the analysis layer to make specific adjustments to the data processing based on the actual data, enabling the processing layer to better integrate the data and further greatly improving the efficiency of data acquisition by the application layer.
[0069] Specifically, under the fourth data processing method, the analysis layer determines the processing method for the acquisition layer based on the evaluation parameters of the discrete data, wherein:
[0070] The first processing method is that the analysis layer re-determines new data packets based on the character values of each discrete data to repackage the data; the first processing method is that the evaluation parameter is less than or equal to the first preset evaluation parameter.
[0071] The second processing method is that the analysis layer determines the processing method for the acquisition layer a second time based on the number of abnormal data; the second processing method is that the evaluation parameter is less than or equal to the second preset evaluation parameter and greater than the first preset evaluation parameter, and the first preset evaluation parameter is less than the second preset evaluation parameter.
[0072] The third processing method involves the analysis layer lowering the screening threshold to a corresponding value based on the parameter difference between the evaluation parameter and the first preset evaluation parameter; the third processing method satisfies the condition that the evaluation parameter is greater than the second preset evaluation parameter.
[0073] The analysis layer calculates the proportion of the length of a blank value in a single remaining data in the discrete data to the total length of the remaining data. The evaluation parameter is the average value of the proportions of each length in the discrete data.
[0074] When the total number of bytes is too large, the processing method for the processing layer is determined based on the proportion of blank values in each data item, i.e., the evaluation parameter of discrete data. When the evaluation parameter is small, it is determined that there is a large amount of data with mismatches. Therefore, it is determined to generate several new data packets to repackage the discrete data, and the packaging of each new data packet is based on the character value of each data item. When the evaluation parameter is large, it is determined that there are a lot of blank values. In this case, due to the excessive number of blank values, the threshold for data filtering is lowered to allow more data to pass through the filtering. Data filling is performed on the passed data to reduce the number of blank values, so that the processing layer can better integrate the data and further improve the efficiency of data acquisition by the application layer.
[0075] Specifically, the analysis layer maps the character values of each discrete data into nodes on a coordinate grid, with the horizontal axis representing the character value. For the same character value, it marks the same horizontal axis coordinate upwards along the vertical axis. The analysis layer delineates several domains based on the clustering of each node, and then packages the data corresponding to each node within a single domain to generate a new data packet.
[0076] Specifically, under the second processing method, the analysis layer determines the data processing method for the filtering threshold of the acquisition layer based on the number of abnormal data, wherein:
[0077] The first data processing method involves the analysis layer lowering the screening threshold to a corresponding value based on the parameter difference between the evaluation parameter and the first preset evaluation parameter; the first data processing method satisfies the condition that the number of abnormal data is less than or equal to the second preset number of abnormal data.
[0078] The second data processing method involves the analysis layer re-determining new data based on the character values of each discrete data to perform secondary packaging of the data; the second data processing method satisfies the requirement that the number of abnormal data is greater than the second preset number of abnormalities.
[0079] Specifically, the analysis layer determines the adjustment method for the screening threshold based on the parameter difference between the evaluation parameter and the first preset evaluation parameter, wherein:
[0080] The first adjustment method involves the analysis layer using a first preset adjustment coefficient to adjust the screening threshold to a corresponding value; the first adjustment method satisfies that the parameter difference is less than or equal to the first preset parameter difference.
[0081] The second adjustment method is that the analysis layer uses a second preset adjustment coefficient to adjust the screening threshold to the corresponding value; the second adjustment method satisfies that the parameter difference is less than or equal to the second preset parameter difference and greater than the first preset parameter difference, and the first preset parameter difference is less than the second preset parameter difference;
[0082] The third adjustment method is that the analysis layer uses a third preset adjustment coefficient to adjust the screening threshold to the corresponding value; the third adjustment method satisfies that the parameter difference is greater than the second preset parameter difference.
[0083] Specifically, the analysis layer determines the domain adjustment method for the preset number of domains based on the amount of abnormal data, wherein:
[0084] The first domain adjustment method involves the analysis layer using a first preset domain adjustment coefficient to adjust the number of preset domains to a corresponding value; the first domain adjustment method satisfies the condition that the number of abnormal data is less than or equal to the first preset number of abnormalities.
[0085] The second domain adjustment method is that the analysis layer uses a second preset domain adjustment coefficient to adjust the number of preset domains to a corresponding value; the second domain adjustment method satisfies that the number of abnormal data is less than or equal to the second preset abnormal number and greater than the first preset abnormal number, and the first preset abnormal number is less than the second preset abnormal number.
[0086] The third domain adjustment method involves the analysis layer using a third preset domain adjustment coefficient to adjust the number of preset domains to a corresponding value; the third domain adjustment method satisfies the condition that the number of abnormal data is greater than the second preset number of abnormal data.
[0087] Under the condition that the analysis layer has adjusted the number of preset domains based on the number of abnormal data, the acquisition layer reprocesses the data based on the adjusted parameters. The analysis layer determines whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data re-output by the processing layer. Under the condition that the number of preset domains is adjusted in the second determination, the redundancy data screening threshold is adjusted to the corresponding value according to the abnormal proportion of blank values to abnormal data.
[0088] Structured data extracted from information resources to describe their characteristics and content (such as topic names, versions, related descriptions, including search points) is used to organize, describe, retrieve, store, and manage information and knowledge resources. Data is an important component of the application layer, while the acquisition layer is the core part for building, managing, maintaining, and using the application layer. The acquisition layer mainly includes: dimension & metric extraction, creating new grouped columns, value mapping, missing value imputation, column splitting, whitespace removal, type conversion, date expressions, data range definition, and data feature values. It filters out redundant data by masking unnecessary fields and supplements missing data; through data governance, modeling can be simplified and problems in the data can be automatically discovered, helping users adjust and customize business datasets as needed, ultimately obtaining the required data, driving powerful visualization capabilities, and discovering new insights faster.
[0089] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0090] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A data management system based on data fusion, characterized in that, include: The acquisition layer is used to acquire data, filter out redundant data, and supplement missing data. A processing layer, which is connected to the acquisition layer, is used to integrate the data output by the acquisition layer to output several data packets and discrete data that does not match each data packet. An analysis layer, which is connected to the acquisition layer and the processing layer respectively, is used to determine whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data output by the processing layer, and to adjust the filtering parameters of the acquisition layer when it is determined that the operating parameters of the acquisition layer do not meet the preset standard. The application layer, which is connected to the analysis layer, is used to realize specific business requirements and application scenarios; The analysis layer obtains the original data through lineage analysis for discrete data, and fills in the missing values of the original data using text rules. The analysis layer lists the filling results for a preset number of domains, and lists the filling results for a preset number of domains. The analysis layer matches each filling result with each data packet to add discrete data to the data packet with the highest correlation. The analysis layer has several matching adjustment methods for a preset enumeration number based on the number of data packets, and the adjustment range of each matching adjustment method for the preset enumeration number is different. The analysis layer maps the character values of each discrete data point onto a coordinate grid in the form of nodes. The horizontal axis of the coordinate grid represents the character value. For the same character value, it is marked sequentially upwards along the vertical axis of the horizontal axis. The analysis layer delineates several domains based on the clustering of each node. The analysis layer then packages the data corresponding to each node in a single domain to generate a new data packet.
2. The data management system based on data fusion according to claim 1, characterized in that, The analysis layer compares the number of blank values in the discrete data with the number of abnormal data, and determines the collection and processing method for the collection layer based on the comparison result. The collection and processing method includes adjusting the preset area quantity to a corresponding value based on the number of abnormal data, or lowering the screening threshold for redundant data of the collection layer to a corresponding value based on the abnormal proportion of the number of blank values to the number of abnormal data; wherein, the number of abnormal data is the difference between the total number of discrete data and the number of blank values.
3. The data management system based on data fusion according to claim 2, characterized in that, The analysis layer has several redundancy adjustment methods for filtering redundancy data based on the abnormal proportion of the number of blank values to the number of abnormal data, and the adjustment range of each redundancy adjustment method for the filtering threshold is different. After adjusting the screening threshold based on the proportion of abnormal data, the acquisition layer reprocesses the data based on the adjusted parameters. The analysis layer determines whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data re-output by the processing layer. Furthermore, under the condition of adjusting the screening threshold for redundant data in a secondary judgment, the preset number of fields is adjusted to the corresponding value based on the number of abnormal data.
4. The data management system based on data fusion according to claim 3, characterized in that, The analysis layer determines the processing method for the processing layer based on the evaluation parameters of discrete data. The processing methods include: re-determining new data packets based on the character values of each discrete data to repackage the data; or, determining the processing method for the acquisition layer a second time based on the number of abnormal data; or, adjusting the screening threshold to the corresponding value based on the parameter difference between the evaluation parameter and the first preset evaluation parameter. The analysis layer calculates the proportion of the length of a blank value in a single remaining data in the discrete data to the total length of the remaining data. The evaluation parameter is the average of the proportions of each length in the discrete data. The discrete data includes several remaining data points.
5. The data management system based on data fusion according to claim 4, characterized in that, The analysis layer has several data processing methods for filtering thresholds of the acquisition layer based on the number of abnormal data, and the adjustment range of each data processing method for the filtering threshold is different.
6. The data management system based on data fusion according to claim 5, characterized in that, The analysis layer has several adjustment methods for the screening threshold based on the parameter difference between the evaluation parameter and the first preset evaluation parameter, and the adjustment range of each adjustment method for the screening threshold is different.
7. The data management system based on data fusion according to claim 6, characterized in that, The analysis layer has several domain adjustment methods based on the number of abnormal data, and the adjustment range of each domain adjustment method for the number of preset domains is different. After the analysis layer adjusts the number of abnormal data to the preset number of regions, the acquisition layer reprocesses the data based on the adjusted parameters. The analysis layer determines whether the operating parameters of the acquisition layer meet the preset standard based on the total number of bytes of discrete data re-output by the processing layer. Under the condition of adjusting the number of preset regions in the second determination, the redundancy data screening threshold is adjusted to the corresponding value according to the abnormal proportion of blank values to abnormal data.
Citation Information
Patent Citations
Network data fusion matching method based on multiple dimensions
CN111897644A
Data screening method and device
CN110109938A
Data warehouse management system based on cloud native and storage and calculation separation
CN117149746A