Dynamic data mapping and adaptation methods and systems for heterogeneous data
By constructing candidate mapping graphs and optimizing the data framework, and combining the correlation strength of semantic tags and business attributes, the problem of inaccurate mapping in ETL tools in heterogeneous data processing is solved, realizing the flexibility and accuracy of data type matching, and adapting to the processing needs of complex or fuzzy business data types.
Patent Information
- Application Number
- CN202511461113.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-14
AI Technical Summary
In existing technologies, ETL tools lack accuracy when processing heterogeneous data and cannot effectively handle differences in the definition of the same concept among different business systems, resulting in inaccurate data mapping results.
By acquiring initial data, extracting core features, constructing candidate mapping graphs, optimizing the data framework, and dynamically adjusting the mapping process by combining the correlation strength of semantic tags and business attributes, time-series prediction reports are generated, achieving dynamic data adaptation.
It improves the flexibility and accuracy of data type matching, enhances the business adaptability of mapping results, and can better meet the needs of multi-source data integration and adapt to changes in data characteristics and business requirements.
Smart Images

Figure CN120950589B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data mapping technology, and in particular to a method and system for dynamic data mapping and adaptation of heterogeneous data. Background Technology
[0002] Currently, data has become a core resource driving business decisions and innovation, especially in industries such as finance and logistics that require the integration of multi-source data. The processing and application of heterogeneous data are therefore particularly crucial. Heterogeneous data refers to data from different systems, formats, or business backgrounds, such as structured databases, semi-structured log files, and unstructured text. Its integration and analysis can significantly improve operational efficiency and decision-making accuracy for enterprises.
[0003] In one existing technology, an ETL (Extract-Transform-Load) tool with predefined mapping rules is used to process heterogeneous data. Users establish the correspondence between source data and target data structure by dragging and dropping components (such as "field mapping" and "format conversion"). For data type differences (such as converting a time field from a string to a timestamp), a script or built-in function needs to be written to perform a forced conversion. The processed data is then loaded into the target database or data warehouse to generate a unified report.
[0004] However, ETL tools map data based solely on field names or data types. Different business systems may define the same concept differently, and the contextual meaning of data within a business scenario can lead to inaccuracies in the mapping results during practical applications. In summary, existing technologies suffer from inaccurate data mapping. Summary of the Invention
[0005] This invention provides a method and system for dynamic data mapping and adaptation of heterogeneous data to solve the problem of inaccuracy in data mapping in the prior art.
[0006] In a first aspect, to address the aforementioned technical problems, the present invention provides a method for dynamic data mapping and adaptation of heterogeneous data, comprising:
[0007] Acquire initial data, extract core features, and obtain an initial data set; the initial data set includes business attributes, type features, and semantic tags.
[0008] If the type feature in the initial data set has a matching degree with the preset type library that is lower than the preset type matching degree threshold, then the initial data set is regrouped to obtain a standard data set.
[0009] Based on the standard data set, the association strength between the semantic tag and the business attribute is calculated, and a candidate mapping map of attachment locations is constructed. The candidate set of attachment locations is obtained by optimizing the mapping map.
[0010] Based on the candidate set of attachment locations, the core features are integrated into the corresponding attachment locations to obtain the initial data framework;
[0011] The initial data frame is formatted to obtain a standard data frame;
[0012] Logical rules are generated based on the standard data framework and integrated with a preset analysis report template to obtain an analysis output report;
[0013] Based on the analysis output report, calculate the matching strength of the semantic tag and the business attribute in the candidate attachment location set. If the matching strength is lower than the preset matching strength threshold, optimize the standard data framework to obtain the optimized data framework.
[0014] The candidate mapping of the attachment location is updated based on the analysis output report, and a time series prediction report is obtained by combining the optimized data framework.
[0015] Secondly, the present invention provides a dynamic data mapping and adaptation system for heterogeneous data, comprising:
[0016] The data acquisition module is used to acquire initial data, extract core features, and obtain an initial data set; the initial data set includes business attributes, type features, and semantic tags.
[0017] The data grouping module is used to regroup the initial data group set to obtain a standard data group set if the matching degree between the type feature and the preset type library in the initial data group set is lower than the preset type matching degree threshold.
[0018] The data matching module is used to calculate the association strength between the semantic tag and the business attribute based on the standard data set, construct a candidate mapping map of attachment positions, and obtain a candidate set of attachment positions by optimizing the mapping map;
[0019] The data integration module is used to integrate the core features into the corresponding attachment positions based on the candidate attachment position set to obtain an initial data framework;
[0020] The data unification module is used to format the initial data framework to obtain a standard data framework.
[0021] The data generation module is used to generate logical rules based on the standard data framework and integrate them with a preset analysis report template to obtain an analysis output report.
[0022] The data verification module is used to calculate the matching strength of the semantic tag and the business attribute in the candidate set of attachment locations based on the analysis output report. If the matching strength is lower than the preset matching strength threshold, the standard data framework is optimized to obtain the optimized data framework.
[0023] The data prediction module is used to update the candidate mapping map of the attachment location based on the analysis output report, and to obtain a time-series prediction report by combining the optimized data framework.
[0024] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the dynamic data mapping and adaptation method for heterogeneous data as described in any one of the above.
[0025] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the dynamic data mapping and adaptation method for heterogeneous data as described in any one of the above.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] (1) By regrouping, the data grouping method is dynamically optimized, improving the flexibility and accuracy of data type matching, effectively solving the data processing problem caused by field type mismatch in existing ETL tools, especially when dealing with complex or ambiguous business data types;
[0028] (2) This invention introduces the correlation strength analysis of business attributes and semantic tags, constructs a candidate mapping graph, integrates semantic understanding into the data mapping process, greatly improves the business adaptability of the mapping results, avoids the problem of inaccurate mapping caused by ignoring data semantics and context, thereby improving the accuracy of matching and better meeting the needs of multi-source data integration.
[0029] (3) This invention has a dynamic optimization mechanism that verifies and optimizes the mapping candidate set based on the analysis output report, and continuously improves the data framework. At the same time, combined with time series prediction reports, it can adapt to changes in data characteristics or business needs in advance, making the entire data processing system more flexible and efficient. Attached Figure Description
[0030] Figure 1 This is a schematic diagram of the dynamic data mapping and adaptation method for heterogeneous data provided in the first embodiment of the present invention;
[0031] Figure 2This is a schematic diagram of the dynamic data mapping and adaptation system structure for heterogeneous data provided in the second embodiment of the present invention. Detailed Implementation
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Reference Figure 1 The first embodiment of the present invention provides a dynamic data mapping and adaptation system for heterogeneous data, including the following steps:
[0034] S11, Obtain initial data, extract core features, and obtain an initial data set; the initial data set includes business attributes, type features, and semantic tags;
[0035] S12, if the matching degree between the type feature and the preset type library in the initial data group set is lower than the preset type matching degree threshold, then the initial data group set is regrouped to obtain a standard data group set;
[0036] S13, Based on the standard data set, calculate the association strength between the semantic tag and the business attribute, construct a candidate mapping map of attachment positions, and obtain a candidate set of attachment positions by optimizing the mapping map;
[0037] S14, Based on the candidate set of attachment positions, integrate the core features into the corresponding attachment positions to obtain the initial data framework;
[0038] S15, the initial data frame is formatted to obtain a standard data frame;
[0039] S16, Generate logical rules according to the standard data framework and integrate them with a preset analysis report template to obtain an analysis output report;
[0040] S17. Based on the analysis output report, calculate the matching strength of the semantic tag and the business attribute in the candidate attachment location set. If the matching strength is lower than the preset matching strength threshold, optimize the standard data framework to obtain the optimized data framework.
[0041] S18, update the candidate mapping map of the attachment location according to the analysis output report, and obtain the time series prediction report by combining the optimized data framework.
[0042] In step S11, initial data is acquired, core features are extracted, and an initial data set is obtained; the initial data set includes business attributes, type features, and semantic tags, including:
[0043] Get initial data;
[0044] Type features, semantic labels, and business attributes are extracted from the initial data to obtain a structured dataset;
[0045] If the completeness of the type features in the structured dataset is higher than the preset type completeness threshold, the data is initially classified to obtain an initial set of classified data groups.
[0046] If the completeness of the type features in the structured dataset is lower than a preset type completeness threshold, then the missing features are filled in using the semantic labels to obtain the initial data set.
[0047] It should be noted that obtaining initial data involves multiple data sources, such as transaction records from online e-commerce platforms, sales data from online self-operated stores, and inventory management systems. Raw data from different systems and formats is collected, covering the entire business chain. This data can be collected through API interfaces (such as order APIs for retail systems and warehousing APIs for logistics systems), direct database connections (MySQL / Oracle structured libraries), and file parsing (log files, Excel reports, PDF reports). Type features, semantic tags, and business attributes are extracted from this initial data. Type features are identified using regular expressions (extracting the time "2025-10-01") and data type detection tools (such as Python's `type()` function). Semantic tags are extracted using NLP tools (such as BERT segmentation and keyword matching).
[0048] Based on a pre-established business dictionary mapping, business attributes are extracted from business scenarios through a precise one-to-one mapping of "fields-attributes". The completeness of type features in the structured dataset is verified. The completeness calculation rule is that the completeness of type features equals the number of "non-empty type features that meet business requirements" in a certain data point divided by the total number of "required type features" in that business scenario, multiplied by 100%. The pre-established business dictionary mapping collects fields related to core business attributes from various business systems, establishes a one-to-one mapping relationship of "field name - business attribute", and stores the sorted "field - business attribute" mapping relationship in the database, thus pre-establishing a business dictionary table. By testing the clustering effect under different completeness rates in historical data, when the completeness rate is ≥80%, the stability of the cluster centers (change in iteration <0.001) meets the standard, and the consistency of type features of samples within the category is high. Therefore, the type completeness threshold is set to 80%.
[0049] When the completeness exceeds a threshold (preset to 80%), a classification method combining a "business attribute dictionary + clustering algorithm (such as K-Means)" is used. Based on the "business dictionary mapping" results, the standardized type feature data is first grouped into major categories according to "business attributes" to reduce the computational load of K-Means. For each major category, K-Means clustering is used to split the data into smaller categories with clear business meanings based on "type feature similarity." Then, a clustering algorithm (based on type feature similarity) is used to further subdivide the smaller categories. K samples are randomly selected from the major category data as initial cluster centers. The "type feature similarity distance" between each sample and the K cluster centers is calculated, and the sample is assigned to the category of the nearest cluster center. Different distance formulas are used to calculate the distance between samples based on the standardized type of the type feature (numerical, time-based, One-Hot character). Similarity (the smaller the distance, the higher the similarity): Euclidean distance is used to calculate similarity for numerical / temporal features, and Manhattan distance is used to calculate similarity for One-Hot character features. The distance results of different types are weighted and summed (e.g., 0.4 for numerical, 0.3 for temporal, and 0.3 for character) to obtain the final similarity distance. For each category, the "type feature mean" of all samples in that category is calculated (mean for numerical / temporal features, mode for One-Hot character features) and used as the new cluster center. Samples are repeatedly assigned and cluster centers are updated until the change in cluster centers is less than a preset threshold (e.g., 0.001) or the number of iterations reaches the upper limit (e.g., 100 times), at which point the iteration stops. Each cluster corresponds to a sub-category of a business sub-scenario, forming a hierarchical classification of "large category → small category", resulting in the initial set of classified data groups.
[0050] When the completeness is below the threshold, a "semantic label-type feature" mapping library is established. For each semantic label, its corresponding "type feature type + specific field" is clearly defined. The rule is that one label can be associated with multiple features, but one feature must correspond to a specific label. If the data is missing a certain type feature (e.g., D003 is missing "payment time"), its semantic label is queried, the corresponding type feature is matched from the mapping library, and historical data is extracted from the historical database. The specific value is completed by combining the historical data (e.g., the payment time for "holiday transactions" is mostly 9:00-22:00. If there are user access logs, it can be further located to 14:00-15:00, and completed as "2025-10-01 14:30"), thus obtaining the initial data set.
[0051] In step S12, if the type feature in the initial data set has a matching degree with a preset type library that is lower than a preset type matching degree threshold, then the initial data set is regrouped to obtain a standard data set, including:
[0052] The initial data set is divided into multiple subsets to obtain the initial data set subsets;
[0053] When the type characteristics of the data group in the initial data group subset match the preset type library with a lower than the preset type matching threshold, the initial data group subset is grouped to obtain a standardized data group subset.
[0054] Extract semantic labels from the standardized data subset to obtain an initial semantic label sequence;
[0055] Analyze the correlation between the initial semantic tag sequence and the business attributes to obtain a highly correlated optimized semantic tag sequence;
[0056] By merging the optimized semantic label sequence and the standardized data set subset, a standard data set is obtained.
[0057] It should be noted that the data element's business affiliation is used as the basis for splitting. This means that large-scale, mixed-business data sets are divided into smaller business scenario units based on the core business scenario (such as transactions, inventory, user behavior, etc.). For each initial data subset, the matching degree with the type library is calculated using a weighted method: the basic type matching rate equals the number of standard basic types contained in the subset divided by the total number of basic types required by the type library, multiplied by 100%; the business feature matching rate equals the number of standard business features contained in the subset divided by the total number of business features required by the type library, multiplied by 100%; the total matching degree equals the product of the basic type matching rate multiplied by 0.6 and the product of the business feature matching rate multiplied by 0.4 (weights can be adjusted according to business needs). The pre-defined type library, based on business scenarios, includes two core requirements: basic type requirements: mandatory basic data types (such as numeric - payment amount, time - transaction time, character - order ID); business feature requirements: scenario-specific business type features (such as transaction types needing to include "payment method (enumerated type)" and "whether it's a promotion (Boolean type)"). If the type characteristics of a data group in the initial data subset match the preset type library with a lower than the preset type matching threshold (80%), the subset is regrouped to correct the type, ultimately resulting in a standardized data subset. The type matching threshold can be determined through historical data sample analysis and testing. Data analysis shows that when the threshold is 80%, the sample meets business requirements, and the type matching threshold is set to 80%.
[0058] The process involves iterating through each data element in the standardized data subset, extracting labeled semantic tags, and sorting them according to their impact on business attributes to form a semantic tag sequence. This process is repeated for all data elements in the subset, removing completely duplicate sequences to obtain an initial set of semantic tag sequences. By analyzing the correlation strength between the initial tag sequences and business attributes, highly correlated sequences are selected. Using the initial set of semantic tag sequences as the processing object, each sequence is treated as a transaction itemset. The semantic tags and corresponding business attributes contained in the sequence are identified, forming a transaction dataset to be analyzed. For the semantic tag sequences and business attributes in the dataset, two core indicators are calculated to quantify the correlation strength: support (the frequency of a semantic tag sequence and a business attribute co-occurring, i.e., the proportion of transactions containing both the sequence and the attribute to the total number of transactions); and confidence (the conditional probability of a corresponding business attribute appearing given a semantic tag sequence, i.e., the ratio of transactions containing both the sequence and the attribute to transactions containing only the sequence). The average of these two indicators is used as the correlation strength. Tag sequences with a correlation strength higher than a threshold (e.g., confidence ≥ 70%) are retained, while weakly correlated sequences are removed, resulting in the "optimized semantic tag sequence". The optimized semantic label sequence is bound to the original data elements in the standardized data set subset to form a complete structure of "data element-type feature-business attribute-optimized label sequence", thus obtaining the standard data set.
[0059] In step S13, based on the standard data set, the association strength between the semantic tag and the business attribute is calculated, and a candidate mapping map of attachment locations is constructed. The candidate set of attachment locations is obtained by optimizing the mapping map, including:
[0060] Get context variables;
[0061] Extract semantic tag sequences from the standard data set and calculate the correlation of the semantic tag sequences to obtain the correlation strength of the semantic tags;
[0062] If the association strength is lower than a preset strength threshold, the semantic tag sequence is grouped to obtain a subset of the grouped semantic tag sequences;
[0063] Construct feature vectors associated with the semantic label sequence subset and the context variables, and obtain an initial attachment position candidate set based on the associated feature vectors;
[0064] The business attributes in the standard data set are transformed to obtain a business attribute vector;
[0065] If the matching degree between the feature vector of the initial attachment location candidate set and the business attribute vector is higher than the preset matching degree threshold, then a candidate mapping map is constructed according to the spatial distribution of the feature vector to obtain a preliminary mapping map set.
[0066] Based on the preliminary mapping graph set, the context variable association weights of the candidate mapping graphs are calculated, and the optimized mapping graph structure is obtained based on the association weights.
[0067] Based on the optimized mapping graph structure, an attachment location candidate table is generated, resulting in the final attachment location candidate set.
[0068] It should be noted that core variables are determined based on business scenarios, including time-related variables (such as holidays, time periods), activity-related variables (such as promotional activities), and status-related variables (such as inventory warning levels, user activity levels). Time-related variables are extracted from the system clock / calendar interface, activity-related variables are read from the business system configuration table, and status-related variables are obtained through real-time calculation by the business system. Variables are converted into a unified encoding format (such as binary identifiers, level quantification values) to obtain context variables. From the standard data set, optimized semantic label sequences that are strongly associated with business attributes are extracted. Cosine similarity is used to calculate the semantic association between sequences. Each sequence is converted into a one-hot vector, and similarity is calculated by the ratio of vector dot product to magnitude (the higher the value, the stronger the association). If the association strength is lower than a preset threshold (which can be obtained through historical data sample analysis and is set to 0.6 through data analysis), the average linking algorithm is used for grouping: initially, each sequence is regarded as an independent cluster, and the average distance between clusters is calculated (the distance is obtained by subtracting the similarity from 1); the clusters with the smallest distance are merged, and this process is repeated until the inter-cluster distance is ≥ the threshold (such as 0.4), forming a subset of grouped semantic label sequences.
[0069] For each subset of semantic label sequences after grouping, its label vector is concatenated with the context variable vector to form a semantic-context fusion feature vector; matching attachment position rule base: the rule base pre-sets the correspondence between "feature vector → attachment position"; generating initial candidate set: all "feature vector - attachment position" are summarized to form an initial attachment position candidate set. One-hot encoding based on a business attribute vocabulary converts business attributes in the standard dataset into vectors. Cosine similarity is used to calculate the matching degree between the "feature vector of the initial attachment position candidate set" and the "business attribute vector" (the higher the value, the better the attachment position matches the business objective). Only candidate positions with a matching degree higher than a preset threshold (which can be obtained through historical data sample analysis and is set to 0.7 through data analysis) are retained. The spatial distribution of feature vectors is transformed into a structured mapping graph, which intuitively presents the relationship between semantic labels, context variables, and attachment positions. The mapping graph structure is defined as follows: three types of nodes are set: semantic label nodes, context variable nodes, and attachment position nodes. The edge weights between semantic labels and context variables can be based on their cosine similarity. The edge weights between context variables and attachment positions can be based on the frequency of occurrence of this combination in the initial candidate set. A candidate mapping graph is constructed. An independent mapping graph is constructed for each grouped subset of semantic label sequences to form a preliminary mapping graph set.
[0070] Adjust the weights of the mapping graph based on the actual impact of context variables on business attributes to make the mapping relationship more aligned with business priorities (e.g., promotions have a higher impact on revenue than normal periods). Define the impact weights of context variables, which can be based on the frequency of the variable's occurrence within the business cycle, divided into 3 levels: High frequency (occurrences more than 100 times per day): impact weight = 0.8; Medium frequency (occurrences less than 100 times per day but more than 50 times per day): impact weight = 0.5; Low frequency (occurrences less than 50 times per day): impact weight = 0.3. For the edges of "context variable - attachment position" in the mapping graph, multiply their original weights (occurrence frequency) by the variable's own weight coefficient and take the average. Replace the original edge weights with the optimized weights to form an "optimized mapping graph structure" (the higher the weight, the more important the variable-position association). The optimized mapping graph is transformed into a candidate table that can be directly called, supporting dynamic adaptation when the context changes. The rule of "semantic label sequence + context variable → attachment position" is extracted from the optimized mapping graph, and the confidence of the rule (i.e. the optimized weight) is recorded. The candidate table structure is constructed, which contains four columns: semantic label sequence, context variable combination, attachment position, and association weight. Records with association weight ≥ preset threshold (e.g., 0.7) are retained to form the final attachment position candidate set.
[0071] In step S14, based on the candidate attachment location set, the core features are integrated into the corresponding attachment locations to obtain an initial data framework, including:
[0072] Retrieve time series data and numeric fields;
[0073] If there are multiple candidate locations in the candidate set of attachment locations, the adaptation score of each candidate location is evaluated to obtain a list of adaptation scores;
[0074] Sort the adaptation score list according to the score from highest to lowest, extract the attachment position of the highest adaptation score, and obtain the attachment position mapping;
[0075] Based on the attachment location mapping, the associated data set is extracted to obtain the initial target structural framework;
[0076] Analyze the correspondence between the attachment positions in the initial target structural framework and the time series data to obtain a time series feature set;
[0077] The numerical fields are parsed based on the time series feature set to obtain a structured data table;
[0078] Check the integrity of fields in the structured data table. If the integrity is higher than the preset field integrity threshold, integrate the data to obtain the initial data framework.
[0079] It should be noted that, from the standard data set, data elements with continuous time labels (such as hourly sales and daily order volume) are extracted, requiring consistent time intervals (such as fixed hourly or daily granularity) and containing explicit timestamps (such as "2025-10-01 08:00"). Non-continuous time data (such as scattered records at random time points) are excluded, and time series data is obtained. Numerical technical features (such as transaction amount, inventory, and conversion rate) are selected from the standard data set, and their associated business attribute labels are retained. Non-numerical or business-meaningless fields are excluded, and numerical fields are extracted. If there are multiple candidate attachment locations in the candidate set, the adaptation score of each candidate location is evaluated. First, the scoring indicators and weights are determined: Association weight (weight 0.4): obtained based on the association weight of "semantic label - context variable - attachment location" in the optimized mapping diagram described in step S13; Data type matching degree (weight 0.3): the matching ratio between the data type of the candidate location (e.g., "numerical + time") and the current core feature type (e.g., matching 2 / 3 of the fields yields 0.67); Historical adaptation success rate (weight 0.3): the percentage of successful processing of similar data in the past at this location (e.g., 90 out of 100 successful processing yields 0.9). The adaptation score is obtained by calculating the product of the association weight, data type matching degree, and historical adaptation success rate with their corresponding weights, and then summing the results. The scores are then sorted from highest to lowest to form a list. The attachment location with the highest adaptation score is selected from the list to establish the correspondence between core feature and attachment location, i.e., the attachment location mapping.
[0080] Based on the attachment location mapping, relevant data is extracted and organized according to business logic to form a structured framework. Each attachment location corresponds to a specific business module, defining the business scope of that location. Data related to this business scope is filtered from the standard data set to obtain associated data elements. This data is then organized according to business process logic to form a multi-level framework (user layer → transaction layer → activity layer). Each layer includes field attributes (name, type, business attributes) and data relationships (e.g., "user ID" associated with "transaction record"). The time series correspondence is analyzed to clarify the dynamic association between attachment locations and time series data. Attach locations in the initial target structural framework (e.g., core category pages) are associated with timestamps of time series data (e.g., "2025-10-01 08:00"), forming "location-time-value" triples. Trends (e.g., data growth / decline), cycles, and abrupt changes (e.g., surges on promotional days) of the time series data are analyzed to form a time series feature set. Numerical fields are structured according to time series features. Based on the time series feature set, numerical fields are associated with their corresponding timestamps. A two-dimensional table is formed using the time axis as the row index and the numerical fields as columns, resulting in a structured data table. Verify data integrity and calculate field integrity: integrity equals the number of non-empty fields that meet business requirements divided by the total number of fields multiplied by 100%. Fields exceeding the preset threshold (which can be determined through historical data sample analysis and is set to 90% through data analysis) are integrated to obtain the initial data framework.
[0081] In step S15, the initial data frame is formatted to obtain a standard data frame, including:
[0082] Numerical fields are extracted from the initial data framework, and the encoding of the numerical fields is adjusted to obtain a formatted data set;
[0083] The standard data framework is obtained by formatting the data set according to the preset field specification table.
[0084] It should be noted that, from the initial data framework, all numeric fields are filtered out, and non-numeric fields such as text and enumeration types are excluded. The encoding method is adjusted, converting string-type numeric values to a unified numeric type (e.g., "1000 yuan" is converted to floating-point 1000.0 after removing the unit). All numeric fields are encoded to UTF-8 to avoid garbled characters during cross-system parsing; ultimately forming a formatted data set. Through preset business specifications, it is ensured that the data meets business requirements in terms of field naming, format, and constraints. The preset field specification table includes the following: basic information, including field names (e.g., "payment amount" instead of synonyms such as "transaction amount" or "payment amount"), business meaning (clearly defining the business attributes corresponding to the field); and technical specifications, including data types and formats. When validating and correcting according to the specification table, synonymous fields are unified with the standard names in the specification table. All field names in the formatted dataset are traversed and fuzzy matched against the synonym list in the specification table. For the matched synonymous fields, they are uniformly replaced with the standard field names in the specification table. This can be done in batches using data processing tools (such as the "field mapping" component of an ETL tool). The field format is checked to see if it conforms to the field specification. If the format is incorrect and cannot be automatically converted and corrected (such as invalid values like "abc" for time fields), it is marked as an abnormal field and filled in. The average time value of the same batch of data can be used to fill in the error and generate a standard data framework.
[0085] In step S16, logical rules are generated according to the standard data framework and integrated with a preset analysis report template to obtain an analysis output report, including:
[0086] Based on the correspondence between time series data and attachment locations in the standard data framework, the mapping graph is grouped to obtain grouped mapping graph data clusters;
[0087] Extract time series data from the grouped mapping data clusters and analyze fluctuation characteristics to obtain a set of fluctuation characteristics;
[0088] The integrity of the set of fluctuation features is determined. If the integrity is higher than a preset set integrity threshold, the fluctuation features are fused using a preset analysis report template to obtain an updated analysis report.
[0089] Key indicators are extracted from the updated analysis report and categorized to obtain a set of categorized indicators;
[0090] Extract the context variables of the classification index set analysis and analyze their correlation with the update analysis report, and generate association rules to obtain an initial association rule set;
[0091] Calculate the coverage of the initial association rule set. If the coverage is higher than a preset first coverage threshold, optimize the context variables in the initial association rule set to obtain an optimized association rule set.
[0092] The update analysis report is optimized based on the optimized association rule set to obtain the analysis output report.
[0093] It should be noted that, based on the correspondence between time series data and attached locations in the standard data framework, the cosine similarity of time series data (such as sales revenue) at different attached locations is calculated to measure the consistency of time series trends (the higher the value, the more similar the trends). The consistency of business attributes (such as "high customer flow", "low customer flow", "inventory-intensive") of attached locations is judged by comparing the business attribute labels of all attached locations in the same grouped mapping data cluster. If the labels belong to the same category (such as all being "high customer flow"), the consistency is considered to be up to standard. If there are cross-category labels (such as some being "high customer flow" and some being "low customer flow"), the consistency is considered to be down to standard. In this case, a clustering algorithm (such as K-means, with k set to 3) is used to aggregate the mapping maps with similar dimensions into one cluster. For example: mapping maps with time series trends of "large fluctuations" and attachment locations of "high passenger flow" are grouped into the "high fluctuation high passenger flow cluster"; mapping maps with time series trends of "stable" and attachment locations of "low passenger flow" are grouped into the "stable low passenger flow cluster"; mapping maps with time series trends of "large fluctuations" and attachment locations of "low passenger flow" are grouped into the "high fluctuation low passenger flow cluster"; mapping maps with time series trends of "stable" and attachment locations of "high passenger flow" are grouped into the "low fluctuation high passenger flow cluster"; ultimately forming multiple "grouped mapping map data clusters", with data within each cluster having similar time characteristics and business attributes. From each grouped mapping data cluster, the corresponding time series data is extracted, and the fluctuation characteristics are analyzed from two dimensions: amplitude characteristics: the mean and range (the difference between the maximum and minimum values in the time series) are calculated to reflect the magnitude of the fluctuation; periodic characteristics: periodic patterns are identified using the moving average method. The length of the moving average window is determined according to the business scenario (e.g., a window size of 7 days when analyzing by weekly cycle). Starting from the beginning of the time series, the moving average is calculated with the window size as the step size, and the average value of the data within each window is calculated. Trend analysis is performed on the smoothed moving average sequence. If a regular repetition is observed (e.g., the weekly window average peaks on Saturday), it is identified as a periodic feature (e.g., "Saturday is the peak sales week"). The features of each cluster are integrated to form a fluctuation feature set. To determine whether the set of fluctuation features covers the core features of "amplitude and cycle", the completeness is calculated by dividing the number of feature types included by 2 and then multiplying by 100%. If the completeness is greater than or equal to the preset threshold (which can be obtained through historical data sample analysis and is set to 80% through data analysis), a preset analysis report template (such as a sales trend analysis template) is called, and the fluctuation features are filled into the corresponding fields of the template. If the completeness is insufficient, the missing features are retrospectively supplemented (such as supplementing the cycle feature through historical data interpolation). After the template integrates the features, a report containing the basic fluctuation pattern is formed.
[0094] Key indicators that directly impact business decisions are extracted from the updated analysis reports and categorized by business function into revenue, efficiency, and customer flow categories, forming a set of categorized indicators. The correlation between these categorized indicators and context variables is analyzed to extract business rules. Context variables related to the categorized indicators are selected, and an association rule algorithm (such as Apriori) is used, with the categorized indicator set as the result and the context variables as the condition, to uncover association patterns. Rules with a confidence level ≥ a preset threshold (80%) are retained, forming an initial set of association rules. The coverage rate of this initial set of association rules is calculated. Coverage rate equals the quotient of the data volume covered by the initial rules divided by the total data volume of the standard data framework, multiplied by 100%. If the coverage rate ≥ a preset first coverage threshold (which can be determined through historical data sample analysis and is set to 0.7 through data analysis), the context variables are refined. If the coverage rate is insufficient, the indicators or variables are adjusted back and re-mined to obtain a more accurate and optimized set of association rules. The optimized association rules are integrated into the report and bound to the core fields of the report. For example, the "Promotion Impact" field is updated from "Promotion increases sales" to "Sales increase by 30% during large promotions (covering 85% of online self-operated malls)". The optimized report contains clear fluctuation characteristics, classification indicators and association rules, forming an analysis output report.
[0095] In step S17, based on the analysis output report, the matching strength of the semantic tag and the business attribute in the candidate attachment location set is calculated. If the matching strength is lower than a preset matching strength threshold, the standard data framework is optimized to obtain an optimized data framework, including:
[0096] Numerical fields are extracted from the analysis output report, and the numerical fields are grouped to obtain grouped field data clusters;
[0097] Calculate the matching strength between the semantic tag and the business attribute in the candidate set of attachment locations. If the matching strength is lower than the strength threshold, calculate the correlation between the grouping field data cluster and the semantic tag to obtain the semantic correlation.
[0098] If the semantic relevance is lower than the preset relevance threshold, it is determined that the format consistency is insufficient, and the numeric field is optimized to obtain the optimized numeric field.
[0099] By integrating the optimized numerical fields and the fluctuation characteristics in the analysis output report, an optimized feature set is generated;
[0100] Based on the optimized feature set, the mapping relationship of the data framework is adjusted to obtain the optimized framework;
[0101] Key fields are extracted from the optimization framework and categorized to obtain a set of key fields;
[0102] Analyze the mapping relationship between the set of key fields and the attachment location. If the coverage of the mapping relationship is higher than the preset second coverage threshold, optimize the mapping relationship to obtain an optimized mapping set.
[0103] The optimized data framework is obtained by updating the field configuration in the optimized framework based on the optimized mapping set.
[0104] It should be noted that, from the analysis output reports, numerical fields with business significance are selected, and the K-means clustering algorithm is used to group the numerical fields based on the similarity of their business attributes, forming grouped field data clusters. Cosine similarity is then used to compare the semantic tag sequence vectors in the candidate attachment location set with the business attribute vectors (based on one-hot encoding of the business attribute vocabulary) to obtain the matching strength. If the matching strength is lower than a preset threshold (0.6), the cosine similarity between the feature vectors and semantic tag vectors of the grouped field data clusters is further calculated to obtain the semantic relevance. If the semantic relevance is lower than a preset relevance threshold (0.7), it is determined to be insufficient format consistency. For null fields, linear interpolation of data from preceding and following time periods is used to fill in the missing values. In optimizing the numerical fields, the time or business sequence position of the missing value is identified, and two adjacent valid data points are taken, denoted as (…). )and( ), where x is the time / series index and y is the corresponding value; based on the linear relationship, missing points ( Estimated value: , calculate Fill in the missing positions; use The rule identifies outliers, based on The rule is to set the threshold range of normal data as "data mean ± 3 times standard deviation". Values outside this range are judged as outliers. Outliers are replaced with the mean of the same cluster field to correct outliers and obtain optimized numerical fields. The fluctuation features associated with the optimized numerical fields are extracted from the analysis output report. The optimized numerical fields and fluctuation features are concatenated according to business attributes (such as "daily sales of online self-operated mall (120,000 yuan) + monthly fluctuation standard deviation (5,000 yuan) + promotion increase (30%)"). The features are then integrated. Business attribute tags are added to each integrated feature (such as "revenue-high volatility-promotion sensitive") to form an optimized feature set.
[0105] Based on the mapping relationship of "numerical field - attached location - business attribute" in the standard data framework, the mapping relationship is redefined based on the business attributes of the optimized feature set. The adjusted framework has consistent logic among "feature - location - business attribute". The optimized framework is output, and core fields directly related to business decisions (such as online self-operated mall ID, optimized sales, attached location type, inventory turnover rate, etc.) are selected and classified according to business function dimensions (such as revenue category including optimized sales and average order value; location category including attached location type and online self-operated mall ID) to form a key field set. The mapping coverage rate is calculated: the coverage rate is equal to the quotient of the number of key fields that can accurately match the attached location divided by the total number of key fields, and then multiplied by 100%. If the coverage rate is higher than the preset second coverage rate threshold (which can be obtained through historical data sample analysis and is set to 0.85 through data analysis), the mapping rules are refined (such as "optimized sales > 100,000 yuan → core category page" is refined to "50,000 yuan < sales ≤ 100,000 yuan → promotional activity page") to obtain the optimized mapping set. Based on the optimized mapping set, the field attributes in the optimized framework are updated to obtain the optimized data framework.
[0106] It should be noted that if the matching strength is not lower than the preset matching strength threshold, the standard data framework is retained as the optimized data framework for prediction in the next step.
[0107] In step S18, the candidate mapping map of the attachment location is updated according to the analysis output report, and a time series prediction report is obtained by combining the optimized data framework, including:
[0108] By analyzing the context variables of the output report data, the scope of event impact and time series characteristics of the data are determined, and an event feature set is obtained.
[0109] If the time series features of the event feature set meet the preset dynamic threshold, then the attachment location candidate is generated according to the scope of promotional influence, and the candidate location set is obtained.
[0110] The updated mapping graph is obtained by updating the node weights and connections of the updated mapping graph based on the candidate location set;
[0111] Node features and time-series data are extracted from the updated mapping graph to obtain the enhanced dataset;
[0112] Time series prediction is performed based on the enhanced dataset to obtain a time series prediction report.
[0113] It should be noted that the variables related to the business event are selected from the analysis output report, including event type (e.g., discounts, promotions), duration, participating entities (e.g., product categories, online self-operated mall scope), and changes in key indicators during the event period (e.g., sales volume, click volume increase). By comparing the data during the event period and the non-event period, the affected spatial range is divided, and the "indicator growth rate" for each position is calculated (the quotient of the average during the event period and the average during the non-event period minus 1). The range is divided into three levels: core impact range: growth rate ≥ 50%; secondary impact range: 20% ≤ growth rate < 50%; no impact range: growth rate < 20%. The time patterns related to the event are analyzed, including the peak time after the event starts, the decline period after the event ends, and the fluctuation range during the event period (e.g., standard deviation). The "event type + impact range + time characteristics" are then integrated. The event feature set is obtained. Based on the dynamism of the event feature set, a dynamism threshold (preset to 30%) is set to determine whether the time series features in the event feature set significantly affect the data. If the dynamism threshold is met, candidate attachment positions are generated according to the scope of the event's impact. The importance and correlation strength of the attachment positions are adjusted according to the candidate position set to make the mapping map accurately reflect the business relationships under the dynamics of the event. The node weight represents the business importance of the attachment position. The indicator increment of each attachment position during the event is calculated (e.g., the sales increase equals the sales during the event period minus the average value during the non-event period). The contribution ratio of a single position is calculated as the quotient of the indicator increment of that position divided by the total indicator increment of all positions, then multiplied by 100%. The contribution ratio is the node weight. The connection relationship reflects the business relationship between positions. The strength is optimized according to the event's impact (e.g., the connection strength is increased from 0.6 to 0.8), and the updated mapping map is output.
[0114] The features of each attachment location are obtained from the updated mapping map. Historical data corresponding to the features of the attachment locations are extracted from the optimized data framework. The node features are bound to the time series data by the attachment location ID (e.g., "Attachment location: core category page - node weight 0.9 - average daily sales of 80,000 yuan in the past 90 days - association strength 0.8") to form an enhanced dataset. From the enhanced dataset, the data is grouped by "attachment location + event type". The time series pattern of similar historical events is calculated (the time point and peak percentage of the average index and statistical index of similar events at each time point are calculated to determine the overall trend and pattern). The historical pattern is adjusted according to the duration and intensity of the current event (e.g., discount range). The adjusted values are output sequentially according to the time series (e.g., the next 7 days) to finally form a time series prediction report.
[0115] It is worth noting that the specific values of each threshold can be calibrated according to the actual business scenario and historical data. The values given in this embodiment of the invention are only examples.
[0116] In summary, this invention discloses a dynamic data mapping and adaptation method for heterogeneous data. Through a regrouping mechanism, it dynamically optimizes data grouping methods, improving the flexibility and accuracy of data type matching. This effectively solves the data processing problems caused by field type mismatches in existing ETL tools, performing particularly well when handling complex or ambiguous business data types. Furthermore, this invention introduces association strength analysis of business attributes and semantic tags to construct candidate mapping graphs, integrating semantic understanding into the data mapping process. This significantly improves the business adaptability of the mapping results, avoiding inaccurate mapping caused by ignoring data semantics and context, thereby improving matching accuracy and better meeting the needs of multi-source data integration.
[0117] Reference Figure 2 The second embodiment of the present invention provides a dynamic data mapping and adaptation system for heterogeneous data, comprising:
[0118] The data acquisition module is used to acquire initial data, extract core features, and obtain an initial data set; the initial data set includes business attributes, type features, and semantic tags.
[0119] The data grouping module is used to regroup the initial data group set to obtain a standard data group set if the matching degree between the type feature and the preset type library in the initial data group set is lower than the preset type matching degree threshold.
[0120] The data matching module is used to calculate the association strength between the semantic tag and the business attribute based on the standard data set, construct a candidate mapping map of attachment positions, and obtain a candidate set of attachment positions by optimizing the mapping map;
[0121] The data integration module is used to integrate the core features into the corresponding attachment positions based on the candidate attachment position set to obtain an initial data framework;
[0122] The data unification module is used to format the initial data framework to obtain a standard data framework.
[0123] The data generation module is used to generate logical rules based on the standard data framework and integrate them with a preset analysis report template to obtain an analysis output report.
[0124] The data verification module is used to calculate the matching strength of the semantic tag and the business attribute in the candidate set of attachment locations based on the analysis output report. If the matching strength is lower than the preset matching strength threshold, the standard data framework is optimized to obtain the optimized data framework.
[0125] The data prediction module is used to update the candidate mapping map of the attachment location based on the analysis output report, and to obtain a time-series prediction report by combining the optimized data framework.
[0126] It should be noted that the dynamic data mapping and adaptation system for heterogeneous data provided in this embodiment of the invention is used to execute all the process steps of the dynamic data mapping and adaptation method for heterogeneous data in the above embodiment. The working principles and beneficial effects of the two are one-to-one, so they will not be described again.
[0127] This invention also provides an electronic device. The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a dynamic data mapping and adaptation program for heterogeneous data. When the processor executes the computer program, it implements the steps in the aforementioned embodiments of the dynamic data mapping and adaptation system and method for heterogeneous data, for example... Figure 1 The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, such as the data prediction module.
[0128] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0129] The electronic device may be a desktop computer, laptop, handheld computer, or smart tablet, etc. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above components are merely examples of electronic devices and do not constitute a limitation on the electronic device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0130] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting all parts of the electronic device via various interfaces and lines.
[0131] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0132] Wherein, if the modules / units integrated in the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0133] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0134] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for dynamic data mapping and adaptation of heterogeneous data, characterized in that, The method comprises the following steps: acquiring initial data, extracting core features, and obtaining an initial data set collection; the initial data set collection contains business attributes, type features, and semantic labels; if the type features in the initial data set collection have a matching degree with a preset type library that is lower than a preset type matching degree threshold, the initial data set collection is regrouped to obtain a standard data set collection; based on the standard data set collection, the association strength of the semantic labels and the business attributes is calculated, a candidate mapping graph of the attached position is constructed, and an attached position candidate set is obtained by optimizing the mapping graph; according to the attached position candidate set, the core features are integrated into the corresponding attached position to obtain an initial data framework; the initial data framework is formatted to obtain a standard data framework; logical rules are generated according to the standard data framework, and a preset analysis report template is fused to obtain an analysis output report; according to the analysis output report, the matching strength of the semantic labels and the business attributes in the attached position candidate set is calculated, if the matching strength is lower than a preset matching strength threshold, the standard data framework is optimized to obtain an optimized data framework; the candidate mapping graph of the attached position is updated according to the analysis output report, and a time sequence prediction report is obtained in combination with the optimized data framework.
2. The method of claim 1, wherein, if the type features in the initial data set collection have a matching degree with a preset type library that is lower than a preset type matching degree threshold, the initial data set collection is regrouped to obtain a standard data set collection, which comprises the following steps: the initial data set collection is divided into multiple subsets to obtain initial data set subsets; when the type features of the data sets in the initial data set subsets have a matching degree with a preset type library that is lower than a preset type matching degree threshold, the initial data set subsets are grouped to obtain standardized data set subsets; the semantic labels in the standardized data set subsets are extracted to obtain an initial semantic label sequence; the correlation of the initial semantic label sequence and the business attributes is analyzed to obtain an optimized semantic label sequence with high correlation; the optimized semantic label sequence and the standardized data set subsets are fused to obtain a standard data set collection.
3. The method for dynamic data mapping and adaptation of heterogeneous data according to claim 1, wherein, based on the standard data set collection, the association strength of the semantic labels and the business attributes is calculated, a candidate mapping graph of the attached position is constructed, and an attached position candidate set is obtained by optimizing the mapping graph, which comprises the following steps: acquiring context variables; extracting the semantic label sequence in the standard data set collection and calculating the correlation of the semantic label sequence to obtain the association strength of the semantic labels; if the association strength is lower than a preset strength threshold, the semantic label sequence is grouped to obtain a grouped semantic label sequence subset; a feature vector associated with the context variables is constructed based on the semantic label sequence subset, and an initial attached position candidate set is obtained based on the associated feature vector; the business attributes in the standard data set collection are converted to obtain a business attribute vector; If the matching degree of the feature vector of the initial attachment position candidate set and the service attribute vector is higher than a preset matching degree threshold, a candidate mapping graph is constructed according to the spatial distribution of the feature vector, and a preliminary mapping graph set is obtained; According to the preliminary mapping graph set, the context variable association weight of the candidate mapping graph is calculated, and the optimized mapping graph structure is obtained according to the association weight mapping; According to the optimized mapping graph structure, an attachment position candidate table is generated, and a final attachment position candidate set is obtained.
4. The method for dynamic data mapping and adaptation of heterogeneous data according to claim 1, wherein, According to the attachment position candidate set, the core feature is integrated into the corresponding attachment position to obtain an initial data framework, including: Obtain time series data and numerical fields; If there are multiple alternative positions in the attachment position candidate set, the adaptation score of each alternative position is evaluated to obtain an adaptation score list; According to the adaptation score list, the attachment positions with the highest adaptation scores are extracted according to the score from high to low, and an attachment position mapping is obtained; According to the attachment position mapping, the associated data set is extracted to obtain an initial target structure framework; Analyze the correspondence between the attachment position in the initial target structure framework and the time series data to obtain a time series feature set; According to the time series feature set, the numerical fields are parsed to obtain a structured data table; Check the integrity of the fields in the structured data table. If the integrity is higher than a preset field integrity threshold, integrate the data to obtain an initial data framework.
5. The method for dynamic data mapping and adaptation of heterogeneous data according to claim 1, wherein, The standard data framework is used to generate logical rules, and a preset analysis report template is fused to obtain an analysis output report, including: According to the correspondence between the time series data and the attachment position in the standard data framework, the mapping graph is grouped to obtain a grouped mapping graph data cluster; Extract the time series data in the grouped mapping graph data cluster and analyze the fluctuation characteristics to obtain a fluctuation feature set; Determine the integrity of the fluctuation feature set. If the integrity is higher than a preset set integrity threshold, fuse the fluctuation features through a preset analysis report template to obtain an updated analysis report; Extract key indicators from the updated analysis report and classify the key indicators to obtain a classified indicator set; Extract the context variables analyzed in the classified indicator set and analyze the relevance with the updated analysis report, and generate association rules to obtain an initial association rule set; Calculate the coverage rate of the initial association rule set. If the coverage rate is higher than a preset first coverage rate threshold, optimize the context variables in the initial association rule set to obtain an optimized association rule set; According to the optimized association rule set, the updated analysis report is optimized to obtain an analysis output report.
6. The method for dynamic data mapping and adaptation of heterogeneous data according to claim 1, wherein, According to the analysis output report, the matching strength of the semantic label and the service attribute in the attachment position candidate set is calculated. If the matching strength is lower than a preset matching strength threshold, the standard data framework is optimized to obtain an optimized data framework, including: Extract the numerical fields from the analysis output report and group the numerical fields to obtain a grouped field data cluster; Calculate the matching strength of the semantic label and the service attribute in the attachment position candidate set, and if the matching strength is lower than the strength threshold, calculate the association degree of the grouping field data cluster and the semantic label to obtain a semantic association degree; If the semantic association degree is lower than a preset association degree threshold, it is determined that the format consistency is insufficient, and an optimized numerical field is obtained by optimizing the numerical field; Fuse the optimized numerical field and the fluctuation characteristics in the analysis output report to generate an optimized feature set; According to the optimized feature set, adjust the mapping relationship of the data framework to obtain an optimized framework; Extract key fields from the optimized framework and classify the key fields to obtain a key field set; Analyze the mapping relationship between the key field set and the attachment position, and if the coverage rate of the mapping relationship is higher than a preset second coverage rate threshold, optimize the mapping relationship to obtain an optimized mapping set; According to the optimized mapping set, update the field configuration in the optimized framework to obtain an optimized data framework.
7. The method for dynamic data mapping and adaptation of heterogeneous data according to claim 1, wherein, The analysis output report is updated according to the analysis output report, and a time series prediction report is obtained in combination with the optimized data framework, including: Analyze the context variables of the analysis output report data to determine the event influence range and time series characteristics of the data to obtain an event feature set; If the time series characteristics of the event feature set meet a preset dynamic threshold, generate an attachment position candidate according to the promotion influence range to obtain a candidate position set; Update the node weight and connection relationship of the mapping graph according to the candidate position set to obtain an updated mapping graph; Extract node features and time series data from the updated mapping graph to obtain an enhanced data set; According to the enhanced data set, time series prediction is performed to obtain a time series prediction report.
8. The method for dynamic data mapping and adaptation of heterogeneous data according to claim 1, wherein, The initial data is obtained, the core features are extracted, and an initial data group set is obtained, including: Obtain initial data; Extract type features, semantic labels and service attributes from the initial data to obtain a structured data set; If the integrity of the type features in the structured data set is higher than a preset type integrity threshold, the data is preliminarily classified to obtain a classified initial data group set; If the integrity of the type features in the structured data set is lower than a preset type integrity threshold, missing features are completed through the semantic label to obtain an initial data group set.
9. The method for dynamic data mapping and adaptation of heterogeneous data according to claim 1, wherein, The initial data framework is formatted to obtain a standard data framework, including: Extract numerical fields from the initial data framework and adjust the encoding mode of the numerical fields to obtain a formatted data set; According to a preset field specification table, the formatted data set is specified to obtain a standard data framework.
10. A dynamic data mapping and adaptation system for heterogeneous data, characterized by, Including: A data acquisition module is configured to acquire initial data, extract core features, and obtain an initial data group set; The initial data group set includes business attributes, type features and semantic labels; A data grouping module is configured to re-group the initial data group set if the matching degree of the type features in the initial data group set with a preset type library is lower than a preset type matching degree threshold, to obtain a standard data group set; The data matching module is configured to calculate the association strength of the semantic label and the business attribute based on the standard data set, and construct a candidate mapping graph of the attached position, and obtain a candidate set of the attached position by optimizing the mapping graph. The data integration module is configured to integrate the core feature into the corresponding attached position according to the candidate set of the attached position, and obtain an initial data framework. The data unification module is configured to format the initial data framework, and obtain a standard data framework. The data generation module is configured to generate a logical rule according to the standard data framework, and fuse a preset analysis report template, and obtain an analysis output report. The data verification module is configured to calculate the matching strength of the semantic label and the business attribute in the candidate set of the attached position according to the analysis output report, and if the matching strength is lower than a preset matching strength threshold, optimize the standard data framework, and obtain an optimized data framework. The data prediction module is configured to update the candidate mapping graph of the attached position according to the analysis output report, and obtain a time sequence prediction report in combination with the optimized data framework.
Citation Information
Patent Citations
Integrated analysis method and system for tobacco multi-source heterogeneous data
CN116541449A
Financial data generation method and system based on rule configuration engine
CN120259003A