Information center data integration method for heterogeneous system data information aggregation
By identifying features and constructing semantic graphs for heterogeneous systems, the problems of data access failure and semantic ambiguity in traditional methods are solved, and efficient, accurate aggregation and semantic unification of heterogeneous system data are achieved.
Patent Information
- Application Number
- CN202511536667.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Traditional methods lack systematic identification and analysis of multi-dimensional characteristics of data sources in heterogeneous system data integration, leading to data access failures, resource waste, and semantic ambiguity, making it difficult to achieve efficient and accurate data aggregation.
By identifying the feature types of heterogeneous systems, constructing cross-system semantic graphs, and calculating multi-dimensional semantic similarity, the driver programs and access strategies of the target data source are accurately matched, achieving feature matching of the transport layer, storage layer, and business layer. Furthermore, through semantic alignment of entities, attributes, and relationships, a semantically unified integrated dataset is output.
It enables efficient and stable access and semantic unification of data information from heterogeneous systems, reduces operation and maintenance costs, and improves the accuracy and efficiency of data aggregation.
Smart Images

Figure CN121009427B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data integration, more particularly, to an information center data integration method for heterogeneous system data information aggregation. BACKGROUND
[0002] In the field of heterogeneous system data integration, due to the differences in design goals and technical architecture, different communication protocols, storage structures, data encoding formats and business logic definitions are often adopted for target data sources of various types of heterogeneous systems, forming independent data islands. Traditional data integration methods lack systematic identification and analysis mechanisms for multi-dimensional characteristics of data sources, making it difficult to accurately match the technical characteristics of different heterogeneous system data sources, resulting in problems such as the inability to establish stable connections efficiently due to inaccurate judgments of data source characteristics when interfacing different system data sources, hindering the smooth flow and aggregation of data between heterogeneous systems.
[0003] In the existing process of heterogeneous system data integration, the selection of drivers and the configuration of preset access strategies are mostly dependent on manual experience judgments, lacking scientific decision-making basis based on data source characteristics. Since the core characteristic dimensions of data sources such as transmission layer, storage layer and business layer are not subdivided and matched for analysis, it is easy to cause incompatibility between drivers and target data sources, and mismatch between access strategies and actual operational requirements of data sources, not only causing data access failure, but also consuming a large amount of system resources due to repeated trial and error, increasing the operation and maintenance cost and time cost of data integration, making it difficult to meet the demand for efficient aggregation of heterogeneous system data.
[0004] Even if part of the traditional methods achieve the initial access of heterogeneous system data sources, there are still significant shortcomings in the data semantic alignment process. Different heterogeneous systems have different naming and definitions for the same business entities, attributes and relationships between entities, and traditional methods lack unified semantic standards and multi-dimensional semantic similarity analysis mechanisms, making it difficult to effectively penetrate such naming and definition differences, resulting in semantic ambiguity in integrated data, making it difficult to form semantically unified data sets. This makes it difficult to fully exploit the value of heterogeneous system data, affecting the accuracy and effectiveness of information center data information aggregation. SUMMARY
[0005] In view of the deficiencies in the prior art, the present application aims to provide an information center data integration method for heterogeneous system data information aggregation.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0007] An information center data integration method for heterogeneous system data information aggregation, comprising
[0008] Step S1: The information center identifies the feature types of the target data sources of various heterogeneous systems, determines the matching degree of the feature type identification result, loads the corresponding target data source driver and the preset access strategy when the matching degree reaches the preset threshold, and triggers the feature learning mechanism to update the preset feature template library when the matching degree does not reach the threshold.
[0009] Step S2: Constructing a cross-system semantic graph for various heterogeneous systems.
[0010] Step S3: Based on the cross-system semantic graph, the multi-dimensional semantic similarity of the accessed target data sources is calculated to complete the semantic alignment between the target data sources, and the integrated data set with unified semantics is output.
[0011] Further, the target data source of each type of heterogeneous system collects three core feature dimensions, including the transmission layer feature dimension, the storage layer feature dimension, and the business layer feature dimension.
[0012] Further, the sub-features under the transmission layer feature dimension include the communication protocol, port number, data encryption method, and connection timeout threshold of the target data source.
[0013] The sub-features under the storage layer feature dimension include the storage structure and data encoding format of the target data source.
[0014] The sub-features under the business layer feature dimension include the update frequency, data volume order of magnitude, and core field type of the target data source.
[0015] Further, the matching degree of the feature type identification result is determined as follows: a protocol template mapping table is established, the actual communication protocol name of the target data source is obtained, the associated scene template group is determined according to the actual communication protocol name and the protocol template mapping table, the optimal possible template is further filtered out from the associated scene template group, the feature matching value of each sub-feature and the optimal possible template is obtained , the matching weight of each sub-feature is obtained , the matching degree of the feature type identification result is calculated by the formula ; ; is the total number of sub-features.
[0016] Further, the optimal possible template is further filtered out from the associated scene template group as follows: the template matching value of each template in the associated scene template group is calculated ; is the number of core feature labels of the associated scene template group, is the weight of the jth core feature label, is the feature matching value of the jth core feature label; and the template matching value The template with the highest value is marked as the preferred possible template.
[0017] Further, the target data sources accessed are calculated for multi-dimensional semantic similarity based on the cross-system semantic graph to complete semantic alignment between the target data sources, and the specific steps are as follows:
[0018] Step one: based on the cross-system semantic graph, the target data sources to be aligned are preprocessed to eliminate the interference of ''format / unit / representation difference'' on the similarity calculation;
[0019] Step two: calculating entity semantic similarity , attribute semantic similarity and relationship semantic similarity , and comprehensively calculating comprehensive semantic similarity according to the entity semantic similarity, the attribute semantic similarity and the relationship semantic similarity, wherein d1, d2 and d3 are weight coefficients ;
[0020] The upper value and the lower value of the comprehensive semantic similarity are set, when the comprehensive semantic similarity is greater than or equal to the upper value of the comprehensive semantic similarity, the entities / attributes / relationships of the data sources are directly mapped to the global semantics of the graph;
[0021] When the comprehensive semantic similarity is less than or equal to the lower value of the comprehensive semantic similarity, manual intervention is triggered to confirm whether it is a new semantic;
[0022] When the comprehensive semantic similarity is between the upper value and the lower value of the comprehensive semantic similarity, the alignment is performed after supplementing the local adaptation rule.
[0023] Further, the entity semantic similarity ; is an entity semantic weight coefficient.
[0024] Further, the attribute semantic similarity ; is an attribute identifier similarity; is attribute value compatibility;
[0025] The attribute identifier similarity ; is an attribute identifier weight coefficient;
[0026] The attribute value compatibility .
[0027] Further, the relationship semantic similarity ; is a relationship semantic weight coefficient.
[0028] Compared with the prior art, the present application has the following beneficial effects:
[0029] The method of the present application accurately analyzes the feature matching values of the multi-dimensional sub-features of the transmission layer, storage layer and service layer, designs differentiated matching rules for the technical attributes of different sub-features, accurately judges the adaptation degree of the target data source and the preferred possible template, and loads the corresponding drive and strategy only when the matching degree meets the standard, thereby avoiding resource waste caused by drive loading errors and ensuring stable and efficient data access through the accurate effect of the preset access strategy, realizing the upgrade from "experience-driven adaptation" to "data-driven accurate adaptation";
[0030] The entity semantic similarity is verified through name and business attribute, and the business essence is locked through the penetration of heterogeneous system entity naming differences. The attribute semantic similarity is powered from two aspects of identification matching and value compatibility, which not only solves the attribute mapping problem of "synonymous different names", but also ensures that the attribute values meet the global standard. The relationship semantic similarity is verified by relying on global relationship mapping and associated entities, which guarantees the consistency of business logic between entities. The three are independently analyzed to solve the core pain points of entity recognition, attribute matching and relationship verification in semantic alignment, and provide accurate single-dimensional support for comprehensive alignment;
[0031] Through the hierarchical processing mechanism, when the similarity is high, automatic mapping is used to improve efficiency, when the similarity is medium, local rule adaptation is used to reduce cost, and when the similarity is low, manual intervention is triggered to ensure accuracy, effectively balancing the efficiency and accuracy of semantic alignment, and finally outputting integrated data sets with unified semantics, providing a reliable and efficient technical path for heterogeneous system data information aggregation. BRIEF DESCRIPTION OF DRAWINGS
[0032] Fig. 1 A method flowchart of an information center data integration method for heterogeneous system data information aggregation;
[0033] Fig. 2 A construction principle diagram of a cross-system semantic graph. DETAILED DESCRIPTION
[0034] Referring to Figs. 1-2 An information center data integration method for heterogeneous system data information aggregation, comprising
[0035] Step S1: The information center identifies the feature types of the target data sources of various heterogeneous systems (collects three core feature dimensions for each type of target data source of a heterogeneous system, including transmission layer feature dimension, storage layer feature dimension, and business layer feature dimension, and each of the three core feature dimensions is subdivided into multiple sub-features), determines the matching degree of the feature type identification result, and when the matching degree reaches a preset threshold value (the preset threshold value is set based on the "adaptation success rate target" (to ensure that when the threshold value is reached, the driving program and the preset access strategy can be loaded to stably access after the driving program and the preset access strategy are loaded, and the success rate needs to be ≥95%) and "misjudgment cost balance" (to avoid too low threshold value leading to frequent loading failure, or too high threshold value leading to excessive triggering of learning), loads the driving program and the preset access strategy corresponding to the target data source (each template corresponds to a driving program and a preset access strategy, the driving program and the preset access strategy corresponding to the optimal possible template are selected, the driving program is automatically loaded: the driving program corresponding to the optimal possible template is called from the "driver resource library" (official drivers of various data sources are pre-stored, such as MySQL-connector-java and MongoDB-driver) to establish a stable connection with the target data source; the preset access strategy takes effect: the access strategy bound by the optimal possible template is loaded, including data reading frequency, data sharding rule, and exception retry mechanism, to ensure the efficiency and stability of data access); when the matching degree does not reach the threshold value, a feature learning mechanism is triggered to update the preset feature template library (sample labeling: the types and key attributes of the feature vectors (transmission / storage / business layer sub-features) of the target data source are labeled by humans, such as decoding rules of new data sources; template training: new templates are trained by algorithms (such as Naive Bayes) combined with historical similar features, to generate standard feature values and weights; library update and reuse: the new templates are written into the library, bound with corresponding drivers and access strategies, and subsequent similar data sources can be directly matched);
[0036] Sub-features under the transmission layer feature dimension: include the communication protocol (such as JDBC, ODBC, MQTT, HTTP / HTTPS) of the target data source, port number, and data encryption method (such as SSL / TLS, plaintext);
[0037] Sub-features under the storage layer feature dimension: include the storage structure (row storage / column storage, key-value pair / document type / time series type) of the target data source and data encoding format (UTF-8, GBK, Base64);
[0038] Sub-features under the business layer feature dimension: include the update frequency (real-time stream / minute level / hour level / day level) of the target data source, data volume order of magnitude (KB level / MB level / GB level), and core field type (numeric type / string type / date type);
[0039] Determine the matching degree of the feature type identification result, as follows: obtain the feature matching value of each sub-feature , synchronously acquire the matching weight of each sub-feature , analyze the matching weight of each sub-feature by AHP method , calculate the matching degree of the feature type recognition result ; k is the total number of sub-features, is the feature matching value of the i-th sub-feature;
[0040] Feature matching value of communication protocol The acquisition method is: establishing a protocol template mapping table, acquiring the actual communication protocol name of the target data source, determining the associated scene template group according to the actual communication protocol name and the protocol template mapping table, further filtering out the preferred possible template from the associated scene template group, if the actual communication protocol name is in the preferred possible template in the preset feature template library (such as "JDBC" in the MySQL template) → ; if the actual communication protocol name is not in the preferred possible template (such as "HTTP" not in the MySQL template) → (No compatible relationship, no partial matching scenario);
[0041] Feature matching value of port number The acquisition method is: scanning the actual listening port of the target data source, if the actual listening port is the "default port" of the preferred possible template → ; if the actual listening port is the "compatible port" of the preferred possible template → ; if the actual listening port is not the "default port" and "compatible port" of the preferred possible template → ; the template is stored in "default port + compatible port" (such as "MySQL template" port = {default 3306, compatible 3307}, "HTTP template" port = {default 80, compatible 8080});
[0042] Feature matching value of data encryption method The acquisition method is: acquiring the data encryption method of the target data source, if the data encryption method is the "main encryption method" of the preferred possible template → ; if the data encryption method is the "compatible encryption method" of the preferred possible template → ; if the data encryption method is not the "main encryption method" and "compatible encryption method" of the preferred possible template → ; the template is stored in "main encryption method + compatible encryption method" (such as "MySQL template" encryption method = {main SSL / TLSv1.3, compatible TLSv1.2}, "HTTP template" encryption method = {plain text, compatible SSL / TLSv1.1});
[0043] Feature matching value of connection timeout threshold Acquisition method: Obtain the actual timeout configuration of the target data source through the connection test tool. If the actual timeout configuration is within the timeout threshold interval of the preferred possible template → ; If the actual timeout configuration is not within the timeout threshold interval of the preferred possible template → ; The template is stored in "time interval", such as "MySQL template" timeout threshold = {5-10 seconds}, "HTTP template" timeout threshold = {6-12 seconds};
[0044] Characteristic matching value of storage structure Acquisition method: Analyze the storage major category and storage subclass of the target data source. If the storage major category and storage subclass of the target data source are both within the preferred possible template → ; If the storage major category of the target data source is within the preferred possible template, and the storage subclass is not within the preferred possible template → ; If the storage major category of the target data source is not within the preferred possible template → ; The template is stored in "hierarchical classification", that is, major category + subclass, such as "MySQL template" storage "major category = row storage, storage subclass = MySQL";
[0045] Characteristic matching value of data encoding format Acquisition method: (through sample data decoding test) Obtain the actual encoding of the target data source. If the actual encoding of the target data source is the "main encoding" of the preferred possible template → ; If the actual encoding of the target data source is the "compatible encoding" of the preferred possible template → ; If the actual encoding of the target data source is neither the "main encoding" nor the "compatible encoding" of the preferred possible template → ; The template is stored in "main encoding + compatible encoding", such as "MySQL template" encoding: main encoding: UTF-8mb4, compatible encoding: GBK;
[0046] Characteristic matching value of update frequency Acquisition method: (through querying the target data source update log or scheduled task configuration) Obtain the actual update frequency interval of the target data source (such as "5-15 minutes / time", "2-4 hours / time"), if the actual update frequency interval of the target data source is completely contained in the update frequency interval of the preferred possible template (such as "5-15 minutes / time" in the minute level 1-30 minute interval) → ; If the actual update frequency interval of the target data source partially overlaps with the update frequency interval of the preferred possible template (such as "25-40 minutes / time" overlapping with the minute level 1-30 minute interval for 5 minutes) → ; If the actual update frequency interval of the target data source does not overlap with the update frequency interval of the preferred possible template → ;
[0047] Data volume order of magnitude feature matching value The acquisition method is as follows: the data volume of a single batch transmission / storage of the statistical target data source is determined to determine the actual magnitude interval (such as “512 KB-1 MB”), and if the actual magnitude interval is completely contained in the data volume magnitude interval of the preferred possible template (such as “512 KB-1 MB” in the MB level 1-1024 MB interval) -> ; if the actual magnitude interval of the target data source partially overlaps with the data volume magnitude interval of the preferred possible template -> ; if the actual magnitude interval of the target data source does not overlap with the data volume magnitude interval of the preferred possible template -> ;
[0048] Core field type feature matching value The acquisition method is as follows: the target data source is parsed to obtain the actual types of the core fields (such as “numeric (amount), string (order number), and string (order time)”) of the target data source, and if the core field types of the target data source are completely consistent with the required core field type list of the preferred possible template -> ; if the core field types of the target data source are partially consistent with the required core field type list of the preferred possible template -> ; if the core field types of the target data source are completely inconsistent with the required core field type list of the preferred possible template -> ; the template is stored in the “required core field type list” (such as “MySQL template” core field type={numeric (amount), string (order number), and date (order time)}).
[0049] Protocol template mapping table, for example: if the actual communication protocol name is JDBC / ODBC, the mapped associated scene template group is the relational database template group (the mapped core logic: JDBC / ODBC is a dedicated access protocol for relational databases (such as MySQL and Oracle), and will not be used for Internet of Things terminals or file systems); if the actual communication protocol name is MQTT / CoAP, the mapped associated scene template group is the Internet of Things terminal template group (the mapped core logic: MQTT / CoAP is a lightweight protocol designed for Internet of Things devices, and does not support SQL interaction of relational databases); if the actual communication protocol name is Kafka protocol, the mapped associated scene template group is the stream data template group (Kafka protocol is only used for real-time data transmission of stream data platforms, and is irrelevant to static storage databases / files);
[0050] Further filter the preferred possible templates from the associated scene template group, as follows: calculate the template matching value of each template in the associated scene template group ; m is the number of core feature labels of the associated scene template group (such as the MySQL template and the Oracle template in the relational database template, the relational database template group m = 3, the core feature labels are determined according to the following steps: first, anchor the scene adaptation core demand (such as the relational database template group requires "stable connection + SQL interaction"); second, filter the features that play a decisive role in the success or failure of adaptation (exclude secondary features); third, ensure that the features are shared by the template group, can distinguish other groups, and can be obtained through lightweight detection; finally, determine the core feature label through historical adaptation cases), is the weight of the jth core feature label The weight distribution of the core feature labels of different associated scene template groups is not the same. For example, in the relational database template group, the weight of the port number is 0.4, the weight of the storage structure is 0.3, and the weight of the SQL support is 0.3. The weight of the port number is the highest because it is a "preliminary decisive feature for connection establishment". The weights of the storage structure and the SQL support are the same because both of them are basic features for data interaction and reading. is the feature matching value of the jth core feature label (the core feature label is included in all the sub-features mentioned above, and the feature matching value is obtained in the same way as the feature matching value of the sub-feature mentioned above); and the template matching value The template with the highest value is marked as the optimal possible template.
[0051] Step S2: Construct a cross-system semantic graph for each type of heterogeneous system.
[0052] The cross-system semantic graph for each type of heterogeneous system is constructed as follows:
[0053] Step 1: Based on the target data source that has been accessed, automatically extract and normalize the core metadata, and clearly define the "basic node information" of the graph: target data source node collection: synchronize the key information of the target data source from the optimal possible template, and mark it as the "data source node" of the graph, including: basic attributes, technical attributes; basic attributes: data source belonging system (such as "e-commerce order system A" and "logistics warehouse system B"), access template type (such as "MySQL template" and "MQTT Internet of Things template"), data update frequency (sub-feature "update frequency" in the business layer feature dimension); technical attributes: data storage location (such as database IP + table name, MQTT topic), access driver version;
[0054] Business entity extraction: Identify the core business entities (entity nodes in the graph) that are common across systems from the target data sources through "rule matching + manual assistance", such as: E-commerce scenario: extract "user", "order", "goods", "logistics order"; Industrial scenario: extract "equipment", "sensor", "production order", "fault record"; Entity definition refers to industry standard ontology (such as e-commerce reference "Electronic Commerce Data Exchange Specification"), to ensure the uniformity of cross-system entity meaning;
[0055] Entity attribute regularization: Reuse the sub-feature "core field type" under the business layer feature dimension to extract the key attributes of the entity (attribute nodes in the graph) and supplement semantic constraints: basic information: attribute name (such as "user ID", "order amount"), data type (reuse "core field type"); Semantic constraints: attribute unit (such as "order amount" in "yuan / minute"), format specification (such as "mobile phone number" in "11 digits"), whether it is mandatory (such as "user ID" is mandatory);
[0056] Step two: Establish a "global semantic dictionary" to map the "synonyms with different names" data of various heterogeneous systems to a unified semantics: Entity semantic unification: For entities with the same meaning but different names in various heterogeneous systems, define a global unified entity name and semantic interpretation, for example: System A's "member", system B's "user", system C's "customer", are unified as global entity "user", and the semantic interpretation is "all registered / consumed individual / enterprise principal in the platform";
[0057] Attribute semantic unification: For "synonyms with different names" attributes of the same entity, define a global unified attribute name, semantic interpretation, and conversion rule, for example: System attribute name is system A "user number", corresponding global unified attribute name is user unique identifier, semantic interpretation is a code that uniquely identifies a user and cannot be repeated, and conversion rule is direct reuse (consistent format); System attribute name is system B "UID", corresponding global unified attribute name is user unique identifier, semantic interpretation is a code that uniquely identifies a user and cannot be repeated, and conversion rule is direct reuse (consistent format);
[0058] Bind the global semantic dictionary to the graph nodes, and each entity / attribute node is labeled with a "global semantic ID" to ensure accurate matching when aligning subsequent data;
[0059] Step three: Through "business rules + data mining", the semantic relationship between the nodes of the graph is determined, and the isolated nodes form a "traceable and associated" network. The core relationship types include: entity-attribute association (hasAttribute): defines the basic relationship of "entity containing attributes", such as "user" hasAttribute "user unique identifier", "name", and "mobile phone number". Entity-entity business association: based on cross-system business logic, defines the core business relationship between entities, such as: "user" and "order": hasOrder (user order generates order); "order" and "logistics order": hasLogistics (order associated with logistics order); "device" and "sensor": hasSensor (device mounts sensor). Cross-system attribute mapping association (mapTo): marks the mapping relationship between different system attributes and "global unified attributes", such as "system A. user number" mapTo "global, user unique identifier", and associates the corresponding conversion rules (such as unit conversion). Data bloodline association (deriveFrom): records the source of integrated data, such as "global order summary table" deriveFrom "system A. order table", "system B. order table", which facilitates subsequent conflict tracing.
[0060] Step S3: Based on the cross-system semantic graph, the multi-dimensional semantic similarity of the target data source is calculated to complete the semantic alignment between the target data sources, and the integrated data set with unified semantics is output.
[0061] Based on the cross-system semantic graph, the multi-dimensional semantic similarity of the target data source is calculated to complete the semantic alignment between the target data sources, and the integrated data set with unified semantics is output.
[0062] Step one: Based on the cross-system semantic graph, the target data sources (such as system A's order data, system B's transaction data) to be aligned are preprocessed to eliminate the interference of "format / unit / representation differences" on similarity calculation:
[0063] Entity name regularization: convert the entity name in the target data source to a candidate expression of "global entity name", such as system A's "member order" → associated graph "order" entity synonym expression list (graph pre-stored "order" synonyms: member order, transaction order, purchase order); attribute format standardization: unify attribute format / unit according to graph "attribute semantic constraints", such as: date attribute: system A "2024-05-20", system B "05 / 20 / 2024" → standardized to "YYYY-MM-DD" specified by the graph; numerical attribute: system A "order amount (1000 points)", system B "OrderAmt (10 yuan)" → converted to "yuan" unit specified by the graph (1000 points → 10 yuan);
[0064] Relationship expression unification: map the entity relationship name in the target data source to the candidate of the graph "global relationship name", for example, system A "user generates order" → the synonymous expression of the "hasOrder (user order)" relationship in the associated graph (generate order, create order, order);
[0065] Step two: calculate entity semantic similarity , attribute semantic similarity and relationship semantic similarity , according to the comprehensive semantic similarity of entity semantic similarity, attribute semantic similarity and relationship semantic similarity , d1, d2, d3 are weight coefficients, , because the entity is the "core carrier" of the data (first determine "is order or commodity", then talk about others), it is the basis of semantic alignment, the attribute is the "landing feature" of the entity (without "order amount / user ID" and other attributes, the entity cannot be actually integrated), it is as important as the entity, the relationship is the "association logic" between entities, which depends on the existence of entities / attributes, and its importance is less than the first two. Therefore, the value of d1 can be 0.4, the value of d2 can be 0.4, and the value of d3 can be 0.2; Set the comprehensive semantic similarity upper value and the comprehensive semantic similarity lower value (the comprehensive semantic similarity upper value is greater than the comprehensive semantic similarity lower value, the comprehensive semantic similarity upper value and the comprehensive semantic similarity lower value are threshold values based on industry experience + historical data verification + business fault tolerance cost), when the comprehensive semantic similarity ≥ comprehensive semantic similarity upper value, directly map the entity / attribute / relationship of the data source to the graph "global semantic", for example, system A "member order number" → global "order unique identifier", system B "transaction number" → global "order unique identifier"; When the comprehensive semantic similarity ≤ comprehensive semantic similarity lower value, trigger manual intervention to confirm whether it is a new semantic (such as system D's "pre-order", which needs to add a "pre-order" entity and associated rules in the graph); When the comprehensive semantic similarity is between the comprehensive semantic similarity upper value and the comprehensive semantic similarity lower value, align after supplementing "local adaptation rules"; "Local adaptation rules" are small rules that only target the details of the current data source (do not affect the global semantic graph, and do not need to be reused for all data sources), the purpose is to modify "details that do not meet the global standard" to "details that meet the standard", for example: global semantic standard: "order amount" needs to "retain 2 decimal places" (such as 10.50 yuan); System C difference: "order amount" only retains 1 decimal place (such as 10.5 yuan); Local adaptation rule: "add 0 to 2 decimal places" (i.e. 10.5 → 10.50);
[0066] Entity semantic similarity ; is the entity semantic weight coefficient, The "directionality" of the semantic of the entity name is stronger than that of the attribute, and the name is preferred to determine the direction, and then the attribute is verified, so The value of the attribute semantic similarity can be 0.6.
[0067] The name similarity is obtained in the following manner: first, the "data source entity name" and "global entity name" are standardized, and the core semantics are extracted: redundant prefixes / suffixes are removed: for example, the data source entity "E-commerce platform 2024 member order table"→the core name "member order" is extracted; the global entity "order (global)"→the core name "order" is extracted; the expression is unified: for example, "order" (typos)→corrected to "order", "User" (English)→corresponding to "user" (global entity Chinese standard name). The core semantics are split into word vectors ("member", "order" vs "order"), and the cosine similarity calculation method is used to calculate the name similarity;
[0068] The business attribute matching degree is obtained in the following manner: the core mandatory attributes (i.e. "the smallest attribute set for determining the entity", defined by industry business logic) of the target global entity are extracted from the "cross-system semantic graph": for example, the mandatory attributes (obtained from the graph) of the global "order" entity: order unique identifier (such as order number), order amount, order time (3 attributes); the mandatory attributes of the global "logistics order" entity: waybill unique identifier, logistics company name, shipping time (3 attributes). From the target data source that has been accessed, the core business attributes of the data source entity are extracted by reusing the "core field type of business layer feature dimension" sub-feature: example 1: the core attributes of the data source entity "member order": member order number (corresponding to "order unique identifier"), order amount, order time (3 attributes); example 2: the core attributes of the data source entity "logistics order": waybill number (corresponding to "waybill unique identifier"), logistics company, shipping time (3 attributes); the score in the 0-1 interval is calculated according to the "number of coinciding attributes / global mandatory attribute number": example 1: "member order" and global "order": the number of coinciding attributes is 3 (order number, amount, time), and the matching degree is 3 / 3=1; example 2: "logistics order" and global "order": the number of coinciding attributes is 0 (no common attributes), and the matching degree is 0 / 3=0.
[0069] The attribute semantic similarity ; The attribute value compatibility The attribute value compatibility The attribute value compatibility
[0070] The attribute semantic similarity ; The attribute semantic similarity The attribute semantic similarity , the mapping rule pre-stored in the graph has high accuracy, the rule is preferred to determine the result, and the name semantic is only used as a supplement, so The value of the attribute semantic similarity can be 0.7.
[0071] The acquisition method of global mapping matching degree: query the pre-stored "data source attribute-global attribute" association relationship in the "cross-system semantic graph", such as the global attribute is the user unique identifier, the associated data source attributes are system A. User number, system B. UID, system C. Member ID; such as the global attribute is the order amount, the associated data source attributes are system A. Order amount, system B. OrderAmt, system D. Transaction amount; if the two data source attributes to be aligned (such as system A "user number", system B "UID") can find the same global attribute (user unique identifier) in the table → global mapping matching degree = 1; if any data source attribute cannot be found in the table (such as newly added system E "customer ID", the graph has not been marked), or two are mapped to different global attributes (such as system A "order amount" → global "order amount", system B "product amount" → global "product amount") → global mapping matching degree = 0.
[0072] The acquisition method of attribute name semantic similarity: first, standardize the "data source attribute name", extract the core semantics, and split the core semantics into word vectors ("member", "order" vs "order"); calculate the attribute name semantic similarity by using the cosine similarity calculation method;
[0073] Attribute value compatibility The acquisition method of compatible value number: extract "global attribute constraints" (such as numerical range, format, unit: for example, the "order amount" constraint is "positive number, 2 decimal places, unit yuan") from the semantic graph; randomly sample the attribute values of the target data source, and judge one by one whether they meet the constraints (such as the sampled value "10.5" → 0 is added to "10.50" to meet 2 decimal places, and it is calculated as compatible; "-5.20" → violates the positive number constraint, and is not calculated); count the total number of samples that meet the constraints, which is the compatible value number; the total number of samples is the sum of the number of samples that meet the constraints and the number of samples that do not meet the constraints;
[0074] Relationship semantic similarity ; is the relationship semantic weight coefficient, Since the global relationship directly corresponds to the business logic, the accuracy is the highest, so The value of can be 0.8;
[0075] The global relationship matching degree is obtained in the following manner: a "global relationship mapping table" is extracted from the semantic graph (a "system relationship-global relationship" association is pre-stored, such as system A "user generates an order" -> global "hasOrder", and system B "member creates a transaction order" -> global "hasOrder"), if two system relationships to be aligned (such as system A "generates" and system B "creates") are mapped to the same global relationship -> global relationship matching degree = 1; if they are mapped to different global relationships (such as system A "generates" -> "hasOrder", and system C "refunds" -> "refundOrder") -> global relationship matching degree = 0.
[0076] The associated entity matching degree is obtained in the following manner: an entity pair associated with a system relationship is extracted (such as the target data source for which the order data of system A and the transaction data of system B are required, and the extracted entity pair associated with the system relationship is system A "user-member order" and system B "member-transaction order"), and the graph is queried to determine whether the entity pairs are mapped to the same "global entity pair" (such as "user-member order" -> global "user-order", and "member-transaction order" -> global "user-order") -> entity pair complete matching -> matching degree = 1; if the entity pairs are different (such as "user-goods" vs. "user-order") -> matching degree = 0.
[0077] The above formulas are all dimensionless values, and the preset parameters in the formulas are set by a person skilled in the art according to actual conditions.
[0078] The above embodiments can be realized wholly or partially by software, hardware, firmware or any combination thereof. When realized by software, the above embodiments can be realized wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD) or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0079] It should be understood that the size of the sequence number of the above processes does not mean the order of execution in various embodiments of the present application, and the execution order of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0080] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0081] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0082] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0083] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0084] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data integration method for information centers used for summarizing data information from heterogeneous systems, characterized in that, include Step S1: The information center identifies the feature types of target data sources from various heterogeneous systems, determines the matching degree of the feature type identification results, and loads the corresponding target data source driver and preset access strategy when the matching degree reaches a preset threshold. When the matching degree does not reach the threshold, the feature learning mechanism is triggered to update the preset feature template library. Step S2: Construct cross-system semantic graphs for various heterogeneous systems; Step S3: Calculate the multi-dimensional semantic similarity of the accessed target data sources based on the cross-system semantic graph to complete the semantic alignment between the target data sources and output a semantically unified integrated dataset; Based on cross-system semantic graphs, multi-dimensional semantic similarity is calculated for the accessed target data sources to achieve semantic alignment between target data sources, as detailed below: Step 1: Based on the cross-system semantic graph, preprocess the target data source to be aligned to eliminate the interference of differences in format, units, and expression on similarity calculation: Step 2: Calculate entity semantic similarity Attribute semantic similarity and relational semantic similarity The comprehensive semantic similarity is calculated based on entity semantic similarity, attribute semantic similarity, and relation semantic similarity. d1, d2, and d3 are all weighting coefficients. ; Set an upper and lower bound for comprehensive semantic similarity. When the comprehensive semantic similarity is greater than or equal to the upper bound, directly map the entities, attributes, and relationships of the data source to the global semantic graph. When the overall semantic similarity is less than or equal to the lower limit of overall semantic similarity, manual intervention is triggered to confirm whether it is a new semantic meaning; When the comprehensive semantic similarity is between the upper and lower values of comprehensive semantic similarity, alignment is performed after supplementing local adaptation rules. Entity semantic similarity ; For entity semantic weight coefficients; Attribute semantic similarity ; To indicate the similarity of attributes; For attribute value compatibility; attribute identifier similarity ; Assign weight coefficients to attributes; attribute value compatibility. ; Relational semantic similarity ; This represents the semantic weight coefficient of the relation.
2. The information center data integration method for summarizing data information from heterogeneous systems according to claim 1, characterized in that, The target data source for each type of heterogeneous system collects three core feature dimensions, namely the transport layer feature dimension, the storage layer feature dimension, and the business layer feature dimension.
3. The information center data integration method for summarizing data information from heterogeneous systems according to claim 2, characterized in that, Sub-features under the transport layer feature dimension include the target data source's communication protocol, port number, data encryption method, and connection timeout threshold. Sub-features under the storage layer feature dimension: including the storage structure and data encoding format of the target data source; Sub-features under the business layer feature dimension include the update frequency of the target data source, the data volume, and the core field types.
4. The information center data integration method for summarizing data information from heterogeneous systems according to claim 3, characterized in that, The matching degree of the feature type identification results is determined as follows: A protocol template mapping table is established, the actual communication protocol name of the target data source is obtained, and the associated scenario template group is determined based on the actual communication protocol name and the protocol template mapping table. From the associated scenario template group, preferred possible templates are further selected, and the feature matching values of each sub-feature and the preferred possible template are obtained. Simultaneously obtain the matching weights of each sub-feature Through formula Calculate the matching degree of the feature type recognition results. k represents the total number of sub-features.
5. The information center data integration method for summarizing data information from heterogeneous systems according to claim 4, characterized in that, Further refine the selection of preferred templates from the associated scene template group, as follows: Calculate the template matching value for each template in the associated scene template group. This represents the number of core feature labels for the associated scene template group. The weight of the j-th core feature label, The feature matching value for the j-th core feature label; the template matching value The template with the highest value is marked as the preferred possible template.
Citation Information
Patent Citations
Cross-modal semantic alignment method based on multi-source heterogeneous data
CN120579144A
Multi-layered knowledge base system and processing method thereof
US20210192372A1