Big data standardization method and system based on data integration
Through the combination of dynamic format adapter and quantum annealing algorithm combined with reinforcement learning Schema algorithm, the format conflict and redundancy problems of heterogeneous data sources are solved, efficient data integration and interoperability are achieved, and the accuracy and efficiency of data analysis are improved.
Patent Information
- Application Number
- CN202510488545.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-18
AI Technical Summary
When dealing with heterogeneous data sources, the prior art faces problems such as format conflict, data redundancy, inconsistency, data loss and analysis accuracy, resulting in increased complexity of data integration and interoperability.
The dynamic format adapter is used to combine the reinforcement learning Schema algorithm to convert the data into an intermediate format; quantum superposition states of fields and knowledge graphs are generated by quantum state encoding, and the best mapping nodes are solved through quantum annealing algorithm, and data standardization is performed by combining semantic constraints and standard formats.
It improves the accuracy and efficiency of data conversion, solves the problem of data heterogeneity, and realizes efficient data integration and interoperability.
Smart Images

Figure CN120336307A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data standardization, and particularly to a big data standardization method and system based on data integration. Background Art
[0002] With the rapid development of information technology, big data has become an important resource in modern society. The core value of big data lies in providing support for decision-making through the analysis and mining of massive data. However, the diversity and complexity of big data also bring many challenges, especially in data integration and interoperability. Data integration refers to integrating data from different data sources to form a unified data view, while data interoperability refers to the ability of different systems to seamlessly exchange and share data. In a big data environment, data sources are usually heterogeneous, including structured data, semi-structured data, and unstructured data, which makes data integration and interoperability particularly complex.
[0003] Traditional data integration methods mainly include ETL (Extract, Transform, Load) tools, which can extract data from different data sources, perform transformation and cleaning, and then load it into the target database. However, with the increasing diversity of data sources and the rapid increase in data volume, traditional ETL tools face many challenges when dealing with heterogeneous data sources.
[0004] The data formats of different data sources vary greatly, resulting in format conflicts easily occurring during the data integration process; and due to the independence between data sources, the same data may exist in different forms in different data sources, leading to data redundancy and inconsistency, which not only increases the cost of data storage and processing, but also may affect the accuracy of data analysis; some data may be lost during the data acquisition process, which will affect the integrity of the data and the reliability of the analysis results. Summary of the Invention
[0005] To overcome the above deficiencies of the prior art, the present invention proposes a big data standardization method based on data integration, including:
[0006] Obtaining the original data of multiple different data sources;
[0007] Using a dynamic format adapter to convert the original data into intermediate format data based on data conversion mapping rules in combination with a reinforcement learning Schema algorithm;
[0008] Based on each field in the intermediate format data, generating a quantum superposition state of the field and a node in the knowledge graph by using quantum state encoding; based on the quantum superposition state, using a quantum annealing algorithm to solve the best mapped node of the field in the knowledge graph;
[0009] Based on the optimal mapping nodes of each field in the intermediate format data, combining the semantic constraints and standard format of the nodes in the knowledge graph, standardize the mapping of the data corresponding to each field in the intermediate format data to obtain standardized data.
[0010] Optionally, the process of constructing the knowledge graph includes:
[0011] Determine the core entities and attributes of the external multi-source data as the nodes of the initial knowledge graph, where the nodes include entity nodes and attribute nodes;
[0012] Based on each sample field in the external multi-source data, use the semantic similarity between the sample field and each node to determine the mapping node of the sample field, and map the sample field to the initial knowledge graph based on the mapping node;
[0013] In the initial knowledge graph that has completed mapping and filling, set the semantic constraints and standard format of the attribute nodes to obtain a knowledge graph for data conversion.
[0014] Optionally, the generation of the quantum superposition state of the field and the nodes in the knowledge graph by using quantum state encoding includes:
[0015] Based on the semantic similarity between the field and each node in the knowledge graph, and in combination with a similarity threshold, select multiple candidate mapping nodes of the field from the knowledge graph;
[0016] Use quantum state encoding to quantify the semantic relationship between the field and each candidate mapping node as a probability amplitude;
[0017] Combine the probability amplitudes between the field and each candidate mapping node to obtain the quantum superposition state of the field and the nodes in the knowledge graph.
[0018] Optionally, based on the quantum superposition state, using the quantum annealing algorithm to solve the optimal mapping node of the field in the knowledge graph includes:
[0019] Utilize the quantum tunneling effect in the quantum degenerate algorithm to iteratively update the probability amplitudes of the candidate mapping nodes in the quantum superposition state multiple times based on the quantum superposition state, a preset temperature parameter, and an energy parameter;
[0020] After each iterative update, the temperature parameter actively decreases according to a preset attenuation coefficient; stop the iterative update until the temperature parameter drops to a threshold, and take the candidate node with the highest probability amplitude in the quantum superposition state when the iterative update stops as the optimal mapping node.
[0021] Optionally, the calculation process of the semantic similarity between the field and each node in the knowledge graph includes:
[0022] Extract the semantic vector of the field by using named entity recognition technology;
[0023] Calculate the cosine similarity between the semantic vector of the field and the semantic vectors of each node in the knowledge graph to obtain the semantic similarity between the field and each node.
[0024] Optionally, the use of the dynamic format adapter to convert the original data into intermediate format data based on the data conversion mapping rule in combination with the reinforcement learning Schema algorithm includes:
[0025] Use the dynamic format adapter to identify the format of the original data and parse the initial field type in the original data;
[0026] Convert the original data into an internal data structure based on the format by using a parser, denoted as the initial parsed data;
[0027] Revise the initial field type based on the initial parsed data through the reinforcement learning schema algorithm to obtain the field type in the original data;
[0028] Convert the original data into intermediate format data based on the data conversion mapping rule and the machine learning algorithm in combination with the field type.
[0029] Optionally, the conversion of the original data into intermediate format data based on the data conversion mapping rule and the machine learning algorithm in combination with the field type includes:
[0030] Based on the field type of the original data and the predefined mapping rule, determine the original data with mapping rules and the original data without mapping rules;
[0031] Convert the original data with mapping rules into intermediate format data based on the field type of the original data in combination with the mapping rule;
[0032] Use the machine learning algorithm to perform semantic inference and conversion on the original data without mapping rules based on the field type of the original data to obtain intermediate format data.
[0033] Optionally, before using the dynamic format adapter to convert the original data into intermediate format data based on the data conversion mapping rule and the machine learning algorithm, it further includes:
[0034] Use edge computing technology to extract and screen the original data from multiple different data sources to obtain the processed original data.
[0035] Based on the same inventive concept, the present invention provides a big data standardization system based on data integration, including:
[0036] A data acquisition module for acquiring raw data from multiple different data sources;
[0037] An intermediate format conversion module for converting the raw data into intermediate format data by using a dynamic format adapter based on data conversion mapping rules in combination with a reinforcement learning Schema algorithm;
[0038] A node solving module for generating a quantum superposition state of each field in the intermediate format data and a node in the knowledge graph by using quantum state encoding based on each field in the intermediate format data; and solving for the best mapping node of the field in the knowledge graph by using a quantum annealing algorithm based on the quantum superposition state;
[0039] A standardization mapping module for performing a standardization mapping on the data corresponding to each field in the intermediate format data based on the best mapping node of each field in the intermediate format data, in combination with the semantic constraints and standard format of the nodes in the knowledge graph, to obtain standardized data.
[0040] Optionally, the node solving module is specifically configured to:
[0041] Determine the core entities and attributes of external multi-source data as the nodes of the initial knowledge graph, where the nodes include entity nodes and attribute nodes;
[0042] Based on each sample field in the external multi-source data, determine the mapping node of the sample field by using the semantic similarity between the sample field and each node, and map the sample field to the initial knowledge graph based on the mapping node;
[0043] In the initial knowledge graph after mapping and filling, set the semantic constraints and standard format of the attribute nodes to obtain a knowledge graph for data conversion.
[0044] Optionally, the node solving module is specifically configured to:
[0045] Based on the semantic similarity between the field and each node in the knowledge graph, and in combination with a similarity threshold, select multiple candidate mapping nodes of the field from the knowledge graph;
[0046] Use quantum state encoding to quantify the semantic relationship between the field and each candidate mapping node into a probability amplitude;
[0047] Combine the probability amplitudes between the field and each candidate mapping node to obtain a quantum superposition state of the field and the nodes in the knowledge graph.
[0048] Optionally, the node solving module is specifically configured to:
[0049] Utilize the quantum tunneling effect in the quantum degradation algorithm to iteratively update the probability amplitudes of the candidate mapping nodes in the quantum superposition state multiple times based on the quantum superposition state, a preset temperature parameter, and an energy parameter;
[0050] After each iterative update, the temperature parameter actively decreases according to a preset attenuation coefficient; stop the iterative update until the temperature parameter drops to a threshold, and use the candidate node with the highest probability amplitude in the quantum superposition state when the iterative update stops as the optimal mapping node.
[0051] Optionally, the node solving module is specifically configured to:
[0052] Extract the semantic vector of the field using named entity recognition technology;
[0053] Calculate the cosine similarity between the semantic vector of the field and the semantic vectors of each node in the knowledge graph to obtain the semantic similarity between the field and each node.
[0054] Optionally, the intermediate format conversion module is specifically configured to:
[0055] Use a dynamic format adapter to identify the format of the original data and parse the initial field type in the original data;
[0056] Based on the format, use a parser to convert the original data into an internal data structure, denoted as the initial parsed data;
[0057] Use the reinforcement learning schema algorithm to correct the initial field type based on the initial parsed data to obtain the field type in the original data;
[0058] Based on the data conversion mapping rules and machine learning algorithms, combine the field type to convert the original data into intermediate format data.
[0059] Optionally, the intermediate format conversion module is specifically configured to:
[0060] Based on the field type of the original data and predefined mapping rules, determine the original data with mapping rules and the original data without mapping rules;
[0061] Based on the field type of the original data and in combination with the mapping rules, convert the original data with mapping rules into intermediate format data;
[0062] Use machine learning algorithms to perform semantic inference and conversion on the original data without mapping rules based on the field type of the original data to obtain intermediate format data.
[0063] Optionally, the system further includes an edge computing module for:
[0064] Using edge computing technology to extract and filter the raw data from the multiple different data sources to obtain processed raw data.
[0065] On the other hand, the present application also provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected by a bus;
[0066] The memory is used to store one or more programs;
[0067] When the one or more programs are executed by the at least one processor, a big data standardization method based on data integration as described above is implemented.
[0068] On the other hand, the present application also provides a computer-readable storage medium with an execution program stored thereon. When the execution program is executed, a big data standardization method based on data integration as described above is implemented.
[0069] Compared with the closest prior art, the beneficial effects of the present invention are as follows:
[0070] A big data standardization method and system based on data integration provided by the present invention includes: obtaining raw data from multiple different data sources; using a dynamic format adapter to convert the raw data into intermediate format data based on data conversion mapping rules combined with a reinforcement learning Schema algorithm; generating a quantum superposition state of each field in the intermediate format data and a node in a knowledge graph based on the quantum state encoding; solving for the best mapping node of the field in the knowledge graph based on the quantum superposition state using a quantum annealing algorithm; based on the best mapping node of each field in the intermediate format data, combining the semantic constraints and standard format of the nodes in the knowledge graph, performing a standardized mapping on the data corresponding to each field in the intermediate format data to obtain standardized data; the dynamic format adapter in the present invention integrates multi-source data and parses it into intermediate format data that the system can process, solving the data heterogeneity problem, and the schema algorithm solves the structural deviation between the raw data and the conversion format by correcting the field type, improving the accuracy of the intermediate format data; the quantum superposition state can simultaneously retain multiple mapping possibilities, and the quantum annealing algorithm can jump out of the local optimum, and through efficient parallel search capabilities, quickly obtain the global optimum solution, improving the accuracy and efficiency of node mapping; performing semantic constraint and standard format mapping in the knowledge graph based on the best mapping node realizes data standardization processing. Description of the Drawings
[0071] Figure 1Flow schematic diagram of a big data standardization method based on data integration provided by the present invention;
[0072] Figure 2 Structural schematic diagram of a big data standardization system based on data integration provided by the present invention;
[0073] Figure 3 Structural schematic diagram of an electronic device provided by the present invention. Detailed implementation manners
[0074] The following further elaborates on the detailed implementation manners of the present invention in conjunction with the accompanying drawings.
[0075] Embodiment 1
[0076] A big data standardization method based on data integration provided by the present invention, as Figure 1 shown, includes:
[0077] S1. Obtain the original data of multiple different data sources;
[0078] S2. Use a dynamic format adapter to convert the original data into intermediate format data based on data conversion mapping rules in combination with a reinforcement learning Schema algorithm;
[0079] S3. Based on each field in the intermediate format data, use quantum state encoding to generate a quantum superposition state of the field and nodes in the knowledge graph; based on the quantum superposition state, use a quantum annealing algorithm to solve for the best mapped node of the field in the knowledge graph;
[0080] S4. Based on the best mapped nodes of each field in the intermediate format data, in combination with the semantic constraints and standard format of the nodes in the knowledge graph, perform a standardized mapping on the data corresponding to each field in the intermediate format data to obtain standardized data.
[0081] In step S1, the original data of different data sources is obtained. It is responsible for extracting data from different data sources, supporting multiple data formats and protocols, and ensuring wide compatibility of the data. Specifically:
[0082] Use multiple data source access devices distributed near the multiple different data sources to obtain the original data of the multiple different data sources; each data source access module corresponds to one or more of the data sources;
[0083] Distribute the data source access devices to edge nodes, and use edge computing technology (a set of multiple edge computing devices) to extract and screen the original data of the multiple different data sources to obtain processed original data, realizing real-time data access and preliminary processing, and reducing the load of the system.
[0084] In the Internet of Things environment, sensors may frequently send similar data points. Multiple data sources are connected to devices and distributed near the data sources. At the edge node, the data is initially processed (such as filtering, aggregating, compressing, etc.), and only key information is sent to the dynamic format adapter, which can significantly reduce the amount of data that needs to be transmitted to the dynamic format adapter, thereby reducing network bandwidth requirements and transmission costs and accelerating transmission efficiency.
[0085] In step S2, the dynamic format adapter converts the original data into intermediate format data based on the data conversion mapping rules in combination with the reinforcement learning Schema algorithm.
[0086] Among them, the dynamic format adapter supports real-time learning and adaptation to new data formats. Through machine learning algorithms, the system can automatically learn the format rules of new data sources and generate corresponding conversion logics. The design process of the dynamic format adapter is as follows:
[0087] Design a modular architecture including a data format recognition module, a conversion logic generation module, and a data conversion module based on the data format conversion requirements of different data sources;
[0088] Introduce machine learning algorithms in the conversion logic generation module, predict the best conversion strategy through learning historical data, and continuously update the machine learning algorithms as more data formats are added to achieve the function that the dynamic format adapter supports real-time learning and adaptation to new data formats.
[0089] S21. Use the dynamic format adapter to recognize the format of the original data and parse the initial field types in the original data; based on the format, convert the original data into an internal data structure using a parser, denoted as the initial parsed data.
[0090] Among them, recognizing the format of the original data is to call the parsers corresponding to different data formats to parse the original data and convert it into an internal data structure (such as a dictionary, a data frame) that can be operated by the program, providing a basis for subsequent Schema inference and field extraction; by parsing the initial field types in the original data, that is, predicting the field types through format features (such as values wrapped in quotes are strings), the initial search space of the reinforcement learning schema algorithm (reinforcement learning architecture algorithm) is reduced, and the data processing efficiency is improved.
[0091] S22. Based on the initial parsed data, correct the initial field types through the reinforcement learning schema algorithm to obtain the field types in the original data.
[0092] By using the enhanced learning schema algorithm to correct the field types and nested relationships in the original data, the structural deviation between the original data and the format to be converted is solved, providing a data structure reference for the subsequent rule engine and machine learning, ensuring that the conversion logic corresponds to the data structure, and improving the accuracy of the intermediate format data.
[0093] S23. Based on the data conversion mapping rules and machine learning algorithms, convert the original data into intermediate format data in combination with the field types, including:
[0094] Based on the field types of the original data and predefined mapping rules, determine the original data with mapping rules and the original data without mapping rules;
[0095] Based on the field types of the original data in combination with the mapping rules, convert the original data with mapping rules into intermediate format data;
[0096] Use machine learning algorithms to perform semantic inference and conversion on the original data without mapping rules based on the field types of the original data to obtain intermediate format data.
[0097] Among them, the data structures of different data sources may vary greatly, and the intermediate format needs to be able to be compatible with the structures of different data sources; if the data volume is large, the intermediate format needs to be designed as an efficient structure to reduce the conversion and transmission overhead.
[0098] For the original data with clear mapping rules, it can be directly converted with high conversion efficiency. For the original data without mapping rules, machine learning algorithms can be used to intelligently parse and convert the data, solve the problems of complex or unstructured formats, and improve the coverage and accuracy of data conversion; combining rules and machine learning can process both structured data and unstructured data simultaneously, adapting to the diversity of multiple data sources.
[0099] Use a dynamic format adapter to extract the original data from different data sources and preliminarily convert it into an intermediate format that the system can process, facilitating subsequent processing. It can solve the problem of data source heterogeneity, such as different protocols, encodings, structures, etc., ensure that the data can be received by the system, improve the compatibility of data access, that is, support multiple source data types, and provide high-quality input for subsequent modules, that is, the obtained intermediate format data can ensure being correctly parsed and processed by subsequent modules, retaining the accuracy of the original data.
[0100] In step S3, the pre-construction process of the knowledge graph is as follows:
[0101] Obtain the semantic information of multiple sample fields in external multi-source data;
[0102] Define core entities and attributes based on the external multi-source data, and determine the core entities and attributes as nodes of the initial knowledge graph, where the nodes include entity nodes and attribute nodes;
[0103] Based on each of the sample fields, obtain the semantic similarity between the sample field and each of the nodes; determine the mapping nodes using the semantic similarity between the sample field and each of the nodes, and map the sample field to the initial knowledge graph based on the mapping nodes;
[0104] And represent the semantic relationship between the field name and the core entities and attributes as edges, and populate the initial knowledge graph based on the nodes and edges;
[0105] In the initial knowledge graph that has completed mapping and population, set the semantic constraints and standard formats of the attribute nodes to obtain a knowledge graph for data transformation.
[0106] The semantic constraints in the knowledge graph are used to ensure that the field values conform to the business logic rules, which do not directly change the format but affect the legality of the transformation;
[0107] For example: verify the data type (e.g., "price" must be numeric, and an error will be reported if the original value is "100 yuan"); or verify the business rule (e.g., "date of birth ≤ current date", and an exception will be marked if it is a future date).
[0108] Only when the field value passes the semantic constraints will it enter the format conversion stage.
[0109] S31. Before performing quantum state encoding, extract multiple fields from the intermediate format data, where the fields include field names and field values. Common types of intermediate format data include:
[0110] JSON (JavaScript Object Notation, a lightweight data interchange format) data: key-value pair structure, the field name is the key, and the field name is obtained by parsing the first-level keys of the JSON data, and the field value is obtained by directly reading the value corresponding to the key.
[0111] XML (eXtensible Markup Language) data: tag structure, the field name is the tag name, and the field name is obtained by parsing the sub-tags under the root node of the XML data, and the text content within the tag is extracted to obtain the field value.
[0112] CSV (Comma-Separated Values, a text file format) data: table structure, the field name is usually in the first row, and the field name is obtained by reading the first row of the file, and the field value is obtained from the corresponding column position in the subsequent rows.
[0113] S32. Generate the quantum superposition state of each field in the intermediate format data and the nodes in the knowledge graph by using quantum state encoding.
[0114] For each of the fields in the intermediate format data:
[0115] a. Calculate the semantic similarity between the field and each node in the knowledge graph. Based on the semantic similarity between the field and each node in the knowledge graph and in combination with a similarity threshold, select multiple candidate mapping nodes of the field from the knowledge graph. This includes:
[0116] Use named entity recognition technology or a word vector model to extract the semantic vector of the field;
[0117] Obtain the semantic similarity between the field and each of the nodes by calculating the cosine similarity between the semantic vector of the field and the semantic vectors of each node in the knowledge graph.
[0118] Select the nodes in the knowledge graph with a semantic similarity greater than the preset similarity threshold as candidate mapping nodes.
[0119] By screening the nodes in the knowledge graph through semantic similarity, the computational amount in the subsequent quantum matching process is significantly reduced, and the data processing efficiency is improved.
[0120] b. Use quantum state encoding to quantify the semantic relationship between the field and each of the candidate mapping nodes into probability amplitudes; combine the probability amplitudes between the field and each of the candidate mapping nodes to obtain the quantum superposition state of the field and the nodes in the knowledge graph.
[0121] By using the quantum superposition state to represent the association probabilities between a field and multiple candidate nodes simultaneously, the inefficiency problem of multiple serial matches required by traditional methods is avoided, and the processing speed in complex semantic scenarios (such as polysemy) is significantly improved.
[0122] S33. Based on the quantum superposition state, use the quantum annealing algorithm to solve for the best mapping node of the field in the knowledge graph. Specifically:
[0123] For each of the fields in the intermediate format data:
[0124] Take the probability amplitudes in the quantum superposition state corresponding to the field as the initial state;
[0125] Utilize the quantum tunneling effect in the quantum annealing algorithm to iteratively update the probability amplitudes of the candidate mapping nodes in the quantum superposition state multiple times based on the quantum superposition state, a preset temperature parameter (initial high temperature parameter), and an energy parameter (semantic similarity calculation rule based on semantic vectors).
[0126] After each iteration update, the temperature parameter actively decreases according to a preset attenuation coefficient; until the temperature parameter drops to a threshold value, the iteration update stops, and the candidate node with the highest probability amplitude in the quantum superposition state when the iteration update stops is used as the optimal mapping node.
[0127] Among them, in each iteration, the probability amplitude of the candidate mapping node is adjusted according to the semantic similarity. After multiple iteration updates, the system will converge to a stable state, and the node with the highest probability amplitude is retained as the optimal mapping node.
[0128] Quantum annealing can simultaneously explore the matching possibilities of multiple candidate mapping nodes through the quantum tunneling effect, avoiding the traditional algorithm from falling into local optimal solutions; combined with the adjustment of the probability amplitude and the energy function (semantic similarity), while retaining the business semantic relevance (such as the polysemy of the same field in different scenarios), it quickly converges to the most logical mapping node.
[0129] In step S4, according to the semantic constraints and standard format corresponding to the optimal mapping node, the formats of multiple fields in the intermediate format data are automatically corrected to obtain standardized data, so that it can be used for data integration and interoperability of multi-source data.
[0130] Among them, semantic constraints are used to ensure that the field values conform to business logic rules. The field values that have passed the semantic constraints are rewritten in the standard format, that is, illegal values can be filtered through semantic constraints. If the field value cannot pass the semantic constraints, the format conversion is terminated and the error is recorded.
[0131] Embodiment 2
[0132] Based on the same inventive concept, the present invention also provides a big data standardization system based on data integration, as Figure 2 shown, including:
[0133] A data acquisition module for acquiring raw data from multiple different data sources;
[0134] An intermediate format conversion module for converting the raw data into intermediate format data by using a dynamic format adapter based on data conversion mapping rules in combination with the enhanced learning Schema algorithm;
[0135] A node solving module for generating a quantum superposition state of each field in the intermediate format data and a node in the knowledge graph by using quantum state encoding based on each field in the intermediate format data; and solving the optimal mapping node of the field in the knowledge graph by using the quantum annealing algorithm based on the quantum superposition state.
[0136] A standard mapping module for performing standard mapping on the data corresponding to each field in the intermediate format data based on the best mapping nodes of each field in the intermediate format data, in combination with the semantic constraints and standard format of the nodes in the knowledge graph, to obtain standardized data.
[0137] In a possible implementation manner, the above node solving module is specifically configured to:
[0138] Determine the core entities and attributes of the external multi-source data as the nodes of the initial knowledge graph, where the nodes include entity nodes and attribute nodes;
[0139] Based on each sample field in the external multi-source data, use the semantic similarity between the sample field and each node to determine the mapping node of the sample field, and map the sample field to the initial knowledge graph based on the mapping node;
[0140] In the initial knowledge graph that has completed mapping and filling, set the semantic constraints and standard format of the attribute nodes to obtain a knowledge graph for data conversion.
[0141] In a possible implementation manner, the above node solving module is specifically configured to:
[0142] Based on the semantic similarity between the field and each node in the knowledge graph, in combination with a similarity threshold, select multiple candidate mapping nodes of the field from the knowledge graph;
[0143] Use quantum state encoding to quantify the semantic relationship between the field and each candidate mapping node as a probability amplitude;
[0144] Combine the probability amplitudes between the field and each candidate mapping node to obtain the quantum superposition state of the field and the nodes in the knowledge graph.
[0145] In a possible implementation manner, the above node solving module is specifically configured to:
[0146] Utilize the quantum tunneling effect in the quantum annealing algorithm to iteratively update the probability amplitudes of the candidate mapping nodes in the quantum superposition state multiple times based on the quantum superposition state, a preset temperature parameter, and an energy parameter;
[0147] After each iterative update, the temperature parameter actively decreases according to a preset attenuation coefficient; stop the iterative update until the temperature parameter drops to a threshold, and use the candidate node with the highest probability amplitude in the quantum superposition state when the iterative update stops as the best mapping node.
[0148] In a possible implementation manner, the above node solving module is specifically configured to:
[0149] Extract the semantic vector of the field using named entity recognition technology;
[0150] By calculating the cosine similarity between the semantic vector of the field and the semantic vectors of each node in the knowledge graph, obtain the semantic similarity between the field and each of the nodes.
[0151] In a possible implementation manner, the above intermediate format conversion module is specifically configured to:
[0152] Use a dynamic format adapter to identify the format of the original data and parse the initial field type in the original data;
[0153] Based on the format, use a parser to convert the original data into an internal data structure, denoted as initial parsed data;
[0154] Based on the initial parsed data through a reinforcement learning schema algorithm, correct the initial field type to obtain the field type in the original data;
[0155] Based on a data conversion mapping rule and a machine learning algorithm, combine the field type to convert the original data into intermediate format data.
[0156] In a possible implementation manner, the above intermediate format conversion module is specifically configured to:
[0157] Based on the field type of the original data and a predefined mapping rule, determine the original data with a mapping rule and the original data without a mapping rule;
[0158] Based on the field type of the original data and in combination with the mapping rule, convert the original data with a mapping rule into intermediate format data;
[0159] Use a machine learning algorithm based on the field type of the original data to perform semantic inference and conversion on the original data without a mapping rule to obtain intermediate format data.
[0160] In a possible implementation manner, the above system further includes an edge computing module for:
[0161] Use edge computing technology to extract and filter the original data of the multiple different data sources to obtain the processed original data.
[0162] Embodiment 3
[0163] As Figure 3As shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and this data can be called and / or modified when the instructions are executed.
[0164] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a big data standardization method based on data integration in the above embodiment.
[0165] Embodiment 4
[0166] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of a big data standardization method based on data integration in the above embodiment can be implemented.
[0167] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0168] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0169] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0170] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0171] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present invention, various changes, modifications, or equivalent replacements can still be made to the specific implementation manners of the application. However, these changes, modifications, or equivalent replacements are all within the scope of the protection of the claims pending for approval of the application.
Claims
1. A big data standardization method based on data integration, characterized in that, Including: Obtaining raw data from multiple different data sources; Using a dynamic format adapter to convert the raw data into intermediate format data based on data conversion mapping rules combined with a reinforcement learning Schema algorithm; Based on each field in the intermediate format data, generating a quantum superposition state of the field and nodes in the knowledge graph using quantum state encoding; based on the quantum superposition state, solving for the best mapping node of the field in the knowledge graph using a quantum annealing algorithm; Based on the best mapping node of each field in the intermediate format data, and combining the semantic constraints and standard format of the nodes in the knowledge graph, performing a standardized mapping on the data corresponding to each field in the intermediate format data to obtain standardized data.
2. The method according to claim 1, wherein The construction process of the knowledge graph includes: Determining the core entities and attributes of external multi-source data as the nodes of the initial knowledge graph, where the nodes include entity nodes and attribute nodes; Based on each sample field in the external multi-source data, determining the mapping node of the sample field using the semantic similarity between the sample field and each node, and mapping the sample field to the initial knowledge graph based on the mapping node; In the initial knowledge graph after mapping filling, setting the semantic constraints and standard format of the attribute nodes to obtain a knowledge graph for data conversion.
3. The method according to claim 1, wherein The generating of the quantum superposition state of the field and nodes in the knowledge graph using quantum state encoding includes: Based on the semantic similarity between the field and each node in the knowledge graph, and in combination with a similarity threshold, selecting multiple candidate mapping nodes of the field from the knowledge graph; Using quantum state encoding to quantify the semantic relationship between the field and each candidate mapping node as a probability amplitude; Combining the probability amplitudes between the field and each candidate mapping node to obtain the quantum superposition state of the field and nodes in the knowledge graph.
4. The method according to claim 3, characterized in that, The solving for the best mapping node of the field in the knowledge graph using a quantum annealing algorithm based on the quantum superposition state includes: Using the quantum tunneling effect in the quantum degradation algorithm, and based on the quantum superposition state, a preset temperature parameter, and an energy parameter, iteratively updating the probability amplitudes of the candidate mapping nodes in the quantum superposition state multiple times; After each iterative update, the temperature parameter actively decreases according to a preset decay coefficient; stop iterative update until the temperature parameter drops to a threshold, and take the candidate node with the highest probability amplitude in the quantum superposition state when stopping iterative update as the best mapping node.
5. The method according to claim 3, characterized in that, The calculation process of the semantic similarity between the field and each node in the knowledge graph includes: Using named entity recognition technology to extract the semantic vector of the field; By calculating the cosine similarity between the semantic vector of the field and the semantic vectors of each node in the knowledge graph, obtaining the semantic similarity between the field and each node.
6. The method according to claim 1, characterized in that, The converting of the raw data into intermediate format data using a dynamic format adapter based on data conversion mapping rules combined with a reinforcement learning Schema algorithm includes: Using a dynamic format adapter to identify the format of the raw data and parse the initial field type in the raw data; Convert the original data into an internal data structure based on the format using a parser, denoted as initial parsed data; Based on the initial parsed data, correct the initial field types through a reinforcement learning schema algorithm to obtain the field types in the original data; Based on data conversion mapping rules and machine learning algorithms, convert the original data into intermediate format data in combination with the field types.
7. The method according to claim 6, wherein The conversion of the original data into intermediate format data based on data conversion mapping rules and machine learning algorithms in combination with the field types includes: Based on the field types of the original data and predefined mapping rules, determine the original data with mapping rules and the original data without mapping rules; Based on the field types of the original data in combination with the mapping rules, convert the original data with mapping rules into intermediate format data; Use machine learning algorithms based on the field types of the original data to perform semantic inference and conversion on the original data without mapping rules to obtain intermediate format data.
8. The method according to claim 1, characterized in that Before converting the original data into intermediate format data using a dynamic format adapter based on data conversion mapping rules and machine learning algorithms, it further includes: Use edge computing technology to extract and filter the original data from multiple different data sources to obtain processed original data.
9. A big data standardization system based on data integration, characterized in that, It includes: A data acquisition module for acquiring the original data from multiple different data sources; An intermediate format conversion module for converting the original data into intermediate format data using a dynamic format adapter based on data conversion mapping rules in combination with a reinforcement learning Schema algorithm; A node solving module for generating a quantum superposition state of each field in the intermediate format data with nodes in the knowledge graph based on quantum state encoding; based on the quantum superposition state, use a quantum annealing algorithm to solve for the best mapped node of each field in the knowledge graph; A standardized mapping module for performing standardized mapping on the data corresponding to each field in the intermediate format data based on the best mapped node of each field in the intermediate format data in combination with the semantic constraints and standard format of the nodes in the knowledge graph to obtain standardized data.
10. The system according to claim 9, wherein The node solving module specifically is used for: Determine the core entities and attributes of external multi-source data as the nodes of the initial knowledge graph, where the nodes include entity nodes and attribute nodes; Based on each sample field in the external multi-source data, determine the mapped node of the sample field using the semantic similarity between the sample field and each node, and map the sample field to the initial knowledge graph based on the mapped node; In the initial knowledge graph after completing mapping and filling, set the semantic constraints and standard format of the attribute nodes to obtain a knowledge graph for data conversion.
Citation Information
Patent Citations
Intrusion detection and response method and system of satellite internet target range
CN119155101A
Digital precise drainage sale management method and system fused with knowledge graph
CN119168028A
Systems, methods, and apparatus for recursive quantum computing algorithms
US20080313114A1
Knowledge graph reasoning systems using self-supervised reinforcement learning and methods thereof
US20240249154A1
Cited By
Large model-based standardized data processing method, electronic equipment, storage medium and computer program product
CN121009134A
Standardized data processing method based on large model, electronic device, storage medium and computer program product
CN121009134B
Intelligent conversion method and system for data units
CN121412399A