Distributed fusion method for multi-source heterogeneous data
By employing a distributed fusion method for multi-source heterogeneous data, using logarithmic ratio transformation, range-based sharding, and the Edge for Cloud model, the preprocessing and consistency issues of large-scale multi-source heterogeneous data are solved, achieving efficient data fusion and real-time processing.
Patent Information
- Application Number
- CN202411035308.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2044-07-31
AI Technical Summary
Existing data fusion technologies suffer from problems such as complex and time-consuming data preprocessing, difficulty in ensuring consistency in a distributed environment, and low efficiency due to large differences in heterogeneous data formats and semantics when dealing with large-scale, multi-source, and heterogeneous data, making it difficult to meet real-time processing requirements.
A distributed fusion approach for multi-source heterogeneous data is adopted, including data preprocessing, data sharding, multi-level caching, and distributed fusion. Data preprocessing defines the data space through logarithmic ratio transformation; data sharding adopts a range-based sharding method; multi-level caching uses the Edge for Cloud model, combined with overhead control mechanisms and parallel processing technology; and heterogeneous data mapping uses machine learning to optimize rules to ensure data consistency.
It achieves data dimension invariance, fast retrieval and classification in multi-source heterogeneous data processing, improves data processing efficiency, meets real-time requirements, and optimizes NP-hard problems in the data fusion process.
Smart Images

Figure CN118916411B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of big data and edge computing, and particularly to a multi-source heterogeneous data distributed fusion method. BACKGROUND
[0002] In the era of big data, the multi-source and heterogeneous nature of data has become a major challenge in data processing and analysis. Data sources are diverse, including structured data, semi-structured data and unstructured data, which are usually distributed across different systems and platforms, with different formats and semantics. The existence of such multi-source and heterogeneous data makes it difficult for traditional data processing methods to cope, especially in application scenarios that require real-time processing and fusion of these data.
[0003] Structured data is usually stored in relational databases, characterized by strict schema, fixed data types and clear relationships. Common sources of structured data include enterprise business systems, financial transaction records, etc. Semi-structured data contains some structural information, but not as strict as structured data, such as XML and JSON format data. They are widely used in web data, API returned data and other scenarios. Unstructured data has no fixed structure, commonly found in text, image, video and other types of data, which are widely available from social media, sensor data, log files, etc.
[0004] In order to effectively utilize these multi-source and heterogeneous data, data fusion technology has emerged. Data fusion aims to integrate data from different sources and formats to form a unified view for subsequent analysis and mining. However, existing data fusion techniques still face many challenges when dealing with large-scale, multi-source, heterogeneous data. First, data preprocessing is a complex and time-consuming process, including data cleaning, format conversion, outlier handling, etc. Second, in a distributed environment, ensuring data consistency and integrity is a key issue for data fusion. Traditional distributed consistency algorithms such as Paxos and Raft can solve consistency problems to some extent, but are inefficient in large-scale data processing and cannot meet real-time processing requirements. Third, the format and semantic differences of heterogeneous data are large, and how to effectively map and convert them is another major difficulty in achieving data fusion. Existing format conversion methods are usually inefficient and difficult to adapt to dynamically changing data sources. SUMMARY
[0005] To address the above problems, the present application aims to provide a multi-source and heterogeneous data distributed fusion method to solve the problems raised in the background.
[0006] The purpose of the present application is achieved by the following technical solutions:
[0007] 1. A multi-source heterogeneous data distributed fusion method, comprising: data preprocessing, data slicing, multi-level caching, and distributed fusion.
[0008] Further, the data preprocessing defines a data space containing m groups of data elements, and simultaneously performs a logarithmic ratio transformation on all components of the m groups of data elements, but the reference is not a single component, but the geometric mean of all components, which is:
[0009]
[0010] where x i is the i-th initial data element, and z i is the i-th data element.
[0011] Further, the data slicing step is specifically defined as follows: the data space is Z, the data space includes m data, each data includes n groups of labels, if the label does not contain any element, the data space is not null, denoted as null, then there are
[0012]
[0013] where z1 is the first group of data elements in the data space Z, z2 is the second group of data elements in the data space Z, and z m is the m-th group of data elements in the data space Z, is the first group of labels of z1, is the second group of labels of z1, is the n-th group of labels of z1, is the first group of labels of z2, is the second group of labels of z2, is the n-th group of labels of z2, is the first group of labels of z m , is the second group of labels of z m , is the n-th group of labels of z m , and a range-based slicing method is proposed to slice the data elements with labels in the data space, which is:
[0014]
[0015] where low_1 is the lower limit of the first group of ranges, up_1 is the upper limit of the first group of ranges, low_2 is the lower limit of the second group of ranges, up_2 is the upper limit of the second group of ranges, low_q is the lower limit of the q-th group of ranges, up_q is the upper limit of the q-th group of ranges, is the data element existing under different range constraints The labels of the data elements belong to the ranges [low_1, up_1], [low_2, up_2], [low_q, up_q] respectively, and there are:
[0016]
[0017] Where, Θ is an empty set, there exists The case, because the data elements have multiple sets of labels, in the range-based sharding mode, the selection of the data elements is blurred, and more attention is paid to the attribute values of the labels, the mapping relationship between the data elements and the labels is one-to-many, and there are:
[0018]
[0019] Further, the multi-level caching step is specifically: setting multi-level caches at each node to accelerate data reading and writing operations, and proposing to use an Edge for Cloud model to cache data. On the basis of range-based sharding of data elements with labels in the data space, q groups of subspaces are formed, each group of spaces contains independent labels, and the labels are cached through distributed edge nodes, each label is regarded as a node, and the data space contains q groups of nodes for caching labels within the range. The data space contains 1 cloud computing power, which mainly provides computing power and storage support for the q edge nodes, and specifically includes the following:
[0020] Under the framework of the q groups of subspaces, there are labels: Where note (1) , note (2) ,…, note (q) are the labels of the first group, the second group, and the qth group of classifications respectively, the computing power set of the edge node is defined as computingpower=
[0021] {computingpower1, computingpower2, …, computingpower q}, where computingpower1, computingpower2, …, computingpower q are the computing power of the first edge node, the computing power of the second edge node, and the computing power of the qth edge node respectively, and the computing power of the cloud is defined as COMPUTINGPOWER CLOUD , and there are
[0022] Further, in order to enhance the competitiveness of data space, an overhead control mechanism is proposed to quantitatively control the cache of data space, and in the overhead control mechanism, it includes:
[0023] (1) Edge node overhead and cloud overhead
[0024] The data attribute of the edge node is defined as:
[0025] {Edge:computing power, storage space, offloading, cost edge ,cost cloud}
[0026] Wherein, Edge is the edge node, computing power is the computing power, storage space is the storage space, offloading is the computing offloading, the label will be calculated offloaded to the designated edge node for caching, cost edge is the edge node overhead, cost cloud is the cloud computing power overhead, the node overhead includes two cases: 1. When the node computing power or storage space supports the cache amount of the current moment of computing offloading, the node overhead belongs to the current edge node; 2. When the node computing power or storage space cannot meet the cache amount of the current moment of computing offloading, it needs to rent the idle edge node of the node computing power or storage space in the data space for caching;
[0027] (2) Service selection of edge node
[0028] There are three ways for service selection of edge node, if the computing power of the current edge node is sufficient but the storage space cannot provide service for the label, the label waits online to complete the service; if the computing power of the current edge node is insufficient but the storage space can provide service for the label, the cloud computing power can be involved in the cache service; if the computing power of the current edge node is insufficient and the storage space cannot provide service for the label, other edge nodes with computing power and storage space are called to cooperate in the cache task;
[0029] For label online waiting, there are:
[0030]
[0031] C total ={1:Load1,2:Load2,…|EMPTY_ROOM}
[0032] Wherein, ρ is the service intensity, μ is the service rate, λ is the offloading arrival rate, is the offloading of label note * , computing power * is the computing power for label note* Edge node of service, Bandwidyh is network communication bandwidth, T cache is cache time, Total execution time of tag online waiting, C total Total storage space of edge node, Load1 is the first unloading amount, Load2 is the second unloading amount, EMPTY_ROOM is the remaining storage space, if Then Enter the waiting service sequence of edge node, complete the cache service in order;
[0033] For collaborative cloud computing power participating in cache service, there are:
[0034]
[0035]
[0036] COST (2) =cost edge +cost cloud
[0037] Among them, Total execution time of collaborative cloud computing power participating in cache service, Storage space of cloud computing power, COST (2) Total cost of collaborative cloud computing power participating in cache service;
[0038] For other edge nodes with computing power and storage space to coordinate cache service, there are:
[0039]
[0040] COST (3) =cost edge +cost edge+
[0041] Among them, edge+ is the edge node participating in the coordination of cache service, computingpower + Computing power of edge+, EMPTYROOM + Remaining space of edge+, Total execution time of other edge nodes with computing power and storage space to coordinate cache service, cost edge+ Cost of renting edge+, COST (3) Total cost of other edge nodes with computing power and storage space to coordinate cache service;
[0042] (3) Parallel processing uses parallel processing technology, each node independently processes the data of the slice, uses multi-thread or multi-process technology to accelerate processing, uses message queue to realize the communication and data synchronization between nodes, processes different data segments at the same time, and shortens the processing time;
[0043] (4) Consistency check After the data processing is completed, the consistency of the data of each node is checked, the Raft protocol is used for consistency check between nodes, and the consistency of the data is ensured.
[0044] Further, the heterogeneous data mapping mechanism dynamically generates a heterogeneous data mapping rule according to the label of the data element, and has:
[0045] W=(z-b)note T (note note T ) -1
[0046]
[0047] Where W is a weight matrix, z is a data element matrix, b is a bias vector, note is a label vector, z i is the i-th group of data elements, note j is the label of z j , b j is the bias of z j , the mapping rule is continuously optimized by learning historical data using machine learning technology.
[0048] Further, the data fusion includes data extraction, data matching and data merging, the data extraction is to extract the data to be fused from each data source, the data matching is to match the extracted data, and the same or similar data items are identified, and the data merging is to merge the matched data items to generate a unified data view.
[0049] The beneficial effects of the application are:
[0050] 1. The application adopts a logarithmic ratio transformation method to preprocess the initial data elements, and the traditional ordinary logarithmic ratio transformation method usually takes the ratio of the elements in a composition vector to another selected reference element, and then takes the natural logarithm of the ratio. This transformation will cause the data to lose the original proportionality, and can handle the scale problem and proportion dependence commonly seen in proportional data. The logarithmic ratio transformation method provided by the application simultaneously performs logarithmic ratio transformation on all components, but the reference is not a single component, but the geometric mean of all components. The advantage is that it maintains the dimensionality of the data unchanged, and does not need to select a particular reference component, so it is more convenient and consistent when processing multivariate composition data.
[0051] 2. The application proposes a range-based data slice, and the data is characterized by a one-to-many mapping relationship, which means that in the data slice, the selection of data elements is not independent, and the selection of labels is unique. In the range-based slicing method, the selection of data elements is blurred, and the attribute value of the label is more valued, which enhances the effect of data classification, can quickly retrieve the required data information within the specified range, and can extract data elements.
[0052] 3. The application proposes to use the Edge for Cloud model to cache data. On the basis of the range-based slicing method for slicing the data elements with labels in the data space, q groups of subspaces are formed, each group of spaces contains independent labels, the labels are cached through distributed edge nodes, an overhead control mechanism is proposed to quantitatively control the caching of the data space, the distributed data is processed in the service selection link of the edge node, the data queuing, insufficient computing power, and insufficient storage space are considered, and a model is provided for online waiting of labels, collaborative cloud computing power to participate in caching services, and other edge nodes with computing power and storage space to coordinate caching services. Through the overhead as a constraint condition, the NP-hard problem is controlled within the range of feasible solution space. BRIEF DESCRIPTION OF DRAWINGS
[0053] The invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the invention. For ordinary skilled in the art, other drawings can be obtained without creative labor on the basis of the following drawings.
[0054] Figure 1 The figure is a schematic diagram of the structure of the application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by ordinary skilled in the art without creative labor belong to the scope of protection of the application.
[0056] Please refer to Figure 1 , the application will be further described with reference to the following examples.
[0057] See Figure 1 , the application aims to provide a multi-source heterogeneous data distributed fusion method, characterized in that it comprises the following steps: data preprocessing, data slicing, multi-level caching and distributed fusion.
[0058] Specifically, the data preprocessing defines a data space containing m groups of data elements, and simultaneously performs a logarithmic ratio transformation on all components of the m groups of data elements, but the reference is not a single component, but the geometric mean of all components, that is,
[0059]
[0060] where x i is the i-th initial data element, z i is the i-th data element.
[0061] Specifically, the data slicing step is specifically: defining the data space as Z, the data space includes m data, each data includes n groups of tags, if the tag does not contain any element, the data space is not default, recorded as null, then
[0062]
[0063] where z1 is the first group of data elements in the data space Z, z2 is the second group of data elements in the data space Z, z m is the mth group of data elements in the data space Z, is the first group of tags of z1, is the second group of tags of z1, is the n-th group of tags of z1, is the first group of tags of z2, is the second group of tags of z2, is the n-th group of tags of z2, is the first group of tags of z m , is the second group of tags of z m , is the n-th group of tags of z m .
[0064] Specifically, a range-based slicing method is proposed to slice the data elements with tags in the data space, that is,
[0065]
[0066] where low_1 is the lower limit of the first group of ranges, up_1 is the upper limit of the first group of ranges, low_2 is the lower limit of the second group of ranges, up_2 is the upper limit of the second group of ranges, low_q is the lower limit of the qth group of ranges, up_q is the upper limit of the qth group of ranges, is the data element existing under different range constraints The labels of the data elements belong to the ranges [low_1, up_1], [low_2, up_2], [low_q, up_q] respectively, and there are:
[0067]
[0068] There are cases, because the data elements have multiple sets of labels, and in the range-based sharding mode, the selection of the data elements is blurred, and more attention is paid to the attribute values of the labels, the mapping relationship between the data elements and the labels is one-to-many, and there are:
[0069]
[0070] Specifically, the multi-level caching step is specifically: setting multi-level caches at each node to accelerate data reading and writing operations, and proposing to cache data using an Edge for Cloud model, on the basis of range-based sharding of data elements with labels in the data space, forming q groups of subspaces, each group of spaces containing independent labels, caching labels through distributed edge nodes, regarding each label as a node, caching q groups of nodes in the data space as labels within the range, and including 1 cloud computing power in the data space, which mainly provides computing power and storage support for q edge nodes, and specifically includes the following:
[0071] Under the framework of the q groups of subspaces, there are labels: Where note (1) ,note (2) ,…,note (q) are the labels of the first group, the second group, and the qth group of classifications respectively, the computing power set of the edge nodes is defined as computingpower=
[0072] {computingpower1,computingpower2,…,computingpower q}, where computingpower1,computingpower2,…,computingpower q are the computing power of the first edge node, the computing power of the second edge node, and the computing power of the qth edge node respectively, and the computing power of the cloud is defined as COMPUTINGPOWER CLOUD , and there are
[0073] Specifically, in order to enhance the competitiveness of data space, an overhead control mechanism is proposed to quantitatively control the cache of data space. In the overhead control mechanism, it includes:
[0074] (1) Edge node overhead and cloud overhead
[0075] The data attribute of the edge node is defined as:
[0076] {Edge:computing power, storage space, offloading, cost edge , cost cloud}
[0077] Wherein, Edge is the edge node, computing power is the computing power, storage space is the storage space, offloading is the computing offloading, the label will be calculated offloaded to the designated edge node for caching, cost edge is the edge node overhead, cost cloud is the cloud computing power overhead;
[0078] Specifically, the node overhead includes two cases: 1. When the node computing power or storage space supports the cache amount of the current moment of computing offloading, the node overhead belongs to the current edge node; 2. When the node computing power or storage space cannot meet the cache amount of the current moment of computing offloading, it needs to rent the idle edge node of the node computing power or storage space in the data space for caching;
[0079] (2) Service selection of edge node
[0080] There are three ways for service selection of edge node. If the computing power of the current edge node is sufficient but the storage space cannot provide service for the label, the label waits online to complete the service; if the computing power of the current edge node is insufficient but the storage space can provide service for the label, the cloud computing power can be involved in the cache service; if the computing power of the current edge node is insufficient and the storage space cannot provide service for the label, other edge nodes with computing power and storage space are called to cooperate in the cache task;
[0081] Specifically, for the label online waiting, there are:
[0082]
[0083] C total ={1:Load1,2:Load2,…|EMPTY_ROOM}
[0084] Wherein, ρ is the service intensity, μ is the service rate, λ is the offloading arrival rate, is the label note *computingpower * is the label ntoe * is the edge node serving the label, Bandwidth is the network communication bandwidth, T cache is the cache time, is the total execution time of the label online waiting, C total is the total storage space of the edge node, Load1 is the first unloading amount, Lodad2 is the second unloading amount, EMPTY_ROOM is the remaining storage space, if then enter the waiting service sequence of the edge node, and complete the cache service in order;
[0085] Specifically, for the cooperative cloud computing power participating in the cache service, there are:
[0086]
[0087] COST (2) = cost edge + cost cloud
[0088] wherein, is the total execution time of the cooperative cloud computing power participating in the cache service, is the storage space of the cloud computing power, COST (2) is the total cost of the cooperative cloud computing power participating in the cache service;
[0089] Specifically, for other edge nodes with computing power and storage space to coordinate cache service, there are:
[0090]
[0091] COST (3) = cost edge + cost edge+
[0092] wherein, edge+ is the edge node participating in the coordinated cache service, computingpower + is the computing power of edge+, EMPTYROOM + is the remaining space of edge+, is the total execution time of other edge nodes with computing power and storage space to coordinate cache service, cost edge+ is the cost of renting edge+, COST (3) is the total cost of other edge nodes with computing power and storage space to coordinate cache service;
[0093] (3) Parallel processing uses parallel processing technology, each node independently processes the data of the slice, uses multi-thread or multi-process technology to accelerate processing, uses message queue to realize the communication and data synchronization between nodes, processes different data segments at the same time, and shortens the processing time;
[0094] (4) Consistency check After the data processing is completed, the consistency of the data of each node is checked, the Raft protocol is used for consistency check between nodes, and the consistency of the data is ensured.
[0095] Further, the heterogeneous data mapping mechanism dynamically generates a heterogeneous data mapping rule according to the label of the data element, and has:
[0096] W=(z-b)note T (note note T ) -1
[0097]
[0098] Where W is a weight matrix, z is a data element matrix, b is a bias vector, note is a label vector, z i is the i-th group of data elements, note j is the label of z j , b j is the bias of z j , the mapping rule is continuously optimized by using machine learning technology to learn historical data.
[0099] Further, the data fusion includes data extraction, data matching and data merging, the data extraction is to extract the data to be fused from each data source, the data matching is to match the extracted data, and the same or similar data items are identified, and the data merging is to merge the matched data items to generate a unified data view.
[0100] The beneficial effects of the application are:
[0101] 1. The application adopts a logarithmic ratio transformation mode to pretreat the initial data elements, and the traditional ordinary logarithmic ratio transformation mode usually takes the ratio of the elements in a composition vector to another selected reference element, and then takes the natural logarithm of the ratio. This transformation will cause the data to lose the original proportional properties, and can handle the scale problem and proportional dependence commonly seen in proportional data. The logarithmic ratio transformation mode provided by the application simultaneously performs logarithmic ratio transformation on all components, but the reference is not a single component, but the geometric mean of all components. The advantage is that it maintains the dimensionality of the data unchanged, and does not need to select a particular reference component, so it is more convenient and consistent when processing multivariate composition data.
[0102] 2. The application proposes a range-based data slice, and the data is characterized by a one-to-many mapping relationship, which means that in the data slice, the selection of data elements is not independent, and the selection of labels is unique. In the range-based slicing method, the selection of data elements is blurred, and more attention is paid to the attribute value of the label, which enhances the effect of data classification, can quickly retrieve the required data information in the specified range, and can extract data elements.
[0103] 3. The application proposes to use the Edge for Cloud model to cache data. On the basis of the range-based slicing method for slicing the data elements with labels in the data space, q groups of subspaces are formed, each group of spaces contains independent labels, the labels are cached through distributed edge nodes, an overhead control mechanism is proposed to quantitatively control the caching of the data space, the distributed data is processed in the service selection link of the edge node, the data queuing, insufficient computing power, and insufficient storage space are considered, and a model is provided for online waiting of labels, cooperative cloud computing power to participate in caching services, and other edge nodes with computing power and storage space to coordinate caching services. Through the overhead as a constraint condition, the NP-hard problem is controlled within the range of feasible solution space.
[0104] Although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features, any modification, equivalent replacement, improvement, etc. within the spirit and principles of the application shall be included in the protection scope of the application.
Claims
1. Distributed fusion methods for multi-source heterogeneous data, including: Data preprocessing, data sharding, multi-level caching, and distributed fusion; Define the data space to include Each data element will Each data element undergoes a logarithmic ratio transformation on all components simultaneously, but the benchmark is not a single component, but rather the geometric mean of all components, as follows: in, For the first One initial data element, For the first One data element; The data sharding step specifically involves: defining the data space as... The data space includes Each data point includes [number] data points. If a tag contains no element and the data space is not defaulted, it is denoted as null. Then there are... ,in For data space The first data element in This is the second data element in data space Z. For data space The first in One data element, for The first tag, for The second tag, for The One tag, for The first tag, for The second tag, for The One tag, for The first tag, for The second tag, for The A range-based partitioning method is proposed to partition labeled data elements in the data space using a single label. in, This is the lower limit of the first range. This is the upper limit of the first range. This is the lower limit of the second range. This is the upper limit of the second range. For the first A lower limit of the range, The upper limit of the q-th range, , , To ensure that data elements exist under different range constraints , , The tags belong to the ranges respectively. , , When data element tags Scope At time, data elements Put in A set, when data elements tags Scope At time, data elements Put in A set, when data elements tags Scope At time, data elements Put in The set of , and we have: in, It is an empty set, and it exists. This situation arises because data elements have multiple tags. In range-based sharding, the selection of data elements becomes blurred, with greater emphasis placed on tag attribute values. The mapping relationship between data elements and tags is one-to-many, and simultaneously: Multi-level caching is implemented at each node to accelerate data read and write operations. The Edge for Cloud model is proposed for data caching. Based on range-based sharding to partition labeled data elements in the data space, a system is formed... Each subspace contains independent labels. These labels are cached via distributed edge nodes, with each label considered a node. The data space contains... Each node caches tags within its range, while the data space contains one cloud computing power unit, whose main function is to... Each edge node provides computing power and storage support, specifically including the following: exist Within the framework of each subspace, there are tags: ,in Let the labels in the first, second, and qth categories be the basis for defining the computing power set of the edge nodes. ,in, Let be the computing power of the 1st edge node, the 2nd edge node, and the qth edge node, respectively. Define the cloud computing power as... ,have A heterogeneous data mapping mechanism is proposed, which dynamically generates heterogeneous data mapping rules based on the labels of data elements, including: Where W is the weight matrix, z is the data element matrix, and b is the bias vector. For the label vector, For the first One data element, for The tag, for The bias is determined by using machine learning techniques to learn from historical data and continuously optimize the mapping rules.
2. The distributed fusion method for multi-source heterogeneous data according to claim 1, in order to enhance the competitiveness of the data space, proposes an overhead control mechanism to quantitatively control the caching of the data space. The overhead control mechanism includes: (1) Edge node overhead and cloud overhead Define the data attributes of the edge nodes as: { } in, For edge nodes, For computing power, For storage space, To facilitate unloading, tags are computed and unloaded to designated edge nodes for caching. For edge node overhead, For cloud computing power overhead, node overhead includes two cases:
1. When the node's computing power or storage space supports the cache amount of the current computation unloading, then the node overhead belongs to the current edge node; 2. When the node's computing power or storage space cannot meet the cache amount of the current computation unloading, it is necessary to lease edge nodes with free computing power or storage space in the data space for caching. (2) Service selection for edge nodes There are three ways to select services for edge nodes: if the current edge node has sufficient computing power but insufficient storage space to provide services for the tag, the tag waits online to complete the service; if the current edge node has insufficient computing power but sufficient storage space to provide services for the tag, it can collaborate with cloud computing power to participate in caching services; if the current edge node has insufficient computing power and insufficient storage space to provide services for the tag, it can call other edge nodes with computing power and storage space to collaborate on caching tasks. Regarding the online waiting period for tags, there are: in, For service intensity, For service rate, For unloading arrival rate, For tags Uninstallation, This is a label service edge nodes, It refers to network communication bandwidth. It's the cache time. The total execution time while the label is online. The total storage space for edge nodes. This is the first uninstallation. This is the second unload volume. For the remaining storage space, if > ,but Enter the waiting service sequence of the edge node and complete the cache service in order; Regarding the use of collaborative cloud computing power in caching services, the following applies: in, To coordinate the total execution time of cloud computing power participating in caching services, Storage space for cloud computing power The total overhead for coordinating cloud computing power to participate in caching services; For coordinating caching services with other edge nodes that have computing power and storage space, the following are options: in, Edge nodes that participate in coordinating caching services. for computing power for The remaining space, Total execution time when coordinating caching services for other edge nodes with computing power and storage space. For leasing Expenses, The total overhead of coordinating caching services for other edge nodes with computing power and storage space; (3) Parallel processing utilizes parallel processing technology, each node independently processes the fragmented data in parallel, uses multi-threading or multi-process technology to accelerate processing, uses message queues to realize communication and data synchronization between nodes, and processes different data fragments at the same time to shorten processing time. (4) Consistency check After the data processing is completed, the data of each node is checked for consistency. The Raft protocol is used to check the consistency between nodes to ensure data consistency.
3. The distributed fusion method for multi-source heterogeneous data according to claim 1, comprising the following through data fusion technology: Data extraction, data matching, and data merging; data extraction involves extracting the data that needs to be merged from various data sources. Data matching involves matching extracted data to identify identical or similar data items; Data merging combines matching data items to generate a unified data view.
Citation Information
Patent Citations
Constructive geochemical collaborative simulation hidden mine exploration risk evaluation method
CN116384747A
Drug information storage method based on distributed edge calculation and multi-modal data
CN118152481A