Multi-source heterogeneous data fusion method, device and equipment and computer readable medium
By preprocessing, quality inspection, and heterogeneous conflict resolution of multi-source heterogeneous data, and combining resource scheduling and multimodal data fusion models, the problems of low efficiency, low quality, and low security in multi-source heterogeneous data fusion are solved, achieving efficient and secure data fusion and storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies fail to effectively consider the correlation and dependence between different modalities when fusing multi-source heterogeneous data, resulting in low data fusion efficiency, low quality, low security, and serious waste of resources.
By preprocessing, quality inspection, and heterogeneous conflict resolution of enterprise heterogeneous datasets from different sources, and by using resource scheduling and multimodal data fusion models for federated learning, a globally fused dataset is generated and then encrypted and transmitted to the data warehouse.
It improves the efficiency and accuracy of data fusion, reduces waste of transmission resources, and enhances data security and storage resource utilization efficiency.
Smart Images

Figure CN121834646A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to methods, apparatus, devices, and computer-readable media for multi-source heterogeneous data fusion. Background Technology
[0002] With the rapid development of computer technology, the volume of multi-source heterogeneous data is increasing, leading to the data silo problem. Data fusion technology, by comprehensively processing data at multiple levels, including raw data and data features, can reduce redundancy and storage resource waste, thus solving the data silo problem. The typical approach to fusion of multi-source heterogeneous data is to use a backpropagation (BP) neural network model to perform simple single-modal concatenation and single-modal decision-making on heterogeneous datasets from multiple data sources on the same server, resulting in a fused heterogeneous dataset, and then store this fused dataset.
[0003] However, in practice, it has been found that when using the above methods to fuse multi-source heterogeneous data, the following technical problems often arise: Since only single-modal splicing and combination of enterprise heterogeneous data is performed, the correlation and dependency between different modalities are not considered. Furthermore, data conflicts and redundancy exist in enterprise heterogeneous datasets from multiple sources, making it impossible to accurately remove redundant and conflicting data. Additionally, fusing multi-source heterogeneous data on the same server requires the transmission of a large amount of redundant and erroneous data, resulting in low data transmission security. Finally, when fusing data from multiple terminal devices, insufficient device resources lead to long data fusion times, resulting in low data fusion efficiency, low data fusion quality, low data security, and a waste of storage and transmission resources.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure provide methods, apparatuses, devices, and computer-readable media for fusing multi-source heterogeneous data to address one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure provide a method for multi-source heterogeneous data fusion, comprising: controlling a set of terminal devices; preprocessing enterprise heterogeneous datasets acquired from different sources to obtain preprocessed enterprise heterogeneous datasets as local enterprise heterogeneous datasets; performing data quality detection on each local enterprise heterogeneous data in the aforementioned local enterprise heterogeneous dataset to generate local data quality values, thereby obtaining a local data quality value set; performing heterogeneous conflict resolution processing on the aforementioned local enterprise heterogeneous dataset based on the aforementioned local data quality value set, thereby obtaining a local heterogeneous conflict resolution dataset; and responding to the detection that there are resource-constrained terminal devices in the aforementioned terminal device set, adjusting the settings of the terminal devices... The system performs resource scheduling on the backup set to obtain a terminal resource scheduling device set; controls the terminal resource scheduling device set to train the corresponding initial heterogeneous multimodal data fusion model set based on the local heterogeneous conflict resolution dataset to obtain a heterogeneous multimodal data fusion model set; performs multimodal federated fusion processing on the model parameter set corresponding to the heterogeneous multimodal data fusion model set to obtain a federated fusion model parameter set; performs global multi-level fusion on the local heterogeneous conflict resolution dataset based on the federated fusion model parameter set and the heterogeneous multimodal data fusion model set to obtain an enterprise global fusion dataset; and encrypts and transmits the enterprise global fusion dataset to the data warehouse for data storage.
[0008] Secondly, some embodiments of this disclosure provide a multi-source heterogeneous data fusion apparatus, comprising: a control unit configured to control a set of terminal devices to preprocess data from different sources of enterprise heterogeneous datasets to obtain preprocessed enterprise heterogeneous datasets as local enterprise heterogeneous datasets; a data quality detection unit configured to perform data quality detection on each local enterprise heterogeneous data in the aforementioned local enterprise heterogeneous dataset to generate local data quality values to obtain a local data quality value set; a heterogeneous conflict resolution unit configured to perform heterogeneous conflict resolution processing on the aforementioned local enterprise heterogeneous dataset based on the aforementioned local data quality value set to obtain a local heterogeneous conflict resolution dataset; and a resource scheduling unit configured to, in response to detecting that there are resource-constrained terminal devices in the aforementioned terminal device set, schedule the aforementioned terminal device set... The system performs resource scheduling to obtain a set of terminal resource scheduling devices; a model training unit is configured to control the terminal resource scheduling device set to train the corresponding initial heterogeneous multimodal data fusion model set based on the local heterogeneous conflict resolution dataset to obtain a heterogeneous multimodal data fusion model set; a multimodal federated fusion unit is configured to perform multimodal federated fusion processing on the model parameter set corresponding to the heterogeneous multimodal data fusion model set to obtain a federated fusion model parameter set; a global multi-level fusion unit is configured to perform global multi-level fusion on the local heterogeneous conflict resolution dataset based on the federated fusion model parameter set and the heterogeneous multimodal data fusion model set to obtain an enterprise global fusion dataset; and an encrypted transmission unit is configured to encrypt and transmit the enterprise global fusion dataset to a data warehouse for data storage.
[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.
[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any of the implementations of the first aspect.
[0011] The above embodiments of this disclosure have the following beneficial effects: The multi-source heterogeneous data fusion method of some embodiments of this disclosure can improve the efficiency of data fusion, shorten the data fusion time, improve data security, and reduce the waste of transmission resources by fusing data located on different terminal devices. Specifically, the reasons for the low data fusion efficiency, low data fusion quality, low data security, and waste of storage and transmission resources are as follows: because only single-modal splicing and combination of enterprise heterogeneous data is performed, the correlation and dependency between different modal data are not considered, and the enterprise heterogeneous datasets from multiple sources have data conflicts and redundancy, making it impossible to accurately remove redundant and conflicting data. Furthermore, multi-source heterogeneous data fusion on the same server requires the transmission of a large amount of redundant and erroneous data, resulting in low data transmission security. Additionally, when fusing data from multiple terminal devices, the data fusion time is long due to insufficient device resources to support data processing, leading to low data fusion efficiency, low data fusion quality, low data security, and waste of storage and transmission resources. Based on this, the multi-source heterogeneous data fusion method of some embodiments of this disclosure can first control the terminal device set to perform data preprocessing on the acquired heterogeneous enterprise datasets from different sources, obtaining preprocessed heterogeneous enterprise datasets as local heterogeneous enterprise datasets. Here, data preprocessing can perform initial processing on the heterogeneous enterprise datasets, removing some erroneous and redundant data, improving data quality and reducing data volume. Second, data quality detection is performed on each local heterogeneous data in the aforementioned local enterprise heterogeneous dataset to generate local data quality values, obtaining a local data quality value set. Here, data quality detection can provide a better understanding of the data from different data sources, ensuring data quality and the quality of subsequent data fusion. Third, based on the aforementioned local data quality value set, heterogeneous conflict resolution processing is performed on the aforementioned local enterprise heterogeneous datasets, obtaining a local heterogeneous conflict resolution dataset. Here, removing inconsistent data can ensure data quality and data integrity. Next, in response to the detection of resource-constrained terminal devices in the aforementioned terminal device set, resource scheduling is performed on the aforementioned terminal device set, obtaining a terminal resource scheduling device set. Here, resource scheduling can improve the efficiency of terminal devices in processing local heterogeneous datasets, ensuring that terminal devices have sufficient resources for data processing, shortening processing time, reducing terminal device failure rates, and improving terminal device security and stability. Subsequently, the aforementioned terminal resource scheduling device set is controlled to train the corresponding initial heterogeneous multimodal data fusion model set based on the aforementioned local heterogeneous conflict resolution dataset, resulting in a heterogeneous multimodal data fusion model set. Here, model training can improve the accuracy of subsequent model-based multi-source heterogeneous data fusion.Next, the model parameter sets corresponding to the aforementioned heterogeneous multimodal data fusion model set are subjected to multimodal federated fusion processing to obtain the federated fusion model parameter set. Here, performing global-local training fusion based on federated learning on the heterogeneous multimodal data fusion model can improve the accuracy of the model output. Then, based on the aforementioned federated fusion model parameter set and the heterogeneous multimodal data fusion model set, the aforementioned local heterogeneous conflict resolution dataset is subjected to global multi-level fusion to obtain the enterprise global fusion dataset. Here, only the local model parameters used for data fusion can be transmitted without transmitting the data itself, which can reduce the amount of data transmitted and improve transmission efficiency, thereby improving the accuracy and efficiency of data fusion and enhancing the security of data transmission. Finally, the aforementioned enterprise global fusion dataset is encrypted and transmitted to the data warehouse for data storage. This improves the data quality and security of heterogeneous enterprise datasets from different sources and reduces the waste of storage resources. Therefore, this multi-source heterogeneous data fusion method can reduce the waste of transmission resources and shorten the transmission time by performing federated learning-based model fusion on data located on different terminal devices. It can reduce the waste of transmission resources and shorten the transmission time by transmitting only the model parameters without transmitting the data itself. Based on model fusion, it can improve the efficiency and accuracy of data fusion, shorten the data fusion time, improve the security and quality of data, and reduce the waste of storage resources. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0013] Figure 1 This is a flowchart of some embodiments of the multi-source heterogeneous data fusion method according to the present disclosure; Figure 2 These are schematic diagrams of the structure of some embodiments of the multi-source heterogeneous data fusion apparatus according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] Figure 1 A flow 100 of some embodiments of a multi-source heterogeneous data fusion method according to the present disclosure is shown. This multi-source heterogeneous data fusion method includes the following steps: Step 101: Control the terminal device set to perform data preprocessing on the acquired heterogeneous enterprise datasets from different sources to obtain preprocessed heterogeneous enterprise datasets, which will serve as local heterogeneous enterprise datasets.
[0021] In some embodiments, the executing entity (e.g., an electronic device) of the above-described multi-source heterogeneous data fusion method can control a set of terminal devices via wired or wireless connections to preprocess the acquired heterogeneous enterprise datasets from different sources, obtaining preprocessed heterogeneous enterprise datasets as local heterogeneous enterprise datasets. The terminal devices in the aforementioned set of terminal devices can be edge servers used to store and process heterogeneous enterprise data from different aspects. The aforementioned data preprocessing may include, but is not limited to, at least one of the following: standardizing data formats, data cleaning, and mean-based data imputation. The heterogeneous enterprise data in the aforementioned heterogeneous enterprise datasets from different sources can be data related to the enterprise's status. For example, the aforementioned heterogeneous enterprise data from different sources may include, but is not limited to, at least one of the following: enterprise financial data, enterprise operating status data, enterprise bond data, data from similar enterprises in the same industry, enterprise business data, enterprise equity structure data, and enterprise financial data.
[0022] Step 102: Perform data quality checks on each piece of heterogeneous data in the local enterprise heterogeneous dataset to generate local data quality values and obtain a set of local data quality values.
[0023] In some embodiments, the aforementioned executing entity may perform data quality checks on each piece of local enterprise heterogeneous data in the aforementioned local enterprise heterogeneous dataset to generate local data quality values, thereby obtaining a set of local data quality values. These local data quality values can be numerical values that evaluate the data quality of heterogeneous enterprise data from different sources, i.e., local enterprise heterogeneous data, in terms of data integrity, accuracy, and consistency.
[0024] In some optional implementations of certain embodiments, the above-mentioned data quality detection of each local enterprise heterogeneous data in the aforementioned local enterprise heterogeneous dataset to generate a local data quality value and obtain a local data quality value set may include the following steps: The first step is to perform the following data quality assessment steps for each piece of heterogeneous local enterprise data in the above heterogeneous local enterprise dataset: Sub-step 1: Determine the data quality assessment dimension set for the aforementioned heterogeneous data from local enterprises. The data quality assessment dimensions in this set can be dimensional data used to evaluate the data quality of heterogeneous data from different aspects. This data quality assessment dimension set may include: completeness, standardization, consistency, timeliness, uniqueness, and accuracy. Completeness characterizes the degree of data completeness of heterogeneous data from local enterprises compared to real data. Standardization characterizes the degree to which heterogeneous data from local enterprises conforms to relevant data specifications and standards. Consistency characterizes the degree to which heterogeneous data from local enterprises is consistent with other heterogeneous data from local enterprises within a context and the degree to which related data conforms to logical rules. Timeliness characterizes whether the collection of heterogeneous data from local enterprises is timely and meets collection time requirements, i.e., the freshness of the recorded data. Uniqueness characterizes the degree of non-duplication in heterogeneous data from local enterprises. Accuracy characterizes the degree to which heterogeneous data from local enterprises accurately represents real-world values.
[0025] Sub-step 2 involves dividing each data quality assessment dimension in the aforementioned data quality assessment dimension set into sub-dimensions to generate a data quality assessment indicator set. This data quality assessment indicator set can be a group of dimension indicators for the same data quality assessment dimension, derived from different sub-dimensions to further evaluate the quality of heterogeneous data from local enterprises. The data quality assessment indicator set may include: attribute completeness, record completeness, numerical completeness, data type standardization, data precision standardization, data order consistency, data logical consistency, timeliness, attribute uniqueness, record uniqueness, range accuracy, and numerical accuracy. Attribute completeness characterizes the degree of attribute loss in heterogeneous data from local enterprises, measured by the ratio of the set of non-missing and non-repeating attributes in the heterogeneous data to the total set of all attributes in the heterogeneous data. Record completeness characterizes the degree of record loss (i.e., data collected by terminal devices at the same point in time constitutes a record) in heterogeneous data from local enterprises, measured by the ratio of the total number of records included in the heterogeneous data to the expected total number of records to be collected. The aforementioned numerical integrity characterizes the degree of missing elements in local enterprise heterogeneous data, measured by the ratio of the number of missing values in the local enterprise heterogeneous data to the number of values in the expected total number of records. The aforementioned data type normalization characterizes the correctness of the data types in local enterprise heterogeneous data, measured by the ratio of the number of attributes (excluding time attributes) that cannot be converted to numeric values to the total number of all attributes. The aforementioned data precision normalization characterizes the degree to which data values in local enterprise heterogeneous data meet the requirement of retaining decimal places. The aforementioned data order consistency characterizes the degree to which local enterprise heterogeneous data conforms to temporal order, measured by the ratio of the number of data points in the subset of local enterprise heterogeneous data that determines the largest monotonically non-decreasing subsequence based on the time attribute to the total number of data points. The aforementioned data logical consistency characterizes the degree to which the logical relationships between different attributes are consistent and non-conflicting, measured by the ratio of data points in local enterprise heterogeneous data that violate attribute value logic to the total number of data points.
[0026] In practice, the implementing entity can first divide each data quality assessment indicator in the aforementioned data quality assessment indicator set into sub-dimensions to obtain an initial sub-dimension indicator set. Then, it determines the relevance set of the sub-dimension indicators within the initial sub-dimension indicator set. Finally, it selects at least one initial sub-dimension indicator from the initial sub-dimension indicator set whose corresponding sub-dimension indicator relevance meets a preset relevance threshold, and uses this as the data quality assessment dimension indicator set. The preset relevance threshold can be a pre-defined critical value for whether or not to remove initial sub-dimension indicators. For example, the preset relevance threshold could be 0.7.
[0027] Sub-step 3 involves determining the evaluation indicator weight values for each data quality evaluation indicator within each data quality evaluation indicator group in the aforementioned data quality evaluation indicator set, thus obtaining a set of evaluation indicator weight values. These evaluation indicator weight values characterize the importance of each data quality evaluation indicator within the data quality evaluation dimension. This determination can be performed using the entropy weight method.
[0028] Sub-step 4 involves determining the weight value of each data quality assessment dimension in the aforementioned data quality assessment dimension set, thus obtaining a set of assessment dimension weight values. The assessment dimension weight values in this set characterize the importance of each data quality assessment dimension. This determination can be achieved through a combination of the expert Delphi method and the analytic hierarchy process.
[0029] Sub-step 5: Based on the aforementioned set of evaluation dimension weight values, the aforementioned set of evaluation indicator weight values, the aforementioned set of data quality evaluation indicators, and the aforementioned set of data quality evaluation dimensions, determine the local data quality value set for the aforementioned heterogeneous data of the local enterprise. In practice, the executing entity can first perform a weighted summation of the aforementioned set of evaluation indicator weight values and the aforementioned set of data quality evaluation indicators to obtain the data dimension value set of the aforementioned data quality evaluation dimension set. Then, it can perform a weighted summation of the aforementioned data dimension value set and the aforementioned set of evaluation dimension weight values to obtain the local data quality value set for the aforementioned heterogeneous data of the local enterprise.
[0030] Step 103: Based on the local data quality numerical set, perform heterogeneous conflict resolution processing on the local enterprise heterogeneous dataset to obtain the local heterogeneous conflict resolution dataset.
[0031] In some embodiments, the aforementioned executing entity can perform heterogeneous conflict resolution processing on the aforementioned local enterprise heterogeneous dataset based on the aforementioned local data quality value set, to obtain a local heterogeneous conflict resolution dataset. This local heterogeneous conflict resolution dataset can be a dataset composed of local enterprise heterogeneous datasets with conflict-free data (data with fusion conflicts resolved) and local enterprise heterogeneous datasets without fusion conflicts. The aforementioned heterogeneous conflict resolution processing can be conflict detection and conflict resolution processing using knowledge graphs.
[0032] In addressing the aforementioned technical challenges by employing technical solutions, the application scenario of cross-platform, multi-source sensitive data fusion often presents the following technical problems: due to variations in data quality and the presence of sensitive data from different local enterprise heterogeneous data sources, as well as data conflicts and missing data, the fused data suffers from low quality, prolonged fusion time, and significant waste of storage resources. Based on the characteristics of this application scenario—multi-source heterogeneous data, data conflicts, and missing data—we have decided to adopt the following solution: In some optional implementations of certain embodiments, the process of performing heterogeneous conflict resolution processing on the local enterprise heterogeneous dataset based on the local data quality value set to obtain a local heterogeneous conflict resolution dataset may include the following steps: The first step is to perform semantic matching processing on the aforementioned heterogeneous datasets of local enterprises to obtain a heterogeneous indicator matching result set. The heterogeneous indicator matching results in this set can be the results of matching local enterprise heterogeneous data within the aforementioned datasets that have the same semantics but different names. For example, the heterogeneous indicator matching results could be the successful matching of different indicators representing the same semantics as tax and financial statement tax payment. In practice, the implementing entity can determine the Pearson correlation coefficient and Spearman correlation coefficient of each local enterprise heterogeneous data in the aforementioned datasets with other local enterprise heterogeneous datasets, obtaining the Pearson correlation coefficient set and the Spearman correlation coefficient set, which serve as the heterogeneous indicator matching result set. The aforementioned other local enterprise heterogeneous datasets can be the set obtained by removing local enterprise heterogeneous data from the aforementioned local enterprise heterogeneous datasets.
[0033] The second step involves performing indicator relationship fitting on the local enterprise heterogeneous data corresponding to each heterogeneous indicator matching result in the aforementioned heterogeneous indicator matching result set, thereby obtaining an indicator relationship fitting function. This indicator relationship fitting function characterizes the degree of mapping correlation between the local enterprise heterogeneous data corresponding to the aforementioned heterogeneous indicator matching results. In practice, the executing entity can first filter out at least one heterogeneous indicator matching result from the aforementioned heterogeneous indicator matching result set that meets preset indicator matching conditions. These preset indicator matching conditions can be conditions where both the Pearson correlation coefficient and the Spearman correlation coefficient are greater than or equal to a preset threshold. The preset threshold can be a pre-set value. For example, the preset threshold could be 0.75. Then, in response to determining that the local enterprise heterogeneous dataset corresponding to at least one heterogeneous indicator matching result is local enterprise heterogeneous data with a linear relationship, the least squares method is used to perform linear relationship fitting on the local enterprise heterogeneous dataset corresponding to at least one heterogeneous indicator matching result, thereby obtaining the indicator relationship fitting function. Finally, in response to determining that the local enterprise heterogeneous dataset corresponding to at least one heterogeneous indicator matching result is local enterprise heterogeneous data with non-linear relationships, an indicator relationship mapping function is used to fit the indicator relationships of the local enterprise heterogeneous dataset corresponding to at least one heterogeneous indicator matching result, resulting in an indicator relationship fitting function. The indicator relationship mapping function can be a deep neural network model that maps indicator relationships to the input local enterprise heterogeneous dataset with non-linear relationships. For example, the indicator relationship mapping function can be an XGBoost (eXtreme Gradient Boosting) model.
[0034] The third step involves performing semantic indicator conflict detection on the local enterprise heterogeneous dataset corresponding to the heterogeneous indicator matching result set, based on the aforementioned indicator relationship fitting function. This yields a semantic conflict detection result set. The semantic conflict detection results in this set characterize the inconsistencies and contradictions existing between the local enterprise heterogeneous datasets.
[0035] As an example, the aforementioned execution entity can perform the following conflict detection steps for each heterogeneous indicator matching result set in the aforementioned heterogeneous indicator matching result set: First, input the first local enterprise heterogeneous data corresponding to the aforementioned heterogeneous indicator matching result set into the aforementioned indicator relationship fitting function to obtain fitted heterogeneous data. Here, the aforementioned first local enterprise heterogeneous data can be one local enterprise heterogeneous data point in the local enterprise heterogeneous dataset corresponding to the heterogeneous indicator matching result. Then, determine the mean squared error and the squared error value of the aforementioned fitted heterogeneous data and the second local enterprise heterogeneous data after multiple iterations. Here, the aforementioned second local enterprise heterogeneous data can be another local enterprise heterogeneous data point in the local enterprise heterogeneous dataset corresponding to the heterogeneous indicator matching result. Afterwards, in response to determining that the aforementioned squared error value is greater than or equal to the product of the mean squared error value and the detection sensitivity coefficient, the local enterprise heterogeneous dataset corresponding to the heterogeneous indicator matching result set is considered to have a semantic conflict, and this is determined as a semantic conflict detection result. Finally, in response to determining that the aforementioned squared error value is less than the product of the mean squared error value and the detection sensitivity coefficient, the local enterprise heterogeneous dataset corresponding to the heterogeneous indicator matching result set does not have a semantic conflict, and this is determined as a semantic conflict detection result.
[0036] Fourth, in response to the determination that a single conflict exists in the semantic conflict detection result set, based on the aforementioned indicator relationship fitting function and the aforementioned heterogeneous indicator matching result set, conflict splitting processing is performed on the local enterprise heterogeneous dataset corresponding to the single conflict conflict detection result set to obtain the first conflict-split heterogeneous dataset. The single conflict conflict detection result can be indicator data in the local enterprise heterogeneous data corresponding to the semantic conflict detection result where only one semantic indicator conflict exists. The first conflict-split heterogeneous data in the first conflict-split heterogeneous dataset can be local enterprise heterogeneous data corresponding to the single conflict conflict detection result split into local enterprise heterogeneous data that does not contain data conflicts, based on the data source.
[0037] As an example, the aforementioned execution entity can perform the following conflict splitting step for each heterogeneous indicator matching result in the aforementioned heterogeneous indicator matching result set: Based on the aforementioned indicator relationship fitting function, perform conflict splitting processing on the local enterprise heterogeneous dataset corresponding to the conflict detection result set with a single conflict, to obtain a first conflict-split heterogeneous data group. The aforementioned first conflict-split heterogeneous data group can be local enterprise heterogeneous data from one source that includes local enterprise heterogeneous data without data conflicts and conflicting data obtained from the indicator association fitting function, and local enterprise heterogeneous data from another source.
[0038] Fifth, in response to the determination that multiple conflicts exist in the semantic conflict detection result set, based on the aforementioned indicator relationship fitting function and the aforementioned heterogeneous indicator matching result set, the local enterprise heterogeneous dataset corresponding to the multiple conflict conflict detection result set is subjected to group conflict splitting processing to obtain a second conflict split heterogeneous dataset. Here, the aforementioned multiple conflict conflict detection results can be indicator data with data conflicts, and the local enterprise heterogeneous data corresponding to the semantic conflict detection results contains only at least two semantic indicator conflicts.
[0039] As an example, the aforementioned execution entity can first extract local enterprise heterogeneous data with single conflicts from the local enterprise heterogeneous dataset corresponding to the conflict detection result set with multiple conflicts, thus obtaining a local enterprise heterogeneous data group. Then, based on the aforementioned indicator relationship fitting function, the aforementioned local data quality numerical set, and the aforementioned heterogeneous indicator matching result set, conflict splitting processing is performed on each local enterprise heterogeneous data in the local enterprise heterogeneous data group to obtain a second conflict-split heterogeneous dataset. The specific implementation method of this step can be referred to the specific implementation method of step four, and will not be repeated here.
[0040] Step 6: Input the aforementioned first conflict-split heterogeneous dataset and the aforementioned second conflict-split heterogeneous dataset into the conflict resolution anomaly detection model to obtain a set of conflict-split anomaly values. The conflict-split anomaly values in this set characterize the degree to which the first or second conflict-split heterogeneous data deviates from normal data. The aforementioned conflict resolution anomaly detection model can be a machine learning model that performs anomaly detection on the input first and second conflict-split heterogeneous datasets. Alternatively, it can be a model that integrates an isolated forest equalization sampling anomaly model, a priori anomaly detection model, and a clustering anomaly detection model using a soft-voting ensemble method. Finally, it can be a model that uses the EasyEnsemble undersampling algorithm to equalize the data before it is input into the isolated forest algorithm, and then inputs this data into the isolated forest for anomaly detection. The aforementioned prior anomaly detection model can first use the Pearson correlation coefficient to select strongly correlated conflict splitting heterogeneous datasets from the first and second conflict splitting heterogeneous datasets as the target conflict splitting heterogeneous dataset. Strong correlation can be defined as a Pearson correlation coefficient greater than or equal to a preset correlation coefficient threshold, which can be a pre-defined value, such as 0.2. Second, the target conflict splitting heterogeneous dataset is input into an indicator prediction model to obtain a target conflict prediction dataset. This indicator prediction model can be a machine learning model that performs linear and non-linear correlation predictions on the input target conflict splitting heterogeneous dataset. It can be a model composed of a parallel linear regression model for predicting linear correlation between indicators and a KNN (K-Nearest Neighbors) regression model for predicting non-linear correlation between indicators. Finally, the set of abnormal indicator values in the target conflict prediction dataset and the aforementioned target conflict splitting heterogeneous dataset is determined as the conflict splitting abnormal value set. The abnormal indicator data can be the ratio of the squared difference to the mean square error between the target conflict prediction dataset and the target conflict splitting heterogeneous dataset. The aforementioned clustering anomaly detection model can be an anomaly detection model that uses the distance values to cluster centers in the improved nearest neighbor propagation clustering algorithm. The aforementioned clustering anomaly detection model can obtain the conflict and split anomaly value set through the following steps: First, determine the initial distance value set between the first conflict and split heterogeneous dataset and the second conflict and split heterogeneous dataset and their nearest cluster centers. The aforementioned cluster centers can be the centers obtained by clustering the target local enterprise heterogeneous dataset after removing the first and second conflict and split heterogeneous datasets.Secondly, the mean distance from the first conflict-split heterogeneous dataset and the second conflict-split heterogeneous dataset to each target cluster center in the target cluster center set is determined to obtain the target distance mean set. The target cluster centers can be data from either the first or second conflict-split heterogeneous datasets. Finally, the difference between each initial distance value in the initial distance value set and the corresponding target distance mean in the target distance mean set is determined as the conflict-split anomaly value set.
[0041] Step 7: Based on the aforementioned conflict and split anomaly value set and the aforementioned local data quality value set, perform data type conflict resolution processing on the aforementioned first conflict and split heterogeneous dataset and the aforementioned second conflict and split heterogeneous dataset to obtain an initial resolved enterprise heterogeneous dataset. The initial resolved enterprise heterogeneous data in the aforementioned initial resolved enterprise heterogeneous dataset can be obtained by replacing the conflicted data included in the first and second conflict and split heterogeneous datasets with normal local enterprise heterogeneous data.
[0042] As an example, the aforementioned executing entity can perform the following resolution steps for each conflict-splitting anomaly pair in the conflict-splitting anomaly value set: First, determine the anomaly difference between the aforementioned conflict-splitting anomaly pairs. Second, in response to determining that the anomaly difference is greater than or equal to a preset conflict anomaly threshold, identify the first conflict-splitting heterogeneous data or the second conflict-splitting heterogeneous data with the larger value in the corresponding local data quality value pair as the initial resolution of enterprise heterogeneous data. The preset conflict anomaly threshold can be a pre-set critical value for determining the conflict resolution method. For example, the preset conflict anomaly threshold could be 0.2. Then, in response to determining that the anomaly difference is less than the preset conflict anomaly threshold, determine the ratio of the weighted sum of each conflict-splitting anomaly value in the conflict-splitting anomaly pair and the corresponding first conflict-splitting heterogeneous data or second conflict-splitting heterogeneous data from different sources to the sum of the conflict-splitting anomaly pair, as the initial resolution of enterprise heterogeneous data.
[0043] Step 8: The missing enterprise heterogeneous dataset included in the initial resolved enterprise heterogeneous dataset is sorted by missing rate to obtain a sequence of missing enterprise heterogeneous data. This sequence can be obtained by sorting the missing enterprise heterogeneous data in ascending order of missing rate. The missing rate can be the proportion of missing data in the initial resolved enterprise heterogeneous dataset.
[0044] Step nine involves performing sequence feature interpolation imputation on the aforementioned missing enterprise heterogeneous data sequences to obtain the imputed enterprise heterogeneous dataset. In practice, the executing entity can first determine the missing data importance set and missing mutual information value set of the target local enterprise heterogeneous dataset with respect to the aforementioned missing enterprise heterogeneous data sequences. The missing data importance represents the degree of contribution of each target local enterprise heterogeneous data point to the missing enterprise heterogeneous data, and can be calculated using the Random Forest Regressor in the Scikit-learn machine learning library. The missing mutual information value represents the strength of the dependency relationship between each target local enterprise heterogeneous data point and each missing enterprise heterogeneous data point, and can be calculated using the mutual_info_regression algorithm in the Scikit-learn machine learning library. The aforementioned target local enterprise heterogeneous dataset can be obtained by removing the aforementioned missing enterprise heterogeneous data sequences from the local enterprise heterogeneous dataset. Then, the target local enterprise heterogeneous dataset that meets the missing imputation conditions is selected from the aforementioned missing data importance set and missing mutual information value set as the missing dependent enterprise heterogeneous dataset. The missing data imputation conditions can be either greater than or equal to a preset data importance threshold and a preset mutual information threshold. Both the preset data importance threshold and the preset mutual information threshold can be pre-set critical values. For example, the preset data importance threshold and the preset mutual information threshold can be 0.25 and 0.75, respectively. Finally, using pre-trained multiple sets of MissForest (random forest imputation) models, the missing data sequences of the aforementioned heterogeneous enterprise datasets are iteratively imputed to obtain the imputed heterogeneous enterprise datasets. The aforementioned multiple sets of MissForest models can include: models with pre-trained KNN numerical index interpolation and categorical index interpolation strategies, and models with pre-trained linear regression numerical index interpolation and categorical index interpolation strategies. The numerical interpolation strategy can be a weighted average of multiple sets of MissForest interpolations. The categorical index interpolation strategy can be an interpolation strategy that uses the mode of multiple sets of MissForest interpolations.
[0045] Step 10: The above-mentioned filled heterogeneous enterprise dataset, the initial resolved heterogeneous enterprise dataset after removing missing heterogeneous enterprise datasets, and the local heterogeneous enterprise dataset without data conflicts are identified as the local heterogeneous conflict resolution dataset. The above-mentioned local heterogeneous conflict resolution dataset is encrypted and transmitted to the data warehouse for data storage.
[0046] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem mentioned in the background art: "Due to the differences in data quality and sensitive data, as well as data conflicts and missing data, the data quality of fused data is low, the fusion time is long, and a large amount of storage resources are wasted because of the heterogeneous data from local enterprises from different sources." The factors leading to low data quality, long fusion time, and wasted storage resources in fused data are often as follows: The differences in data quality and sensitive data, as well as data conflicts and missing data, lead to low data quality, long fusion time, and wasted storage resources in fused data. Solving these factors can improve the data quality of fused data, reduce fusion time, and reduce the waste of storage resources. To achieve this effect, this disclosure firstly uses Pearson correlation coefficient and Spearman correlation coefficient for semantic matching of indicators. Multi-correlation calculation can improve the accuracy of semantic matching and avoid the false positive rate of single correlation and the large amount of computational data required for subsequent conflict resolution. Secondly, semantic indicator conflict detection is performed through an indicator relationship fitting function. Adjusting the conflict rate through dynamic thresholds can reduce oversensitivity in low-conflict scenarios, reduce the false positive rate of conflict, and improve the accuracy of conflict detection. Subsequently, conflict splitting is performed on single and multiple conflicts to avoid interference between different types of conflict resolution. Then, the conflict splitting anomaly sets of the first and second conflict splitting heterogeneous datasets are identified and conflict resolution is performed. The accuracy of identifying conflict splitting anomaly values is ensured by using an integrated conflict resolution anomaly detection model, thereby improving the accuracy of conflict resolution in different scenarios using local data quality sets and conflict splitting anomaly sets, and improving the quality of the resolved data. Next, the initial resolved heterogeneous enterprise datasets are sorted and then filled with sequence feature interpolation. By considering the correlation between missing data and sorting for gap filling, the number of interpolation operations is reduced, improving interpolation efficiency and accuracy, reducing interpolation time, and improving the quality of the interpolated data. Finally, the local heterogeneous conflict resolution dataset is encrypted and transmitted to the data warehouse for data storage, which improves the security of sensitive data included in the local heterogeneous conflict resolution dataset and reduces the waste of storage resources.
[0047] Step 104: In response to the detection that there are resource-constrained terminal devices in the terminal device set, resource scheduling is performed on the terminal device set to obtain the terminal resource scheduling device set.
[0048] In some embodiments, the execution entity may, in response to detecting that there are resource-constrained terminal devices in the terminal device set, perform resource scheduling on the terminal device set to obtain a terminal resource scheduling device set. The terminal resource scheduling devices in the terminal resource scheduling device set may be devices that, after resource scheduling, have sufficient resources for processing. The resource-constrained terminal devices may be devices whose memory, CPU, power consumption, and other resources exceed preset limits, making them insufficient to support data processing.
[0049] In addressing the aforementioned technical problems in the application scenario—resource balancing and scheduling in data fusion based on edge terminal devices—the following technical issues often arise: Edge terminal devices have limited energy, and training models on them requires significant energy and resources, leading to reduced accuracy in data fusion, increased energy consumption, longer latency in resource balancing adjustments, lower precision in resource balancing, increased failure rate, and reduced stability. Based on the characteristics of this application scenario—edge terminal devices, limited terminal device resources and energy consumption, multi-source heterogeneity of data, deep neural network model training, and low latency—we have decided to adopt the following solution: In some optional implementations of certain embodiments, the above-mentioned resource scheduling of the terminal device set to obtain a terminal resource scheduling device set may include the following steps: The first step, based on the set of terminal devices, is to perform the following resource scheduling steps: Sub-step 1 involves identifying resource-constrained devices within the aforementioned set of terminal devices to obtain a set of scheduled terminal devices. These scheduled terminal devices can be those whose resources are insufficient to complete the initial heterogeneous multimodal data fusion model training during the current communication round when the cloud server (i.e., the executing entity) is engaged in communication training.
[0050] Sub-step 2 involves determining the data transmission rate set, data transmission delay set, data transmission energy consumption set, and device-assisted model training delay set for the scheduling terminal device set and the edge server set, as the first device resource scheduling set. The aforementioned edge server set consists of servers used to assist the scheduling terminal devices in completing federated training. The edge servers in this set can be servers capable of covering the target terminal set, possessing stronger computing and storage performance compared to the terminal devices, stable power supply, and communicating with the terminal devices using orthogonal frequency division multiple access (OFDMA) technology. The edge servers and scheduling terminal devices can have short data links and sufficient bandwidth resources. One edge server can provide resource support for at least one scheduling terminal device. The data transmission rate in the aforementioned data transmission rate set characterizes the data transmission rate between the scheduling terminal device and the corresponding edge server during current communication. The data transmission delay in the aforementioned data transmission delay set characterizes the transmission delay of data that needs to be uploaded to the edge server when the scheduling terminal device cannot complete the initial heterogeneous multimodal data fusion model. The data transmission energy consumption in the aforementioned data transmission energy consumption set characterizes the resource energy consumption consumed by the target terminal device in transmitting data during current communication between the scheduling terminal device and the corresponding edge server. The device-assisted model training latency in the aforementioned set of latency characteristics can represent the training latency generated when the data transmitted by the scheduling terminal device during its current communication with the corresponding edge server is used for model training on the edge server. The data transmission rate, data transmission latency, data transmission energy consumption, and device-assisted model training latency can be expressed as follows: .
[0051] in, Indicates the first The first communication The first dispatch terminal device and the first Data transfer rate between edge servers. Indicates data transmission delay. This indicates the energy consumption for data transmission. This indicates the latency of the device-assisted model training. Indicates the first The first communication The first dispatch terminal device and the first The number of units of bandwidth used for communication between edge servers. Indicates the first The bandwidth of an edge server. Indicates the first The total bandwidth per unit of edge server. A target terminal device can be allocated multiple units of bandwidth, and each unit of bandwidth can only be allocated to one target terminal device. Indicates the first The first edge server coverage area The transmission power of each dispatch terminal device. Indicates the first The first edge server coverage area Gain of the transmission channel between scheduling terminal devices. Indicates noise power. Indicates the first The total amount of data that a scheduling terminal device needs to train the model in each communication. Indicates the first The number of CPU cycles required for each scheduling terminal device to process the total amount of data. , This indicates the number of CPU cycles required to process each bit of data. Indicates the first The first communication The dispatch terminal device from the first The computing frequency allocated to each edge server.
[0052] Sub-step 3: Determine the local training latency set, local training energy consumption set, device upload latency set, and device upload energy consumption set of the target terminal device set as the second device resource scheduling function set. The target terminal device set is a collection of terminal devices that autonomously complete federated training. Specifically, the local training latency in the local training latency set represents the training latency of the target terminal device completing model training locally. The local training energy consumption in the local training energy consumption set represents the energy resources consumed by the target terminal device in completing model training locally. The device upload latency in the device upload latency set represents the upload latency of the target terminal device uploading model parameters to the cloud server. The device upload energy consumption in the device upload energy consumption set represents the energy resources consumed by the target terminal device in uploading model parameters to the cloud server. The local training latency, local training energy consumption, device upload latency, and device upload energy consumption can be represented as follows: .
[0053] in, Indicates that it is located at the th The first edge server coverage area Local training latency of each target terminal device. Indicates that it is located at the th The first edge server coverage area Local training energy consumption of each target terminal device. Indicates that it is located at the th The first edge server coverage area Device upload latency of each target terminal device. Indicates that it is located at the th The first edge server coverage area The device upload power consumption of each target terminal device. Indicates that it is located at the th The first edge server coverage area The computing frequency of each target terminal device. This refers to the effective switched capacitors within the target terminal device, determined by the chip structure. Indicates the first The sparsity of the communication model. This represents the amount of model parameter data in the initial heterogeneous multimodal data fusion model. Indicates the first During the second communication, it was located at the first The first edge server coverage area The data transmission rate at which a target terminal device transmits data to a cloud server. Indicates the current number of communications. Indicates the total number of communications.
[0054] Sub-step 4: Generate a device resource scheduling delay optimization function based on the first device resource scheduling function set and the second device resource scheduling function set. This device resource scheduling delay optimization function can be a set of delay optimization functions and communication constraint functions that minimize the maximum delay when the target terminal device set and the scheduling terminal device set complete the current model training communication. The device resource scheduling delay optimization function can be expressed as: .
[0055] in, Indicates the first During the second communication, it was located at the first The first edge server coverage area Model training latency for a scheduling terminal device or a target terminal device. Indicates the first During the second communication, it was located at the first The first edge server coverage area The energy resources consumed by the training of a scheduling terminal device or a target terminal device model. Indicates the first During the second communication, it was located at the first The first edge server coverage area Whether a scheduling terminal device or a target terminal device receives auxiliary training from an edge server. This represents the latency optimization function that minimizes the latency when the target terminal device set and the scheduling terminal device set complete the current model training communication. This indicates the computing frequency of the edge server. Indicates the first The first edge server coverage area The initial energy resources of a target terminal device or a scheduling terminal device. This represents a constraint function that ensures the sum of the computing frequencies allocated to the aforementioned set of scheduling terminal devices does not exceed the computing power of the edge server set. This represents a constraint function that ensures the bandwidth allocated to the aforementioned set of scheduling terminal devices does not exceed the bandwidth of the edge server. This is a constraint function that indicates the energy consumption of the scheduling terminal device and the target terminal device does not exceed the energy resources they possess. This indicates that model training is completed on the target terminal device or edge server. This indicates that the computation frequency allocated to the scheduling terminal equipment is non-negative. This indicates that the bandwidth allocated to the scheduling terminal device is non-negative.
[0056] Sub-step 5: Based on the aforementioned device resource scheduling latency optimization function, perform resource balancing scheduling on the aforementioned set of terminal devices to obtain an initial resource scheduling device set. The initial resource scheduling devices in this initial resource scheduling device set can be terminal devices obtained by providing resource assistance to the scheduling terminal device set through an edge server set.
[0057] As an example, the aforementioned execution entity can utilize a heuristic optimization algorithm to perform resource balancing scheduling on the aforementioned set of terminal devices based on the aforementioned device resource scheduling delay optimization function, thereby obtaining an initial resource scheduling device set. The aforementioned heuristic optimization algorithm can be one or a combination of particle swarm optimization, gray wolf optimization, simulated annealing, and genetic algorithms.
[0058] Sub-step 6 involves determining the current sparsity of the initial global heterogeneous multimodal data fusion model corresponding to the initial heterogeneous multimodal data fusion model set located on the initial resource scheduling device set, after a certain number of executions. This current sparsity characterizes the degree of unstructured pruning performed on the global heterogeneous multimodal data fusion model, i.e., the initial heterogeneous multimodal data fusion model located on the cloud server, to further reduce the energy consumption of the initial heterogeneous multimodal data fusion model on various terminal devices. The global heterogeneous multimodal data fusion model can be a global model located on the cloud server, used to fuse and update the model parameters sent by the initial heterogeneous multimodal data fusion model set on the initial resource scheduling device set. In practice, the executing entity can use the model adaptive sparsity calculation formula to determine the current sparsity of the initial global heterogeneous multimodal data fusion model corresponding to the initial heterogeneous multimodal data fusion model set located on the initial resource scheduling device set, after a certain number of executions. The model adaptive sparsity calculation formula can be expressed as: .
[0059] in, This represents the sparsity of the current fusion model, determined by the adaptive sparse computation formula. This represents the sparsity control factor of the model. When the sparsity of the initial global heterogeneous multimodal data fusion model is within the preset range, the value is 1; otherwise, it is 0. This indicates the final sparsity ratio. This represents the initial sparsity ratio, i.e., the sparsity of the initial global heterogeneous multimodal data fusion model during initial communication. Indicates sparse frequency. Indicates the number of times the operation has been performed. This represents the total number of communications between the initial resource scheduling device set and the cloud server. This represents the rate of change of model sparsity in the initial global heterogeneous multimodal data fusion model.
[0060] Sub-step 7: Based on the sparsity of the current fusion model, perform unstructured iterative pruning on the initial heterogeneous multimodal data fusion model set to obtain a sparse heterogeneous multimodal fusion model set. The sparse heterogeneous multimodal fusion models in the aforementioned sparse heterogeneous multimodal fusion model set can be models obtained by deleting redundant connections or neurons from the initial heterogeneous multimodal data fusion model.
[0061] As an example, the aforementioned execution entity can first generate a global model mask matrix based on the current sparsity of the fusion model. This global model mask matrix represents the mask matrix obtained by comparing the absolute values of the parameter weights of each model parameter in the initial global heterogeneous multimodal data fusion model with the current sparsity of the fusion model; a value of 0 is taken when the comparison result is less than the current sparsity, and a value of 1 is taken when the comparison result is greater than or equal to the current sparsity. The order of the elements in the global model mask matrix is obtained by sorting the parameter weights from smallest to largest. Then, using the local model gradient update formula, based on the current sparsity of the fusion model, unstructured iterative pruning is performed on the initial heterogeneous multimodal data fusion model set to obtain a sparse heterogeneous multimodal fusion model set. The local model gradient update formula can be expressed as: .
[0062] in, This represents the initial resource scheduling device determined by the local model gradient update formula. The updated model parameters of the initial heterogeneous multimodal data fusion model deployed on the platform. Indicates the initial resource scheduling device The model parameters before updating the initial heterogeneous multimodal data fusion model deployed on the platform. This represents the model learning rate. Indicates the initial resource scheduling device The local gradient of the initial heterogeneous multimodal data fusion model deployed on it. This represents the global model mask matrix.
[0063] Sub-step 8: In response to determining that the global heterogeneous multimodal data fusion model corresponding to the sparse heterogeneous multimodal fusion model set has completed training, the initial resource scheduling device set deploying the sparse heterogeneous multimodal fusion model set is determined as the terminal resource scheduling device set. Here, "training completed" can be defined as the global heterogeneous multimodal data fusion model completing training when the loss value of the global heterogeneous multimodal data fusion model meets the condition of being less than or equal to a preset loss threshold. The preset loss threshold can be a pre-set value, which needs to be determined according to specific circumstances and is not limited here.
[0064] The second step, in response to the determination that the global heterogeneous multimodal data fusion model has not been trained, is to determine the initial resource scheduling device set that deploys the sparse heterogeneous multimodal data fusion model set as the terminal device set, and to determine the sparse heterogeneous multimodal data fusion model set and the global heterogeneous multimodal data fusion model as the initial heterogeneous multimodal data fusion model set and the initial global heterogeneous multimodal data fusion model, respectively, so as to execute the above resource scheduling steps again.
[0065] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem of "reducing the accuracy of model data fusion, increasing the energy consumption of terminal devices, lengthening the delay in resource balancing adjustment, reducing the accuracy of resource balancing scheduling, increasing the damage rate of terminal devices, and reducing stability." Factors leading to reduced accuracy of model data fusion, increased energy consumption of terminal devices, lengthening the delay in resource balancing adjustment, reducing the accuracy of resource balancing scheduling, increasing the damage rate of terminal devices, and reducing stability are often as follows: Edge terminal devices have limited energy and power, and training models on edge terminal devices requires a large amount of energy and resources, leading to reduced accuracy of model data fusion, increased energy consumption of terminal devices, lengthening the delay in resource balancing adjustment, reducing the accuracy of resource balancing scheduling, increasing the damage rate of terminal devices, and reducing stability. Solving these factors can improve the accuracy of model data fusion, reduce the energy consumption of terminal devices, shorten the delay in resource balancing adjustment, improve the accuracy of resource balancing scheduling, reduce the damage rate of terminal devices, and improve stability. To achieve this effect, this disclosure first identifies the set of scheduling terminal devices in each communication round and constructs a three-layer architecture consisting of the terminal device set, edge servers, and cloud servers. This effectively reduces the transmission latency and duration of direct communication between terminal devices and cloud servers. Furthermore, resource load balancing via edge servers closer to the terminal device set shortens the resource scheduling distance, improves scheduling efficiency, and reduces scheduling time. Secondly, a device resource scheduling latency optimization function is constructed by considering transmission latency and energy consumption, as well as model training latency and energy consumption. This quantifies the problem of terminal devices falling behind due to resource and energy consumption during federated learning. Constructing the function from different perspectives effectively avoids resource waste and shortages caused by over-allocation and ensures that the constraint function better matches the federated learning scenario of the terminal devices, reducing resource adjustment latency. Finally, an initial resource scheduling device set is obtained by solving the device resource scheduling latency optimization function using a heuristic algorithm. This improves the accuracy and efficiency of resource scheduling and reduces resource scheduling latency. Finally, after adaptive sparsity processing of the global heterogeneous multimodal data fusion model, unstructured pruning is performed on the initial heterogeneous multimodal data fusion model set located on the aforementioned terminal devices to obtain the terminal devices with the pruned models. As a set of terminal resource scheduling devices, redundant connections or neurons in the model can be deleted. This can reduce the resources and energy required by the terminal devices while ensuring the accuracy of the model, reduce the computational pressure on the terminal devices running the model, improve the stability of the terminal devices, and reduce the damage rate of the terminal devices.
[0066] In addressing the aforementioned technical problems in the application scenario—resource balancing scheduling based on edge terminal devices—the following technical issues often arise: Due to the limited resources and energy consumption of terminal devices, when performing multi-source heterogeneous data fusion based on federated learning through these devices, terminal devices may fail to receive and send local model parameters in a timely manner, leading to issues such as terminal device lag and uneven resource allocation. This results in higher terminal device failure rates and energy consumption, lower accuracy of resource scheduling, waste of terminal device resources, and longer resource balancing scheduling times. Based on the characteristics of this application scenario—federated learning, edge terminal devices, limited terminal device resources and energy, multi-source heterogeneous data, and low latency—we have decided to adopt the following solution: Optionally, the above-mentioned resource scheduling of the terminal device set to obtain the terminal resource scheduling device set may include the following steps: The first step is to generate a device resource scheduling latency optimization function based on the aforementioned set of terminal devices. It should be noted that the specific implementation of this step can refer to the implementation methods of steps one through five in some optional implementation methods of the preceding embodiments.
[0067] The second step involves transforming the aforementioned device resource scheduling delay optimization function to obtain a resource scheduling Markov decision model. This model includes a resource scheduling state space, a resource scheduling action space, and a resource scheduling reward function. The resource scheduling state space can include: the remaining battery power of the terminal device in the current communication round, the allocated sample data transmission bandwidth, the allocated computing frequency, and whether auxiliary training is received. The resource scheduling action space can be a set of resource allocation strategies and auxiliary training strategies that the terminal device can adopt during resource scheduling. This action space can include: changes in the allocated transmission bandwidth, changes in the allocated computing frequency, and changes in the auxiliary training flag in the current communication round. The resource scheduling reward function can be a reward / penalty function for the terminal device performing actions corresponding to those in the resource scheduling action space. This reward function can maximize the long-term cumulative reward.
[0068] The resource scheduling reward function described above can be expressed as: .
[0069] The third step involves inputting the current resource scheduling state information corresponding to the resource scheduling state space included in the aforementioned resource scheduling Markov decision model into the initialization resource scheduling policy network included in the initialization resource scheduling deep reinforcement model, thereby obtaining the target resource scheduling action information. The aforementioned initialization resource scheduling deep reinforcement model further includes: initializing the resource scheduling value network, initializing the target policy network, and initializing the target value network. The aforementioned initialization resource scheduling deep reinforcement model is a model obtained by initializing the model parameters of the resource scheduling deep reinforcement model. This resource scheduling deep reinforcement model can be a deep reinforcement learning agent based on a value network and policy network architecture used to solve the aforementioned resource scheduling Markov decision model. For example, the aforementioned resource scheduling deep reinforcement model can be a model based on DDPG (Deep Deterministic Policy Gradient) that includes a value network and a policy model. The aforementioned current resource scheduling state information can be the remaining battery power of the terminal device set, the allocated sample data transmission bandwidth, the allocated computing frequency, and whether auxiliary training information is received. The aforementioned target resource scheduling action information can be the action information in the resource scheduling action space determined by the initialization resource scheduling policy network.
[0070] The fourth step involves generating the next resource scheduling state information and resource scheduling reward information based on the aforementioned target resource scheduling action information. The next resource scheduling state information can be the state information of the terminal device obtained after the deep agent executes the aforementioned target resource scheduling action information. The resource scheduling reward information can be the reward function obtained by inputting the model training delay obtained through the device resource scheduling delay optimization function after the deep agent executes the aforementioned target resource scheduling action information back into the aforementioned resource scheduling reward function. As an example, the executing agent can first add random noise to the aforementioned target resource scheduling action information to obtain target noise action information. The random noise can be Gaussian noise. Then, the deep reinforcement agent is controlled to execute the aforementioned target noise action information to generate the next resource scheduling state information and resource scheduling reward information.
[0071] The fifth step involves determining the current resource scheduling status information, target resource scheduling action information, resource scheduling reward information, and next resource scheduling status information as a resource scheduling quadruple, and inputting this resource scheduling quadruple into the initialization resource scheduling buffer pool to obtain the target resource scheduling buffer pool. The initialization resource scheduling buffer pool can be a data pool used to store the quadruple of {resource scheduling status, resource scheduling action, resource scheduling reward, next resource scheduling status}.
[0072] Step 6: In response to determining that the target resource scheduling buffer pool meets the preset storage conditions, perform batch gradient sampling processing on the target resource scheduling buffer pool to obtain a device resource scheduling sample set. The preset storage conditions can be that the amount of data included in the target resource scheduling buffer pool is greater than or equal to a preset data amount. The preset data amount can be 0.8 times the total data amount of the target resource scheduling buffer pool. The device resource scheduling sample set can be a sample set obtained by sampling from the target resource scheduling buffer pool using a mini-batch gradient descent algorithm.
[0073] Step 7: Based on the above equipment resource scheduling sample set, update the model of the initial resource scheduling value network to obtain the updated resource scheduling value network.
[0074] As an example, the aforementioned execution entity can utilize the model loss gradient function to update the initial resource scheduling value network based on the aforementioned device resource scheduling sample set, thereby obtaining the updated resource scheduling value network. The aforementioned model loss function can be expressed as: .
[0075] in, This represents the loss function for initializing the resource scheduling value network, i.e., the mean squared error function based on temporal difference. This represents the immediate reward obtained after executing the target resource scheduling action, calculated by the resource scheduling reward function. This indicates the current resource scheduling status information for the current communication round. This indicates the next resource scheduling status information. This represents the output of the initialization target strategy, which is the optimal action under the next resource scheduling state information. This represents the set of parameters used to initialize the resource scheduling value network. This indicates the resource scheduling action to be executed under the current resource scheduling status information, and its value is data in the resource scheduling action space. The Q value represents the predicted action value of the initial target value network. This represents the predicted Q-value of the initial resource scheduling value network. This indicates the initialization of the target value network. This represents the model parameters for initializing the target value network. This represents the output of the initialization target policy network, indicating that in The optimal action under the given conditions. This indicates the initialization of the target policy network. This represents the model parameters for initializing the target policy network. This represents the target value of DDPG, which is used to evaluate the "long-term value of the current action," combining the predicted value of the immediate reward and the target value network. express. It represents the mathematical expectation, which is the average of the equipment resource scheduling sample set, reflecting the expected level of loss. This represents the sample set of equipment resource scheduling. This represents the discount factor, used to calculate the discount on long-term value. Indicates the first Instant rewards for each sample of device resource scheduling. Indicates the first Next resource scheduling status information corresponding to each device resource scheduling sample. Indicates the first Resource scheduling status information corresponding to each device resource scheduling sample. Indicates the first Resource scheduling actions corresponding to each device resource scheduling sample. This represents the sample index of the equipment resource scheduling sample in the equipment resource scheduling sample set. Indicates the first The predicted Q-value of the initial resource scheduling value network corresponding to each device resource scheduling sample. Indicates the first The predicted Q-value of the initial target value network corresponding to each device resource scheduling sample. This represents the gradient of the initial value network with respect to its own parameters.
[0076] Step 8: Based on the aforementioned equipment resource scheduling sample set, perform a delayed update on the initial resource scheduling policy network to obtain the updated resource scheduling policy network. Also, perform a soft update on the initial target policy network and the initial target value network to obtain the updated target value network and the updated target policy network. Repeat steps 3 through 8 until a preset number of training rounds are reached. The delayed update can be performed once after every preset number of updates to the initial resource scheduling value network to avoid fluctuations in the initial resource scheduling value network updates affecting the policy optimization updates. This preset number can be pre-set and can be determined based on specific circumstances; it will not be elaborated further here.
[0077] As an example, the aforementioned execution entity can first use a deterministic gradient update formula to perform a delayed update on the initial resource scheduling policy network based on the target resource scheduling buffer pool, thereby obtaining the updated resource scheduling policy network. The deterministic gradient update formula can be expressed as: .
[0078] in, This represents the objective function for initializing the resource scheduling policy network, and the policy is represented by... The long-term cumulative benefits generated from interactions within the environment. express The approximate gradient is the basis for updating the network parameters of the initial resource scheduling strategy. Indicates the first The output of the initialization target policy network corresponding to each device resource scheduling sample under the current resource scheduling status information. Representation Strategy The mathematical expectation. These represent instant rewards at different times. This represents the gradient of the initial resource scheduling policy network with respect to its own parameters, reflecting the degree to which changes in the initial resource scheduling policy network parameters affect the output action.
[0079] Finally, the initial target policy network and the initial target value network are soft-updated using the soft update formula to obtain the updated target value network and the updated target policy network. The soft update formula can be expressed as: .
[0080] in, This represents the soft update formula for initializing the target policy network. This parameter represents the speed at which the model is updated. This represents the soft update formula for initializing the target value network.
[0081] Step 9: Using the updated resource scheduling strategy network, the resource scheduling Markov decision model is solved in a reinforced manner to obtain device resource scheduling information. Based on the device resource scheduling information, the terminal device set is subjected to resource balancing scheduling to obtain the initial resource scheduling device set.
[0082] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the second technical problem mentioned in the background: "high damage rate and energy consumption of terminal devices, low accuracy of resource scheduling, waste of terminal device resources, and long resource balancing scheduling time." The factors leading to high damage rate and energy consumption of terminal devices, low accuracy of resource scheduling, waste of terminal device resources, and long resource balancing scheduling time are often as follows: Due to the limited resources and energy consumption of terminal devices, when performing multi-source heterogeneous data fusion based on federated learning through terminal devices, there are issues such as terminal devices failing to receive and send local model parameters in a timely manner, resulting in terminal device lag and uneven resource allocation. This leads to high damage rate and energy consumption of terminal devices, low accuracy of resource scheduling, waste of terminal device resources, and long resource balancing scheduling time. Solving these factors can reduce the damage rate and energy consumption of terminal devices, improve the accuracy of resource scheduling, reduce waste of terminal device resources, and shorten the long resource balancing scheduling time. To achieve this effect, this disclosure first transforms the generated device resource scheduling delay optimization function to obtain a resource scheduling Markov decision model. This constructs the resource allocation problem as a learnable Markov decision process to learn the optimal strategy in a continuous action space and provides basic input data for subsequent solutions. Second, an initialization resource scheduling deep reinforcement learning model, including an initialization resource scheduling policy network, an initialization resource scheduling value network, an initialization target value network, and an initialization target policy network, is used to obtain target resource scheduling action information. This initialization model avoids the volatility of Q-value estimation. Next, by having the deep reinforcement agent interact with the surrounding environment and adding random noise, the exploration range is increased, generating the next resource scheduling state information and resource scheduling reward information. This information, along with the current resource scheduling state information and target resource scheduling action information, is stored and randomly sampled. An experience replay mechanism reduces the temporal correlation of samples, effectively avoiding model training oscillations and improving the prediction and evaluation accuracy of the initialization resource scheduling deep reinforcement learning model. Finally, different updates are performed on each network included in the initialization resource scheduling deep reinforcement learning model to avoid instability in update training caused by sudden changes in model parameters, ensuring the asymptotic nature of value evaluation and policy optimization. Ultimately, by generating device resource scheduling information through the updated resource scheduling strategy network and performing resource balancing scheduling, the accuracy of resource balancing scheduling can be improved, the scheduling problem of terminal devices in federated learning can be reduced, the processing latency and energy consumption of terminal devices can be reduced, the resource waste and damage rate of terminal devices can be reduced, and the resource balancing scheduling time can be shortened.
[0083] Step 105: Control the terminal resource scheduling device set, and train the corresponding initial heterogeneous multimodal data fusion model set according to the local heterogeneous conflict resolution dataset to obtain the heterogeneous multimodal data fusion model set.
[0084] In some embodiments, the aforementioned execution entity can control the aforementioned set of terminal resource scheduling devices to train a model on the corresponding initial heterogeneous multimodal data fusion model set based on the aforementioned local heterogeneous conflict resolution dataset, thereby obtaining a heterogeneous multimodal data fusion model set. The aforementioned initial heterogeneous multimodal data fusion model can be a deep neural network model located on each terminal device for fusing data from the input local heterogeneous conflict resolution dataset. The aforementioned initial heterogeneous multimodal data fusion model can be an M3amba model (CLIP-driven Mamba Model for Multi-modal Remote Sensing Classification). The heterogeneous multimodal data fusion models in the aforementioned heterogeneous multimodal data fusion model set can be models after initial model training using the local heterogeneous conflict resolution dataset. The aforementioned model training can be performed using a combination of supervised loss and unsupervised consistency loss. The aforementioned supervised loss can be loss optimization by comparing the predicted probability distribution with the true labels. The aforementioned unsupervised consistency loss can be loss optimization by calculating the cosine similarity between different modal outputs to further enhance the consistency between modalities.
[0085] Step 106: Perform multimodal federated fusion processing on the model parameter set corresponding to the heterogeneous multimodal data fusion model set to obtain the federated fusion model parameter set.
[0086] In some embodiments, the execution entity may perform multimodal federated fusion processing on the model parameter set corresponding to the heterogeneous multimodal data fusion model set to obtain a federated fusion model parameter set. The federated fusion model in the federated fusion model parameter set may be global model parameters determined through a federated learning algorithm.
[0087] Step 107: Based on the federated fusion model parameter set and the heterogeneous multimodal data fusion model set, perform global multi-level fusion on the local heterogeneous conflict resolution dataset to obtain the enterprise global fusion dataset.
[0088] In some embodiments, the executing entity may perform global multi-level fusion of the local heterogeneous conflict resolution dataset based on the federated fusion model parameter set and the heterogeneous multimodal data fusion model set to obtain an enterprise global fusion dataset. The enterprise global fusion dataset may be data obtained by fusing local heterogeneous conflict resolution datasets located on a set of terminal resource scheduling devices.
[0089] As an example, the aforementioned executing entity can first encrypt and transmit the aforementioned federated fusion model parameter set to the aforementioned heterogeneous multimodal data fusion model set and the control terminal resource scheduling device set. After decrypting the received encrypted federated fusion model parameter set, the entity can then retrain the heterogeneous multimodal data fusion model set to obtain a trained heterogeneous multimodal data fusion model set. Then, the local heterogeneous conflict resolution dataset is input into the trained heterogeneous multimodal data fusion model set to obtain the enterprise-wide fusion dataset.
[0090] In some optional implementations of certain embodiments, the process of performing global multi-level fusion of the local heterogeneous conflict resolution dataset based on the federated fusion model parameter set and the heterogeneous multimodal data fusion model set to obtain an enterprise global fusion dataset may include the following steps: The first step involves filtering the aforementioned local heterogeneous conflict resolution dataset to obtain an enterprise text dataset and an enterprise image dataset. The enterprise text data in the enterprise text dataset can be data representing the local heterogeneous conflict resolution data in text form. Similarly, the enterprise images in the enterprise image dataset can be data representing the local heterogeneous conflict resolution data in image form.
[0091] The second step involves performing word embedding encoding on the aforementioned enterprise text dataset to obtain a set of word embedding vectors. The word embedding vectors in this set can represent the characters included in the enterprise text dataset in vector form. In practice, the executing entity can first perform keyword extraction processing on the enterprise text dataset to obtain a set of enterprise text keywords. Then, it can perform word embedding encoding on the enterprise text keyword set to obtain a set of word embedding vectors.
[0092] The third step involves determining the positional feature vector and textual feature vector for each word embedding vector in the aforementioned word embedding vector set, resulting in a positional feature vector set and a textual feature vector set. The positional feature vector represents the location information of the keyword corresponding to the word embedding vector within the aforementioned enterprise text data. The textual feature vector represents the semantic information of the keyword corresponding to the word embedding vector.
[0093] The fourth step is to perform feature concatenation on the above-mentioned word embedding vector set, position feature vector set, and text feature vector set to obtain the text concatenation feature vector set.
[0094] The fifth step involves inputting the aforementioned set of concatenated text feature vectors into a text encoding model to obtain a set of text feature vectors. These text feature vectors represent the complete semantic information of the keywords within the context of the aforementioned enterprise text. The text encoding model can be a Transformer encoder.
[0095] The sixth step involves inputting the aforementioned set of text feature vectors into a residual network to obtain a set of text residual feature vectors. This residual network can be a deep neural network model that extracts local and global features from the input set of text feature vectors to output the set of text residual feature vectors. For example, the residual network could be a ResNet34 model.
[0096] Step 7: For each enterprise image in the above enterprise image set, perform the following convolutional embedding step: Sub-step 1 involves segmenting the aforementioned enterprise images into blocks for encoding, resulting in an image feature vector group. The image feature vectors in this group can represent the feature information of each segmented enterprise image block. The executing entity can first segment the aforementioned enterprise images into blocks to obtain a set of enterprise image blocks. Each enterprise image block can have 16 pixels. 16. Then, the above enterprise images are input into the linear projection layer to obtain image feature vector groups.
[0097] Sub-step 2 involves concatenating the aforementioned image feature vector group and the preset image block position feature vector group to obtain a concatenated image feature vector group. The preset image block position feature vector in the preset image block position feature vector group can be the position information of the image block corresponding to the image feature vector within the enterprise image.
[0098] Sub-step 3 involves performing overlapping convolution on the stitched image feature vector group to obtain an image spatial feature vector group. The image spatial feature vectors in this group can characterize the positional relationships of pixels in the enterprise image and local structural information, including texture, edges, and shape. This overlapping convolution can be performed using a fully convolutional network.
[0099] Sub-step 4: Input the above image spatial feature vector group into the multi-head attention mechanism layer to obtain the image weight feature vector group.
[0100] Sub-step 5 involves performing convolutional embedding on the aforementioned image weight feature vector group to obtain an image semantic feature vector group. The image semantic feature vectors in this group can represent high-level semantic information about the object category, image content, and event logic of the enterprise image. This convolutional embedding process can be performed using the U-Net model.
[0101] Sub-step 6 involves fusing the aforementioned text residual feature vector set and the obtained image semantic feature vector set to obtain the enterprise-wide fused dataset. In practice, the executing entity can first align the text residual feature vector set and the image semantic feature vector set to obtain an aligned feature vector set. Then, a weighted sum is performed on the aligned feature vector set to obtain the enterprise-wide fused dataset.
[0102] Optionally, the above-mentioned feature data fusion of the text residual feature vector set and the obtained image semantic feature vector set to obtain the enterprise global fusion dataset may include the following steps: The first step is to perform text feature encoding on the above set of text residual feature vectors to obtain a set of text feature vectors. This text feature encoding can be performed using a Transformer model.
[0103] The second step involves image feature encoding of the aforementioned set of image semantic feature vectors to obtain a set of image feature vectors. This image feature encoding can be performed using a Vision Transformer.
[0104] The third step is to perform channel concatenation on the above-mentioned text feature vector set and the above-mentioned image feature vector set to obtain the concatenated feature vector.
[0105] The fourth step is to perform tensor decomposition on the spliced feature vectors to obtain the height feature vector, width feature vector, and channel feature vector.
[0106] The fifth step is to perform feature multiplication on the height feature vector and the width feature vector to obtain the spatial attention feature vector.
[0107] The sixth step is to perform feature multiplication on the spatial attention feature vector and the channel feature vector to obtain the three-dimensional attention feature vector.
[0108] Step 7: Determine the sum of the above three-dimensional attention feature vector and the above concatenated feature vector as the semantic feature vector.
[0109] Step 8: Perform global average pooling on the above semantic feature vectors to obtain the pooled feature vectors.
[0110] The ninth step involves inputting the pooled feature vectors into a feature fusion network to obtain the enterprise-wide fused dataset. The feature synthesis network can be a multilayer perceptron with four layers for feature fusion.
[0111] Step 108: Encrypt and transmit the enterprise global fusion dataset to the data warehouse for data storage.
[0112] In some embodiments, the executing entity may encrypt and transmit the aforementioned enterprise-wide fused dataset to a data warehouse for data storage. The data warehouse may be a database used to store the enterprise-wide fused dataset. The encrypted transmission may only encrypt sensitive data included in the enterprise-wide fused dataset, reducing the amount of encrypted data, improving encryption efficiency, enhancing the security of the enterprise-wide fused dataset, and reducing transmission resource consumption.
[0113] In some optional implementations of certain embodiments, encrypting and transmitting the aforementioned enterprise-wide fused dataset to a data warehouse for data storage may include the following steps: The first step is to obtain the public key for encrypting heterogeneous data. This public key can be the Shamir secret shared key.
[0114] The second step involves classifying the aforementioned enterprise-wide fusion dataset into three data types: a first heterogeneous fusion dataset, a second heterogeneous fusion dataset, and a third heterogeneous fusion dataset. The first heterogeneous fusion dataset can contain structured data from the enterprise-wide fusion dataset. The second heterogeneous fusion dataset can contain unstructured data from the enterprise-wide fusion dataset. The third heterogeneous fusion dataset can contain semi-structured data from the enterprise-wide fusion dataset.
[0115] The third step involves performing empirical mode decomposition on the aforementioned first heterogeneous fusion dataset to obtain an enterprise modality fusion data sequence. The enterprise modality fusion data in this sequence can be stable and interpretable feature pattern data extracted from the aforementioned first heterogeneous fusion dataset, and each enterprise modality fusion data point is independent and non-redundant.
[0116] The fourth step is to perform homomorphic encryption on the above-mentioned heterogeneous data encryption public key to obtain the first heterogeneous fusion dataset after encryption.
[0117] The fifth step involves segmenting the aforementioned second heterogeneous fusion dataset to obtain a segmented second heterogeneous fusion dataset. The segmented second heterogeneous fusion data in this dataset can be obtained by segmenting data according to the acquisition time sequence or logical relationship of the second heterogeneous fusion dataset (e.g., character order in text, time flow of enterprise data). The segment length of the segmented second heterogeneous fusion data can meet the length requirements of the sequence prediction model input.
[0118] Step 6: Perform frequency domain transformation on the segmented second heterogeneous fused data set to obtain the transformed second heterogeneous fused data set. The frequency domain transformation can be performed by normalizing the segmented second heterogeneous fused data set, then performing a Fourier transform to convert it from the time domain to the frequency domain, followed by tensor adaptation processing. The tensor adaptation processing can be a process of converting the two-dimensional tensor transformed to the frequency domain into a three-dimensional tensor.
[0119] Step 7: Input the transformed second heterogeneous fused data set into the sequence prediction model to obtain the time-series heterogeneous fused data sequence. The sequence prediction model can be a deep neural network model that extracts complex sequence association features from the input transformed second heterogeneous fused data set to output the heterogeneous sequence fused data sequence. The sequence prediction model can be a bidirectional long short-term memory neural network model.
[0120] Step 8 involves performing data mapping processing on the aforementioned heterogeneous temporal fusion data sequence to obtain a heterogeneous byte fusion data sequence. The heterogeneous byte sequence fusion data in the heterogeneous byte sequence fusion dataset can be a byte sequence that meets the requirements of subsequent forward and backward feedback encryption. In practice, the executing entity can first randomly select a decimal place from the Logistic chaotic mapping system to obtain a chaotic random number. Then, the product of the chaotic random number and the aforementioned heterogeneous temporal fusion data sequence is determined as the chaotic temporal fusion data sequence. Finally, the chaotic temporal fusion data sequence is multiplied by a factor and the remainder after dividing by 256 is taken to obtain the heterogeneous byte fusion data sequence.
[0121] Step 9: Based on the aforementioned heterogeneous byte fusion data sequence, perform forward and backward feedback encryption on the aforementioned temporal heterogeneous fusion data sequence to obtain the encrypted second heterogeneous fusion dataset. In practice, the aforementioned execution entity can first select data located at the initial and final positions from the aforementioned heterogeneous byte fusion data sequence and the aforementioned temporal heterogeneous fusion data sequence, as the starting heterogeneous byte fusion data, the starting temporal heterogeneous fusion data, the ending heterogeneous byte fusion data, and the ending temporal heterogeneous fusion data. Secondly, after XORing the starting heterogeneous byte fusion data and the starting temporal heterogeneous fusion data, add them to the starting heterogeneous byte fusion data and determine the remainder after dividing by 256, which is used as the initial forward encryption fusion data. For the temporal heterogeneous fusion data following the starting temporal heterogeneous fusion data, determine the XOR result of the temporal heterogeneous fusion data and the corresponding heterogeneous byte fusion data, add it to the previously encrypted forward encryption fusion data, and determine the remainder after dividing by 256 to obtain the forward encryption fusion data sequence. Next, backward feedback encryption is performed on the initial forward encrypted fused data and the forward encrypted fused data sequence to obtain the encrypted second heterogeneous fused dataset. The backward feedback encryption method is the same as the forward feedback encryption method, only in reverse direction.
[0122] Step 10: Perform third data encryption on the above-mentioned third heterogeneous fusion dataset to obtain the encrypted third heterogeneous fusion dataset.
[0123] Step 11: Store the encrypted first heterogeneous fusion dataset, the encrypted second heterogeneous fusion dataset, and the encrypted third heterogeneous fusion dataset into the data warehouse.
[0124] Optionally, the above-mentioned third heterogeneous fusion dataset is subjected to third data encryption to obtain an encrypted third heterogeneous fusion dataset, which may include the following steps: The first step is to generate a heterogeneous chaotic sequence for enterprises. This heterogeneous chaotic sequence can be a two-dimensional sequence used for data encryption of a third heterogeneous fused dataset. In practice, the executing entity can utilize a two-dimensional discrete chaotic system to generate the heterogeneous chaotic sequence for enterprises.
[0125] The second step involves performing a nonlinear transformation on the aforementioned heterogeneous chaotic sequence to obtain a heterogeneous chaotic key stream. This heterogeneous chaotic key stream can be a chaotic key stream obtained by inputting the heterogeneous chaotic sequence into a nonlinear transformation function. The nonlinear transformation function can be a function of the sum of the product of 0.4 and the square of the x-coordinate in the heterogeneous chaotic sequence, and the product of 0.5 and the square of the y-coordinate in the heterogeneous chaotic sequence.
[0126] The third step is to perform analog-to-digital conversion on the above heterogeneous chaotic key stream to obtain the enterprise heterogeneous key.
[0127] The fourth step involves grouping the aforementioned third heterogeneous fusion dataset into separate groups, resulting in grouped third heterogeneous data sets. Each group within a grouped third heterogeneous data set contains data of the same length. For example, a grouped third heterogeneous data set could contain 8 identical characters of third heterogeneous data.
[0128] Fifth, based on the third heterogeneous data set after grouping, perform the following transformation steps: Sub-step 1 involves grouping the aforementioned heterogeneous enterprise keys to obtain grouped heterogeneous enterprise key sets. The grouped heterogeneous enterprise keys in these sets can be 32-bit keys.
[0129] Sub-step 2 involves performing XOR encryption on the aforementioned grouped third heterogeneous data set and the aforementioned grouped enterprise heterogeneous key set to obtain an encrypted third heterogeneous data set. The XOR encryption process can be performed using round-based encryption rules in IDEA (International Data Encryption Algorithm).
[0130] Sub-step 3 involves shifting the heterogeneous enterprise key to obtain the shifted heterogeneous enterprise key. This shifting process can be a cyclic left shift of 25 bits.
[0131] Sub-step 4: In response to determining that the number of times the above transformation steps have been executed is greater than or equal to a preset execution threshold, the encrypted third heterogeneous data set is transformed according to the shifted enterprise heterogeneous key to obtain the encrypted third heterogeneous fusion dataset. The preset execution threshold can be a pre-defined maximum number of times the transformation steps are executed. For example, the preset execution threshold can be 8. In practice, the executing entity can first group the shifted enterprise heterogeneous keys to obtain target heterogeneous key groups. Then, each encrypted third heterogeneous data in each encrypted third heterogeneous data group in the encrypted third heterogeneous data set is multiplied and added to the corresponding target heterogeneous key in the target heterogeneous key group to obtain the encrypted third heterogeneous fusion dataset. The corresponding target heterogeneous key can be the target heterogeneous key with the same position in the set as the encrypted third heterogeneous data. The multiplication and addition process can be multiplying the data in the first and fourth positions and adding the data in the second and fourth positions.
[0132] Step 6: In response to determining that the number of executions has been less than the preset execution threshold, the shifted heterogeneous enterprise key is identified as the heterogeneous enterprise key, and the above transformation steps are executed again.
[0133] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a multi-source heterogeneous data fusion device, which are similar to... Figure 1 Corresponding to the method embodiments shown, this multi-source heterogeneous data fusion device can be specifically applied to various electronic devices.
[0134] like Figure 2As shown, a multi-source heterogeneous data fusion device 200 includes: a control unit 201, a data quality detection unit 202, a heterogeneous conflict resolution unit 203, a resource scheduling unit 204, a model training unit 205, a multimodal federated fusion unit 206, a global multi-level fusion unit 207, and an encrypted transmission unit 208. The control unit 201 is configured to control a set of terminal devices to preprocess the acquired heterogeneous enterprise datasets from different sources, obtaining preprocessed heterogeneous enterprise datasets as local heterogeneous enterprise datasets. The data quality detection unit 202 is configured to perform data quality detection on each local heterogeneous data in the aforementioned local heterogeneous enterprise dataset to generate local data quality values, obtaining a local data quality value set. The heterogeneous conflict resolution unit 203 is configured to perform heterogeneous conflict resolution processing on the aforementioned local heterogeneous enterprise dataset based on the aforementioned local data quality value set, obtaining a local heterogeneous conflict resolution dataset. Resource scheduling unit 204 is configured to: in response to detecting resource-constrained terminal devices in the aforementioned terminal device set, perform resource scheduling on the aforementioned terminal device set to obtain a terminal resource scheduling device set. Model training unit 205 is configured to: control the aforementioned terminal resource scheduling device set to train the corresponding initial heterogeneous multimodal data fusion model set based on the aforementioned local heterogeneous conflict resolution dataset to obtain a heterogeneous multimodal data fusion model set. Multimodal federated fusion unit 206 is configured to: perform multimodal federated fusion processing on the model parameter set corresponding to the aforementioned heterogeneous multimodal data fusion model set to obtain a federated fusion model parameter set. Global multi-level fusion unit 207 is configured to: perform global multi-level fusion on the aforementioned local heterogeneous conflict resolution dataset based on the aforementioned federated fusion model parameter set and the heterogeneous multimodal data fusion model set to obtain an enterprise global fusion dataset. Encrypted transmission unit 208 is configured to: encrypt and transmit the aforementioned enterprise global fusion dataset to the data warehouse for data storage.
[0135] It is understandable that the various units and references described in the multi-source heterogeneous data fusion device 200 Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the multi-source heterogeneous data fusion device 200 and the units contained therein, and will not be repeated here.
[0136] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0137] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0138] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0139] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0140] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0141] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0142] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: control the set of terminal devices to preprocess the acquired heterogeneous enterprise datasets from different sources to obtain preprocessed heterogeneous enterprise datasets as local heterogeneous enterprise datasets; perform data quality checks on each local heterogeneous data in the aforementioned local heterogeneous enterprise dataset to generate local data quality values, obtaining a local data quality value set; perform heterogeneous conflict resolution processing on the aforementioned local heterogeneous enterprise dataset based on the aforementioned local data quality value set, obtaining a local heterogeneous conflict resolution dataset; and respond to the detection of resource-constrained terminals in the aforementioned set of terminal devices. The device performs resource scheduling on the aforementioned set of terminal devices to obtain a terminal resource scheduling device set; controls the aforementioned terminal resource scheduling device set to train the corresponding initial heterogeneous multimodal data fusion model set based on the aforementioned local heterogeneous conflict resolution dataset to obtain a heterogeneous multimodal data fusion model set; performs multimodal federated fusion processing on the model parameter set corresponding to the aforementioned heterogeneous multimodal data fusion model set to obtain a federated fusion model parameter set; performs global multi-level fusion on the aforementioned local heterogeneous conflict resolution dataset based on the aforementioned federated fusion model parameter set and the heterogeneous multimodal data fusion model set to obtain an enterprise global fusion dataset; and encrypts and transmits the aforementioned enterprise global fusion dataset to a data warehouse for data storage.
[0143] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0145] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a control unit, a data quality detection unit, a heterogeneous conflict resolution unit, a resource scheduling unit, a model training unit, a multimodal federated fusion unit, a global multi-level fusion unit, and an encrypted transmission unit. The names of these units do not necessarily limit the specific unit; for example, a control unit may also be described as "a unit that controls a set of terminal devices to preprocess enterprise heterogeneous datasets from different sources to obtain preprocessed enterprise heterogeneous datasets, which serve as local enterprise heterogeneous datasets."
[0146] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0147] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A multi-source heterogeneous data fusion method, comprising: controlling a terminal device set, respectively performing data preprocessing on acquired different source enterprise heterogeneous data sets to obtain preprocessed enterprise heterogeneous data sets as local enterprise heterogeneous data sets; performing data quality detection on each local enterprise heterogeneous data in the local enterprise heterogeneous data set to generate local data quality values and obtain a local data quality value set; performing heterogeneous conflict resolution processing on the local enterprise heterogeneous data set according to the local data quality value set to obtain a local heterogeneous conflict resolution data set; in response to detecting that there is a resource-limited terminal device in the terminal device set, performing resource scheduling on the terminal device set to obtain a terminal resource scheduling device set; controlling the terminal resource scheduling device set, performing model training on a corresponding initial heterogeneous multi-modal data fusion model set according to the local heterogeneous conflict resolution data set to obtain a heterogeneous multi-modal data fusion model set; performing multi-modal federated fusion processing on a model parameter set corresponding to the heterogeneous multi-modal data fusion model set to obtain a federated fusion model parameter set; performing global multi-level fusion on the local heterogeneous conflict resolution data set according to the federated fusion model parameter set and the heterogeneous multi-modal data fusion model set to obtain an enterprise global fusion data set; encrypting and transmitting the enterprise global fusion data set to a data warehouse for data storage.
2. The method of claim 1, wherein, The data quality detection on each local enterprise heterogeneous data in the local enterprise heterogeneous data set to generate local data quality values and obtain a local data quality value set comprises: for each local enterprise heterogeneous data in the local enterprise heterogeneous data set, the following data quality evaluation steps are performed: determining a data quality evaluation dimension set of the local enterprise heterogeneous data; performing sub-dimension division on each data quality evaluation dimension in the data quality evaluation dimension set to generate a data quality evaluation index group and obtain a data quality evaluation index group set; determining the evaluation index weight values of each data quality evaluation index included in each data quality evaluation index group in the data quality evaluation index group set to obtain an evaluation index weight value group set; determining the evaluation dimension weight values of each data quality evaluation dimension in the data quality evaluation dimension set to obtain an evaluation dimension weight value set; determining the local data quality value set of the local enterprise heterogeneous data according to the evaluation dimension weight value set, the evaluation index weight value group set, the data quality evaluation index group set, and the data quality evaluation dimension set.
3. The method of claim 1, wherein, The global multi-level fusion on the local heterogeneous conflict resolution data set according to the federated fusion model parameter set and the heterogeneous multi-modal data fusion model set to obtain an enterprise global fusion data set comprises: performing screening on the local heterogeneous conflict resolution data set to obtain an enterprise text data set and an enterprise image set; performing word embedding coding on the enterprise text data set to obtain a word embedding vector group set; determining the position feature vector and the text feature vector of each word embedding vector in the word embedding vector group set to obtain a position feature vector group set and a text feature vector group set; concatenate the word embedding vector set, the position feature vector set and the text feature vector set to obtain a text concatenated feature vector set; input the text concatenated feature vector set into a text encoding model to obtain a text feature vector set; input the text feature vector set into a residual network to obtain a text residual feature vector set; for each enterprise image in the enterprise image set, the following convolutional embedding steps are performed: block coding is performed on the enterprise image to obtain an image feature vector set; the image feature vector set and a preset image block position feature vector set are concatenated to obtain a concatenated image feature vector set; the concatenated image feature vector set is subjected to overlapping convolution processing to obtain an image spatial feature vector set; the image spatial feature vector set is input into a multi-head attention mechanism layer to obtain an image weight feature vector set; the image weight feature vector set is subjected to convolutional embedding processing to obtain an image semantic feature vector set; the text residual feature vector set and the obtained image semantic feature vector set are subjected to feature data fusion to obtain an enterprise global fusion data set.
4. The method of claim 3, wherein, the text residual feature vector set and the obtained image semantic feature vector set are subjected to feature data fusion to obtain an enterprise global fusion data set, including: text feature coding is performed on the text residual feature vector set to obtain a text feature vector set; image feature coding is performed on the image semantic feature vector set to obtain an image feature vector set; the text feature vector set and the image feature vector set are subjected to channel concatenation to obtain a concatenated feature vector; the concatenated feature vector is subjected to tensor decomposition to obtain a height feature vector, a width feature vector and a channel feature vector; the height feature vector and the width feature vector are subjected to feature multiplication to obtain a spatial attention feature vector; the spatial attention feature vector and the channel feature vector are subjected to feature multiplication to obtain a three-dimensional attention feature vector; the sum of the three-dimensional attention feature vector and the concatenated feature vector is determined as a semantic feature vector; the semantic feature vector is subjected to global average pooling processing to obtain a pooled feature vector; the pooled feature vector is input into a feature fusion network to obtain an enterprise global fusion data set.
5. The method of claim 1, wherein, the enterprise global fusion data set is encrypted and transmitted to a data warehouse for data storage, including: an heterogeneous data encryption public key is obtained; the enterprise global fusion data set is subjected to data type classification to obtain a first heterogeneous fusion data set, a second heterogeneous fusion data set and a third heterogeneous fusion data set; the first heterogeneous fusion data set is subjected to empirical mode decomposition to obtain an enterprise modal fusion data sequence; the enterprise modal fusion data sequence is subjected to data homomorphic encryption according to the heterogeneous data encryption public key to obtain an encrypted first heterogeneous fusion data set; the second heterogeneous fusion data set is subjected to data segmentation processing to obtain a segmented second heterogeneous fusion data group set; the segmented second heterogeneous fusion data group set is subjected to frequency domain conversion to obtain a converted second heterogeneous fusion data group set; inputting the converted second isomerism fusion data set into a sequence prediction model to obtain a time series isomerism fusion data sequence; performing data mapping processing on the time series isomerism fusion data sequence to obtain an isomerism byte fusion data sequence; performing forward and backward feedback encryption on the time series isomerism fusion data sequence according to the isomerism byte fusion data sequence to obtain an encrypted second isomerism fusion data set; performing third data encryption on the third isomerism fusion data set to obtain an encrypted third isomerism fusion data set; storing the encrypted first isomerism fusion data set, the encrypted second isomerism fusion data set and the encrypted third isomerism fusion data set into a data warehouse.
6. The method of claim 5, wherein, The third data encryption on the third isomerism fusion data set to obtain an encrypted third isomerism fusion data set comprises: generating an enterprise isomerism chaotic sequence; performing nonlinear transformation on the enterprise isomerism chaotic sequence to obtain an isomerism chaotic key stream; performing analog-digital conversion on the isomerism chaotic key stream to obtain an enterprise isomerism key; performing grouping processing on the third isomerism fusion data set respectively to obtain a grouped third isomerism data set; based on the grouped third isomerism data set, the following transformation steps are performed: performing grouping processing on the enterprise isomerism key to obtain a grouped enterprise isomerism key; performing XOR encryption processing on the grouped third isomerism data set and the grouped enterprise isomerism key to obtain an encrypted third isomerism data set; performing shift processing on the enterprise isomerism key to obtain a shifted enterprise isomerism key; in response to determining that the number of times of execution of the transformation step is greater than or equal to a preset execution threshold, performing transformation processing on the encrypted third isomerism data set according to the shifted enterprise isomerism key to obtain an encrypted third isomerism fusion data set; in response to determining that the number of times of execution is less than the preset execution threshold, determining the shifted enterprise isomerism key as the enterprise isomerism key to execute the transformation step again.
7. A multi-source isomerism data fusion device, comprising: a control unit configured to control a terminal device set, and perform data preprocessing on enterprise isomerism data sets of different sources obtained respectively to obtain preprocessed enterprise isomerism data sets as local enterprise isomerism data sets; a data quality detection unit configured to perform data quality detection on each local enterprise isomerism data in the local enterprise isomerism data sets to generate local data quality values to obtain a local data quality value set; an isomerism conflict resolution unit configured to perform isomerism conflict resolution processing on the local enterprise isomerism data sets according to the local data quality value set to obtain a local isomerism conflict resolution data set; a resource scheduling unit configured to perform resource scheduling on the terminal device set in response to detecting that there is a resource-limited terminal device in the terminal device set to obtain a terminal resource scheduling device set; a model training unit configured to control the terminal resource scheduling device set, and perform model training on a corresponding initial isomerism multi-modal data fusion model set according to the local isomerism conflict resolution data set to obtain an isomerism multi-modal data fusion model set; The multi-modal federated fusion unit is configured to perform multi-modal federated fusion processing on the model parameter sets corresponding to the set of heterogeneous multi-modal data fusion models to obtain a set of federated fusion model parameters; The global multi-level fusion unit is configured to perform global multi-level fusion on the set of local heterogeneous conflict resolution data according to the set of federated fusion model parameters and the set of heterogeneous multi-modal data fusion models to obtain a set of enterprise global fusion data. The encrypted transmission unit is configured to perform encrypted transmission of the set of enterprise global fusion data to a data warehouse for data storage. 8.An electronic device, comprising: one or more processors; a memory device having stored thereon one or more programs, when the one or more programs are executed by the one or more processors, cause the one or more processors to carry out the method of any one of claims 1-6.
9. A computer readable medium having stored thereon a computer program, wherein, The computer program, when executed by a processor, carries out the method of any one of claims 1-6.