A method and system for exchanging and sharing multi-source data information resources

By distinguishing authenticity, complete reconstruction and semantic merging of business information resources, and using multiple data fusion technologies, the information island problem of multi-source heterogeneous massive data is solved, and efficient integration and optimized utilization of information resources are achieved.

CN119697250BActive Publication Date: 2025-07-18CHINESE PEOPLES LIBERATION ARMY UNIT 92942
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411821650.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-07-18
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

In the process of information resource exchange and sharing, the information island problems caused by multi-source, heterogeneous and massive data are difficult to achieve efficient use of information resources across applications, multi-dimensional and multi-level.

Method used

By obtaining business information resources, authenticity identification, completeness reconstruction and semantic merging are carried out, and data fusion and matrix decomposition are used to obtain methods such as tensor expansion, fractal learning, covariance cross algorithm, support vector machine and D-S evidence theory, and a unified semantic description specification is established to achieve standardization and efficient integration of information resources.

Benefits of technology

It realizes efficient use of information resources across applications, multi-dimensional and multi-level, breaks information silos, and promotes the optimization of information resources and maximizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119697250B_ABST
    Figure CN119697250B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for exchanging and sharing multi-source data information resources. The method includes obtaining business information resources, and further includes the following steps: integrating the business information resources according to business requirements to obtain shared information resources; and pushing the shared information resources according to a preset pushing strategy. The present invention proposes a method and system for exchanging and sharing multi-source data information resources, which can effectively improve the utilization rate of information resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data information processing, and particularly to a method and system for exchanging and sharing multi-source data information resources. Background Art

[0002] The exchange and sharing of information resources refer to the exchange and common use of information and information products between information systems at different levels and in different departments. It is to share this kind of resource, which is becoming increasingly important in the Internet era, with others in order to more reasonably achieve resource allocation and save costs. Specifically, the exchange and sharing of information resources is a process that enables information resources to be exchanged and shared between information systems at different levels and in different departments based on information system technology and transmission technology, which can improve the utilization rate of information resources and avoid duplicate waste in information collection, storage, and management.

[0003] In the process of information resource exchange and sharing, information resources usually have characteristics such as multi-source, heterogeneous, and massive, which are likely to cause the problem of "information islands" and are not conducive to the efficient utilization of information resources across applications, multi-dimensions, and multi-levels. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention proposes a method and system for exchanging and sharing multi-source data information resources, which can effectively improve the utilization rate of information resources.

[0005] The first object of the present invention is to provide a method for exchanging and sharing multi-source data information resources, including obtaining business information resources, and further including the following steps:

[0006] Integrating and integrating the business information resources according to business requirements to obtain shared information resources;

[0007] Pushing the shared information resources according to a preset pushing strategy.

[0008] Preferably, the business information resources are data generated by various terminal devices.

[0009] In any of the above solutions, preferably, the integrating and integrating of the business information resources includes the following sub-steps:

[0010] Identifying the authenticity of the business information resources;

[0011] Reconstructing the completeness of the business information resources;

[0012] Merging and reorganizing the information resources according to the semantics of the information resources.

[0013] In any of the above solutions, preferably, the identifying the authenticity of the business information resources includes:

[0014] Perform tensor expansion on the business information resource to obtain an expansion matrix;

[0015] Perform eigenvalue decomposition on the expansion matrix to obtain eigenvalues;

[0016] Conduct correlation analysis based on the eigenvalues to detect the authenticity of the data;

[0017] Perform tensor filling on the business information resource to obtain the restored business information resource.

[0018] Preferably, in any of the above solutions, the first completeness reconstruction of the business information resource includes the following sub-steps:

[0019] Calculate the fractal dimension of the business information resource;

[0020] Obtain a fractal learning model, and extract the potential relationships between business information resources from the fractal dimension through the fractal learning model;

[0021] Perform initial clustering based on the fractal dimension to generate clustering members;

[0022] Use the voting method to re-cluster the clustering members to generate a clustering result;

[0023] Fuse the business information resources according to the potential relationships between the business information resources and / or the clustering result to obtain the business information resource after the first completeness reconstruction.

[0024] Preferably, in any of the above solutions, the fractal dimension is an estimate of the freedom degree of the dataset distribution. For a dataset with an embedding dimension of n, embed it into an n-dimensional lattice with a cell side length of γ (γ ∈ (γ1, γ2)), where (γ1, γ2) is the change range of the edge length measurement with fractal characteristics of the dataset. Calculate the number of data points falling into the i-th cell, denoted as The fractal dimension D q Is defined as follows:

[0025]

[0026] where q is a parameter affecting the calculation of the fractal dimension, and D2 is the probability that the distance between two randomly selected points is less than an eigenvalue.

[0027] Preferably, in any of the above solutions, the second completeness reconstruction of the business information resource after the first completeness reconstruction further includes the following sub-steps:

[0028] Construct a Gaussian mixture model of the business information resource;

[0029] The Gaussian mixture model is estimated using the covariance intersection algorithm to obtain the business information resources after the second completeness reconstruction.

[0030] Preferably, in any of the above solutions, the observation model is set as: X = Hx* + n, where X is the observed value, x* is the true value, H is the degradation matrix, and n is the noise;

[0031] Given that the fusion sources (x1, x2, ···, x N ) are N random vectors, the covariance intersection algorithm is used for estimation:

[0032]

[0033] Among them, And

[0034] Preferably, in any of the above solutions, the weight ω i is solved by minimizing tr(P 00 );

[0035]

[0036] Among them, p ii is the matrix involved in solving the weight ω i .

[0037] Preferably, in any of the above solutions, for the business information resources after the second completeness reconstruction, a third completeness reconstruction is performed, which further includes the following sub-steps:

[0038] The support vector machine is used to process the business information resources to obtain the output result of the support vector machine;

[0039] The basic probability assignment function in the D-S evidence theory is used to fuse the business information resources according to the output result to obtain the business information resources after the third completeness reconstruction.

[0040] Preferably, in any of the above solutions, for the business information resources after the third completeness reconstruction, a fourth completeness reconstruction is performed, which further includes the following sub-steps:

[0041] A relationship matrix, a constraint matrix, and an objective function are constructed according to the business information resources;

[0042] The relationship matrix is decomposed, and the process of matrix decomposition is regularized by the constraint matrix, and the matrix decomposition is iterated k times. When k = k max , the iteration stops, and the objective function is minimized to obtain the decomposition matrix, where k max is the maximum number of iterations;

[0043] Fuse the service information resources according to the decomposition matrix to obtain the service information resources after final completeness reconstruction.

[0044] Preferably, in any of the above solutions, for different types of data sources e and f with dimension n e , n f , the relationship between them is represented by a sparse matrix R ef ; for the same type of data source e with dimension n e , the relationship between the data sources is represented by a constraint matrix Θ e . Decompose the relationship matrix R ef through the triple penalty matrix decomposition method, and perform regularization constraints on the decomposition process through the constraint matrix, that is, obtain Construct an objective function. The optimization objective function of this matrix decomposition is the sum of the relationship matrix reconstruction error and the constraint matrix reconstruction trace:

[0045]

[0046] where G is a certain potential variable matrix, G e is a certain potential feature matrix, S ef is a certain weight or relationship matrix between data sources e and f during the conversion process, is the transpose matrix of a certain potential feature, t is the number of iterations, maxt e is the objective function, tr is the trace of the matrix, G T is the conversion matrix, Θ (t) is the relationship matrix after t iterations.

[0047] Preferably, in any of the above solutions, merging and reorganizing the information resources according to the semantics of the information resources includes the following sub-steps:

[0048] Define a unified semantic description specification, and establish a global information resource directory according to the semantic description specification;

[0049] Merge and reorganize the information resources according to the global information resource directory.

[0050] The second object of the present invention is to provide a multi-source data information resource exchange and sharing system, including an information resource pre-exchange module and an information resource integration module,

[0051] The information resource pre-exchange module is used to obtain service information resources.

[0052] The information resource pre-exchange module is located between the information resource integration module and the service application system,

[0053] The information resource integration module is used to fuse and integrate the service information resources according to service requirements to obtain shared information resources.

[0054] The system adopts the method described in the first objective to realize the exchange and sharing of multi-source data information resources.

[0055] Preferably, the information resource pre-exchange module 401 includes: an adapter management unit, an exchange control unit, a data extraction unit, a data push unit, and a data exchange security control unit.

[0056] In any of the above solutions, preferably, the information resource integration module includes a task monitoring unit, a data processing component management unit, an integration task scheduling unit, a data integration and processing unit, an integration task scheduling unit, an integration processing unit, an inter-station data push unit, and an inter-station transmission security control unit.

[0057] In any of the above solutions, preferably, the multi-source data information resource exchange and sharing system further includes: a platform management module 403, a data management module 404, an application integration support module 405, an operation and maintenance monitoring module 406, a data sharing service module 407, and a public information service module 408.

[0058] In any of the above solutions, preferably, the platform management module is used to integrally manage the information resource directory system, the system operation mode, the access rules, and the system configuration.

[0059] In any of the above solutions, preferably, the data management module is used to centrally store and manage the shared information resources, including: a distributed file storage engine unit, a distributed database engine unit, a data organization and management unit, a data storage security control unit, a data disaster recovery and backup unit, a public basic database, a public business database, and a special service database.

[0060] In any of the above solutions, preferably, the application integration support module is used to provide basic function call support.

[0061] In any of the above solutions, preferably, the operation and maintenance monitoring module is used to monitor the operation status, record the usage logs, and manage the computing resources, storage resources, network resources, and application resources, including: a platform resource management unit, a platform operation status monitoring unit, and a data usage log management unit.

[0062] In any of the above solutions, preferably, the data sharing service module is used to provide information resource sharing services in the form of application access control, including: a principal directory service unit, a data query service unit, a file service unit, and an application access control service unit.

[0063] Preferably, in any of the above solutions, the public information service module is used to provide information resource sharing services in the form of thematic applications, including: a public service portal unit, a thematic management unit, a thematic information organization unit, and a thematic information display unit.

[0064] The present invention proposes a method and system for exchanging and sharing multi-source data information resources, which can integrate and integrate the service information resources according to business requirements, so as to standardize and normalize and analyze and process data from different sources, with different specifications, and of different qualities, integrate fragmented information resources and explore the knowledge contained in the information resources, obtain integrated and shared information resources, realize the efficient utilization of information resources across applications, multi-dimensions, and multi-levels, and promote the optimization of information resources and the maximization of resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 FIG. is a flowchart of a preferred embodiment of the multi-source data information resource exchange and sharing method according to the present invention.

[0066] Figure 2 FIG. is a schematic flowchart of an embodiment for discriminating the authenticity of information resources in the multi-source data information resource exchange and sharing method according to the present invention.

[0067] Figure 3 FIG. is a schematic flowchart of an embodiment for information resource fusion based on support vector machine and D-S evidence theory in the multi-source data information resource exchange and sharing method according to the present invention.

[0068] Figure 4 FIG. is a schematic diagram of a preferred embodiment of the multi-source data information resource exchange and sharing system according to the present invention.

[0069] Figure 5 FIG. is a schematic structural diagram of an embodiment of a computer device in the multi-source data information resource exchange and sharing method according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0071] Embodiment 1

[0072] This embodiment provides an information resource exchange and sharing method that can be executed by a server / terminal device. As Figure 1 shown, in step 100, service information resources are obtained.

[0073] Business information resources are usually data generated by business application systems. Business application systems can be various terminal devices, such as mobile phones, computers, monitoring devices, etc. The diversification of business application systems often leads to the diversification of the data formats generated by business application systems. It can be seen that business information resources are usually a kind of heterogeneous data.

[0074] The acquisition of business information resources is realized by the information resource pre-exchange module. The information resource pre-exchange module is deployed in the network environment where the information source is located and is responsible for docking with business application systems. The information resource pre-exchange module orderly integrates various information resources from various data sources, completes the extraction of the original information resources, and after data exchange, cleaning, and integration, pushes them to the information resource integration module.

[0075] Step 200: Integrate and fuse the business information resources according to business requirements to obtain shared information resources.

[0076] Fusion integration is to standardize and normalize data from different sources, with different specifications and different qualities, and perform analysis and processing, to integrate fragmented information resources and explore the knowledge contained in information resources, so as to obtain highly complete shared information resources and realize the efficient utilization of information resources. For example, customer information may be scattered in different business units. In order to obtain an overall portrait of the customer, it is necessary to integrate relevant data into a platform for fusion integration.

[0077] In the application scenarios of cross-organizational collaboration and auxiliary decision-making, due to the fact that decision-making models, business models, and processes are to a large extent quite flexible and comprehensive, "single-domain" information services are far from meeting the needs. It is necessary to base on the public basic database and public business database, and through data fusion, refine it into comprehensive information, so as to better provide information resource services as needed during the cross-organizational collaboration of various departments. Based on the above analysis, information resource integration can complete the overall integration and sharing of information resources according to a unified information resource catalog, exchange strategy, and integration strategy.

[0078] The fusion integration of the business information resources can be realized by the information resource integration module. The information resource integration module receives various information resources from each information resource pre-exchange module, and according to the various comprehensive theme business requirements configured dynamically, processes the data from each business application system through methods such as integration and processing, constructs the business databases required by business applications, enriches the public service database, enhances the value of the data, and realizes the transformation from data to information. The information resource integration module also pushes the shared information resources to each information resource pre-exchange module according to the information push strategy.

[0079] The integration of the business information resources includes at least one of Step 210, Step 220, and Step 240. There is a certain degree of independence among Step 210, Step 220, and Step 240. In specific implementation, one or more of these steps can be selected according to needs to achieve the integration of business information resources.

[0080] Step 210: Authenticity discrimination of the business information resources.

[0081] When integrating the business information resources, the authenticity of the data in the business information can be discriminated, which is beneficial to filling in missing data and improving the performance of data analysis.

[0082] In view of the characteristics of multi - dimensional attributes of multi - source data, tensor representation methods and high - order decomposition techniques (high - order singular value decomposition or PARAFAC decomposition) are adopted to conduct more reasonable and scientific analysis of multi - source data, including: the generalized tensor expansion method that can significantly improve the system recognition ability; the tensor filling method for estimating missing data.

[0083] As Figure 2 shown, the authenticity discrimination of the business information resources includes Step 211 to Step 214.

[0084] Step 211: Tensor expansion of the business information resources to obtain an expansion matrix. That is, by converting the business information resources into matrix form, an expansion matrix is obtained.

[0085] Step 212: Eigenvalue decomposition of the expansion matrix to obtain eigenvalues. For example, eigenvalue decomposition is performed on the γ - mode expansion matrix of the tensor, and thus the sample eigenvalues in the γ - mode are obtained.

[0086] Step 213: Correlation analysis based on the eigenvalues to detect the authenticity of the data. Utilizing the tensor structure of multi - dimensional data can significantly improve the recognition ability of the detection method, and thus low - rank recovery can be performed on data with a relatively high subspace rank. The authenticity of the data can be judged according to the correlation of the eigenvalues. For example, data with a correlation higher than the threshold is determined as real data, and the data with a correlation higher than the threshold is retained, thereby completing the authenticity discrimination of the business information resources.

[0087] Steps 211 to 213 are for generalized tensor expansion of the business information resources, adopting a tensor expansion method based on the (γ1, γ2, ···, γ n ) - mode. Utilizing the tensor structure of multi - dimensional data can significantly improve the recognition ability, which is beneficial to the authenticity discrimination of the business information resources.

[0088] Step 214: Perform tensor filling on the business information resource to obtain the restored business information resource.

[0089] In the process of data acquisition, missing data records due to region, time and industry, as well as serious noise interference, will lead to data loss, thus affecting the integrity of the data. In order to restore important data for subsequent data processing and feature extraction, the matrix filling method can be used to estimate the missing data, which is conducive to solving the problem of missing data. Moreover, in the process of matrix filling of matrix data, the intrinsic tensor structure of massive data can be used to improve the estimation performance.

[0090] Keyi directly performs tensor filling on the business information resources to complete the missing data in the business information resources. It can also first identify the authenticity of the business information resources, remove the false data and retain the true data in the business information resources, and then perform tensor filling on the business information resources after removing the false data and retaining the true data to complete the fusion and integration of business information resources.

[0091] In order to enhance the robustness to impulse noise or outliers, this embodiment adopts a tensor filling method based on the high-dimensional structure of multivariate data, and uses the tensor trace norm and lp(0< <p<1)范数来提高张量填充理论的鲁棒性。此外,为了提高算法的收敛速度和降低计算复杂度,利用松弛手段对目标函数进行合理变换,利用迭代重加权方法进行快速求解。

[0092] Step 220: Completely reconstruct the business information resources.

[0093] Since business information resources come from databases and data monitoring points of different systems and platforms, they are often defective and incomplete, and cannot directly support the construction of the underlying data resource directory system. Therefore, this embodiment mines the correlation between multi-source data to solve the problem of data defect, realize the complete reconstruction of data, and build a complete basic data resource library.

[0094] The key methods used to reconstruct the completeness of data include: using machine learning methods based on fractal theory and fractal dimension clustering (FC) algorithm to explore the potential relationship between multi-source data; using covariance crossover algorithm, support vector machine (SVM), (Dempster-Shafer, DS) evidence theory and matrix decomposition methods to fuse data to overcome data uncertainty and missing problems, thereby providing scientific and objective decision support for data integration.

[0095] The completeness reconstruction of the business information resources includes steps 221 to 225.

[0096] Step 221: Calculate the fractal dimension of the business information resources. For example, use the box-counting method, random-walk method, frequency-domain method, etc. to calculate the fractal dimension of the business information resources.

[0097] Step 222: Obtain a fractal learning model, and extract the potential relationships between business information resources from the fractal dimension through the fractal learning model.

[0098] By calculating the fractal dimension of business information resources, relevant rules are extracted and corresponding machine learning models are established, forming a fractal learning process to provide guidance for future behaviors; and the feedback information of behaviors can be used to correct and update rules or models, thereby continuously optimizing the performance of machine learning models.

[0099] Among them, the fractal dimension is an estimate of the freedom degree of the distribution of a data set, reflecting the distribution characteristics of the data in a multi-dimensional space and the ability to fill the space, and plays a very important role in fractal theory. For a data set with an embedding dimension of n, embed it into an n-dimensional grid with the side length of each cell being γ (γ ∈ (γ1, γ2)), where (γ1, γ2) is the change range of the side length measurement with fractal characteristics of the data set. Calculate the number of data points falling into the i-th cell, denoted as Fractal dimension D q is defined as follows:

[0100]

[0101] Among them, different values of q can be used to calculate different fractal dimension values, and these fractal dimension values describe the characteristics of the data set from different angles. When q = 0, the calculated D0 is the Hausdorff fractal dimension; when q approaches 1, the calculated D1 is the information dimension; when q = 2, the calculated D2 is the correlation dimension. D2 characterizes the probability that the distance between two randomly selected points is less than a characteristic value, reflecting the distribution characteristics of the data.

[0102] Steps 221 and 222 are steps of a machine learning method based on fractal theory. Applying the fractal idea to the research of machine learning methods, using the self-similarity of the data resource management system as the basic basis for reasoning and learning, quantitatively describing the system self-similarity with the fractal dimension, and proposing a learning method that uses the self-similarity of the whole and the part to understand things, which can extract hidden information from high-dimensional and massive data.

[0103] Step 223: Perform initial clustering based on the fractal dimension to generate clustering members. For example, use a clustering algorithm based on the fractal dimension to perform initial clustering and generate clustering members. The clustering algorithm can be the K-means clustering algorithm, a network-based clustering algorithm, or a model-based method.

[0104] Step 224: Re-cluster the clustering members using the voting method to generate a clustering result. Re-"clustering" the clustering members using the voting method or the like can maximize the mutual information of the existing clustering results and obtain a result that is superior to that of a single clustering algorithm.

[0105] Step 225: Fuse the business information resources according to the potential relationships between the business information resources and / or the clustering result to obtain the business information resources after completeness reconstruction. During the fusion process, integrate the business information resources with potential relationships and / or belonging to the same clustering category to complete the completeness reconstruction of the business information resources.

[0106] To overcome the influence of poor clustering and clustering inconsistency on subsequent data fusion steps, a selective clustering algorithm based on the fractal dimension can be used, and incremental clustering can be used to discover clusters of any shape. The mutual information is used to calculate weights and select high-quality clusters to achieve clustering fusion.

[0107] Step 223, Step 224, and Step 225 are steps based on the fractal dimension clustering fusion algorithm. By applying the fractal method to the clustering analysis of multi-source data and statistical data, the robustness of data fusion can be enhanced. Specifically, a clustering algorithm based on the fractal dimension is used to perform clustering according to the self-similarity difference between data rather than the distance, which can be unrestricted by the shape of any data cluster and can handle the situation where the internal density of the data set is uneven. At the same time, as data points are added, the algorithm can dynamically describe the characteristics of the data set, is suitable for data stream clustering, can effectively overcome the problems of uneven internal density and non-adjacency of clusters, and the real-time monitoring data has the typical characteristics of data streams. Data stream clustering can filter out random noise in the data and increase the accuracy and reliability of the data.

[0108] A data stream clustering method for the non-stationary characteristics of monitoring data, combined with the idea of density-based data clustering, realizes the real-time processing of monitoring data streams, improves the data stream clustering efficiency and clustering accuracy, and is suitable for clustering data of any distribution shape.

[0109] The completeness reconstruction of the business information resources further includes: Step 226 and Step 227.

[0110] Step 226: Construct a Gaussian mixture model for business information resources. Using a Gaussian mixture model to represent the state of business information resources is conducive to better describing the state of business information resources. Specifically, through a sampling-based Gaussian mixture model solving algorithm, the mathematical closed solution of the Gaussian mixture model is solved, and an approximation algorithm is used to accelerate the solving speed and ensure the approximation accuracy, so as to obtain a fusion method of multiple Gaussian mixture models in a probability framework.

[0111] Step 227: Use the covariance intersection algorithm to estimate the Gaussian mixture model to obtain the business information resources after completeness reconstruction. The method of covariance intersection fusion (CI fusion) can realize the fusion of business information resources without calculating the cross-covariance.

[0112] Steps 226 and 227 are the steps of the fusion algorithm based on covariance intersection. For the uncertainty problem of massive and multi-source business information resources, the estimation theory can be used to realize data fusion, that is, regarding fusion as an estimation problem. Assume the observation model is:

[0113] X = Hx* + n

[0114] where X is the observed value, x* is the true value, H is the degradation matrix, and n is the noise.

[0115] Given that the fusion sources (x1, x2, ···, x N ) are N random vectors, in order to obtain a reliable fusion result x0, it is proposed to use the covariance intersection algorithm for estimation:

[0116]

[0117] where, and

[0118] The covariance intersection algorithm gives a consistent estimate in the case where P ij (i≠j) is unknown. Therefore, the lowest accuracy of the fusion estimate is not lower than the lowest accuracy of any fusion source estimate. The weight ω i in the estimation formula is controllable and can be solved by minimizing tr(P 00 ):

[0119]

[0120] The covariance intersection algorithm can well fuse the business information resources to obtain the business information resources after completeness reconstruction.

[0121] As Figure 3 shown, the completeness reconstruction of the business information resources further includes steps 228 and 229.

[0122] Step 228: Process the business information resources using a support vector machine (SVM) to obtain the SVM output result. The SVM method is based on the VC dimension theory and the principle of structural risk minimization in statistical learning theory. It seeks the best compromise between the complexity of the model and the learning ability according to limited sample information, in order to obtain the best output result.

[0123] Step 229: Use the basic probability assignment function in D-S evidence theory to fuse the business information resources according to the output result, and obtain the business information resources after completeness reconstruction.

[0124] This embodiment provides a soft decision output method. Since the output (f1, f2, ···, f n ) of the SVM is a relative distance, which contains the relative support degree information for the category, a probability output model is established based on the lower limit of the SVM accuracy, and the basic probability assignment (BPA) function in D-S evidence theory is used as the fusion output data.

[0125] Steps 228 and 229 are steps of the fusion method based on the support vector machine and D-S evidence theory. For some business data and statistical information with the characteristics of small samples, non-linearity, and multi-dimensions, using SVM and D-S evidence theory to fuse this type of data can improve the accuracy of data fusion.

[0126] The completeness reconstruction of the business information resources further includes steps 230 to 232. For various types of heterogeneous data sets, the relationships between the data are complex and unknown, which brings great difficulties to data fusion. Based on this, aiming at the problems of multiple target objects, multi-level relationships, multi-type data, and the lack of mutual relationships of state-type data involved in multi-source data fusion, a matrix decomposition method is used to fuse multiple heterogeneous data sets.

[0127] Step 231: Construct a relationship matrix, a constraint matrix, and an objective function according to the business information resources. For example, for different types of data sources i and j with dimension n i , n j , the relationship between them is represented by a sparse matrix R ij ; for data sources i of the same type with dimension n i , the relationship between the data sources is represented by a constraint matrix Θ i . The relationship matrix R ij is decomposed by the triple penalty matrix decomposition method, and the decomposition process is regularized by the constraint matrix, that is, Construct an objective function. The optimization objective function of this matrix decomposition is the sum of the relationship matrix reconstruction error and the constraint matrix reconstruction trace:

[0128]

[0129] Among them, R ij is a sparse matrix, Θ i is a constraint matrix.

[0130] Step 232: Perform matrix decomposition on the relationship matrix, and perform regularization constraints on the process of matrix decomposition through the constraint matrix. Iterate the matrix decomposition multiple times to reduce the objective function, and obtain a decomposed matrix. The minimization of the objective function value or making it less than a preset threshold is achieved through iterative calculation, that is, alternately updating matrices G and S until convergence, and obtaining the final decomposed matrix.

[0131] Step 233: Integrate the service information resources according to the decomposed matrix, and obtain the service information resources after completeness reconstruction. The data integration method based on matrix decomposition can well integrate multi-source, heterogeneous, and complex state data. The constrained matrix decomposition algorithm is used to achieve the integration of multiple heterogeneous data sources in the middle layer (between the data layer and the decision layer). By decomposing the relationship matrix composed of multiple relationships and multiple types of objects, and constructing an objective function related to minimizing the approximate error and relationship constraints, the decomposition form of the multi-source data matrix is calculated, and the relationship between service information resources is detected from the factors of matrix decomposition, and then the service information resources are integrated.

[0132] Steps 230 to 232 are the steps of the heterogeneous data fusion method based on matrix decomposition, which uniformly represents multiple heterogeneous data sets through a single model, converts them into a structured prediction problem, and uses an inference algorithm to solve it.

[0133] Aiming at the requirements of real-time performance of multi-source data fusion and the mixture of dynamic and static data, an approximate fast algorithm for matrix decomposition in the data flow scenario is established. The combination of pre-decomposition of the static data matrix and approximate fast decomposition of the dynamic data is beneficial to solving the problem of real-time calculation in data processing.

[0134] Step 240: Merge and reorganize the information resources according to the semantics of the information resources. By performing semantic analysis on the information resources, it is possible to achieve the orderly integration and organic fusion of multi-source heterogeneous data according to the semantics of the information resources, and it is also possible to form data services required for semantic-consistent common basic data services, common business data services, and other business services. Specifically, the semantic-based information exchange and sharing method includes steps 241 and 242.

[0135] Step 241: Define a unified semantic description specification and establish a global information resource directory based on the semantic description specification. To address the problem of structural matching among multi-source heterogeneous data, a unified semantic description specification is defined, based on which the data structure for exchange and sharing is defined, and a global information resource directory is established.

[0136] Step 242: Merge and reorganize the information resources according to the global information resource directory. The original information resources from different application systems are merged and reorganized according to the unified semantics defined in the global information resource directory, realizing the organic integration of multi-source heterogeneous information resources at the semantic level, and then realizing information exchange and sharing among different networks.

[0137] This embodiment aims at service quality business data analysis and data decision-making, ensuring that during the process of data collection, processing, sharing, and exchange, the business meaning of the data is not distorted, confused, or duplicated. On this basis, the basic capabilities of data services are formed, providing effective data support for data analysis.

[0138] Execute Step 300: Push the shared information resources according to the preset push strategy.

[0139] After the shared information resources are generated, they need to be pushed to users to assist users in making decisions. Among them, the preset push strategy can be a push strategy set by users according to their needs, which can realize pushing the corresponding shared information resources according to user requirements.

[0140] Pushing the shared information resources is implemented by the information resource pre-exchange module. The information resource pre-exchange module receives the shared information resources pushed by the information resource integration system and pushes them to the demand business unit.

[0141] In the information resource exchange and sharing method of this embodiment, the business information resources can be integrated according to business requirements, so as to standardize and normalize the data from different sources, with different specifications, and of different qualities, and perform analysis and processing, integrating the fragmented information resources and exploring the knowledge contained in the information resources, obtaining the integrated and shared information resources, realizing the efficient utilization of information resources across applications, multi-dimensions, and multi-levels, and promoting the optimization of information resources and the maximization of resource utilization.

[0142] Step 300 is to push the shared information resources into the corresponding business application systems to complete the sharing and exchange of information resources. The preset push strategy can be set according to the needs and levels of departments, so as to provide the corresponding shared information resources for departments, which is conducive to realizing the efficient utilization of information resources.

[0143] This embodiment comprehensively applies methods for storing, transmitting, fusing, and analyzing multiple data to achieve resource sharing of multi-source, heterogeneous, and massive quality data. Based on the business meanings of unified specifications, it explores the knowledge contained in information resources, and conducts real-time collection, organic integration, and fusion processing on various information resources such as business information, inherent data, and status data, breaks various regional and systematic barriers, and establishes a complete information resource exchange database to address the common needs of basic data resources in different industries, departments, and regions, avoid the "information island" problem caused by the differences in information identifiers, realize the efficient utilization of information resources across applications, multi-dimensions, and multi-levels, promote the optimization of information resources and the maximization of resource utilization, and support relevant decisions in information management.

[0144] Embodiment Two

[0145] As Figure 4 shown, this embodiment provides an information resource exchange and sharing system 40, including: an information resource pre-exchange module 401 and an information resource integration module 402.

[0146] The information resource pre-exchange module 401 is used to obtain business information resources.

[0147] The information resource pre-exchange module 401 is located between the information resource integration module 402 and the business application system, and is used to realize the information resource exchange between the information resource integration module 402 and the business application system, including obtaining business information resources from the business application system. Among them, the information resource pre-exchange module 401 isolates the information resource integration module 402 from the business application system, ensuring the independence of the application unit's business information database and the business application system.

[0148] The information resource pre-exchange module 401 includes: an adapter management unit, an exchange control unit, a data extraction unit, a data push unit, a data exchange security control unit, etc.

[0149] The information resource integration module 402 is used to fuse and integrate the business information resources according to business requirements to obtain shared information resources. The information resource integration module 402 is used to complete the overall integration and sharing of information resources according to a unified information resource catalog, exchange strategy, and integration strategy. The information resource integration module 402 includes: a task monitoring unit, a data processing component management unit, an integration task scheduling unit, a data integration and processing unit, an integration task scheduling unit, an integration processing unit, an inter-station data push unit, an inter-station transmission security control unit, etc.

[0150] The information resource pre-exchange module 401 is also used to push the shared information resources according to a preset push strategy. The information resource pre-exchange module 401 can push the shared information into the business application system to realize the sharing and exchange of information resources.

[0151] In this embodiment, the information resource exchange system can achieve information exchange and sharing between different networks finally through the information resource pre-exchange module 401 and the information resource integration module 402.

[0152] In one of the implementation manners, the information resource exchange and sharing system 40 further includes: a platform management module 403, a data management module 404, an application integration support module 405, an operation and maintenance monitoring module 406, a data sharing service module 407, and a public information service module 408.

[0153] The platform management module 403 is used for integrally managing the information resource directory system, the system operation mode, the access rules, and the system configuration.

[0154] The platform management module 403 integrally manages the information resource directory system, the operation mode, the access rules, and the system configuration of the entire data platform, ensuring the efficient, stable, safe, and orderly operation of the quality data information resource exchange prototype system, and giving full play to the supporting role of the data center for the business system. The platform management module 403 includes: a main body directory management unit, a main body metadata management unit, a business application management unit, a sharing service management unit, an exchange strategy management unit, an integration strategy management unit, a system configuration management unit, etc.

[0155] The data management module 404 is used for centrally storing and managing the shared information resources.

[0156] The information resource exchange system will be the hub for the transmission, exchange, sharing, and storage of public information resources, quickly processing a large number of structured and unstructured data with high concurrency requirements. To meet the data exchange and sharing requirements, the data management module 404 in the information resource exchange system adopts distributed data storage methods such as HADOOP, and can use the parallel computing capabilities of data distributed cluster technologies (MapReduce, YARN, etc.), the high-speed processing capabilities of real-time in-memory database technologies (SPARK, etc.), and the data storage capabilities of distributed file storage technologies (HDFS, etc.) to achieve centralized storage management of all quality resource perception data and business information, and realize high-performance exchange, high-speed computing and processing, mass storage, and reliable services under high concurrency of information resources.

[0157] The data management module 404 includes: a distributed file storage engine unit, a distributed database engine unit, a data organization and management unit, a data storage security control unit, a data disaster recovery and backup unit, a public basic database, a public business database, a special topic service database, etc.

[0158] The Application Integration Support Module 405 is used to provide support for basic function calls. The Application Integration Support Module 405 provides support for basic function calls for various applications in the Comprehensive Service Platform and the Decision Support Platform. According to the Service-Oriented Architecture concept, common functions in environmental protection applications such as organizational structure management, user management, unified authentication, and permission management are integrated into the support platform in the form of common components or public services, and relevant API interfaces are provided to reduce the coupling degree between systems, facilitating system expansion and deployment. The Application Integration Support Module 405 includes: Organizational Structure Management Unit, Unit User Management Unit, Unified Identity Authentication Unit, Unified Permission Management Unit, etc.

[0159] The Operation and Maintenance Monitoring Module 406 is used to monitor the running status, record usage logs, and manage computing resources, storage resources, network resources, and application resources. The Operation and Maintenance Monitoring Module 406 monitors the operation of the system. By deeply mining the working status of the system, it uniformly manages the computing resources, storage resources, network resources, and application resources of the data platform, comprehensively and real-time monitors the running status of the platform, records the platform data usage logs, and realizes the global, proactive, centralized, and transparent whole-process management of the data platform. The Operation and Maintenance Monitoring Module 406 faces application requirements such as real-time data monitoring, data processing, data analysis, and comprehensive display, supports the whole-process management of task planning, task preparation, process monitoring, effect display, and retrospective analysis, and provides functions such as real-time monitoring, processing, integration, management, and analysis of various types of data; supports unified monitoring of multiple regions, cross-platforms, and various devices, ensures real-time monitoring of data without affecting tasks, collects important parameters, and quickly stores, analyzes, and processes them. The Operation and Maintenance Monitoring Module 406 has characteristics such as stability, security, real-time performance, and portability.

[0160] The Operation and Maintenance Monitoring Module 406 includes: Platform Resource Management Unit, Platform Running Status Monitoring Unit, Data Usage Log Management Unit, etc.

[0161] The Data Sharing Service Module 407 is used to provide information resource sharing services in the form of application access control.

[0162] The Data Sharing Service Module 407 ensures the legitimacy of users (business application systems) using the data platform through application access control. Specifically, it provides a unified information resource sharing (access) service to legitimate application systems through methods such as subscription push and comprehensive data services. The Data Sharing Service Module 407 provides various information resource sharing services including directory services, data services, file services, etc., realizes business collaboration and public data opening, and provides data fusion application services for users.

[0163] The Data Sharing Service Module 407 includes: Subject Directory Service Unit, Data Query Service Unit, File Service Unit, Application Access Control Service Unit, etc.

[0164] The public information service module 408 is used to provide information resource sharing services in the form of topic applications. The public information service module 408 provides integrated services and data perspective services covering information resources in all business fields in the form of topic applications, such as providing public information services such as public service portals, topic management, topic information organization, and topic information display.

[0165] The public information service module 408 includes: a public service portal unit, a topic management unit, a topic information organization unit, a topic information display unit, etc.

[0166] The information resource exchange and sharing system 40 in this embodiment is composed of a platform management module 403, an information resource pre-exchange module 401, an information resource integration module 402, a data management module 404, a data sharing service module 407, a public information service module 408, an application integration support module 405, and an operation and maintenance monitoring module 406. Among them, the application integration support module 405 and the operation and maintenance monitoring module 406 are basic support modules that provide the integration, secure operation, and unified management capabilities of the entire quality resource application module; the information resource pre-exchange module 401 and the information resource integration module 402 are key background modules for completing information resource integration. The information resource pre-exchange module 401 is deployed in the network environment where the information source is located, responsible for docking with business application modules, extracting original information resources, and pushing shared information resources to business modules. The information resource integration module 402 is responsible for the integration and distribution of business information and perception data, and completes the integration and processing of data of multiple information entities across business fields. The data sharing service module 407 is responsible for providing resource services for various applications of quality resources; the public information service module 408 provides integrated services covering information resources in all business fields.

[0167] Embodiment III

[0168] This application embodiment also provides a computer device, such as Figure 5 shown in the basic structural block diagram of the computer device in this embodiment.

[0169] The computer device 50 includes a memory 501, a processor 502, and a network interface 503 that are communicatively connected to each other via a system bus. It should be noted that only the computer device 50 with components 501-503 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of this technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0170] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.

[0171] The memory 501 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 501 can be an internal storage unit of the computer device 5, such as the hard disk or memory of the computer device 50. In other embodiments, the memory 501 can also be an external storage device of the computer device 50, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 6. Of course, the memory 501 can also include both the internal storage unit of the computer device 50 and its external storage device. In this embodiment, the memory 501 is generally used to store the operating system and various application software installed on the computer device 5, such as computer-readable instructions of the information resource exchange and sharing method. In addition, the memory 501 can also be used to temporarily store various data that have been output or will be output.

[0172] In some embodiments, the processor 502 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 502 is generally used to control the overall operation of the computer device 50. In this embodiment, the processor 502 is used to run the computer-readable instructions stored in the memory 501 or process data, such as running the computer-readable instructions of the information resource exchange and sharing method.

[0173] The network interface 503 may include a wireless network interface or a wired network interface. The network interface 503 is generally used to establish a communication connection between the computer device 50 and other electronic devices.

[0174] Embodiment 4

[0175] This application also provides a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the information resource exchange and sharing method as described above.

[0176] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of this application.

[0177] To better understand the present invention, the above has been described in detail in conjunction with specific embodiments of the present invention, but it is not a limitation of the present invention. Any simple modification made to the above embodiments based on the technical essence of the present invention still belongs to the scope of the technical solution of the present invention. Each embodiment in this specification focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference can be made to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

Claims

1. A method for exchanging and sharing multi-source data information resources, including obtaining business information resources, characterized in that, It further includes the following steps: Integrate and integrate the business information resources according to business requirements to obtain shared information resources. The integration and integration of the business information resources include the following sub-steps: Identify the authenticity of the business information resources, including: Perform tensor expansion on the business information resources to obtain an expansion matrix; Perform eigenvalue decomposition on the expansion matrix to obtain eigenvalues; Perform correlation analysis based on the eigenvalues to detect the authenticity of the data; Perform tensor filling on the business information resources to obtain the restored business information resources; Perform completeness reconstruction on the business information resources. The first completeness reconstruction of the business information resources includes the following sub-steps: Calculate the fractal dimension of the business information resources; Obtain a fractal learning model and extract the potential relationships between business information resources from the fractal dimension through the fractal learning model; Perform initial clustering based on the fractal dimension to generate clustering members; Use the voting method to cluster the clustering members again to generate a clustering result; Fuse the business information resources according to the potential relationships between the business information resources and / or the clustering result to obtain the business information resources after the first completeness reconstruction; Merge and reorganize the information resources according to the semantics of the information resources; Push the shared information resources according to the preset push strategy.

2. The multi-source data information resource exchange and sharing method according to claim 1, characterized in that, The fractal dimension is an estimate of the degree of freedom of the distribution of a data set. For a data set with an embedding dimension of n, it is embedded in an n-dimensional lattice with each cell side length of γ (γ ∈ (γ1, γ2)), where (γ1, γ2) is the range of variation of the edge length measurement for which the data set has fractal characteristics. Calculate the number of data points falling into the i-th cell, denoted as the fractal dimension D q is defined as follows: Where q is a parameter affecting the calculation of the fractal bit number, and D2 is the probability that the distance between two randomly selected points is less than an eigenvalue.

3. The multi-source data information resource exchange and sharing method according to claim 2, characterized in that, The completeness reconstruction of the business information resources after the first completeness reconstruction further includes the following sub-steps: Construct a Gaussian mixture model of the business information resources; Estimate the Gaussian mixture model using the covariance intersection algorithm to obtain the business information resources after the second completeness reconstruction.

4. The multi-source data information resource exchange and sharing method according to claim 3, wherein Set the observation model as: X = Hx* + n, where X is the observed value, x* is the true value, H is the degradation matrix, and n is the noise; Given that the fusion sources (x1, x2, ···, x N ) are N random vectors, the covariance intersection algorithm is used for estimation: Among them, and 5. The multi-source data information resource exchange and sharing method according to claim 4, wherein, Weight ω i Solve by minimizing tr(P 00 ) Among them, p ii is the matrix involved when calculating the weight ω i .

6. The multi-source data information resource exchange and sharing method according to claim 5, wherein The completeness reconstruction of the business information resources after the second completeness reconstruction further includes the following sub-steps: Process the business information resources using a support vector machine to obtain the output result of the support vector machine; Fuse the business information resources according to the output result using the basic probability assignment function in the D-S evidence theory to obtain the business information resources after the third completeness reconstruction.

7. A multi-source data information resource exchange and sharing system, including an information resource pre-exchange module, characterized in that, It further includes an information resource integration module, The information resource pre-exchange module is used to obtain business information resources, The information resource pre-exchange module is located between the information resource integration module and the business application system, The information resource integration module is used to integrate and integrate the business information resources according to business requirements to obtain shared information resources, The system realizes the exchange and sharing of multi-source data information resources by using the method described in claim 1.

Citation Information

Patent Citations

  • Method of constructing big data platform of government affairs

    CN106855962A