Method and device for acquiring data resources, equipment, medium and program product
By acquiring and integrating data resources through a data integration server, and constructing a virtual data ontology and metadata mapping table, the problems of high barriers to entry and single mode of data products in the data network system are solved, and the flexible integration and supply-demand matching of ad-hoc data products are realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE COMM LTD RES INST
- Filing Date
- 2025-12-15
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional data network systems suffer from problems such as high barriers to entry for data products, limited usage models, and an immature ecosystem of platform service providers, leading to difficulties in connecting supply and demand.
The data integration server acquires primary data resources and their metadata related to the target business, integrates them, constructs a virtual data ontology and a primary metadata mapping table, and transforms them into secondary data resources to meet business needs.
Significantly lowers the barrier to entry for data products, enables flexible integration of multi-source heterogeneous data, solves the difficulties in connecting supply and demand sides, and provides ad-hoc data products that meet business needs.
Smart Images

Figure CN121901159A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a method, apparatus, device, medium, and program product for acquiring data resources. Background Technology
[0002] The Data Service Shared Network (DSSN) system is a next-generation trusted data circulation infrastructure provided by domestic telecommunications operators. It achieves secure data circulation that is usable but invisible through a single-point access and network-wide reach. The DSSN system includes data providers (Data Service Nodes, DSNs), data requesters (Data Request Nodes, DRNs), platform service providers (Data Service Platforms), and capability providers.
[0003] The traditional data product distribution scheme based on the data network system is as follows: the data provider builds the data product; the data provider publishes the data product on the platform service provider; the data demander subscribes to the data product on the platform service provider; and the data demander uses the successfully subscribed data product. However, the traditional data product distribution scheme based on the data network system suffers from problems such as excessively high barriers to entry for data products, limited usage models, and difficulties in connecting supply and demand sides due to the immature ecosystem of the platform service provider. Summary of the Invention
[0004] This application provides a method, apparatus, equipment, medium, and program product for acquiring data resources, which addresses the problems of high barriers to entry for data products, limited usage modes, and difficulties in connecting supply and demand sides caused by the immature ecosystem of platform service providers in traditional data product circulation schemes based on data network systems.
[0005] Firstly, this application provides a method for acquiring data resources, applied to a data integration server, the method comprising: Acquire primary data resources related to the target business. The primary data resources include metadata corresponding to the data ontology. The data ontology is used to describe the characteristics of multiple data objects in multiple dimensions, and the metadata is used to describe the attributes of the data ontology. Based on the metadata, the primary data resources are fused to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology. The virtual data ontology is used to describe the characteristics of the fused data objects in multiple dimensions, and the first metadata mapping table is used to describe the attributes of the virtual data ontology. Based on the first metadata mapping table and the virtual data ontology, secondary data resources for implementing the target business are obtained.
[0006] Secondly, this application provides an apparatus for acquiring data resources, applied to a data integration server, the apparatus comprising: The first acquisition module is used to acquire primary data resources related to the target business. The primary data resources include metadata corresponding to the data ontology. The data ontology is used to describe the characteristics of multiple data objects in multiple dimensions, and the metadata is used to describe the attributes of the data ontology. The fusion module is used to fuse the first-level data resources according to the metadata to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology. The virtual data ontology is used to describe the characteristics of the fused data objects in multiple dimensions, and the first metadata mapping table is used to describe the attributes of the virtual data ontology. The second acquisition module is used to acquire secondary data resources for implementing the target business based on the first metadata mapping table and the virtual data ontology.
[0007] Thirdly, this application provides an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the method for acquiring data resources described in the first aspect.
[0008] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method for acquiring data resources described in the first aspect.
[0009] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method for acquiring data resources described in the first aspect.
[0010] This application acquires primary data resources and their metadata related to the target business through a data integration server, merges the primary data resources based on the metadata to construct a virtual data ontology and a first metadata mapping table, and obtains secondary data resources oriented towards the target business accordingly. The method of this application has at least the following technical effects: First, addressing the issue of excessively high barriers to entry for data products in existing technologies, this application utilizes an integration server to transform primary data resources (basic data products) into secondary data resources (ad hoc data products). This transforms the originally unstandardized primary data products into pre-packaged ad hoc data products, allowing data requesters to use them directly without investing significant resources in secondary processing. Thus, it significantly lowers the barriers to entry for data products.
[0011] Second, in response to the problem of the single usage mode of data products in the existing technology, this application can logically aggregate and integrate the primary data resources provided by multiple different data providers in multiple dimensions by constructing a virtual data ontology and a first metadata mapping table. This breaks the limitation of the existing technology that only supports privacy alignment of two-party data, and can realize the flexible integration and multiple applications of multi-source heterogeneous data, effectively solving the problem of the single usage mode of data products in the existing technology.
[0012] Third, in response to the difficulty of connecting supply and demand in existing technologies, this application introduces a data integration server as an intermediary. Based on a clear target business, it selects and reconstructs primary data resources (basic data products), which can transform primary data resources that are detached from the application scenario into secondary data resources (ad hoc data products) that meet business needs, effectively solving the problem of difficulty in connecting supply and demand. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a diagram illustrating the architecture of a trusted data circulation infrastructure for the Internet of Things, as shown in an embodiment of this application. Figure 2 This is a flowchart illustrating a method for acquiring data resources according to an embodiment of this application; Figure 3 This is a schematic diagram illustrating a data resource fusion process according to an embodiment of this application; Figure 4 This is a structural block diagram of an apparatus for acquiring data resources, as shown in an embodiment of this application; Figure 5 This is a schematic diagram of the physical structure of an electronic device as shown in an embodiment of this application. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0016] Figure 1 This is a diagram illustrating the architecture of a trusted data circulation infrastructure for the Internet of Things, as shown in an embodiment of this application. Combined with... Figure 1 The traditional data product circulation scheme based on the Internet of Things system is as follows: (I) Building a Data Product. Step 1: The data provider logs in to the DSN; Step 2: The data provider selects the dataset to be shared, specifies the access address of the dataset, and sets the name, version, category, upload method, etc.; Step 3: The data provider prepares metadata for the dataset, such as a data dictionary and instructions for use; Step 4: The data provider completes a legal compliance statement and a source statement for the dataset; Step 5: The data provider packages the above information to generate a locally unpublished data product.
[0017] (II) Publishing Data Products. Step 1: The data provider logs into the DSN; Step 2: The data provider publishes the data product to the DSP on the DSN and sets the status of the data product to "Published"; Step 3: The platform administrator logs into the DSP, reviews the data product, and if the review is approved, sets the status of the data product to "Listed"; Step 4: The data requester can find the listed data product on the DSP.
[0018] (III) Ordering Data Products. Step 1: The data requester logs into the DSP; Step 2: The data requester filters data products through the data product catalog on the DSP; Step 3: After selecting a data product, the data requester submits a purchase application to the DSP. The platform service provider will review the data requester's qualifications to ensure that they meet the usage conditions of the data product and have the corresponding data security and compliance protection capabilities. The review may include the legality of the entity's identity, compliance of the usage scenario, and data security protection measures; Step 4: After the platform service provider approves the application, the supply and demand parties enter the transaction negotiation stage to negotiate details such as the scope of use, term, price, delivery method, after-sales support, privacy and security protection requirements of the data product. After reaching an agreement, the two parties will sign a formal contract based on the data transaction contract template; Step 5: After the contract takes effect, the data delivery stage begins. The data provider will deliver the data product to the data requester in accordance with the method agreed in the contract (e.g., API interface, data package download, sandbox environment access, etc.). During the delivery process, both parties will follow the pre-agreed data security standards (e.g., de-identification, encrypted transmission, etc.) to ensure data security. After the data product is delivered, the data requester will make settlement and payment in accordance with the contract.
[0019] (iv) Using data products. For example, one usage method is as follows: Step 1: The data requester logs into the DRN; Step 2: The data requester selects data, including local data and subscribed data; Step 3: The privacy alignment algorithm is selected to perform privacy alignment on the local data and subscribed data to obtain the intersection data; Step 4: The data cleaning algorithm is selected to perform data cleaning operations such as eliminating outliers, deduplication, and filling gaps on the intersection data; Step 5: The feature selection algorithm is selected to calculate the correlation of feature values on the cleaned data and provide a feature value importance assessment report; Step 6: The federated learning algorithm is selected to train the above-governed data to obtain an inference model.
[0020] However, traditional data product distribution solutions based on data network systems have the following problems: First, the barriers to using data products are too high. The data products provided by data providers are rudimentary, lacking standardization and ready-to-use features. Data requesters cannot use the data products directly after purchase and still need to invest a lot of time and resources in secondary processing.
[0021] Second, the usage model of data products is limited. Data requesters only support the use of local data and subscribed data after privacy alignment (privacy alignment refers to the secure fusion of data from different sources without exposing the original plaintext data). In practice, data requesters often need to fuse data from multiple data providers. For example, a data requester may not have its own data but needs to combine multiple data products through privacy alignment to form a logically private dataset for model training.
[0022] Third, the immature ecosystem of platform service providers makes it difficult to connect supply and demand. Specifically, platform service providers currently only offer basic data products from data providers, but these data providers do not understand the application needs of data requesters, making it difficult to accurately match supply and demand.
[0023] To address the aforementioned issues, this application provides a method for acquiring data resources, applied to a data integration server. The data resources in this application are the data products mentioned earlier. Figure 2 This is a flowchart illustrating a method for acquiring data resources according to an embodiment of this application. (Refer to...) Figure 2 The method for obtaining data resources in this application includes the following steps: Step S101: Obtain primary data resources related to the target business. The primary data resources include metadata corresponding to the data ontology. The data ontology is used to describe the characteristics of multiple data objects in multiple dimensions, and the metadata is used to describe the attributes of the data ontology.
[0024] The method provided in this application can be applied to trusted data circulation infrastructure in the Internet of Things. In this implementation scenario, the device used by the data integrator is a data integration server, the device used by the data provider is a data source server (DSN), the device used by the data requester is a data request server (DRN), and the device used by the platform manager is a platform server (DSP).
[0025] In this embodiment, the data integrator can pre-calculate the business needs of different data requesters and design various typical business processes based on the statistical results. The target business can be any one of these typical business processes. For example, the target business could be a model for predicting disease risk (e.g., diabetes risk, heart disease risk), a model for predicting weather, or a model for predicting price trends, etc. This embodiment does not specifically limit the type of business process.
[0026] In this embodiment, the primary data resource is the initial data product published by the data provider on the platform server. The primary data resource does not contain the data subject, but rather includes metadata, a legality and compliance statement, and a source statement corresponding to the data subject. Each primary data resource corresponds to a data ontology.
[0027] The data body can be a table with multiple rows and columns. Each row represents a data object, and each column represents a feature of a dimension. Dimensions can be set according to actual needs, such as name, age, weight, height, address, and phone number.
[0028] Primary data resources related to the target business refer to the initial data products that may be used when executing the target business. For example, if the target business is to train a model to predict the risk of diabetes, primary data resources could be physical examination data products released by hospitals that include users' blood glucose test values and past medical history, or lifestyle behavior data products released by health management platforms that include users' dietary preferences and daily steps. Metadata is used to explain the attributes of the data subject. The metadata in this application mainly includes: business metadata, management metadata, and technical metadata.
[0029] Business metadata includes a data dictionary and usage instructions. The data dictionary should at least describe the field name, type, length, value range, business meaning, and privacy level of each column in the data body. Technical metadata should at least describe the access address, data format, update frequency, and upload method of the data ontology. Management metadata should at least describe the name, version, category, owner information, and security level of the data ontology. Of course, the attributes of the data body can be set according to actual needs, and this application does not impose any restrictions on this.
[0030] The access address of the data subject can only be used by the data source server, and external devices (including platform servers) cannot read the plaintext information of the data subject through this access address.
[0031] Step S102: Based on the metadata, the primary data resources are fused to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology. The virtual data ontology is used to describe the characteristics of the fused data objects in multiple dimensions, and the first metadata mapping table is used to describe the attributes of the virtual data ontology.
[0032] In this embodiment, the number of primary data resources acquired in step S101 can be one or more. The data integration server can merge the various primary data resources according to a pre-designed fusion method to obtain a virtual data ontology. Furthermore, the data integration server can also obtain a first metadata mapping table corresponding to the virtual data ontology based on the metadata corresponding to each primary data resource.
[0033] Similarly, the virtual data ontology is a multi-row, multi-column table. Each row represents a data object, and each column represents a feature of a dimension. A cell in the table represents a feature of a data object in one dimension, and the value of that feature (i.e., the value of that cell) is a non-plaintext value, meaning that the data integration server cannot determine the true value of a data object's feature in one dimension.
[0034] In this embodiment, the first metadata mapping table is obtained by organizing the metadata corresponding to each primary data resource. The first metadata mapping table can be used to record the metadata corresponding to each cell (i.e., the features of each data object in each dimension) in the virtual data ontology within its respective primary data resource. Therefore, based on the first metadata mapping table, the data integration server can determine the metadata corresponding to the features of each data object in each dimension within the virtual data ontology within its respective primary data resource.
[0035] Step S103: Obtain secondary data resources for implementing the target business based on the first metadata mapping table and the virtual data ontology.
[0036] In this embodiment, the data integration server can process the virtual data ontology according to the first metadata mapping table to improve its quality. Then, based on the improved virtual data ontology, secondary data resources are obtained. These secondary data resources, compared to primary data resources, are better able to achieve the target business objectives.
[0037] In this embodiment, the secondary data resource is an ad-hoc data product, which is a data product that can be used directly through a standardized interface without secondary processing, such as being directly used for model training.
[0038] This application acquires primary data resources and their metadata related to the target business through a data integration server, merges the primary data resources based on the metadata to construct a virtual data ontology and a first metadata mapping table, and obtains secondary data resources oriented towards the target business accordingly. The method of this application has at least the following technical effects: First, addressing the issue of excessively high barriers to entry for data products in existing technologies, this application utilizes an integration server to transform primary data resources (basic data products) into secondary data resources (ad hoc data products). This transforms the originally unstandardized primary data products into pre-packaged ad hoc data products, allowing data requesters to use them directly without investing significant resources in secondary processing. Thus, it significantly lowers the barriers to entry for data products.
[0039] Second, in response to the problem of the single usage mode of data products in the existing technology, this application can logically aggregate and integrate the primary data resources provided by multiple different data providers in multiple dimensions by constructing a virtual data ontology and a first metadata mapping table. This breaks the limitation of the existing technology that only supports privacy alignment of two-party data, and can realize the flexible integration and multiple applications of multi-source heterogeneous data, effectively solving the problem of the single usage mode of data products in the existing technology.
[0040] Third, in response to the difficulty of connecting supply and demand in existing technologies, this application introduces a data integration server as an intermediary. Based on a clear target business, it selects and reconstructs primary data resources (basic data products), which can transform primary data resources that are detached from the application scenario into secondary data resources (ad hoc data products) that meet business needs, effectively solving the problem of difficulty in connecting supply and demand.
[0041] In conjunction with the above embodiments, in one implementation, step S102 may include: Step S1021: If the data ontology corresponding to the primary data resource includes a data ontology that is allowed to leave the domain, then, based on the metadata corresponding to the data ontology that is allowed to leave the domain, the data ontology that is allowed to leave the domain is merged in the sandbox environment to obtain the first fusion result.
[0042] In this embodiment, "out-of-domain" refers to the transmission of data from the local security trust domain to the external environment. For example, if the data ontology corresponding to a primary data resource X is allowed to leave the domain, it means that the data ontology corresponding to the primary data resource X can be transmitted from its data source server to the external environment; if the data ontology corresponding to the primary data resource X is prohibited from leaving the domain, it means that the transmission of the data ontology corresponding to the primary data resource X from its data source server to the external environment is prohibited.
[0043] Since there can be multiple primary data resources, and each primary data resource corresponds to a data ontology, there can be multiple data ontologies corresponding to a primary data resource. Among all the data ontologies, those that are allowed to leave the domain can first be loaded into a secure sandbox environment by the data integration server. Then, based on the metadata corresponding to these data ontologies allowed to leave the domain, the server merges them to obtain the first fusion result. The first fusion result is essentially a multi-row, multi-column table, where each row represents a data object and each column represents a feature of one dimension.
[0044] In one implementation, the sandbox environment can be a sandbox environment based on Trusted Execution Environment (TEE) technology.
[0045] In this embodiment, the sandbox environment can be constructed using hardware-level isolated Trusted Execution Environment (TEE) technology. This involves establishing an independent, secure area on the computing platform, protected by both hardware and software, as the sandbox environment. This ensures that the confidentiality and integrity of data loaded into the sandbox are independent of the main operating system. In this way, when data is allowed to leave the domain and is transferred to the sandbox environment, computation only occurs within the protected memory of the TEE. The platform server or system administrator cannot access the plaintext, thus effectively mitigating the risk of data leakage after leaving the domain.
[0046] Step S1022: If the data ontology corresponding to the primary data resource includes a data ontology that is prohibited from leaving the domain, the data ontology that is prohibited from leaving the domain is fused according to the identification information of the data objects in the encrypted data ontology that is prohibited from leaving the domain, to obtain a second fusion result; or, the data ontology that is prohibited from leaving the domain is fused according to the metadata corresponding to the data ontology that is prohibited from leaving the domain, to obtain a second fusion result.
[0047] In this embodiment, for those data ontologies that are prohibited from leaving the domain, the data integration server can obtain the second fusion result using either of the following two methods: Method 1: Obtain the identification information of data objects in the encrypted prohibited data ontology, and fuse the prohibited data ontology according to the identification information of the encrypted data objects to obtain the second fusion result.
[0048] Method 2: Based on the metadata corresponding to the prohibited data ontology, merge the prohibited data ontology to obtain the second fusion result.
[0049] The second fusion result is essentially a multi-row, multi-column table, where each row represents a data object and each column represents a feature of a dimension.
[0050] Step S1023: Obtain the virtual data ontology based on the first fusion result and / or the second fusion result.
[0051] In this embodiment, if all data ontologies are allowed to leave the domain, then fusing the primary data resources will only yield a first fusion result, which is then used as the virtual data ontology. If all data ontologies are prohibited from leaving the domain, then fusing the primary data resources will only yield a second fusion result, which is then used as the virtual data ontology. If all data ontologies contain both allowed and prohibited data ontologies, then both the first and second fusion results can be obtained simultaneously. In this case, the first and second fusion results can be fused again to obtain the virtual data ontology.
[0052] In this embodiment, sandboxed physical aggregation is used for data ontology that is allowed to leave the domain, fully leveraging the efficiency of centralized secure computing. Encrypted logical aggregation is used for data ontology that is prohibited from leaving the domain, strictly adhering to the rule of data not leaving the domain, thus ensuring absolute privacy and security of the data ontology. This hybrid aggregation architecture breaks down the barriers to data flow between different security levels, enabling the fusion of data resources from different sources with zero privacy leakage, providing a guarantee for the acquisition of subsequent ad-hoc data products.
[0053] In conjunction with the above embodiments, in one implementation, step S1022 may include: Send encryption policies to the data source servers where each prohibited data ontology resides. The encryption policies are used by the data source servers to encrypt the identification information of data objects in the prohibited data ontology. Receive the encrypted identifier information of data objects in the data ontology that is prohibited from leaving the domain, sent by the data source server; Based on the identification information of data objects in the encrypted prohibited outbound data ontology, the prohibited outbound data ontology is fused to obtain a second fusion result.
[0054] In this embodiment, one of the multiple dimensions is used to represent the identification information of the data object. For example, if the data object is a user, and the multiple dimensions include ID card number, gender, and home address, the ID card number can be used as the identification information of the data object. Different data objects have different identification information.
[0055] In this embodiment, the data integration server can send a request to the data source server containing each prohibited outbound data ontology to read the identification information of the data objects. This request carries an encryption policy. Upon receiving the request, the data source server encrypts the identification information of the data objects in the prohibited outbound data ontology according to the encryption policy, obtaining encrypted identification information of the data objects, such as a hash value. Then, the data source server sends the encrypted identification information of the data objects to the data integration server. Finally, the data integration server merges the prohibited outbound data ontologs based on the encrypted identification information of the data objects, obtaining a second fusion result.
[0056] In this embodiment, by implementing an encryption strategy using the data source server, it is ensured that the original plaintext data and its identification information never leave the local security trust domain; only the encrypted identification information is transmitted, which greatly improves data privacy and security. Furthermore, this embodiment uses the encrypted identification information to fuse data resources, avoiding the transmission of redundant data unrelated to the target business and significantly improving the efficiency of multi-source heterogeneous data aggregation and processing.
[0057] In conjunction with the above embodiments, in one implementation, for step S1022, if the second fusion result is obtained using method one, then the specific method is as follows: Based on the identification information of data objects in the encrypted prohibited outbound data ontology, the prohibited outbound data ontology is fused to obtain a second fusion result, including: Based on the identification information of data objects in the encrypted prohibited outbound data ontology, a feature alignment operation is performed on the prohibited outbound data ontology to obtain a second fusion result. The feature alignment operation is used to expand the number of feature dimensions.
[0058] In this embodiment, the data integration server controls each data source server to encrypt the identification information of data objects according to the same encryption strategy. Therefore, the encrypted identification information corresponding to the same data object in different data subjects is also the same. By identifying the type of the identification information of all encrypted data objects, the data integration server can determine how many types of data objects there are.
[0059] For example, the data integration server receives encrypted identification information XXX1, XXX2, and XXX3 from data source server 1 (corresponding to data subject 1), encrypted identification information XXX2, XXX3, and XXX4 from data source server 2 (corresponding to data subject 2), and encrypted identification information XXX2, XXX3, and XXX5 from data source server 3 (corresponding to data subject 3). The data integration server can then determine that there are a total of 5 different data objects. Next, the data integration server identifies the common data objects in data subjects 1 through 3, namely, data object 2 corresponding to XXX2 and data object 3 corresponding to XXX3. Finally, the data integration server concatenates the feature dimensions of data objects 2 and 3 in different data subjects to expand the feature dimensions. For example, if the dimensions of data object 2 in data subject 1 are ID number, name, and gender; in data subject 2, ID number and age; and in data subject 3, ID number, name, and home address, then after concatenating the feature dimensions, the second fusion result will have the feature dimensions of data object 2 as ID number, name, gender, age, and home address. The second fusion result contains only data object 2 and data object 3.
[0060] In this embodiment, when obtaining the second fusion result using Method 1, if there is only one data subject to be fused, then that data subject is directly output as the second fusion result. If the data integration server determines that there are no common data objects among the multiple data subjects, it can output a default data subject as the second fusion result according to pre-set rules, or apply to obtain the second fusion result using Method 2. The default data subject can be a data subject that meets the first preset characteristics, such as having the most feature dimensions or the most data objects, etc. The second preset characteristics can be set according to actual needs.
[0061] Accordingly, if method two is used to obtain the second fusion result, the specific method is as follows: Based on the metadata corresponding to the prohibited outbound data ontology, the prohibited outbound data ontology is fused to obtain a second fusion result, including: Based on the metadata corresponding to the prohibited data ontology, a sample alignment operation is performed on the prohibited data ontology to obtain the second fusion result. The feature alignment operation is used to expand the number of data objects.
[0062] In this embodiment, the data integration server can determine the meaning of each column in each data ontology based on the metadata of the data ontology, that is, it can determine what the dimensions of multiple features of data objects in the data body are. During sample alignment, the data integration server first determines whether there are any data bodies among all data bodies where the dimensions of the features of their data objects are completely identical. For example, if the data integration server determines that multiple features of data objects in data body 1 and data body 2 are completely identical (i.e., ID card number, name, and home address), then the data integration server can directly logically merge the data bodies 1 and 2 based on the number of data objects without obtaining the identification information of the data objects in data body 1 and data body 2, obtaining a second fusion result. The obtained second fusion result includes all data objects in data body 1 and data body 2.
[0063] In this embodiment, when obtaining the second fusion result using method two, if there is only one data subject to be fused, then that data subject is directly output as the second fusion result. If the data integration server determines that there are no data subjects among all data subjects whose data objects have completely identical feature dimensions, it can output a default data subject as the second fusion result according to pre-set rules, or apply to obtain the second fusion result using method one. The default data subject can be a data subject that meets a second preset characteristic, such as having the most feature dimensions or the most data objects, etc., and the second preset characteristic can be set according to actual needs.
[0064] In this embodiment, the data integration server can choose to use either method one or method two to obtain the second fusion result according to actual needs. This embodiment does not impose any specific restrictions on this.
[0065] In this embodiment, firstly, vertical fusion is achieved using feature alignment operations, which associates different attributes of the same data object scattered across different data sources without exposing privacy, effectively expanding the number of feature dimensions. Secondly, horizontal fusion is achieved using sample alignment operations, which aggregates similar data records scattered across different data sources, effectively expanding the number of data objects. In this way, this application overcomes the problem of the single usage mode of traditional data products, enabling the final ad-hoc data product to simultaneously meet diverse business needs such as multi-dimensional deep analysis and large-scale model training.
[0066] In one embodiment, in step S1021, when fusing data ontologs that are allowed to go out of domain, feature alignment or sample alignment can be selected according to actual needs. If there is only one data subject to be fused, that data subject can be directly output as the first fusion result. When using sample alignment, if the data integration server determines that there are no data subjects among all data subjects whose features are completely identical in dimension, it can output a default data subject as the first fusion result according to a pre-set rule, or apply to use feature alignment to obtain the first fusion result. When using feature alignment, if the data integration server determines that there are no common data objects among multiple data subjects, it can output a default data subject as the first fusion result according to a pre-set rule, or apply to use sample alignment to obtain the first fusion result.
[0067] In conjunction with the above embodiments, in one implementation, step S1023 may include: If the primary data resources include data ontology that is allowed to leave the domain and data ontology that is prohibited from leaving the domain, determine the common data objects in the first fusion result and the second fusion result; The data corresponding to the common data objects in the first fusion result and the data corresponding to the common data objects in the second fusion result are fused to obtain the virtual data ontology.
[0068] In this embodiment, if the first fusion result and the second fusion result include common data objects, then the data corresponding to the common data objects in the first fusion result and the data corresponding to the common data objects in the second fusion result are logically encapsulated to obtain a virtual data ontology. Here, logical encapsulation refers to establishing a unified application programming interface (API).
[0069] For example, if the first fusion result includes 1000 records (one row in the virtual data ontology represents one record), and the second fusion result includes 2000 records, then the data integration server can establish a unified application programming interface (API) (named RESTful API) for both the first and second fusion results. When the data request server calls this API, it appears to the outside world as a table with 3000 records.
[0070] In this embodiment, under the premise of following different data out-of-domain policies, the data allowed to leave the domain stored in the sandbox environment and the data prohibited from leaving the domain stored locally are accurately associated through intersection matching. This can break the isolation between the physical storage environment and the security boundary, not only to achieve complementary feature dimensions across security domains, but also to construct a logically unified view (virtual data ontology) that takes into account both privacy compliance and data richness, thereby maximizing the utilization value of multi-source mixed data.
[0071] In conjunction with the above embodiments, in one implementation, step S1023 may include: The first fusion result and the second fusion result are fused to obtain the virtual data ontology.
[0072] In this embodiment, if the first fusion result and the second fusion result do not include common data objects, then the first fusion result and the second fusion result can be logically encapsulated to obtain a virtual data ontology.
[0073] In this embodiment, the process of fusing primary data resources can be referred to Figure 3 As shown. Figure 3 This is a schematic diagram illustrating a data resource fusion process in an embodiment of this application.
[0074] In this embodiment, the fusion method of the first fusion result and the second fusion result can be set according to actual needs. In this way, the problem of the single usage mode of traditional data products can be overcome, enabling ad-hoc data products to simultaneously meet diverse business needs such as multi-dimensional in-depth analysis and large-scale model training, thereby maximizing the value of data resources.
[0075] In conjunction with the above embodiments, in one implementation, in the virtual data ontology, the feature values of the data object in each dimension are non-plaintext values, including the access method in the data source server to which it belongs and the encrypted feature values.
[0076] In this embodiment, the data integration server can send data acquisition requests to the data source servers involved in the virtual data ontology, and populate the virtual data ontology based on the results returned by the data source servers. After population, the feature values of the data objects in each dimension in the virtual data ontology are non-plaintext values. In this embodiment, the non-plaintext values can be the access methods in the respective data source servers, or they can be encrypted feature values.
[0077] For example, a virtual data ontology can be represented as follows: HASH-001 [Ptr->XXX1] [Ptr->XXX3] …… [Ptr->XXYX1] HASH-002 [Ptr->XXX2] [Ptr->XXX4] …… [Ptr->XXYX2] HASH-003 [Ptr->XXX3] [Ptr->XXX5] …… [Ptr->XXYX3] In the table above, the columns containing HASH-001, HASH-002, and HASH-003 represent ID card numbers, and HASH-001, HASH-002, and HASH-003 are all encrypted values of the ID card numbers. Except for the first column, the values in the other columns represent access methods, which can only be recognized by the data source server. Of course, the table above is only an example and does not represent an actual virtual data subject.
[0078] In this embodiment, data privacy and security are ensured by replacing plaintext with access methods or encrypted feature values. Secondly, establishing a mapping relationship between virtual space and physical storage allows the data request server to accurately locate and retrieve data resources or initiate computing tasks from the data source server based on the access method. This significantly improves the execution efficiency of the target business.
[0079] In conjunction with the above embodiments, in one implementation, step S103 may include: Step S1031: Optimize the virtual data ontology according to the first metadata mapping table to obtain the optimized virtual data ontology.
[0080] In this embodiment, the virtual data ontology can be optimized according to a pre-designed optimization strategy, so that the optimized data ontology has higher data quality.
[0081] Step S1032: Obtain the second metadata mapping table corresponding to the optimized virtual data ontology. The second metadata mapping table is used to describe the attributes of the optimized virtual data ontology.
[0082] In this embodiment, a second metadata mapping table that can accurately describe the attributes of the optimized virtual data ontology can be obtained by combining the optimized virtual data ontology and the first metadata mapping table.
[0083] Step S1033: Obtain the secondary data resources used to implement the target business according to the second metadata mapping table.
[0084] In this embodiment, the data integration server can encapsulate ad-hoc data products (secondary data resources) based on the second metadata mapping table. Subsequently, the data request server can directly use the ad-hoc data products to execute target business operations.
[0085] In this embodiment, the data integration server can significantly improve the quality of the final secondary data resources by optimizing the virtual data ontology.
[0086] In conjunction with the above embodiments, in one implementation, step S1031 may include: Based on the first metadata mapping table, perform at least one of the following operations on the virtual data ontology: data cleaning, data transformation, and data optimization. Data cleaning includes at least one of duplicate value handling and consistency checking; data transformation includes at least one of format conversion and feature encoding; and data optimization includes at least one of data compression, data query acceleration, and privacy protection.
[0087] In this embodiment, data cleaning refers to correcting errors and inconsistencies in the virtual data ontology while adhering to privacy constraints, thereby ensuring the quality of the virtual data ontology. Since the virtual data ontology stores access methods or encrypted feature values pointing to the data source server, the data integration server cannot directly read the plaintext for data cleaning. Therefore, a strategy-based approach involving local computation and ciphertext comparison can be adopted for data cleaning.
[0088] For handling duplicate values, the data integration server can employ encrypted comparison algorithms (such as hash-based anonymization matching techniques). Assuming the virtual data ontology includes data objects from data source server A and data source server B, to determine if they contain the same data objects, the data integration server sends a unified hash strategy to both data source servers A and B. Data source servers A and B, according to the unified hash strategy, perform irreversible hashing of the data object's identification information (e.g., ID card number) locally and return digests. The data integration server compares the returned digests; if identical digests are found, it determines that there are identical data objects and duplicate records. Then, the data integration server performs a merging operation on the records corresponding to the identical data objects, achieving deduplication without exposing the original ID card numbers.
[0089] For consistency checks (i.e., checking whether business rules are followed), the data integration server can utilize a rule engine. For example, given that blood type values should be one of A, B, AB, or O, and cannot be C or X, the rule engine first calculates the hash value set for A, B, AB, or O, for example, a set of hash values named Set={Hash(A), Hash(B), Hash(AB), Hash(O)}. Next, the data integration server obtains the ciphertext Hash(X) of the blood type column from each data source server in the virtual data ontology. The rule engine then determines whether Hash(X) exists in the Set. If a Hash(X) is not in the set, it is determined that the value of the feature corresponding to that Hash(X) does not conform to the business rules.
[0090] In this embodiment, data transformation refers to adjusting the values of heterogeneous and non-standardized features in the virtual data ontology to a format suitable for subsequent analysis or model training, while adhering to data security and privacy protection. Since the virtual data ontology stores access methods or encrypted feature values pointing to various data source servers, the data integration server can implement secure transformation through mechanisms such as rule distribution, remote execution, and view updates.
[0091] For format conversion, features of the same dimension (e.g., blood glucose concentration) from different data source servers in the virtual data ontology may be expressed in different units of measurement, such as mmol / L and mg / dL. The data integration server can use template-based conversion tools to generate standardized rules (e.g., setting conversion formulas to uniformly convert mg / dL to mmol / L) and send these rules as calculation instructions to each data source server. Each data source server performs numerical conversion operations locally and establishes a standardized temporary view. Subsequently, the data integration server updates the corresponding access method in the virtual data ontology to a pointer to this standardized temporary view, thereby achieving unit normalization of cross-source data at the logical level.
[0092] When encoding features, the data integration server can transform a feature of a certain dimension into an algorithm-understandable vector through a sandbox environment or cryptographic methods. For example, for the feature of "occupation" in the virtual data ontology (existing in the form of encrypted values), the data integration server can call the secure computing interface within the sandbox environment to perform one-hot encoding. The sandbox maps "engineer" to a vector [1, 0, 0] within a trusted memory region, and then re-encrypts each element in the vector. After completion, the data integration server updates the structure of the virtual data ontology, expanding the original "occupation" feature column into multiple feature columns to adapt it to the input requirements of models such as neural networks.
[0093] Furthermore, the data integration server can perform feature derivation operations to uncover deeper data value. For example, for a feature representing birth date in a virtual data ontology, in order to perform trend analysis without revealing the precise age, the data integration server can send feature derivation rules (such as age grouping) to the data source server. The data source server calculates feature labels such as youth and middle-aged locally, and the data integration server adds a new column representing the derived features for age grouping in the virtual data ontology, pointing to the newly generated feature labels in the data source server.
[0094] In this embodiment, data optimization refers to further improving the balance between storage efficiency, data query response speed, and data privacy risks of the virtual data ontology after data transformation, thereby building a high-performance ad-hoc data product.
[0095] For data compression, the data integration server can employ privacy-friendly compression algorithms and encrypted indexing techniques. For example, for the values of large-scale encrypted features in a virtual data ontology, the data integration server can send compression commands to the data source server, or invoke homomorphic encryption compression algorithms or privacy-friendly compression tools within the sandbox environment to compress the encrypted feature values, thereby significantly reducing storage space usage.
[0096] To accelerate data queries, such as for the educational level dimension in a virtual data ontology, in order to solve the problem of low efficiency of full table scans in encrypted state, the data integration server can instruct the data source server to build an encrypted index locally and establish an index mapping relationship in the metadata of the virtual data ontology. This allows subsequent data request servers to directly locate the target data row through the index, thereby speeding up the query.
[0097] Regarding privacy protection, for example, if a feature of a certain dimension in the virtual data ontology is a home address, the data integration server can update the access method in the virtual data ontology, redirecting the access method to anonymized views pre-generated locally by the data source server (e.g., truncated generalized data down to the street level), thereby implementing anonymization at the virtual level. Furthermore, for numerical features such as salaries, the data integration server can update the corresponding metadata in the virtual data ontology, dynamically adjusting the noise level of differential privacy (e.g., adjusting the Epsilon noise parameter of the Laplace mechanism from a high-privacy, low-efficiency 0.1 to a balanced 1.0 based on the dataset size). This ensures that ad-hoc data products generated based on this virtual data ontology, when invoked by the data request server, can output statistically significant analytical results while effectively resisting differential attacks.
[0098] In this embodiment, when optimizing the virtual data ontology, the data integration server can also record each optimization operation in detail. By persistently storing the configuration snapshots and processing logs generated by each optimization operation, this application can accurately trace the evolution path of the virtual data ontology while ensuring efficient resource utilization. This provides tamper-proof data processing credentials for regulatory agencies or third-party audits, and also supports quickly locating problems and rolling back to a historical stable version when data quality anomalies occur.
[0099] In this embodiment, the data integration server can significantly improve the quality of the final secondary data resources by optimizing the virtual data ontology, thereby improving the execution efficiency and effectiveness of the target business.
[0100] In conjunction with the above embodiments, in one implementation, step S1033 may include: Based on the second metadata mapping table, obtain the computing power distribution information, privacy compliance statement information, and value assessment information corresponding to the optimized virtual data ontology; Based on the second metadata mapping table, computing power distribution information, privacy compliance statement information, and value assessment information, secondary data resources are obtained to achieve the target business.
[0101] The second metadata mapping table includes business metadata, management metadata, and technical metadata. Business metadata describes the data dictionary and usage instructions of the optimized data ontology. Management metadata describes the name, version, category, owner information, and security level of the optimized data ontology. Technical metadata describes the access address, data format, update frequency, and upload method of the optimized data ontology.
[0102] Computing power distribution information refers to the location mapping relationship of physical computing nodes (such as data source servers or security sandboxes) corresponding to the features of each dimension in the virtual data ontology, as well as the scheduling configuration file of the available computing power of each node. This information is used to accurately distribute logical computing tasks to the corresponding physical nodes for execution when responding to application interface calls.
[0103] Privacy compliance statement information refers to: compliance qualification certificates proving that the construction and processing process of ad-hoc data products comply with relevant privacy protection standards (such as ISO 27701 certification) or legal and regulatory requirements, and is used to ensure the legal security of data circulation.
[0104] Value assessment information refers to a comprehensive report generated after quantitatively assessing secondary data resources, used to measure the quality of the data.
[0105] In this embodiment, encapsulating computing power distribution information, privacy compliance statement information, and value assessment information in secondary data resources can effectively improve the quality of secondary data resources, thereby improving the efficiency and effectiveness of the data requester in executing the target business.
[0106] In conjunction with the above embodiments, in one implementation, after acquiring secondary data resources for implementing the target business, the method of this application may further include: Convert the access permission verification policy corresponding to the secondary data resources into access permission verification code, and deploy the access permission verification code on the authorization server; The authorization server is used to run access permission verification code when it receives a call request for the application interface corresponding to the secondary data resource. The access permission verification code is used to verify the access permission of the server calling the application interface to the secondary data resource.
[0107] In this embodiment, legal regulations and business constraints can be transformed into automatically executable code and deployed on an authorization server (e.g., an API gateway) to implement access control. Specifically, privacy rule templates can be preset based on laws and regulations such as GDPR (General Data Protection Regulation) or CCPA (California Consumer Privacy Act), and complex strategies can be formally described using industry-standard terms such as ODRL (Open Digital Rights Language) or DPV (Data Privacy Glossary) to obtain access verification strategies. The data integration server, based on the aforementioned industry standards, logically encodes fine-grained data privacy usage policies involving who, when, where, and why, as well as data provider-defined extended policies (including access frequency control, usage restrictions, data retention periods, etc.), converting them into machine-readable and executable access verification code. Subsequently, the access verification code is deployed on the authorization server, which serves as a unified service entry point.
[0108] When a data request server initiates a call request to a secondary data resource through an application programming interface (RESTful API), the authorization server intercepts the request in real time and runs access permission verification code to perform comprehensive compliance checks on the caller's identity credentials, environment context, and request intent. The authorization server will only allow the request if all verification conditions are met; otherwise, it will directly block access.
[0109] In this embodiment, the authorization server performs real-time interception and comprehensive verification of API call requests before they reach the underlying data. Based on fine-grained policies such as who, when, and where, it automatically blocks unauthorized access, ensuring strict compliance and privacy security of ad-hoc data products throughout the standardized circulation process and effectively avoiding the risk of data abuse.
[0110] In conjunction with the above embodiments, in one implementation, after step S103, the method of this application further includes: Determine the target data source servers involved in the optimized virtual data ontology; Determine the value score of the data provided by each target data source server for the optimized virtual data ontology; Based on the optimized virtual data ontology, determine the number of data objects and dimensions provided by each target data source server during the execution of the target business. Based on the number of data objects provided, the number of dimensions provided, and the value score, the contribution level of each target data source server to the target business is determined, and the contribution level is used to allocate reward resources to the target data source server.
[0111] In this embodiment, after the data requester uses the secondary data resources, the data integration server, based on the optimized virtual data ontology and the secondary metadata mapping table, can determine all target data source servers constituting the secondary data resources. Next, the data integration server determines the value score of the data provided by each target data source server to the optimized virtual data ontology, and determines the number of data objects (effective sample size) and the number of dimensions (effective feature size) provided by each target data source server during the execution of the target business. Then, for each target data source server, the contribution level is calculated using the following formula: Contribution Level = (Effective Sample Size * Effective Feature Size * Value Score) / Total Effective Sample Size * Effective Feature Size * Value Score of All Data Source Servers
[0112] In this embodiment, the reward resources allocated to a single data source server = total reward resources × the contribution level of the data source server.
[0113] In this embodiment, the data integration server, operating under strict privacy protection, can invoke the algorithm tools provided by the data network system to conduct a comprehensive value assessment of the data provided by various target data source servers involved in the optimized virtual data ontology. The assessment dimensions include intrinsic quality, feature utility, feature importance, and feature-label relevance. Intrinsic quality assessment is the fundamental dimension for evaluating the quality of the data itself, typically including indicators such as accuracy, coverage, timeliness, and scarcity. Feature utility assessment utilizes historical log data to evaluate the usefulness of data in specific business scenarios. Feature importance assessment, within the framework of federated learning models, quantifies the marginal contribution of each feature to the final model's prediction performance by calculating Shapley values or based on the model's output (e.g., federated random forests); the greater the contribution, the higher the value. Feature-label relevance assessment analyzes the statistical correlation between features and the prediction target (e.g., correlation coefficients and mutual information calculated through federation); features with strong correlations generally have higher value.
[0114] After value assessment, value quantification is performed. Indicators corresponding to all the aforementioned dimensions are determined, and the assessment result for each indicator (e.g., accuracy, coverage, Shapley score) is converted into a quantified score (e.g., a score between 0 and 1). Next, weights are assigned to each indicator, with different weights assigned to different indicators depending on the data application scenario. For example, for a predictive model, feature importance may have the highest weight; while for a demographic report, coverage and accuracy may have higher weights. Weights can be set by the data network system or determined through negotiation among the transaction parties. Finally, a weighted average or other comprehensive algorithm is used to calculate the value score of the data provided by each data source server. The formula is: Value Score = (Indicator 1 * W1 + Indicator 2 * W2 + Indicator 3 * W3 + Indicator 4 * W4 + ...), where W1, W2, W3, W4, etc., are the weights of each indicator.
[0115] In the above calculation process, in order to ensure fairness and transparency in multi-party cooperation, the data source server can adopt tamper-proof technical means (such as blockchain notarization) or introduce a credible third-party institution to strictly record and audit the core parameters involved in the entire contribution assessment process (such as effective sample size, effective feature quantity, value score, etc.), and generate an assessment report that can be verified and checked by all participants at any time, thereby ensuring that the settlement basis is true and credible.
[0116] In this embodiment, by combining the effective sample size, effective feature size, and value score of the data provided by each data source server in the virtual data ontology, the contribution level of each data source server can be accurately calculated, thus ensuring the fair allocation of subsequent reward resources.
[0117] In one implementation, in conjunction with the above embodiments, the data integration server communicates with the platform server in the data network system to obtain primary data resources related to the target business, including: Among the multiple data resources displayed on the platform server, obtain the primary data resources related to the target business. These multiple data resources are published on the platform server by different data source servers.
[0118] The method described in this application can be applied to data network infrastructure scenarios, effectively addressing the challenges of supply-demand matching and product rudimentary issues within the platform service provider ecosystem. By establishing communication with the platform server, the data integration server can act as a professional ecosystem intermediary, accurately identifying and acquiring data resources matching the target business from the massive heterogeneous data resources centrally published on the platform server by different data source servers (DSNs). This approach can bridge the cognitive gap between data providers and data requesters, significantly improving the efficiency of data resource allocation and the accuracy of supply-demand matching in the data network.
[0119] In conjunction with the above embodiments, in one implementation, after acquiring the secondary data resources used to implement the target business, the method of this application further includes: Publish the secondary data resources to the platform server.
[0120] In this embodiment, publishing secondary data resources (ad hoc data products) to the platform server can solve the problems of high usage threshold, single mode, and difficulty in connecting supply and demand sides caused by the existing platform server only providing primary data products from the data source server (DSN).
[0121] In summary, the solution presented in this application has at least the following technical effects: First, it breaks through the traditional point-to-point data circulation model, forming a three-tier architecture of data provider-data integrator-data requester, thus solving the problem of high integration costs caused by differences in industry knowledge between the supply and demand sides. Data integrators can design vertical domain data products (such as medical image datasets and financial risk control feature libraries) based on industry experience, significantly improving data reuse rates.
[0122] Second, it lowers the barrier to entry for using data products. Data requesters do not need to worry about the various cumbersome algorithms and parameter settings in the data governance process. They can simply select the algorithm provided by the data product and the platform server to obtain the calculation results through privacy-preserving computation, which greatly reduces the barrier to entry for using data products.
[0123] Third, it enriches the supply model of data products, moving from the existing primary data products only available through data providers to ad-hoc data products based on primary data products provided by data integrators, thus meeting the diverse data application needs of data requesters. This application also provides an apparatus for acquiring data resources, applied to a data integration server. The apparatus for acquiring data resources provided in the embodiments of this application is described below, and the apparatus described below corresponds to the method for acquiring data resources described above.
[0124] Figure 4 This is a structural block diagram of an apparatus for acquiring data resources, as illustrated in an embodiment of this application. (Refer to...) Figure 4 The data resource acquisition device 400 provided in this application embodiment may include: The first acquisition module 401 is used to acquire primary data resources related to the target business. The primary data resources include metadata corresponding to the data ontology. The data ontology is used to describe the characteristics of multiple data objects in multiple dimensions, and the metadata is used to describe the attributes of the data ontology. The fusion module 402 is used to fuse the first-level data resources according to the metadata to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology. The virtual data ontology is used to describe the characteristics of the fused data object in multiple dimensions, and the first metadata mapping table is used to describe the attributes of the virtual data ontology. The second acquisition module 403 is used to acquire secondary data resources for implementing the target business based on the first metadata mapping table and the virtual data ontology.
[0125] According to the data resource acquisition apparatus 400 provided in this application, the fusion module 402 is configured to: if the data ontology corresponding to the primary data resource includes a data ontology that is allowed to leave the domain, fuse the data ontology that is allowed to leave the domain in a sandbox environment based on the metadata corresponding to the data ontology that is allowed to leave the domain, to obtain a first fusion result; if the data ontology corresponding to the primary data resource includes a data ontology that is prohibited from leaving the domain, fuse the data ontology that is prohibited from leaving the domain based on the identification information of the data objects in the encrypted data ontology that is prohibited from leaving the domain, to obtain a second fusion result; or, fuse the data ontology that is prohibited from leaving the domain based on the metadata corresponding to the data ontology that is prohibited from leaving the domain, to obtain a second fusion result; and obtain the virtual data ontology based on the first fusion result and / or the second fusion result.
[0126] According to the data resource acquisition device 400 provided in this application, the fusion module 402 is specifically used for: sending an encryption strategy to the data source server where each prohibited data body is located, the encryption strategy being used by the data source server to encrypt the identification information of data objects in the prohibited data body; receiving the encrypted identification information of data objects in the prohibited data body sent by the data source server; and fusing the prohibited data bodies according to the encrypted identification information of data objects in the prohibited data body to obtain the second fusion result.
[0127] According to the data resource acquisition apparatus 400 provided in this application, the fusion module 402 is specifically used for: performing a feature alignment operation on the data ontology that is prohibited from leaving the domain based on the identification information of the data objects in the encrypted prohibited data ontology, to obtain the second fusion result, wherein the feature alignment operation is used to expand the number of feature dimensions; and performing a sample alignment operation on the data ontology that is prohibited from leaving the domain based on the metadata corresponding to the prohibited data ontology, to obtain the second fusion result, wherein the feature alignment operation is used to expand the number of data objects.
[0128] According to the data resource acquisition device 400 provided in this application, the fusion module 402 is specifically used for: if the primary data resource includes a data ontology that is allowed to leave the domain and a data ontology that is prohibited from leaving the domain, determining the common data object in the first fusion result and the second fusion result; fusing the data corresponding to the common data object in the first fusion result and the data corresponding to the common data object in the second fusion result to obtain the virtual data ontology.
[0129] According to the data resource acquisition apparatus 400 provided in this application, in the virtual data ontology, the feature values of the data object in each dimension are non-plaintext values, and the non-plaintext values include the access method in the data source server to which it belongs and the encrypted feature values.
[0130] According to the data resource acquisition apparatus 400 provided in this application, the second acquisition module 403 is configured to: optimize the virtual data ontology according to the first metadata mapping table to obtain an optimized virtual data ontology; acquire a second metadata mapping table corresponding to the optimized virtual data ontology, the second metadata mapping table being used to describe the attributes of the optimized virtual data ontology; and acquire secondary data resources for implementing the target business according to the second metadata mapping table.
[0131] According to the data resource acquisition device 400 provided in this application, the second acquisition module 403 is specifically used to: perform at least one of the following operations on the virtual data ontology in sequence: data cleaning, data conversion, and data optimization, based on the first metadata mapping table; wherein, the data cleaning includes at least one of duplicate value processing and consistency checking, the data conversion includes at least one of format conversion and feature encoding, and the data optimization includes at least one of data compression, data query acceleration, and privacy protection.
[0132] According to the data resource acquisition device 400 provided in this application, the second acquisition module 403 is specifically used for: acquiring computing power distribution information, privacy compliance statement information, and value assessment information corresponding to the optimized virtual data ontology according to the second metadata mapping table; and obtaining secondary data resources for realizing the target business according to the second metadata mapping table, the computing power distribution information, the privacy compliance statement information, and the value assessment information.
[0133] According to the data resource acquisition device 400 provided in this application, the second metadata mapping table includes business metadata, management metadata, and technical metadata; the business metadata is used to describe the data dictionary and usage instructions of the optimized data ontology; the management metadata is used to describe the name, version, category, owner information, and security level of the optimized data ontology; and the technical metadata is used to describe the access address, data format, update frequency, and upload method of the optimized data ontology.
[0134] The device 400 for acquiring data resources according to this application further includes: a conversion module, configured to: convert the access permission verification policy corresponding to the secondary data resource into access permission verification code, and deploy the access permission verification code on an authorization server; wherein, the authorization server is configured to run the access permission verification code when receiving a call request for the application interface corresponding to the secondary data resource, and the access permission verification code is configured to verify the access permission of the server calling the application interface to the secondary data resource.
[0135] The data resource acquisition apparatus 400 provided in this application further includes: an evaluation module, configured to: determine the target data source servers involved in the optimized virtual data ontology; determine the value score of the data provided by each target data source server to the optimized virtual data ontology; determine, based on the optimized virtual data ontology, the number of data objects and the number of dimensions provided by each target data source server during the execution of the target business; and determine the degree of contribution of each target data source server to the target business based on the number of data objects provided, the number of dimensions provided, and the value score, wherein the degree of contribution is used to allocate reward resources to the target data source servers.
[0136] According to the data resource acquisition device 400 provided in this application, the data integration server is communicatively connected to the platform server in the data network system. The first acquisition module 401 is specifically used to: acquire first-level data resources related to the target business from multiple data resources displayed on the platform server. The multiple data resources are published on the platform server by different data source servers.
[0137] The device 400 for acquiring data resources provided in this application further includes: a publishing module, used to: publish the secondary data resources to the platform server.
[0138] Figure 5 This is a schematic diagram of the physical structure of an electronic device as illustrated in an embodiment of this application. Figure 5As shown, the electronic device may include a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call a computer program in the memory 530 to execute steps of a method for acquiring data resources, such as: acquiring primary data resources related to a target business, wherein the primary data resources include metadata corresponding to a data ontology, the data ontology describing the characteristics of multiple data objects in multiple dimensions, and the metadata describing the attributes of the data ontology; fusing the primary data resources according to the metadata to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology, the virtual data ontology describing the characteristics of the fused data objects in multiple dimensions, and the first metadata mapping table describing the attributes of the virtual data ontology; and acquiring secondary data resources for implementing the target business according to the first metadata mapping table and the virtual data ontology.
[0139] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0140] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of a method for acquiring data resources provided in the above embodiments, such as: acquiring primary data resources related to a target business, wherein the primary data resources include metadata corresponding to a data ontology, the data ontology being used to describe the characteristics of multiple data objects in multiple dimensions, and the metadata being used to describe the attributes of the data ontology; fusing the primary data resources according to the metadata to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology, wherein the virtual data ontology is used to describe the characteristics of the fused data objects in multiple dimensions, and the first metadata mapping table is used to describe the attributes of the virtual data ontology; and acquiring secondary data resources for implementing the target business according to the first metadata mapping table and the virtual data ontology.
[0141] On the other hand, embodiments of this application also provide a processor-readable storage medium storing a computer program. The computer program is used to cause a processor to execute the steps of a method for acquiring data resources provided in the above embodiments. For example: acquiring primary data resources related to a target business, wherein the primary data resources include metadata corresponding to a data ontology, the data ontology describing the characteristics of multiple data objects in multiple dimensions, and the metadata describing the attributes of the data ontology; fusing the primary data resources according to the metadata to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology, the virtual data ontology describing the characteristics of the fused data objects in multiple dimensions, and the first metadata mapping table describing the attributes of the virtual data ontology; and acquiring secondary data resources for implementing the target business according to the first metadata mapping table and the virtual data ontology.
[0142] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for acquiring data resources, characterized in that, Applied to a data integration server, the method includes: Acquire primary data resources related to the target business. The primary data resources include metadata corresponding to the data ontology. The data ontology is used to describe the characteristics of multiple data objects in multiple dimensions, and the metadata is used to describe the attributes of the data ontology. Based on the metadata, the primary data resources are fused to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology. The virtual data ontology is used to describe the characteristics of the fused data objects in multiple dimensions, and the first metadata mapping table is used to describe the attributes of the virtual data ontology. Based on the first metadata mapping table and the virtual data ontology, secondary data resources for implementing the target business are obtained.
2. The method for acquiring data resources according to claim 1, characterized in that, The step of fusing the primary data resources based on the metadata to obtain a virtual data ontology includes: If the data ontology corresponding to the first-level data resource includes a data ontology that is allowed to leave the domain, the data ontology that is allowed to leave the domain is fused in a sandbox environment based on the metadata corresponding to the data ontology that is allowed to leave the domain, to obtain a first fusion result; If the data ontology corresponding to the first-level data resource includes a data ontology that is prohibited from leaving the domain, the prohibited data ontology is fused according to the identification information of the data objects in the encrypted prohibited data ontology to obtain a second fusion result; or, the prohibited data ontology is fused according to the metadata corresponding to the prohibited data ontology to obtain a second fusion result. The virtual data ontology is obtained based on the first fusion result and / or the second fusion result.
3. The method for acquiring data resources according to claim 2, characterized in that, The step of fusing the prohibited data ontology based on the identifier information of the data objects in the encrypted prohibited data ontology to obtain a second fusion result includes: An encryption policy is sent to the data source server where each prohibited data ontology resides. The encryption policy is used by the data source server to encrypt the identification information of the data objects in the prohibited data ontology. Receive the identifier information of the data objects in the encrypted, out-of-domain prohibited data body sent by the data source server; Based on the identification information of the data objects in the encrypted prohibited outbound data ontology, the prohibited outbound data ontology is fused to obtain the second fusion result.
4. The method for acquiring data resources according to claim 2, characterized in that, The step of fusing the prohibited data ontology based on the identifier information of the data objects in the encrypted prohibited data ontology to obtain a second fusion result includes: Based on the identification information of data objects in the encrypted prohibited outbound data ontology, a feature alignment operation is performed on the prohibited outbound data ontology to obtain the second fusion result. The feature alignment operation is used to expand the number of feature dimensions. The step of fusing the prohibited outbound data ontology based on the metadata corresponding to the prohibited outbound data ontology to obtain a second fusion result includes: Based on the metadata corresponding to the prohibited outbound data ontology, a sample alignment operation is performed on the prohibited outbound data ontology to obtain the second fusion result. The feature alignment operation is used to expand the number of data objects.
5. The method for acquiring data resources according to claim 2, characterized in that, The step of obtaining the virtual data ontology based on the first fusion result and / or the second fusion result includes: If the primary data resources include data ontology that is allowed to leave the domain and data ontology that is prohibited from leaving the domain, determine the common data objects in the first fusion result and the second fusion result; The data corresponding to the common data object in the first fusion result and the data corresponding to the common data object in the second fusion result are fused to obtain the virtual data ontology.
6. The method for acquiring data resources according to claim 2, characterized in that, In the virtual data ontology, the feature values of the data object in each dimension are non-plaintext values, including the access method in the data source server and the encrypted feature values.
7. The method for acquiring data resources according to claim 1, characterized in that, The step of obtaining secondary data resources for implementing the target business based on the first metadata mapping table and the virtual data ontology includes: Based on the first metadata mapping table, the virtual data ontology is optimized to obtain the optimized virtual data ontology; Obtain the second metadata mapping table corresponding to the optimized virtual data ontology, the second metadata mapping table being used to describe the attributes of the optimized virtual data ontology; Based on the second metadata mapping table, obtain the secondary data resources used to implement the target business.
8. The method for acquiring data resources according to claim 7, characterized in that, The optimization of the virtual data ontology based on the first metadata mapping table includes: Based on the first metadata mapping table, at least one of the following operations—data cleaning, data transformation, and data optimization—is sequentially performed on the virtual data ontology. The data cleaning includes at least one of duplicate value processing and consistency checking; the data conversion includes at least one of format conversion and feature encoding; and the data optimization includes at least one of data compression, data query acceleration, and privacy protection.
9. The method for acquiring data resources according to claim 7, characterized in that, The step of obtaining secondary data resources for implementing the target business based on the second metadata mapping table includes: Based on the second metadata mapping table, obtain the computing power distribution information, privacy compliance statement information, and value assessment information corresponding to the optimized virtual data ontology; Based on the second metadata mapping table, the computing power distribution information, the privacy compliance statement information, and the value assessment information, secondary data resources for realizing the target business are obtained.
10. The method for acquiring data resources according to claim 7, characterized in that, The second metadata mapping table includes business metadata, management metadata, and technical metadata; The business metadata describes the data dictionary and usage instructions of the optimized data ontology; the management metadata describes the name, version, category, owner information, and security level of the optimized data ontology; and the technical metadata describes the access address, data format, update frequency, and upload method of the optimized data ontology.
11. The method for acquiring data resources according to claim 1, characterized in that, After acquiring the secondary data resources used to implement the target business, the method further includes: The access permission verification policy corresponding to the secondary data resources is converted into access permission verification code, and the access permission verification code is deployed on the authorization server; The authorization server is used to run the access permission verification code when it receives a call request for the application interface corresponding to the secondary data resource. The access permission verification code is used to verify the access permission of the data request server that calls the application interface to the secondary data resource.
12. The method for acquiring data resources according to claim 9, characterized in that, After acquiring the secondary data resources used to implement the target business, the method further includes: Determine the target data source server involved in the optimized virtual data ontology; Determine the value score of the data provided by each of the target data source servers for the optimized virtual data ontology; Based on the optimized virtual data ontology, determine the number of data objects and the number of dimensions provided by each of the target data source servers during the execution of the target service; Based on the number of data objects provided, the number of dimensions provided, and the value score, the contribution level of each target data source server to the target business is determined, and the contribution level is used to allocate reward resources to the target data source server.
13. The method for acquiring data resources according to any one of claims 1-12, characterized in that, The data integration server communicates with the platform server in the data network system, and the acquisition of primary data resources related to the target business includes: Among the multiple different data resources displayed on the platform server, the primary data resources related to the target business are obtained. These multiple different data resources are published on the platform server by different data source servers.
14. The method for acquiring data resources according to any one of claims 1-12, characterized in that, After acquiring the secondary data resources used to implement the target business, the method further includes: The secondary data resources are published to the platform server.
15. An apparatus for acquiring data resources, characterized in that, The device, used in a data integration server, includes: The first acquisition module is used to acquire primary data resources related to the target business. The primary data resources include metadata corresponding to the data ontology. The data ontology is used to describe the characteristics of multiple data objects in multiple dimensions, and the metadata is used to describe the attributes of the data ontology. The fusion module is used to fuse the first-level data resources according to the metadata to obtain a virtual data ontology and a first metadata mapping table corresponding to the virtual data ontology. The virtual data ontology is used to describe the characteristics of the fused data objects in multiple dimensions, and the first metadata mapping table is used to describe the attributes of the virtual data ontology. The second acquisition module is used to acquire secondary data resources for implementing the target business based on the first metadata mapping table and the virtual data ontology.
16. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for acquiring data resources as described in any one of claims 1 to 14.
17. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for acquiring data resources as described in any one of claims 1 to 14.
18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for acquiring data resources as described in any one of claims 1 to 14.