Quality infrastructure service data integration method based on cross-hierarchy multi-link
Through the cross-level multi-link data integration method, heterogeneous data, real-time processing, scalability and security issues in the integration of quality infrastructure service data are solved, and efficient data integration and secure processing are achieved.
Patent Information
- Application Number
- CN202510392191.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has difficulties in integrating heterogeneous data, insufficient real-time data processing capabilities, poor system scalability and data security problems when integrating quality infrastructure service data.
The data integration method across hierarchical multi-links is adopted. By designing a format compatible with different data sources, the data acquisition layer, data processing layer, data storage layer and data application layer are built to collect and process data in real time, and the embedding model is used to generate vector fields for data integration and secure processing.
It realizes effective integration of heterogeneous data, improves real-time data processing capabilities, enhances system scalability, and strengthens data security and privacy protection.
Smart Images

Figure CN120179722A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data integration solution, and more specifically to a method for integrating quality infrastructure service data based on cross - level multi - link. Background Art
[0002] Quality infrastructure service data usually involves various data generated during the service process. When integrating this data, compared with other types of data, there are problems such as data heterogeneity, real - time issues, data volume problems, data privacy and security issues, etc. Based on these problems, using traditional methods has the following defects: 1. It is unable to effectively integrate heterogeneous data; 2. The real - time data processing ability is insufficient, resulting in decision - making delays; 3. The system scalability is poor and it is difficult to cope with the rapid growth of data volume; 4. Data security cannot be effectively guaranteed. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a method for integrating quality infrastructure service data based on cross - level multi - link that can effectively solve the above - mentioned defects.
[0004] To achieve the above - mentioned purpose, the present invention provides the following technical solution: A method for integrating quality infrastructure service data based on cross - level multi - link, characterized by including the following steps: Step 1: Input a format that can be compatible with different data sources, and then use this format to integrate heterogeneous data; Step 2: Construct a cross - level multi - link data integration framework, including a data collection layer, a data processing layer, a data storage layer, and a data application layer; Step 3: Collect and process the data of quality infrastructure services in real - time; Step 4: Store the data collected and processed in Step 3 into es8.9, and use an embedding model to generate vector fields for the data as the characteristics of the service institution to complete the integration of service data.
[0005] As a further improvement of the present invention, the data collection and processing method in Step 3 is: using a distributed collection method to connect to each data source, cleaning, deduplicating, and integrating the data during the collection process, converting it into the format with the service institution as the main body designed in Step 1, and at the same time quantifying five influencing factors of service cycle, service price, user evaluation, customer service response efficiency, and data maintenance frequency for each service institution in the data through an algorithm.
[0006] As a further improvement of the present invention, the specific formula for quantifying the five influencing factors in Step 3 is as follows: Wherein, represents the quantization value of the i - th influencing factor, is the weight of the j-th data source for the i-th impact factor, is the feature extraction function of the j-th data source, is the original data of the j-th data source.
[0007] As a further improvement of the present invention, the feature extraction function of the j-th data source is calculated as follows: Use an embedding model to convert the service data corresponding to a certain factor of an institution in jsonl format into a vector representation as features.
[0008] As a further improvement of the present invention, the weight of the j-th data source for the i-th impact factor is calculated by the following formula: where is the attention score of the j-th data source for the i-th impact factor.
[0009] As a further improvement of the present invention, the attention score of the j-th data source for the i-th impact factor is calculated by the following formula: where is the weight matrix for converting the input feature into another space, is the weight vector for calculating the attention score, is the bias term for adjusting the output.
[0010] As a further improvement of the present invention, the data acquisition layer constructed in step two includes: Data sources, which are composed of relational databases, log files, API interfaces, loT devices, and third-party platforms; The acquisition module is used for real-time acquisition and regular acquisition; The protocol adapter stores HTTP / HTTPS protocols, JDBC / ODBC protocols, WebSocket protocols, and FTP / SFTP protocols in its memory to adapt to communication protocols; The constructed data processing layer includes: The stream processing engine and the batch processing engine are used for a kind of business flow processing during data processing and processing the data step by step; The data cleaning module is used for deduplication, format verification, and anomaly detection of data; The data conversion module is used for data conversion and aggregation calculation; The constructed data storage layer includes: Relational database for storing structured data; MongoDB database for storing unstructured data; Cache layer for caching data.
[0011] Advantages of the present invention: Achieve effective integration of heterogeneous data: The present invention designs a cross-level multi-link data integration framework that can be compatible with data from different sources, formats, and structures, thereby improving data integration efficiency.
[0012] Improve real-time data processing capabilities: The present invention adopts distributed computing and stream processing technologies, and processes data based on entity resolution algorithms. The quality infrastructure service data is quantified into five influencing factors: service cycle, service price, user evaluation, customer service response efficiency, and data maintenance frequency, providing a basis for subsequent data analysis, thereby significantly reducing decision-making latency and effectively improving the performance during the operation of the server.
[0013] Enhance system scalability: The present invention adopts a modular design and can dynamically adjust system resources according to the growth of data volume to ensure that the system can handle the increment of massive data.
[0014] Strengthen data security and privacy protection: The present invention introduces encryption technology and access control mechanisms to ensure the security of data during transmission and storage, while protecting user privacy. Description of the Drawings
[0015] Figure 1 It is a module block diagram of a cross-level multi-link data integration framework. Detailed Embodiment
[0016] The following will further elaborate on the present invention with the given embodiments.
[0017] A quality infrastructure service data integration method based on cross-level multi-link in this embodiment includes the following steps: Data Structure Design Design a format compatible with different data sources and use es8.9 to integrate heterogeneous data (without considering multi-modal) The following format is provided in this embodiment: {"orgID": "Service organization ID","orgName": "Service organization name","orgCode": "Service organization code","orgAddress": "Service organization address","areaCode": "Administrative division code","areaName": "Administrative division name","orgLinkTel": "Service organization contact information","orgType": "Service organization type","orgBusiness": "Service scope of the service organization", "extendInfo": [ / / Extended information {"key": "Extended information key","value":"Extended information value"}, {"key": "Extended information key","value":"Extended information value"}, … , "orgCapabilityList": [ / / Organization capability data {Organization capability data object, specific structure omitted}, … , "orgOrderList": [ / / Organization order data {Organization order data object, specific structure omitted}, … , "orgServiceRecordList": [ / / Organization service record data {Organization service record data object, specific structure omitted}, … , "orgEvaluationList": [ / / Organization evaluation data {Organization evaluation data object, specific structure omitted}, … , "orgFactorList": [ / / Organization comprehensive evaluation influencing factor data {"key": "Influencing factor name","value":"Factor value"}, … , "vector": [] / / Vector field, dimension 1024} Construction of Cross - level and Multi - link Data Integration Framework Design a cross - level and multi - link data integration framework, including a data collection layer, a data processing layer, a data storage layer, and a data application layer, where: The data collection layer includes: Data sources, which are composed of relational databases, log files, API interfaces, IoT devices, and third - party platforms; A collection module for real - time collection and regular collection; A protocol adapter that stores HTTP / HTTPS protocols, JDBC / ODBC protocols, WebSocket protocols, and FTP / SFTP protocols to adapt communication protocols; The data processing layer includes: A stream processing engine and a batch processing engine for processing business flows during data processing and step - by - step data processing; A data cleaning module for deduplication, format verification, and anomaly detection of data; A data conversion module for data conversion and aggregation calculation; The data storage layer includes: A relational database for storing structured data; A mongodb database for storing unstructured data; A cache layer for caching data; The data application layer includes: BI tools are used for visual management of processed data by directory and realizing real - time application of data through a monitoring dashboard. Thus, in the cross - level and multi - link data integration framework based on the above - constructed data collection layer, data processing layer, data storage layer, and data application layer, the following functions can be achieved: Real - time data collection and processing Using a distributed collection method, connect to each data source, clean, deduplicate, and integrate data during the collection process, and convert it into the format with service institutions as the main body designed in the first step. At the same time, through a self - developed algorithm (HDQ_IFAN algorithm), quantify five influencing factors for each service institution: service cycle, service price, user evaluation, customer service response efficiency, and data maintenance frequency. The formula is expressed as: Where represents the quantified value of the i - th influencing factor, is the weight of the j - th data source for the i - th influencing factor, is the feature extraction function of the j - th data source, is the original data of the j - th data source.
[0018] For , we use the embedding model to convert the service data corresponding to a certain factor of the institution in jsonl format into vector representations, as features.
[0019] For , we use a method based on the attention mechanism to learn, in order to reflect the contribution degrees of different data sources to each influencing factor, and its formula is expressed as: Among them, is the attention score of the j-th data source to the i-th influencing factor, which can be calculated by the following formula: Among them, is the weight matrix, which is used to convert the input feature into another space. is the weight vector, which is used to calculate the attention score (T is the transpose), is the bias term, which is used to adjust the output, and these parameters are updated through the backpropagation algorithm during the training process to optimize the performance of the model.
[0020] An example of its partial Python code implementation is as follows: # Influence factor quantization model class HDQ_IFAN(nn.Module): def __init__(self, num_factors, num_sources, embedding_dim): super(HDQ_IFAN, self).__init__() self.num_factors = num_factors self.num_sources = num_sources self.embedding_dim = embedding_dim # Initialize weights and biases self.W = nn.Parameter(torch.randn(num_factors, embedding_dim)) self.v = nn.Parameter(torch.randn(num_factors, 1)) self.b = nn.Parameter(torch.randn(num_factors, 1)) # Initialize the embedding model self.embedding_models = nn.ModuleList([EmbeddingModel(num_embeddings=1024, embedding_dim=embedding_dim) for _ in range(num_sources)]) def forward(self, sources): # sources: Data from different data sources embedded_features = torch.stack([embedding_model(source) for embedding_model, source in zip(self.embedding_models, sources)]) # Calculate attention scores s = torch.tanh(torch.mm(self.W, embedded_features) + self.b) # Calculate weights weights = F.softmax(torch.mm(self.v, s), dim=0) # Calculate the influence factor IF = torch.sum(weights * embedded_features, dim=0) return IF Finally, five influence factors are obtained, laying a data foundation for subsequent rapid analysis.
[0021] Data storage and management Store the processed data in es8.9, and use the embedding model to generate vector fields for the data, which serve as the features of this service organization Data analysis and display es8.9 supports vector recall and hybrid retrieval. It can also classify data through the vector field using clustering algorithms to obtain more accurate analysis results compared to traditional technologies At the data application layer, display the analysis results to users through visualization technology to support users in making decisions.
[0022] In summary, the method for integrating quality infrastructure service data based on cross - level multi - link in this embodiment can achieve effective integration of heterogeneous data, improve real - time data processing capabilities, further enhance system scalability, and finally strengthen the privacy protection of data security.
[0023] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above - mentioned embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A quality infrastructure service data integration method based on cross-level multi-links, characterized by: The steps include: Step 1: Input a format that is compatible with different data sources, and then use the format to integrate heterogeneous data; Step 2: Build a cross-level multi-link data integration framework, including data collection layer, data processing layer, data storage layer and data application layer; Step 3: Real-time collection and processing of data from quality infrastructure services; Step 4: Store the data collected and processed in step 3 into es8.9, and use the embedding model to generate vector fields for the data as the characteristics of the service organization to complete the integration of service data.
2. The quality infrastructure service data integration method based on cross-level multi-link according to claim 1 is characterized by: The data collection and processing method in step three is: use a distributed collection method to connect various data sources, clean and deduplicate the data and integrate them during the collection process, convert them into the format based on the service agency designed in step one, and at the same time use an algorithm to quantify the five influencing factors of service cycle, service price, user evaluation, customer service response efficiency, and data maintenance frequency for each service agency in the data.
3. The quality infrastructure service data integration method based on cross-level multi-link according to claim 2 is characterized by: The specific formula for quantifying the five influencing factors in step 3 is as follows: in, represents the quantitative value of the i-th impact factor, is the weight of the j-th data source on the i-th influencing factor, is the feature extraction function of the jth data source, is the original data of the jth data source.
4. The quality infrastructure service data integration method based on cross-level multi-link according to claim 3 is characterized by: The feature extraction function of the j-th data source The calculation method is as follows: Use the embedding model to convert the service data corresponding to a factor of an institution in jsonl format into a vector representation, as characteristics.
5. The method for quality infrastructure service data integration based on cross-level multi-links according to claim 3 or 4, characterized in that: The weight of the j-th data source on the i-th influencing factor Calculated by the following formula: in, is the attention score of the j-th data source on the i-th influencing factor.
6. The method for quality infrastructure service data integration based on cross-level multi-links according to claim 5 is characterized by: The attention score of the j-th data source on the i-th influencing factor Calculated by the following formula: in, is a weight matrix used to transform the input features Transformed into another space, is the weight vector used to calculate the attention score, is the bias term used to adjust the output.
7. The method for quality infrastructure service data integration based on cross-level multi-link according to any one of claims 1 to 6, characterized in that: The data collection layer constructed in step 2 includes: Data sources, which consist of relational databases, log files, API interfaces, IoT devices, and third-party platforms; A collection module, used for real-time collection and periodic collection; The protocol adapter has HTTP / HTTPS protocol, JDBC / ODBC protocol, WebSocket protocol and FTP / SFTP protocol in its memory to adapt the communication protocol; The constructed data processing layer includes: Stream processing engines and batch processing engines are used for business flow processing when processing data, processing data in steps; Data cleaning module, used to remove duplicate data, format check and detect anomalies; Data conversion module, used to convert and aggregate data; The data storage layer constructed includes: Relational databases, used to store structured data; Mongodb database, used to store unstructured data; Cache layer, used to cache data.