Multi-source data management method and device, equipment and storage medium

Through multi-source data management methods and knowledge graph technology, the data island problem in the panel furniture industry has been solved, and the organic integration and intelligent application of data has been realized, data management efficiency and accuracy have been improved, and enterprises have been helped to reduce costs and increase efficiency.

CN120256640APending Publication Date: 2025-07-04ZHONGKE HUIYUAN VISUAL TECHNOLOGY (LUOYANG) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510246108.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The data independence between various business systems in the panel furniture industry leads to information islands, resulting in inefficient data storage and utilization, affecting the timeliness and accuracy of data, and limiting the ability to optimize production and improve quality.

Method used

Through the multi-source data management method, multiple data sources are used to obtain structural and non-structural data, and then preprocess it. After data fusion is used to use the knowledge graph data fusion model to be used to store it in a hierarchical data warehouse, providing data services and query functions.

Benefits of technology

It realizes the organic integration of data in different business systems, improves data management efficiency and accuracy, supports intelligent applications, and helps enterprises reduce costs and improve efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256640A_ABST
    Figure CN120256640A_ABST
Patent Text Reader

Abstract

The invention provides a multi-source data management method, apparatus and device, and a storage medium. The method comprises the steps of obtaining multi-source data of a target service by using a plurality of data sources; the multi-source data comprises structural data and non-structural data; performing first preprocessing on the structural data to obtain first data, and performing second preprocessing on the non-structural data to obtain second data; fusing the first data and the second data by using a pre-constructed knowledge graph data fusion model to obtain a fused data knowledge graph; storing the fused data knowledge graph into a data warehouse with a layered architecture; the data warehouse can perform data service, service packaging and service engine operation on the fused data knowledge graph for the data platform to query; according to the technical scheme provided by the invention, multi-source heterogeneous data can be adaptively fused, intelligent application is realized, quality management and closed-loop control levels are improved, and enterprises are helped to achieve the purposes of reducing cost and increasing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a multi-source data management method, device, equipment and storage medium. Background Art

[0002] With the rapid development of artificial intelligence and big data technology, the panel furniture industry is facing an important opportunity for digital transformation. Manufacturing companies are eager to create business value, reduce costs and improve efficiency through the integrated application of digital technology. However, despite the accumulation of massive data in the industry, the relative independence of various business systems has led to information islands, resulting in repeated data storage and inefficient utilization. This not only seriously affects the timeliness and accuracy of data, but also limits the ability of companies to optimize production and improve quality. Summary of the invention

[0003] The present disclosure provides a multi-source data management method, device, equipment and storage medium to at least solve the above technical problems existing in the prior art.

[0004] According to a first aspect of the present disclosure, a multi-source data management method is provided, the method comprising:

[0005] Utilize multiple data sources to obtain multi-source data of the target business; the multi-source data includes structured data and unstructured data;

[0006] Performing a first preprocessing on the structured data to obtain first data, and performing a second preprocessing on the non-structured data to obtain second data;

[0007] Using a pre-built knowledge graph data fusion model to fuse the first data and the second data to obtain a fused data knowledge graph;

[0008] The fused data knowledge graph is stored in a data warehouse with a layered architecture; the data warehouse can perform data services, service encapsulation and service engine operations on the fused data knowledge graph for data platform query.

[0009] In one possible implementation manner, the acquiring multi-source data of the target business using multiple data sources includes:

[0010] Identify each target data source;

[0011] Based on the data characteristics of each target data source, the data in each target data source is collected using a corresponding collection method;

[0012] The collected data is standardized to obtain multi-source data.

[0013] In an implementable embodiment, the first preprocessing of the structural data to obtain first data includes:

[0014] Performing data cleaning, data dimensionality reduction, and data filling operations on the structural data to obtain first data.

[0015] In an implementable embodiment, the second preprocessing of the unstructured data to obtain second data includes:

[0016] Performing word segmentation on the unstructured data based on the sequence labeling method to obtain multiple word segments;

[0017] Projecting the multiple word segments into a mathematical dimensional space to obtain word vectors, and using the word vectors as second data.

[0018] In an implementable embodiment, using the pre-constructed knowledge graph data fusion model to fuse the first data and the second data to obtain a fused data knowledge graph, includes:

[0019] Identifying data entities and entity relationships in the first data and the second data; the data entities include production batches, equipment parameters, board part codes, inspection determination results, and defect information;

[0020] Performing entity alignment on the obtained data entities to obtain aligned entities;

[0021] Obtaining fused data according to the aligned entities and the corresponding entity relationships;

[0022] Constructing a fused data knowledge graph with the aligned entities as nodes and the corresponding entity relationships as edges.

[0023] In an implementable embodiment, performing entity alignment on the obtained data entities includes:

[0024] Determining a first entity and a second entity for which similarity needs to be calculated to obtain an entity pair;

[0025] Calculating the attribute similarity and structural similarity of the entity pair using a similarity function based on the edit distance;

[0026] Performing weighted calculation on the attribute similarity and the structural similarity to obtain a final similarity;

[0027] Comparing the final similarity with a preset threshold, and screening the entity pair according to the comparison result to complete entity alignment.

[0028] In an implementable embodiment, the following method is used to calculate the attribute similarity of the entity pair,

[0029] sim(E1,E2)=(1-α)sim Attr(E1, E2) + αsim Stru (E1, E2)

[0030] where sim Attr (E1, E2) is the attribute similarity function of the corresponding entity pair, and sim Stru (E1, E2) corresponds to the structure similarity function of the entity pair, α is the adjustment parameter, E1 is the first entity, and E2 is the second entity.

[0031] In an implementable manner, the structure similarity of the entity pair is calculated in the following way;

[0032]

[0033] where Stru(E1) and Stru(E2) represent the common neighbor sets of the entities.

[0034] In an implementable manner, the hierarchical architecture of the data warehouse includes: an operational data store layer, a data detail layer, a data summary layer, and an application data layer; storing the fused data knowledge graph into a data warehouse with a hierarchical architecture includes:

[0035] Storing the fused data through the operational data store layer;

[0036] Performing standardization processing on the fused data through the data detail layer to obtain standard data;

[0037] Performing data summarization on the standard data through the data summary layer based on production key indicators;

[0038] Providing data services to each data platform through the application data layer.

[0039] In an implementable manner, providing data services to each data platform through the application data layer includes:

[0040] Performing attribute classification on the obtained fused data, encapsulating each type of classified data, and connecting to the data platform using a data engine to perform data services.

[0041] According to the second aspect of the present disclosure, a multi-source data management device is provided. The device includes:

[0042] An acquisition module for acquiring multi-source data of a target service using multiple data sources; the multi-source data includes structured data and unstructured data;

[0043] A preprocessing module for performing first preprocessing on the structured data to obtain first data, and performing second preprocessing on the unstructured data to obtain second data;

[0044] A fusion module, configured to fuse the first data and the second data by using a pre-constructed knowledge graph data fusion model to obtain a fused data knowledge graph;

[0045] A storage module, configured to store the fused data knowledge graph in a data warehouse with a hierarchical architecture; the data warehouse can perform data services, service encapsulation, and service engine operations on the fused data knowledge graph for query by a data platform.

[0046] In an implementable manner, the acquisition module includes:

[0047] A determination unit, configured to determine each target data source;

[0048] An acquisition unit, configured to collect data in each target data source by using a corresponding acquisition method based on the data characteristics of each target data source;

[0049] A processing unit, configured to perform standardization processing on the collected data to obtain multi-source data.

[0050] In an implementable manner, the preprocessing module includes:

[0051] A first preprocessing unit, configured to perform data cleaning, data dimensionality reduction, and data filling operations on the structured data to obtain first data.

[0052] In an implementable manner, the preprocessing module includes:

[0053] A second preprocessing unit, configured to perform word segmentation on the unstructured data based on a sequence annotation method to obtain a plurality of word segments; project the plurality of word segments into a mathematical dimension space to obtain word vectors, and use the word vectors as second data.

[0054] In an implementable manner, the fusion module includes:

[0055] An identification unit, configured to identify data entities and entity relationships in the first data and the second data; the data entities include production batches, equipment parameters, board part codes, detection determination results, and defect information;

[0056] An alignment unit, configured to perform entity alignment on the obtained data entities to obtain aligned entities;

[0057] A fusion unit, configured to obtain fused data according to the aligned entities and the corresponding entity relationships;

[0058] A construction unit, configured to construct a fused data knowledge graph with the aligned entities as nodes and the corresponding entity relationships as edges.

[0059] In an implementable manner, the alignment unit includes:

[0060] A first calculation subunit, configured to determine a first entity and a second entity for which similarity needs to be calculated, and obtain an entity pair;

[0061] A second calculation subunit, configured to calculate the attribute similarity and the structural similarity of the entity pair by using a similarity function based on the edit distance;

[0062] A weighting subunit, configured to perform weighted calculation on the attribute similarity and the structural similarity to obtain a final similarity;

[0063] A screening subunit, configured to compare the final similarity with a preset threshold, and screen the entity pair according to the comparison result to complete entity alignment.

[0064] In an implementable manner, the attribute similarity of the entity pair is calculated in the following way

[0065] sim(E1,E2) = (1 - α)sim Attr (E1,E2) + αsim stru (E1,E2)

[0066] wherein, sim Attr (E1,E2) is an attribute similarity function corresponding to the entity pair, sim Stru (E1,E2) corresponds to a structural similarity function of the entity pair, α is an adjustment parameter, E1 is the first entity, and E2 is the second entity.

[0067] In an implementable manner, the structural similarity of the entity pair is calculated in the following way;

[0068]

[0069] wherein, Stru(E1) and Stru(E2) represent the common neighbor sets of the entities.

[0070] In an implementable manner, the storage module includes:

[0071] An operation data storage layer, configured to store fusion data;

[0072] A data details layer, configured to perform standardization processing on the fusion data to obtain standard data;

[0073] A data summary layer, configured to perform data summary on the standard data based on production key indicators;

[0074] An application data layer, configured to provide data services to each data platform.

[0075] In an implementable manner, the application data layer includes:

[0076] A service unit for classifying the attributes of the obtained fusion data, encapsulating various types of classified data, and connecting to a data platform through a data engine for data services.

[0077] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0078] At least one processor; and

[0079] A memory communicatively connected to the at least one processor; wherein,

[0080] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the present disclosure.

[0081] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the present disclosure.

[0082] The multi-source data management method, apparatus, device and storage medium of the present disclosure can adaptively fuse multi-source heterogeneous data from different business systems, improve data quality, enhance the quality management and closed-loop control level of furniture enterprises, and help enterprises achieve the purpose of cost reduction and efficiency improvement.

[0083] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, wherein:

[0085] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0086] Figure 1 Shows the implementation process schematic of the multi-source data management method according to the embodiment of the present disclosure Figure 1 ;

[0087] Figure 2 Shows the implementation process schematic of the multi-source data management method according to the embodiment of the present disclosure Figure 2 ;

[0088] Figure 3 Shows the construction schematic diagram of the knowledge graph according to the embodiment of the present disclosure;

[0089] Figure 4 Shows the schematic diagram of the data warehouse structure according to an embodiment of the present disclosure;

[0090] Figure 5 Shows the schematic diagram of the data warehouse function according to an embodiment of the present disclosure;

[0091] Figure 6 Shows the schematic diagram of the structure of the multi-source data management device according to an embodiment of the present disclosure;

[0092] Figure 7 Shows the schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0093] To make the objectives, features, and advantages of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.

[0094] As Figure 1 shown, the present application provides a multi-source data management method, and the method includes:

[0095] S101, obtaining multi-source data of a target service by using a plurality of data sources; the multi-source data includes structured data and unstructured data;

[0096] The technical solution provided by the present application can be applied to different industries, such as panel furniture enterprises. The obtained data includes data in various links such as production, finance, and quality management of panel furniture enterprises; data from different business systems, equipment operation data, external text data, handwritten data, etc. These data sources have different formats and sources, including structured data and unstructured data, thus forming multi-source heterogeneous data. The data sources of the multi-source data can include data obtained from the Internet, data obtained through manual surveys, etc.

[0097] As Figure 2 shown, the multi-source data includes structured data such as machine operation data, detection data, and enterprise organizational structure data, and unstructured data such as equipment historical maintenance records, equipment SOP files, and handwritten texts.

[0098] S102, performing a first preprocessing on the structured data to obtain first data, and performing a second preprocessing on the unstructured data to obtain second data;

[0099] Since the multi-source data obtained has various formats, it is necessary to first standardize the data from different sources. Structured data is data that is logically expressed and implemented through a two-dimensional table structure, strictly following the data format and length specifications, and is mainly stored and managed through a relational database. Opposite to structured data is unstructured data, including various formats of office documents, XML, HTML, various reports, pictures, audio, video information, etc. Because the attributes of structured data and unstructured data are different, the preprocessing methods for structured data and unstructured data are different. The first preprocessing is performed on structured data to obtain the first data, and the second preprocessing is performed on unstructured data to obtain the second data. For structured data, this application uses a unified table structure for storage to ensure that the fields and data types are consistent; for unstructured data, it is converted into structured records through a parsing tool, and the encoded data uniformly adopts the internal enterprise standard to ensure the consistency of device and product information between different systems.

[0100] S103, using the pre-constructed knowledge graph data fusion model to fuse the first data and the second data to obtain a fused data knowledge graph;

[0101] After constructing the knowledge graph data fusion model, using the entity alignment and relationship matching technologies of the knowledge graph, effectively integrate the data from different sources, solve the data heterogeneity problem, and achieve the unified representation of data.

[0102] S104, storing the fused data knowledge graph in a data warehouse with a hierarchical architecture; the data warehouse can perform data services, service encapsulation, and service engine operations on the fused data knowledge graph for query by the data platform.

[0103] Finally, adopting a hierarchical architecture, store the fused data in a unified data warehouse, and through metadata management and data standardization, ensure the unity, availability, and consistency of the data.

[0104] The multi-source data management method provided by this application effectively fuses the multi-source heterogeneous data from different business systems, combines entity alignment and relationship matching technologies, solves problems such as inconsistent data formats and information silos between systems, realizes the organic integration of data, and improves the efficiency and accuracy of data management. In addition, adopting a hierarchical architecture for data storage, stores the processed multi-source heterogeneous data in a unified data warehouse. Through metadata management and data standardization, ensures the consistency and availability of the data, and provides efficient data service support for subsequent complex queries and intelligent applications.

[0105] In addition, the data warehouse provided by this application can provide services such as data services, service encapsulation, and service engines, support complex queries and multi-dimensional analysis, and ensure the efficient flow and sharing of data between various business modules. Then, based on a unified data platform, it provides a variety of applications such as powerful data query, report display, data analysis, data dashboards, and business decision-making support, thereby helping enterprises to comprehensively achieve digital and intelligent transformation while reducing costs and improving efficiency.

[0106] In some embodiments, the obtaining of multi-source data of a target business by using multiple data sources includes:

[0107] Determine each target data source;

[0108] Based on the data characteristics of each target data source, adopt corresponding acquisition methods to acquire the data in each target data source;

[0109] Perform standardization processing on the acquired data to obtain multi-source data.

[0110] It should be noted that this application first determines various data sources involved in panel furniture enterprises to ensure coverage of all business processes; among them, business processes include data from different business systems such as quality manufacturing execution systems, production manufacturing management systems, and business process management systems; it can also include equipment operation data such as the operating status, energy consumption data, and operating time of edge banding machines, electronic saws, etc., external data such as equipment historical maintenance data, SOP standard operating procedures, and expert knowledge bases; then, according to the characteristics of different data sources, different data acquisition methods are designed, specifically including: automatically capturing operation logs, production data, etc. from internal business systems through interfaces such as APIs or database connections; using Internet of Things sensors to real-time collect equipment operation data such as energy consumption, temperature, and vibration and collect it to the data platform; batch uploading unstructured data such as external maintenance records and SOP documents through file import tools; manually inputting empirical data or data that cannot be automatically collected in the expert knowledge base to ensure data integrity.

[0111] In some embodiments, the first preprocessing of the structured data to obtain first data includes:

[0112] Perform data cleaning, data dimensionality reduction, and data filling operations on the structured data to obtain first data.

[0113] Among them, in this application, data cleaning is achieved by screening out data that does not conform to the rules, so as to achieve the purpose of accurate data. In the production system of panel furniture, due to diverse data sources and complex structures, the collected data cannot be guaranteed to be completely accurate, and incorrect and irregular data are widespread. These irregular data will have a very adverse impact on subsequent analysis and fusion. The purpose of data cleaning is to detect, through means such as detection, the incorrect and irregular data existing in multi-source heterogeneous data, and to eliminate the incorrect and irregular data through methods such as screening and repair, thereby improving the quality of data.

[0114] Data dimensionality reduction is to reduce the number of features in a dataset, remove redundant or unnecessary information, and reduce data complexity to improve the efficiency of subsequent data processing. In the production management of panel furniture, the data dimension is usually relatively large, and may include multiple aspects such as equipment status, production process, energy consumption, etc. To ensure the calculation efficiency and processing speed of the model, data dimensionality reduction uses algorithms such as principal component analysis (PCA) and linear discriminant analysis (LDA) to retain the most valuable features for the model, reduce irrelevant or redundant dimensions, thereby improving data processing efficiency and the prediction ability of the model.

[0115] Time series data filling is to use specific interpolation algorithms to fill in the missing data when dealing with incomplete or uneven time series data caused by different data collection frequencies. In the production process of a panel furniture enterprise, due to different data collection frequencies of sensors on each device, the density of time series data will show differences. To make the data density consistent, this application uses the Lagrange interpolation method to fill in the time series data of the device. The Lagrange polynomial of the time series data is expressed as:

[0116]

[0117] Among them, F scada (t) is the Lagrange interpolation function corresponding to the device time series data; l i (t) is the interpolation basis function; t i is the time series data; t m is the acquisition time corresponding to the time series data m. Then, the time series data in the production process is filled based on the above content, and noise reduction processing is performed during the filling process. The following method is used for noise reduction processing:

[0118]

[0119] Among them, r j is the actual value of the time series data collected at time j; v j is the smoothed value of the first step of the time series data at time j, used to represent the short-term trend of data change; β1 and β2 are respectively the trend smoothing parameters for the step size; sj is the smoothed value of the secondary step length of the time series data at time j, used to represent the long-term trend of data changes; t j represents the time index of the time series data; h is the prediction; v j+h represents the first smoothed value v at time j j plus the prediction step length.

[0120] The processed data is obtained by filling and denoising the time series data, thus providing a good foundation for the fusion of multi-source heterogeneous data.

[0121] In some embodiments, the second preprocessing of the unstructured data to obtain the second data includes:

[0122] Performing word segmentation on the unstructured data based on the sequence annotation method to obtain multiple word segments;

[0123] Projecting the multiple word segments into a mathematical dimensional space to obtain word vectors, and using the word vectors as the second data.

[0124] Specifically, the present application performs preprocessing operations on unstructured text data to ensure the smooth progress of subsequent analysis and processing. The text preprocessing operations mainly include two core steps: Chinese word segmentation and constructing a word vector model. First, since Chinese text is a continuous character sequence, it is difficult to process directly, so word segmentation is required to divide the continuous text into multiple meaningful word segments. The present invention uses a sequence annotation-based method for word segmentation. By injecting a sorted industry corpus, a model is built for word segmentation to predict the distribution and probability of word segmentation and output accurate word segmentation results; second, construct a word vector model, project the text data into a mathematical dimensional space, and transform it into digital features that can be processed by a computer, that is, word vectors. This processing step provides basic data support for subsequent text information extraction and knowledge graph construction.

[0125] In some embodiments, the fusion of the first data and the second data using the pre-constructed knowledge graph data fusion model to obtain a fused data knowledge graph includes:

[0126] Identifying data entities and entity relationships in the first data and the second data; the data entities include production batches, equipment parameters, board part codes, inspection and determination results, and defect information;

[0127] Performing entity alignment on the obtained data entities to obtain aligned entities;

[0128] Obtaining fused data according to the aligned entities and the corresponding entity relationships;

[0129] Constructing a fused data knowledge graph with the aligned entities as nodes and the corresponding entity relationships as edges.

[0130] Specifically, this application first performs text information extraction. Specifically, by taking advantage of the advantages of a rule-based method, which has a precise scope and high accuracy, certain rules are written to match a small number of precise objects as the subsequent corpus import. Then, the sequence labeling method is used to continue extracting key information, and part of the results (such as 70%) output from the previous step can be used as training corpus to replace the process of manually injecting corpus. Next, using the training corpus output from the previous step and based on open-source algorithms, a knowledge graph is modeled. Finally, the model from the above steps is used to judge the remaining part of the corpus. If the model judgment result shows that the model does not meet the standard, more corpus is returned for supplementation until the model automatically judges to meet the standard, at which point the iteration is exited, and the most recently generated model is used as the final model.

[0131] Among them, when constructing the knowledge graph data fusion model, based on the neural network model and deep learning technology, a knowledge graph framework dedicated to multi-source heterogeneous data fusion is constructed. This framework realizes the effective integration and association between different data sources by identifying key entities and relationships in production data. The construction of the knowledge graph is mainly divided into three core steps: entity recognition, entity relationship extraction, and entity alignment.

[0132] First, as Figure 3 shown, in the entity recognition task, this application introduces the BERT BiLSTM-CRF model. BERT, with its pre-trained language model, can deeply understand the semantics of text related to production quality, thus accurately extracting key entities such as production batches, equipment parameters, board codes, inspection judgment results, and defect information. At the same time, the BiLSTM-CRF model further enhances the sequence labeling ability of these entities, ensuring the efficient and accurate extraction of core data for production and quality management in a complex production environment.

[0133] Secondly, in the entity relationship extraction task, this invention adopts the open-source TextCNN model. As a classic text classification model based on convolutional neural networks, TextCNN can process sentences of different lengths through convolutional operations and extract effective features, identify and extract the semantic relationships existing between entities from unstructured text data, and convert unstructured data into structured information. By using the TextCNN model, the multi-dimensional relationships between different entities in the production process can be accurately extracted, ensuring the establishment of effective associations between data in different business systems, device data, and external data. The identification of these relationships provides a solid data foundation for subsequent analysis of quality changes in the production process and the impact of equipment on quality.

[0134] In addition, since the data comes from different business systems, there is a certain degree of duplication in the data information between different data carriers, and there will be certain differences in the representation of entities. Therefore, it is necessary to clean and integrate the data, that is, to eliminate the ambiguity of concepts through knowledge fusion, and to remove redundant and incorrect data, so as to ensure the quality of the final knowledge. The present invention adopts a feature matching algorithm based on a similarity function to perform knowledge disambiguation and fusion. By calculating the similarity of multi-source data such as device information, detection data, and personnel information, the same entity is uniformly represented, and data redundancy or conflicts existing in different systems are eliminated. Especially for the defect data (such as defect names, sizes, positions, etc.) uploaded by different detection devices, this algorithm can automatically identify similar defects and perform merging processing, thereby ensuring the consistency and accuracy of the data.

[0135] In some embodiments, the entity alignment of the obtained data entities includes:

[0136] Determine the first entity and the second entity for which the similarity needs to be calculated to obtain an entity pair;

[0137] Adopt a similarity function based on the edit distance to calculate the attribute similarity and the structure similarity of the entity pair;

[0138] Perform a weighted calculation on the attribute similarity and the structure similarity to obtain the final similarity;

[0139] Compare the final similarity with a preset threshold, and screen the entity pair according to the comparison result to complete the entity alignment.

[0140] Specifically, for two entities E1 and E2 in the entity alignment process, define their similarity function as:

[0141] sim(E1,E2)=(1-α)sim Attr (E1,E2)+αsim Stru (E1,E2)

[0142] Wherein, sim Attr (E1,E2) is the attribute similarity function of the corresponding entity pair, and sim Stru (E1,E2) corresponds to the structure similarity function of the entity pair, and α is an adjustment parameter.

[0143] For the attribute set U of entity E1 and the attribute set V of entity E2, for the attributes u and v that both entities contain, a similarity function based on the edit distance is used to calculate the similarity. This function measures the similarity between two strings by the number of minimum single-character edit operations required to convert one string into the other. The basic edit operations usually include insertion, deletion, replacement, etc. The specific solution process is as follows: Initialize a matrix A of size (|u| + 1) × (|v| + 1), and denote the element in the i-th row and j-th column of A as A i,j , and the value of A is:

[0144]

[0145] After the values in matrix A are calculated, A |u|,|v| is the calculated edit distance. Calculate the weighted average of the edit distances obtained for the attribute pairs of the two entities to obtain sim Attr (E1, E2).

[0146] Structural similarity is an important part of the similarity measurement of entity pairs. The structural similarity of entities is measured by using the Jaccard correlation coefficient of the common neighbors of the entities. The calculation formula is as follows:

[0147]

[0148] Among them, Stru(E1) and Stru(E2) represent the sets of common neighbors of the entities.

[0149] Obtain the similarity of the two entities, screen the entity pairs according to the set threshold, retain the entity pairs within the threshold range, and perform manual verification on the entity pairs exceeding the threshold to confirm whether they are different expressions of the same entity, thus completing entity alignment.

[0150] Finally, the constructed knowledge graph will be stored in the Neo4j graph database and visualized and searched through the Cypher query language.

[0151] In some embodiments, the hierarchical architecture of the data warehouse includes: an operational data store layer, a data detail layer, a data summary layer, and an application data layer; storing the fused data knowledge graph into a data warehouse with a hierarchical architecture includes:

[0152] Store the fused data through the operational data store layer;

[0153] Standardize the fused data through the data detail layer to obtain standard data;

[0154] Perform data summarization on the standard data through the data summary layer based on production key indicators;

[0155] Provide data services to each data platform through the application data layer.

[0156] Specifically, as Figure 4 shown, in this application, data storage adopts a hierarchical architecture to achieve efficient management, mainly including the ODS layer (Operational Data Store), the DWD layer (Data Warehouse Detail), the DWS layer (Data Warehouse Summary), and the ADS layer (Application DataStore). The ODS layer obtains multi-source data such as raw production, equipment, and logistics. The DWD layer cleans, transforms, and standardizes the data to ensure data consistency and accuracy. The DWS layer aggregates key production indicators and conducts data summarization to support operation monitoring. Finally, the ADS layer provides real-time and accurate data services for various business applications and intelligent analysis, improving the intelligent level of management decisions.

[0157] The multi-source data management method provided by this application records and traces the data sources, definitions, and change information of each link in the production process of panel furniture through a complete metadata management system. Metadata management not only ensures data consistency between different systems and modules but also provides transparency and traceability for the integration of multi-source heterogeneous data, making data processing in the digital intelligent management process more standardized and efficient, and ensuring the accurate transmission and application of data in every link of production, equipment, and logistics management. For the multi-source heterogeneous data in panel furniture production, data standardization ensures the consistency of data formats, coding, and naming rules in each business system. By unifying data formats, standardizing product and equipment coding, and standardizing naming, the incompatibility problem of heterogeneous data in the integration process is solved, improving data consistency and availability, providing a stable and reliable data foundation for the intelligent management system of panel furniture enterprises, and facilitating the digital transformation and optimization of the production process.

[0158] In some embodiments, providing data services to each data platform through the application data layer includes:

[0159] Classify the attributes of the obtained fusion data, encapsulate each type of classified data, and use a data engine to connect to the data platform for data services.

[0160] In this application, a data middle platform can be established in the data warehouse, and the data middle platform can implement data services, service encapsulation, and service engines. Specifically, as Figure 5As shown, the data middle platform uniformly encapsulates and manages various types of data (such as production, inventory, equipment status, etc.), integrates data from different sources into standardized data services, and supports applications such as data query, statistical analysis, intelligent warning, and process optimization feedback in enterprise production management. Through the data query service, enterprises can quickly obtain real-time data in multiple links such as production and finance to support the dynamic monitoring of business; the data statistical analysis function helps enterprises conduct in-depth analysis of historical data, generate reports and trend forecasts; the intelligent warning service issues abnormal warnings in a timely manner by monitoring key indicators to ensure the stable operation of the production process; the process optimization and feedback function can put forward optimization suggestions through data analysis to further improve production efficiency.

[0161] Among them, service encapsulation ensures the security and flexibility of data services through modular and standardized designs. First, through the process orchestration function, various data processing steps are automated, manual intervention is reduced, and full-process automated management is achieved; then, the service gateway is responsible for managing the access rights of all data services to ensure the security and compliance of data during the call process; secondly, configuration management allows enterprises to dynamically adjust service configurations according to different needs to improve the adaptability of the system; finally, the service management module provides unified service registration, release, monitoring, and maintenance functions to ensure the high availability and maintainability of the data middle platform services.

[0162] The service engine can support multiple scenarios such as data query, indicator calculation, model prediction, and decision generation. The query engine ensures fast response to complex query requests in the scenario of massive data by optimizing query performance; the indicator engine integrates multi-dimensional business data through a large wide-table design to accelerate the calculation and generation of indicators; the model engine is responsible for intelligent prediction and analysis to support intelligent applications such as demand prediction and quality monitoring; the decision engine combines data analysis and business rules to automatically generate management suggestions to provide strong support for the business decisions of enterprises.

[0163] This application can provide functions such as data query, report display, and data dashboard to other data platforms through the data middle platform, monitor the production quality and equipment operation status in real time, and conduct multi-dimensional analysis through data mining to discover potential problems and optimize production plans and resource allocation; this application integrates intelligent warning and root cause location, automatically detects quality problems and equipment failures through data analysis, provides real-time warnings, and accurately locates the root cause of problems based on the causal analysis model and provides improvement suggestions; this application can also generate key business indicators and dynamic dashboards through comprehensive analysis of production, quality, cost, etc. data, provide optimization suggestions for business strategies, help enterprises improve efficiency, reduce costs, and enhance market competitiveness.

[0164] The multi-source data management method provided by this application utilizes knowledge graph technology to effectively integrate multi-source heterogeneous data from different business systems. By combining entity alignment and relationship matching technologies, it solves problems such as inconsistent data formats and information silos between systems, realizes the organic integration of data, and improves the efficiency and accuracy of data management. This application adopts a hierarchical architecture for data storage, storing the processed multi-source heterogeneous data in a unified data warehouse. Through metadata management and data standardization, the consistency and availability of data are ensured, and efficient data service support is provided for subsequent complex queries and intelligent applications. The data warehouse provided by this application can also provide services such as data services, service encapsulation, and service engines, support complex queries and multi-dimensional analysis, and ensure the efficient flow and sharing of data between business modules. Then, based on a unified data platform, it provides various applications such as powerful data query, report display, data analysis, data dashboards, and business decision-making support, thereby helping enterprises to comprehensively achieve digital transformation while reducing costs and improving efficiency.

[0165] As Figure 6 shown, this application provides a multi-source data management device, and the device includes:

[0166] An acquisition module 601, configured to acquire multi-source data of a target business by using multiple data sources; the multi-source data includes structured data and unstructured data;

[0167] A preprocessing module 602, configured to perform a first preprocessing on the structured data to obtain first data, and perform a second preprocessing on the unstructured data to obtain second data;

[0168] A fusion module 603, configured to fuse the first data and the second data by using a pre-constructed knowledge graph data fusion model to obtain a fused data knowledge graph;

[0169] A storage module 604, configured to store the fused data knowledge graph in a data warehouse with a hierarchical architecture; the data warehouse can perform data services, service encapsulation, and service engine operations on the fused data knowledge graph for query by a data platform.

[0170] In some embodiments, the acquisition module includes:

[0171] A determination unit, configured to determine each target data source;

[0172] An acquisition unit, configured to acquire data in each target data source by using a corresponding acquisition method based on the data characteristics of each target data source;

[0173] A processing unit, configured to perform standardization processing on the acquired data to obtain multi-source data.

[0174] In some embodiments, the preprocessing module includes:

[0175] A first preprocessing unit for performing data cleaning, data dimensionality reduction, and data filling operations on the structural data to obtain first data.

[0176] In some embodiments, the preprocessing module includes:

[0177] A second preprocessing unit for performing word segmentation on the unstructured data based on the sequence annotation method to obtain multiple word segments; projecting the multiple word segments into a mathematical dimensional space to obtain word vectors, and using the word vectors as second data.

[0178] In some embodiments, the fusion module includes:

[0179] An identification unit for identifying data entities and entity relationships in the first data and the second data; the data entities include production batches, equipment parameters, board part codes, inspection determination results, and defect information;

[0180] An alignment unit for performing entity alignment on the obtained data entities to obtain aligned entities;

[0181] A fusion unit for obtaining fusion data according to the aligned entities and the corresponding entity relationships;

[0182] A construction unit for constructing a fusion data knowledge graph with the aligned entities as nodes and the corresponding entity relationships as edges.

[0183] In some embodiments, the alignment unit includes:

[0184] A first calculation subunit for determining a first entity and a second entity for which similarity needs to be calculated to obtain an entity pair;

[0185] A second calculation subunit for calculating the attribute similarity and structural similarity of the entity pair using a similarity function based on the edit distance;

[0186] A weighting subunit for performing weighted calculation on the attribute similarity and the structural similarity to obtain a final similarity;

[0187] A screening subunit for comparing the final similarity with a preset threshold, and screening the entity pair according to the comparison result to complete entity alignment.

[0188] In some embodiments, the attribute similarity of the entity pair is calculated in the following manner

[0189] sim(E1,E2)=(1 - α)sim Attr (E1,E2)+αsim Stru(E1, E2)

[0190] Among them, sim Attr (E1, E2) is the attribute similarity function of the corresponding entity pair, and sim Stru (E1, E2) corresponds to the structural similarity function of the entity pair, α is the adjustment parameter, E1 is the first entity, and E2 is the second entity.

[0191] In some embodiments, the structural similarity of the entity pair is calculated in the following manner;

[0192]

[0193] Among them, Stru(E1) and Stru(E2) represent the set of common neighbors of the entities.

[0194] In some embodiments, the storage module includes:

[0195] An operation data storage layer for storing fusion data;

[0196] A data details layer for standardizing the fusion data to obtain standard data;

[0197] A data summary layer for summarizing the standard data based on production key indicators;

[0198] An application data layer for providing data services to each data platform.

[0199] In some embodiments, the application data layer includes:

[0200] A service unit for classifying the attributes of the obtained fusion data, encapsulating each type of classified data, and connecting to the data platform using a data engine for data services.

[0201] It should be noted that for the multi-source data management device in the embodiments of the present application, since the principle of the multi-source data management device for solving problems is similar to the aforementioned multi-source data management method, the implementation process, implementation principle, and beneficial effects of the multi-source data management device can all refer to the description of the implementation process, implementation principle, and beneficial effects of the aforementioned method, and repeated parts will not be elaborated.

[0202] The embodiments of the present application provide an electronic device, including:

[0203] At least one processor; and

[0204] A memory communicatively connected to the at least one processor; wherein,

[0205] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any of the above embodiments.

[0206] An embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described in any of the above embodiments.

[0207] According to an embodiment of the present application, the present application further provides an electronic device and a readable storage medium.

[0208] Figure 7 A schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0209] As Figure 7 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0210] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0211] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the multi-source data management method. For example, in some embodiments, the multi-source data management method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the multi-source data management method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the multi-source data management method by any other suitable means (e.g., by means of firmware).

[0212] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0213] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0214] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0215] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0216] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0217] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0218] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0219] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this disclosure, "a plurality of" means two or more, unless otherwise specifically defined.

[0220] As described above, the above are only specific embodiments of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed in this disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. A multi-source data management method, characterized in that, The method includes: Obtaining multi-source data of a target business by using multiple data sources; the multi-source data includes structured data and unstructured data; Performing first preprocessing on the structured data to obtain first data, and performing second preprocessing on the unstructured data to obtain second data; Using a pre-constructed knowledge graph data fusion model to fuse the first data and the second data to obtain a fused data knowledge graph; Storing the fused data knowledge graph in a data warehouse with a hierarchical architecture; the data warehouse can perform data services, service encapsulation, and service engine operations on the fused data knowledge graph for query by a data platform.

2. The method according to claim 1, characterized in that, The obtaining multi-source data of a target business by using multiple data sources includes: Determining each target data source; Based on the data characteristics of each target data source, collecting the data in each target data source by using a corresponding collection method; Performing standardization processing on the collected data to obtain multi-source data.

3. The method according to claim 1, wherein The performing first preprocessing on the structured data to obtain first data includes: Performing data cleaning, data dimensionality reduction, and data filling operations on the structured data to obtain first data.

4. The method according to claim 1, wherein The performing second preprocessing on the unstructured data to obtain second data includes: Performing word segmentation processing on the unstructured data based on a sequence annotation method to obtain multiple word segments; Projecting the multiple word segments into a mathematical dimensional space to obtain word vectors, and using the word vectors as second data.

5. The method according to claim 1, characterized in that, The using a pre-constructed knowledge graph data fusion model to fuse the first data and the second data to obtain a fused data knowledge graph includes: Identifying data entities and entity relationships in the first data and the second data; the data entities include production batches, equipment parameters, board part codes, inspection determination results, and defect information; Performing entity alignment on the obtained data entities to obtain aligned entities; Obtaining fused data according to the aligned entities and corresponding entity relationships; Constructing a fused data knowledge graph with the aligned entities as nodes and the corresponding entity relationships as edges.

6. The method according to claim 5, characterized in that, The performing entity alignment on the obtained data entities includes: Determining a first entity and a second entity for which similarity needs to be calculated to obtain an entity pair; Calculating the attribute similarity and structural similarity of the entity pair by using a similarity function based on edit distance; Performing weighted calculation on the attribute similarity and the structural similarity to obtain a final similarity; Comparing the final similarity with a preset threshold, and screening the entity pair according to the comparison result to complete entity alignment.

7. The method according to claim 6, wherein Calculating the attribute similarity of the entity pair in the following manner sim(E1, E2) = (1 - α)sim Attr (E1, E2) + αsim Stru (E1, E2) Among them, sim Attr (E1, E2) is the attribute similarity function of the corresponding entity pair, sim Stru (E1, E2) corresponds to the structure similarity function of the entity pair, α is the adjustment parameter, E1 is the first entity, and E2 is the second entity.

8. A multi-source data management device, characterized in that, The device includes: An obtaining module, configured to obtain multi-source data of a target business by using multiple data sources; the multi-source data includes structured data and unstructured data; A preprocessing module, configured to perform first preprocessing on the structured data to obtain first data, and perform second preprocessing on the unstructured data to obtain second data; A fusion module, configured to use a pre-constructed knowledge graph data fusion model to fuse the first data and the second data to obtain a fused data knowledge graph; A storage module for storing the fused data knowledge graph in a data warehouse with a hierarchical architecture; the data warehouse can perform data services, service encapsulation, and service engine operations on the fused data knowledge graph for query by a data platform.

9. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-7.