A data processing method for irregular data collection, analysis, and organization.
By employing multi-protocol access, semantic extraction, and dynamic knowledge graph technologies, the problems of efficient collection, intelligent analysis, and accurate retrieval of irregular data have been solved, enabling efficient and intelligent data processing and retrieval, and improving the flexibility and accuracy of the data processing system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI BEIKE INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-06-30
AI Technical Summary
Traditional data processing methods suffer from problems such as low collection efficiency, cumbersome preprocessing, insufficient intelligent analysis, and poor classification, clustering, and retrieval results when dealing with irregular data, making it difficult to meet the requirements of modern big data applications for real-time performance, accuracy, and intelligence.
By building a data acquisition gateway that supports multi-protocol access, noise reduction and format normalization are performed. Pre-trained semantic extraction models are used to automatically extract data features, construct dynamic knowledge graphs for autonomous updating and classification, achieve unsupervised adaptive clustering and accurate retrieval, and combine intelligent indexes for multi-dimensional archiving and full-link traceability.
It significantly improves the accuracy and efficiency of data processing, enhances the flexibility and scalability of the system, supports the intelligent processing and application of irregular data, and improves the convenience and accuracy of data retrieval.
Smart Images

Figure CN122311206A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of big data processing, knowledge graphs, and artificial intelligence, specifically a data processing method for the collection, analysis, and organization of irregular data. Background Technology
[0002] In today's information age, data has become a key resource driving social progress and economic development. With the rapid development of technologies such as the Internet, the Internet of Things, and big data, the rate of data generation is exploding, and data sources are becoming increasingly diverse, including social media, sensor networks, enterprise logs, and public datasets. Much of this data is unstructured, heterogeneous, and raw, containing rich information and value, but traditional data processing methods struggle to directly and effectively analyze and utilize it.
[0003] Traditional data processing methods often suffer from several limitations when dealing with irregular data. First, traditional methods typically lack flexibility and scalability in the data acquisition phase, struggling to efficiently handle diverse and heterogeneous data sources, leading to incomplete data collection or significant redundancy. Second, in the data preprocessing stage, traditional methods often rely on manually defined rules and templates for data cleaning and format conversion, which is not only time-consuming and labor-intensive but also ill-suited to the diversity and variability of data structures. Third, traditional data analysis methods often lack intelligent feature extraction and semantic understanding capabilities when processing irregular data, making it difficult to uncover the deeper meanings and relationships behind the data, thus limiting the depth and breadth of data analysis. Finally, traditional methods also suffer from inefficiency and insufficient accuracy in data classification, clustering, and retrieval, failing to meet the real-time, accurate, and intelligent requirements of modern big data applications.
[0004] In summary, traditional data processing methods suffer from drawbacks such as low collection efficiency, cumbersome preprocessing, insufficient intelligent analysis, and poor classification, clustering, and retrieval results when dealing with irregular data. To overcome these limitations, this invention proposes a data processing method specifically for the collection, analysis, and organization of irregular data, which is therefore of paramount importance. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a data processing method for the collection, analysis, and organization of irregular data. It achieves efficient collection and standardized preprocessing of multi-source, heterogeneous, and irregular data by building a collection gateway that supports multi-protocol access. Utilizing a pre-trained semantic extraction model, it automatically extracts entities, attributes, and their relationships from the data, generating feature triples and semantic representation vectors. Based on dynamic knowledge graph technology, it realizes dynamic updates of the knowledge graph ontology architecture and automatic iteration of the classification system, supporting adaptive clustering and accurate retrieval of irregular data. This invention not only significantly improves the accuracy and efficiency of data processing but also significantly enhances the flexibility and scalability of the data processing system, providing strong support for the intelligent processing and application of irregular data.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a data processing method for the collection, analysis, and organization of irregular data, the specific steps of which are as follows:
[0007] S1. Irregular data acquisition and standardization preprocessing: Collect multi-source heterogeneous irregular data, perform noise reduction and format normalization preprocessing on the collected irregular data to obtain a standardized irregular dataset to be processed.
[0008] S2. Automatic extraction of data entities and semantic features: Based on a pre-trained semantic extraction model, the entity, attribute, relationship and core semantic features of each data in the standardized irregular dataset are automatically extracted to generate feature triples and semantic representation vectors for each data.
[0009] S3. Autonomous Evolutionary Dynamic Knowledge Graph Construction and Ontology Update: The initial knowledge graph is constructed based on the extracted feature triples. At the same time, based on the semantic features and entity distribution of the newly added data, the knowledge graph ontology architecture is dynamically updated, the classification system is automatically iterated, and the relationship is completed autonomously, without the need for manual preset ontology rules and classification system.
[0010] S4. Adaptive Clustering and Tag Generation of Irregular Data Based on Dynamic Knowledge Graph: Using the ontology architecture of dynamic knowledge graph as the classification benchmark, and based on the semantic representation vector of data and entity association, unsupervised adaptive clustering of irregular data is completed. At the same time, based on the core semantics and entity attributes of the clusters, a hierarchical tag system for the corresponding data is automatically generated.
[0011] S5. Hierarchical archiving and intelligent index construction of irregular data: Based on the automatically generated hierarchical tag system and entity association relationship of knowledge graph, multi-dimensional hierarchical archiving of irregular data is completed. At the same time, combined with data semantic features, entity attributes and archiving path, an intelligent index library for semantic retrieval is constructed.
[0012] S6. Full-chain traceability and semantic retrieval of archived data: Based on the association links of dynamic knowledge graph, the entire lifecycle of archived data, including its source, flow, modification, and archiving, is traceable. At the same time, based on the constructed intelligent index library, users can complete accurate semantic retrieval of archived data through natural language.
[0013] Furthermore, the specific implementation process of step S1 is as follows: build a collection gateway that supports multi-protocol access, and set up dedicated collection channels for irregular data from different sources; perform coarse filtering on the collected raw data to remove invalid, redundant, damaged and illegal data; perform full format normalization on the filtered valid data, and unify the encoding, timestamp, segment identifier and key-value pair naming convention; generate a globally unique identifier for each normalized data, attach collection metadata and bind it to the data content for storage, and generate a standardized irregular dataset.
[0014] Furthermore, the specific implementation process of step S2 is as follows: A dual-branch collaborative semantic extraction model is constructed. The model is based on a domain-adaptive fine-tuned pre-trained language model as its backbone, with a shared underlying semantic encoding layer and mutually independent entity relation extraction and semantic feature extraction branches. Standardized data is input into the shared semantic encoding layer to generate a contextual semantic encoding matrix. The entity relation extraction branch extracts entities, attributes, and relationships based on the encoding matrix, generating <entity-relationship-entity> feature triples and removing duplicates. The semantic feature extraction branch generates fixed-dimensional semantic representation vectors based on the encoding matrix. The outputs of the two branches are cross-validated, erroneous and redundant triples are removed, and the core features of the semantic vectors are strengthened. Finally, the validated feature triples, semantic representation vectors, and unique identifiers of the corresponding data are bound and stored.
[0015] Furthermore, in step S3, during the autonomous update of the ontology of the dynamic knowledge graph, the compatibility judgment between newly added entities and existing ontology concepts is calculated using the following formula: ,in The comprehensive fit between the newly added entity and the target ontology concept, with a value range of [0, 1]; For the newly added entity semantic vector semantic vector of the ontology concept center Cosine similarity; For the newly added entity association set Set of entities associated with the ontology concept Matching density; To add entity time series heat Baseline of popularity of ontology concept The goodness of fit; , , For the weighting coefficients, satisfying The initial value is , , The system dynamically adjusts the classification accuracy feedback based on each round of ontology updates; the fit threshold is set to 0.75, and when the fit between the new entity and all existing ontology concepts is lower than the threshold, the new ontology concept generation process is automatically triggered to complete the unsupervised update of the ontology architecture and classification system.
[0016] Furthermore, in step S4, the neighborhood radius of the unsupervised adaptive clustering is dynamically adjusted using the following calculation formula: ,in The adaptively adjusted clustering neighborhood radius; The base neighborhood radius is set to an initial value of 0.2; The dynamic weight of the corresponding central node of the ontology has a value range of [0, 1] and is calculated by comprehensively considering the number of entities, access frequency, and number of associations under the ontology. The semantic dispersion of the target dataset, with a value range of [0.5, 2], is calculated from the standard deviation of the data semantic vectors; The local normalized density of the target dataset, with a value range of [0.1, 10], is calculated by normalizing the number of data points in a unit semantic space. After adaptive clustering based on the adjusted neighborhood radius, a three-level hierarchical label system is generated for each cluster, and the label is bound to the unique identifier of the corresponding data and stored.
[0017] Furthermore, the specific implementation process of step S5 is as follows: A multi-level archiving directory architecture is built based on a hierarchical tagging system, with the first-level ontology concept tag as the root directory for archiving, the second-level entity attribute tags as the second-level subdirectories, and the third-level core semantic tags as the prefix for archiving file identification; data is archived to the corresponding directory according to the bound tags, and for cross-topic data, a soft link mapping method is used for multi-directory archiving, storing only one original file; a dual-structure intelligent index library is constructed for the archived data, which combines an inverted index and a semantic vector index. The inverted index uses entities, tags, and core keywords as index items and adopts incremental updates, while the semantic vector index uses data semantic representation vectors as index items and adopts batch updates; a unified retrieval interface and a disaster recovery backup mechanism are set up for the index library.
[0018] Furthermore, in the semantic retrieval process of step S6, the comprehensive ranking of the retrieval results is calculated using the following formula: ,in The overall ranking score of the search results, with a value range of [0, 1]; The semantic vector matching degree between the retrieval statement and the archived data; To determine the correlation between the retrieved entity and the archived data entity; The business value of archived data is calculated by comprehensively considering the length of the data association link, the frequency of reuse, and the degree of business matching. Normalize the historical access frequency of archived data; , , , For the weighting coefficients, satisfying The initial value is , , , The algorithm dynamically adjusts based on user search click feedback. During retrieval, semantic parsing and entity extraction are completed first, and an initial result set is obtained through index matching. Then, the ranking score is calculated using the above formula, and the results are ranked. Simultaneously, related archived data is pushed based on the knowledge graph.
[0019] Furthermore, the specific implementation process of the full-link traceability in step S6 is as follows: In the full-process data processing, a corresponding operation log is generated for each step of each data operation, the log is strongly bound to the unique data identifier, and stored in the corresponding data node of the knowledge graph; a full-process directed acyclic traceability link from collection to archiving is constructed for each data node, as well as the association link with other entities and ontology nodes; when a traceability request is received, the corresponding node is located through the unique data identifier, the full amount of operation logs and link information are retrieved, a visualized full-link traceability graph and a standardized traceability audit report are generated, and an early warning is automatically triggered for abnormal operations found in the traceability.
[0020] Furthermore, the method also includes a privacy protection mechanism covering the entire data lifecycle. The specific implementation process is as follows: In the data collection and preprocessing stage, sensitive data is identified and marked through a preliminary screening module, and a dedicated processing channel is set up; in the feature extraction stage, sensitive semantic units are accurately located and divided into high, medium, and low levels according to leakage risk, and anti-tampering semantic digital watermarks bound to the core semantics are embedded for different levels of sensitive units; in the clustering and archiving stage, adaptive differential privacy desensitization processing is performed on data of different sensitivity levels, injecting corresponding level noise only into sensitive units while preserving the core semantics and relationships of the data; in the retrieval and tracing stage, hierarchical access control for sensitive data is set up, all sensitive data access operations are recorded, access logs are generated, and the data is incorporated into the full-link tracing system.
[0021] Furthermore, the method also includes a dynamic knowledge graph self-optimization and anomaly handling closed-loop mechanism. The specific implementation process is as follows: a monthly optimization window is set to evaluate the effectiveness of all ontology concept nodes, invalid ontology nodes with no matching data for two consecutive periods are taken offline, and redundant classification entries are cleaned up; the effectiveness of all associations in the knowledge graph is verified, low-intensity invalid associations are removed, and high-intensity potential associations are completed; an abnormal data isolation area is set to isolate and store abnormal data with extremely low adaptability and semantic contradictions; a fixed-period secondary analysis window is set to restore the normal processing flow for abnormal data that can be matched with the updated ontology, and to trigger a new ontology generation evaluation for abnormal data that cannot be matched continuously; a full backup and version rollback mechanism for the knowledge graph is set up, and a backup is generated after each ontology update, and a rollback to a stable version is quickly performed when an anomaly occurs.
[0022] Compared with existing technologies, this data processing method for collecting, analyzing, and organizing irregular data has the following advantages:
[0023] I. This invention constructs an initial knowledge graph based on extracted feature triples and can autonomously update the knowledge graph ontology architecture, automatically iterate the classification system, and complete the association relationships according to the semantic features and entity distribution of newly added data. This feature enables the knowledge graph to continuously evolve and keep pace with the data environment. At the same time, by using the ontology architecture of the dynamic knowledge graph as a classification benchmark, it achieves unsupervised adaptive clustering of irregular data and automatically generates a hierarchical label system. Combined with the construction of an intelligent index library, it supports users to complete accurate semantic retrieval of archived data through natural language, greatly improving the convenience and accuracy of data retrieval.
[0024] Second, by building a data acquisition gateway that supports multi-protocol access, this invention can efficiently collect multi-source heterogeneous irregular data and perform preprocessing steps such as noise reduction and format normalization to generate a standardized irregular dataset. Subsequently, a pre-trained semantic extraction model is used to automatically extract entities, attributes, relationships, and core semantic features from the data, generating feature triples and semantic representation vectors. This process not only significantly reduces manual intervention but also improves the accuracy and efficiency of data processing, laying a solid foundation for subsequent knowledge graph construction and data analysis.
[0025] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0027] Figure 1 This is a flowchart illustrating the workflow of a data processing method for collecting, analyzing, and organizing irregular data.
[0028] Figure 2 This is a flowchart of a two-branch collaborative automatic extraction process for data entities and semantic features in a data processing method for irregular data collection, analysis, and organization.
[0029] Figure 3 This is a flowchart of a dynamic knowledge graph ontology self-updating and adaptive data clustering process for data processing methods targeting irregular data collection, analysis, and organization. Detailed Implementation
[0030] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0031] Example 1: Step 1, as follows Figure 1 As shown, this involves the collection and standardized preprocessing of irregular data. A multi-protocol collection gateway supporting HTTP, FTP, WebSocket, and dedicated protocols for the government intranet was built. Dedicated collection channels were set up for irregular data from different sources, such as government approvals, hotline petitions, public services, and grid events, enabling parallel collection of data across departments and systems. The collected raw government data underwent coarse filtering to remove invalid null values, duplicate and redundant reported data, invalid data with corrupted files, and illegal data that did not conform to government data specifications. The filtered valid government data underwent full format normalization, unifying the data encoding format, timestamp standard, data segmentation identifier, and key-value pair naming convention, eliminating format barriers between different departments. A globally unique government data identifier was generated for each normalized government data entry, and collection metadata such as collection source, collection time, and responsible department were attached. The metadata was then strongly bound to the data content and stored to generate a standardized irregular government dataset to be processed.
[0032] The second step, as Figure 2As shown, data entities and semantic features are automatically extracted. A dual-branch collaborative semantic extraction model is constructed. The model is based on a pre-trained language model that is adaptively fine-tuned in the government domain. It sets up a shared underlying semantic encoding layer and two independent branches: entity relationship extraction branch and semantic feature extraction branch. Each data point in the standardized government dataset is input into the shared semantic encoding layer to generate a contextual semantic encoding matrix for the corresponding data. Based on the generated semantic encoding matrix, the entity relationship extraction branch extracts entities, attributes, and relationships such as the handling subject, application object, matter type, and handling result from the government data, generating <entity-relationship-entity> feature triples and performing deduplication. Based on the semantic encoding matrix, the semantic feature extraction branch generates a fixed-dimensional semantic representation vector for each government data point. The outputs of the two branches are cross-validated to remove erroneous and redundant feature triples, while strengthening the core semantic features of the semantic vectors. Finally, the validated feature triples and semantic representation vectors are bound and stored with the globally unique identifier of the corresponding government data.
[0033] The third step, as Figure 3 As shown, a dynamic knowledge graph can be constructed and its ontology updated autonomously. Based on the extracted feature triples, an initial knowledge graph for the government service domain is constructed. Simultaneously, based on the semantic features and entity distribution of newly added government data, the knowledge graph ontology architecture is dynamically updated autonomously, the government classification system is automatically iterated, and entity relationships are completed without the need for manual pre-setting of ontology rules and classification systems. During the autonomous update of the dynamic knowledge graph ontology, the compatibility between new entities and existing ontology concepts is assessed using a comprehensive compatibility calculation formula. Based on the assessment results, it is determined whether the new entity should be included in the existing ontology concept. If the compatibility does not meet the preset threshold, the automatic generation process of a new ontology concept is triggered, simultaneously completing the automatic adjustment of the ontology level and the iterative optimization of the classification system. At the same time, the potential relationships are intelligently completed based on entity association features.
[0034] The fourth step involves adaptive clustering and label generation of irregular data based on a dynamic knowledge graph. Using the dynamically updated ontology architecture of the government service knowledge graph as the classification benchmark, and based on the semantic representation vectors and entity relationships of the government data, unsupervised adaptive clustering of irregular government data is completed. During the clustering process, an adaptive clustering neighborhood radius calculation formula is used to optimize the neighborhood radius parameter in real time, adapting to the data distribution characteristics under different ontology classifications. After clustering, a hierarchical labeling system is automatically generated based on the core semantics and entity attributes of each cluster. The labeling system covers first-level ontology concept labels, second-level entity attribute labels, and third-level core semantic labels, achieving multi-dimensional labeling representation of government data.
[0035] The fifth step involves hierarchical archiving and intelligent index construction for unstructured data. A multi-level archiving directory architecture for government data is built based on an automatically generated hierarchical tagging system. The first-level ontology concept tags serve as the root directory, second-level entity attribute tags as second-level subdirectories, and third-level core semantic tags as the identifier prefix for archived files. Government data is archived to the corresponding directories according to the bound hierarchical tags. For cross-departmental and cross-topic government data, soft link mapping is used to achieve multi-directory archiving, storing only one original data file to avoid redundant data storage. For archived government data, a dual-structure intelligent index library is constructed, combining an inverted index and a semantic vector index. The inverted index uses government entities, hierarchical tags, and core keywords as index items and adopts an incremental update mode, while the semantic vector index uses the semantic representation vector of government data as index items and adopts a batch update mode. A unified government data retrieval interface and a multi-copy disaster recovery backup mechanism are set up for the intelligent index library to ensure its stable operation.
[0036] Step 6: End-to-End Traceability and Semantic Retrieval of Archived Data. Based on the entity association links of a dynamic knowledge graph, the entire lifecycle of archived government data—from collection and transfer to modification and archiving—is traceable. During the entire data processing flow, corresponding operation logs are generated for each step of each piece of government data. These logs are strongly bound to the globally unique identifier of the data and stored in the corresponding data node of the knowledge graph. A directed acyclic traceability link from collection to archiving is constructed for each data node, along with association links with other entities and ontology nodes. Upon receiving a government data traceability request, the corresponding node is located using the globally unique identifier of the data. All operation logs and link information are retrieved, generating a visualized end-to-end traceability graph and a standardized traceability audit report. Simultaneously, automatic alerts are triggered for abnormal operations discovered during traceability. In the semantic retrieval stage, government staff can perform accurate semantic retrieval of archived data using natural language. For the comprehensive ranking of search results, a quantitative scoring formula is used to calculate the comprehensive ranking. Search results are then sorted from highest to lowest based on the comprehensive score, ensuring a high degree of matching between search results and search requirements.
[0037] This embodiment simultaneously implements a privacy protection mechanism covering the entire data lifecycle. In the data collection and preprocessing stage, a preliminary screening module identifies and marks sensitive data such as personal identity information and corporate trade secrets in government data, and sets up a dedicated processing channel. In the feature extraction stage, sensitive semantic units are accurately located and divided into three levels—high, medium, and low—according to leakage risk. Anti-tampering semantic digital watermarks bound to the core semantics are embedded for sensitive units of different levels. In the clustering and archiving stage, adaptive differential privacy desensitization processing is performed on government data of different sensitivity levels, injecting corresponding level noise only into sensitive units while preserving the core semantics and relationships of government data. In the retrieval and tracing stage, hierarchical access control for sensitive data is set up, all access operations for sensitive data are recorded, access logs are generated, and the data is incorporated into the full-chain tracing system. Simultaneously, a dynamic knowledge graph self-optimization and anomaly handling closed-loop mechanism is set up. A monthly optimization window is set up to evaluate the effectiveness of all ontology concept nodes of the government service knowledge graph, remove invalid ontology nodes with no matching data for two consecutive periods, and clean up redundant classification entries; the effectiveness of all associations in the knowledge graph is verified, low-intensity invalid associations are removed, and high-intensity potential associations are added; an abnormal data isolation zone is set up to isolate and store abnormal government data with extremely low adaptability and semantic contradictions; a fixed-period secondary analysis window is set up to restore the normal processing flow for abnormal data that can be matched with the updated ontology, and trigger the generation and evaluation of a new ontology for abnormal data that cannot be matched; a full backup and version rollback mechanism for the knowledge graph is set up to generate a backup after each ontology update, and quickly roll back to a stable version when an anomaly occurs.
[0038] Example 2: This example applies to the full-process processing of irregular operational data from industrial equipment in a smart factory for discrete manufacturing. The data to be processed in this scenario includes real-time sensor data from over 300 CNC machining, testing, and logistics equipment within the factory; unstructured inspection records from equipment maintenance personnel; equipment fault reporting and repair text data; semi-structured inspection reports from product quality inspection processes; and fragmented operational log data from the production system. The data types include time-series fragmented data, unstructured text, semi-structured forms, equipment operation logs, and other types of irregular and heterogeneous data. The specific implementation process is as follows:
[0039] First step, such as Figure 1As shown, this involves the acquisition and standardized preprocessing of irregular data. A multi-protocol acquisition gateway supporting Modbus, OPCUA, MQTT industrial protocols and HTTP and FTP general protocols was built. Dedicated acquisition channels were set up for irregular data from different sources, such as CNC equipment operation data, equipment inspection records, fault repair data, product quality inspection data, and production system logs, enabling real-time parallel acquisition of industrial data across devices and systems. The acquired raw industrial data underwent coarse filtering to remove invalid null values, duplicate and redundant data, damaged data generated by offline equipment, and illegal data that did not conform to industrial data specifications. The filtered valid industrial data underwent full format normalization to unify the data encoding format, timestamp standard, data segmentation identifier, and key-value pair naming convention, eliminating format barriers between data from different devices and systems. A globally unique industrial data identifier was generated for each normalized industrial data, and acquisition metadata such as acquisition device number, acquisition time, production station, and production line were attached. The metadata was strongly bound to the data content and stored to generate a standardized irregular industrial dataset to be processed.
[0040] The second step, as Figure 2 As shown, data entities and semantic features are automatically extracted. A dual-branch collaborative semantic extraction model is constructed. The model is based on a pre-trained language model that is adaptively fine-tuned in the industrial manufacturing field. It sets up a shared underlying semantic encoding layer and two independent branches: entity relation extraction branch and semantic feature extraction branch. Each data point in the standardized industrial dataset is input into the shared semantic encoding layer to generate a contextual semantic encoding matrix for the corresponding data. Based on the generated semantic encoding matrix, the entity relation extraction branch extracts entities, attributes, and relationships such as equipment number, workstation information, fault type, maintenance personnel, and quality inspection results from the industrial data, generating <entity-relationship-entity> feature triples and performing deduplication. Based on the semantic encoding matrix, the semantic feature extraction branch generates a fixed-dimensional semantic representation vector for each piece of industrial data. The outputs of the two branches are cross-validated to remove erroneous and redundant feature triples, while strengthening the core semantic features of the semantic vectors. Finally, the validated feature triples and semantic representation vectors are bound and stored with the globally unique identifier of the corresponding industrial data.
[0041] The third step, as Figure 3As shown, a dynamic knowledge graph can be constructed and its ontology updated autonomously. Based on the extracted feature triples, an initial knowledge graph for the industrial equipment operation and maintenance domain is constructed. Simultaneously, based on the semantic features and entity distribution of newly added industrial data, the knowledge graph ontology architecture is dynamically updated autonomously, the industrial equipment classification system is automatically iterated, and entity association relationships are completed without the need for manual pre-setting of ontology rules and classification systems. During the autonomous update of the dynamic knowledge graph ontology, the compatibility judgment between new entities and existing ontology concepts is performed. A comprehensive compatibility calculation formula between new entities and target ontology concepts is used to complete the quantitative evaluation. Based on the evaluation results, it is determined whether the new entity should be included in the existing ontology concept. If the compatibility does not meet the preset threshold, the automatic generation process of new ontology concepts is triggered, and the automatic adjustment of the ontology level and the iterative optimization of the classification system are completed simultaneously. At the same time, based on entity association features, the potential association relationships between equipment faults, operation and maintenance actions, and quality inspection results are intelligently completed.
[0042] The fourth step involves adaptive clustering and label generation of irregular data based on a dynamic knowledge graph. Using the dynamically updated ontology architecture of the industrial equipment knowledge graph as the classification benchmark, and based on the semantic representation vectors and entity relationships of the industrial data, unsupervised adaptive clustering of irregular industrial data is completed. During the clustering process, an adaptive clustering neighborhood radius calculation formula is used to optimize the neighborhood radius parameter in real time, adapting to the distribution characteristics of industrial data under different ontology classifications. After clustering, a hierarchical labeling system is automatically generated based on the core semantics and entity attributes of each cluster. This labeling system covers first-level ontology concept labels, second-level entity attribute labels, and third-level core semantic labels, achieving multi-dimensional tagged representation of industrial equipment data.
[0043] The fifth step is the hierarchical archiving and intelligent indexing of irregular data. A multi-level archiving directory architecture for industrial data is built based on an automatically generated hierarchical tagging system. The first-level ontology concept tags serve as the root directory, second-level entity attribute tags as second-level subdirectories, and third-level core semantic tags as the identifier prefix for archived files. Industrial data is archived to the corresponding directories according to the bound hierarchical tags. For industrial data spanning multiple devices, processes, and themes, a soft link mapping method is used to achieve multi-directory archiving, storing only one original data file to avoid redundant data storage. For the archived industrial data, a dual-structure intelligent index library is constructed, combining an inverted index and a semantic vector index. The inverted index uses industrial device entities, hierarchical tags, and core keywords as index items and adopts an incremental update mode, while the semantic vector index uses the semantic representation vectors of industrial data as index items and adopts a batch update mode. A unified industrial data retrieval interface and a multi-copy disaster recovery backup mechanism are set up for the intelligent index library to ensure its stable operation in industrial production environments.
[0044] Step 6: Full-chain traceability and semantic retrieval of archived data. Based on a dynamic knowledge graph-based entity association chain, this system achieves full-chain traceability of archived industrial data throughout its entire lifecycle, from collection and transfer to modification and archiving. During the entire data processing flow, a corresponding operation log is generated for each step of each industrial data operation. These logs are strongly bound to the data's globally unique identifier and stored in the corresponding data node of the knowledge graph. This constructs a directed, acyclic traceability chain for each data node, from collection to archiving, as well as association chains with other equipment entities and entity nodes. Upon receiving an industrial data traceability request, the system locates the corresponding node using the data's globally unique identifier, retrieves all operation logs and chain information, and generates a visualized full-chain traceability graph and a standardized traceability audit report. Simultaneously, it automatically triggers alerts for abnormal data operations and equipment parameter tampering discovered during traceability. In the semantic retrieval stage, it supports factory maintenance, process, and quality inspection personnel in performing accurate semantic retrieval of archived data using natural language. For the comprehensive ranking of search results, a quantitative scoring formula is used to quantify the results, and the results are output in descending order of comprehensive score, ensuring a high degree of matching between search results and search requirements.
[0045] This embodiment simultaneously implements a privacy protection mechanism covering the entire data lifecycle. In the data acquisition and preprocessing stage, a preliminary screening module identifies and marks sensitive data such as core process parameters, technical secrets, and production and operation data in industrial data, and sets up a dedicated processing channel. In the feature extraction stage, sensitive semantic units are accurately located and divided into three levels—high, medium, and low—according to leakage risk. Anti-tampering semantic digital watermarks bound to the core semantics are embedded for sensitive units of different levels. In the clustering and archiving stage, adaptive differential privacy desensitization processing is performed on industrial data of different sensitivity levels, injecting corresponding level noise only into sensitive units while preserving the core semantics and relationships of the industrial data. In the retrieval and traceability stage, hierarchical access control for sensitive data is set up, all access operations for sensitive data are recorded, access logs are generated, and the data is incorporated into the full-chain traceability system. Simultaneously, a dynamic knowledge graph self-optimization and anomaly handling closed-loop mechanism is set up. A monthly optimization window is set up to evaluate the effectiveness of all ontology concept nodes of the industrial equipment operation and maintenance knowledge graph, and invalid ontology nodes with no matching data for two consecutive periods are taken offline, and redundant classification entries are cleaned up. The effectiveness of all associations in the knowledge graph is verified, low-intensity invalid associations are removed, and high-intensity potential associations are added. An abnormal data isolation area is set up to isolate and store abnormal industrial data with extremely low adaptability and semantic contradictions. A fixed-period secondary analysis window is set up to restore the normal processing flow for abnormal data that can be matched with the updated ontology, and to trigger the generation and evaluation of a new ontology for abnormal data that cannot be matched. A full backup and version rollback mechanism for the knowledge graph is set up. After each ontology update, a backup is generated, and in case of an anomaly, it is quickly rolled back to a stable version.
[0046] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A data processing method for collecting, analyzing, and organizing irregular data, characterized in that, The specific steps of this method are as follows: S1. Irregular data acquisition and standardization preprocessing: Collect multi-source heterogeneous irregular data, perform noise reduction and format normalization preprocessing on the collected irregular data to obtain a standardized irregular dataset to be processed. S2. Automatic extraction of data entities and semantic features: Based on a pre-trained semantic extraction model, the entity, attribute, relationship and core semantic features of each data in the standardized irregular dataset are automatically extracted to generate feature triples and semantic representation vectors for each data. S3. Autonomous Evolutionary Dynamic Knowledge Graph Construction and Ontology Update: The initial knowledge graph is constructed based on the extracted feature triples. At the same time, based on the semantic features and entity distribution of the newly added data, the knowledge graph ontology architecture is dynamically updated, the classification system is automatically iterated, and the relationship is completed autonomously, without the need for manual preset ontology rules and classification system. S4. Adaptive Clustering and Tag Generation of Irregular Data Based on Dynamic Knowledge Graph: Using the ontology architecture of dynamic knowledge graph as the classification benchmark, and based on the semantic representation vector of data and entity association, unsupervised adaptive clustering of irregular data is completed. At the same time, based on the core semantics and entity attributes of the clusters, a hierarchical tag system for the corresponding data is automatically generated. S5. Hierarchical archiving and intelligent index construction of irregular data: Based on the automatically generated hierarchical tag system and entity association relationship of knowledge graph, multi-dimensional hierarchical archiving of irregular data is completed. At the same time, combined with data semantic features, entity attributes and archiving path, an intelligent index library for semantic retrieval is constructed. S6. Full-chain traceability and semantic retrieval of archived data: Based on the association links of dynamic knowledge graph, the entire lifecycle of archived data, including its source, flow, modification, and archiving, is traceable. At the same time, based on the constructed intelligent index library, users can complete accurate semantic retrieval of archived data through natural language.
2. The data processing method for irregular data collection, analysis, and organization according to claim 1, characterized in that, The specific implementation process of step S1 is as follows: build a collection gateway that supports multi-protocol access, and set up dedicated collection channels for irregular data from different sources; perform coarse filtering on the collected raw data to remove invalid, redundant, damaged and illegal data; perform full format normalization on the filtered valid data, and unify the encoding, timestamp, segment identifier and key-value pair naming convention; generate a globally unique identifier for each normalized data, attach collection metadata and bind it to the data content for storage, and generate a standardized irregular dataset.
3. The data processing method for irregular data collection, analysis, and organization according to claim 1, characterized in that, The specific implementation process of step S2 is as follows: build a dual-branch collaborative semantic extraction model. The model is based on a domain-adaptive fine-tuned pre-trained language model as its backbone, and sets up a shared underlying semantic encoding layer, as well as mutually independent entity relationship extraction branches and semantic feature extraction branches. Standardized data is input into a shared semantic coding layer to generate a contextual semantic coding matrix; The entity-relation extraction branch extracts entities, attributes, and relationships based on the encoding matrix, generating <entity-relation-entity> feature triples and removing duplicates; The semantic feature extraction branch generates a fixed-dimensional semantic representation vector based on the encoding matrix; the outputs of the two branches are cross-validated to remove erroneous and redundant triples and strengthen the core features of the semantic vector; finally, the validated feature triples, semantic representation vectors and corresponding data unique identifiers are bound and stored.
4. The data processing method for irregular data collection, analysis, and organization according to claim 1, characterized in that, In step S3, during the autonomous update of the ontology of the dynamic knowledge graph, the compatibility judgment between newly added entities and existing ontology concepts is calculated using the following formula: ,in To ensure comprehensive compatibility between the newly added entities and the target ontology concept; For the newly added entity semantic vector semantic vector of the ontology concept center Cosine similarity; For the newly added entity association set Set of entities associated with the ontology concept Matching density; To add entity time series heat Baseline of popularity of ontology concept The goodness of fit; , , These are the weighting coefficients.
5. The data processing method for irregular data collection, analysis, and organization according to claim 1, characterized in that, In step S4, the neighborhood radius of the unsupervised adaptive clustering is dynamically adjusted using the following formula: ,in The adaptively adjusted clustering neighborhood radius; The basic neighborhood radius; The dynamic weight corresponding to the central node of the ontology; The semantic dispersion of the target data set; This represents the local normalized density of the target data set.
6. The data processing method for irregular data collection, analysis, and organization according to claim 1, characterized in that, The specific implementation process of step S5 is as follows: A multi-level archive directory architecture is built based on a hierarchical tag system, with the first-level ontology concept tag as the archive root directory, the second-level entity attribute tag as the second-level subdirectory, and the third-level core semantic tag as the archive file identifier prefix; data is archived to the corresponding directory according to the bound tags, and for cross-topic data, a soft link mapping method is used for multi-directory archiving, storing only one original file; a dual-structure intelligent index library is constructed for the archived data, with inverted index and semantic vector index working together. The inverted index uses entities, tags, and core keywords as index items and adopts incremental updates, while the semantic vector index uses data semantic representation vectors as index items and adopts batch updates; a unified retrieval interface and disaster recovery backup mechanism are set up for the index library.
7. The data processing method for irregular data collection, analysis, and organization according to claim 1, characterized in that, In the semantic retrieval process of step S6, the comprehensive ranking of the retrieval results is calculated using the following formula: ,in The overall ranking score of the search results, with a value range of [0, 1]; The semantic vector matching degree between the retrieval statement and the archived data; To determine the correlation between the retrieved entity and the archived data entity; To assess the business value of archived data; Normalize the historical access frequency of archived data; , , , These are the weighting coefficients.
8. The data processing method for irregular data collection, analysis, and organization according to claim 1, characterized in that, The specific implementation process of the full-link traceability in step S6 is as follows: In the full-process data processing, a corresponding operation log is generated for each step of each data operation. The log is strongly bound to the unique data identifier and stored in the corresponding data node of the knowledge graph. A full-process directed acyclic traceability link from collection to archiving is constructed for each data node, as well as the association link with other entities and ontology nodes. When a traceability request is received, the corresponding node is located through the unique data identifier, the full operation log and link information are retrieved, a visualized full-link traceability graph and a standardized traceability audit report are generated, and an early warning is automatically triggered for abnormal operations found in the traceability.
9. A data processing method for collecting, analyzing, and organizing irregular data according to claim 1, characterized in that, The method also includes a privacy protection mechanism covering the entire data lifecycle. The specific implementation process is as follows: In the data collection and preprocessing stage, sensitive data is identified and marked through a preliminary screening module, and a dedicated processing channel is set up; In the feature extraction stage, sensitive semantic units are accurately located and divided into three levels—high, medium, and low—according to the risk of leakage, and anti-tampering semantic digital watermarks bound to the core semantics are embedded for sensitive units of different levels; In the clustering and archiving stage, adaptive differential privacy desensitization processing is performed on data of different sensitivity levels, injecting corresponding level noise only into sensitive units, while preserving the core semantics and correlation of the data. In the retrieval and tracing process, sensitive data access permissions are controlled at different levels. All access operations to sensitive data are recorded, access logs are generated, and the data is incorporated into the full-chain tracing system.
10. A data processing method for collecting, analyzing, and organizing irregular data according to claim 1, characterized in that, The method also includes a dynamic knowledge graph self-optimization and anomaly handling closed-loop mechanism. The specific implementation process is as follows: a monthly optimization window is set to evaluate the effectiveness of all ontology concept nodes, invalid ontology nodes with no matching data for two consecutive periods are taken offline, and redundant classification entries are cleaned up; the effectiveness of all associations in the knowledge graph is verified, low-intensity invalid associations are removed, and high-intensity potential associations are completed; an abnormal data isolation area is set to isolate and store abnormal data with extremely low adaptability and semantic contradictions; a fixed-period secondary analysis window is set to restore the normal processing flow for abnormal data that can be matched with the updated ontology, and to trigger the generation and evaluation of a new ontology for abnormal data that cannot be matched continuously; a full backup and version rollback mechanism for the knowledge graph is set up, and a backup is generated after each ontology update, and a rollback to a stable version is quickly performed when an anomaly occurs.