Classification tagging management method and platform for full life cycle data

By embedding feature-aware nodes and performing cross-stage correlation analysis in the entire life cycle of the data, the adaptive classification label set is generated, and the problem of lack of coherent monitoring between different life cycle stages is solved, data consistency and comparability are achieved, and data management efficiency and quality are improved.

CN120217054APending Publication Date: 2025-06-27LINGSHU TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510347456.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The lack of coherent monitoring mechanisms in the prior art between different life cycle stages leads to inconsistent metadata and poor comparability, and the level of data quality and data security management is uneven.

Method used

By embedding feature perception nodes in the entire life cycle of the data, obtaining feature perception data for each stage, determining time-sensitive metadata, and conducting cross-stage correlation analysis, setting classification label mapping rules under the time constraint layer, combining data semantic similarity and context dependence relationship for label conflict dissolution analysis, and generating an adaptive classification label set.

Benefits of technology

The consistency and comparability of data between different stages are realized, and data management is refined through multimodal tag classification, errors and inconsistencies in the data are identified and corrected, data quality is improved, data retrieval and efficient utilization is realized, and data management efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217054A_ABST
    Figure CN120217054A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field related to data management, in particular to a full-life-cycle data classification tagging management method and platform, and the method comprises the steps: embedding a feature sensing node in a full life cycle of data, obtaining feature data of each stage, setting a classification tag mapping rule under aging constraint, and resolving tag conflicts. According to the method, the technical problems of metadata inconsistency, poor comparability and uneven data quality and data security management level caused by lack of a coherent monitoring mechanism among different life cycle stages are solved, feature sensing nodes are embedded in the full life cycle of data, and the security of the data is improved. According to the method, the consistency and comparability of the data in different stages are ensured, fine data management is carried out through multi-modal label classification, errors and inconsistency in the data are recognized and corrected, the data quality is improved, and the technical effects that the data are quickly retrieved and efficiently utilized, and the data management efficiency is improved are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and particularly relates to a classification and tagging management method and platform for full-life cycle data. Background Art

[0002] Data has become one of the core assets of enterprises, running through all aspects of enterprise operation, decision-making, innovation, etc. The scale and complexity of data are constantly increasing. The wide application of advanced technologies such as big data, cloud computing, and artificial intelligence has brought the collection, storage, processing, and analysis capabilities of data to an unprecedented level. In the current data technology era, the data full life cycle refers to the entire process from data generation, storage, integration, presentation and use, analysis and application to final archiving and destruction. During this process, data goes through different stages, and each stage has its specific management requirements and technical challenges. At the same time, the current data full life cycle management still faces many challenges. On the one hand, the sharp growth of data volume has brought huge pressure to data storage and processing; on the other hand, the issues of data security and privacy protection are becoming increasingly prominent. In addition, the data standards, data quality, and data security management levels vary widely among different industries and enterprises, which also restricts the effective implementation of data full life cycle management.

[0003] In summary, there are technical problems in the prior art that there is a lack of a coherent monitoring mechanism between different life cycle stages, resulting in inconsistent metadata and poor comparability, and uneven data quality and data security management levels. Summary of the Invention

[0004] The present application provides a classification and tagging management platform for full-life cycle data, aiming to solve the technical problems in the prior art that there is a lack of a coherent monitoring mechanism between different life cycle stages, resulting in inconsistent metadata and poor comparability, and uneven data quality and data security management levels.

[0005] In view of the above problems, the technical solution of the present application is as follows:

[0006] On the one hand, the present application provides a method for classifying and labeling the management of full - life - cycle data. The method includes: embedding feature - sensing nodes in the full life cycle of data, obtaining feature - sensing data in the data generation stage, data transfer stage, data storage stage, and data destruction stage, and determining time - sensitive metadata; performing cross - stage correlation analysis on the time - sensitive metadata, setting classification label mapping rules under the time - effect constraint layer, and performing label conflict resolution analysis by combining data semantic similarity and context - dependence relationship to generate an adaptive classification label set; identifying abnormal classification nodes and confidence deviation values in the full life cycle of data according to the classification label mapping rules and the adaptive classification label set; configuring a label management instruction set including label weight adjustment parameters, classification rule iteration strategies, and life - cycle stage remapping schemes through the abnormal classification nodes and confidence deviation values; and using the label management instruction set to perform multi - modal label classification management in the full life cycle of data.

[0007] On the other hand, the present application provides a platform for classifying and labeling the management of full - life - cycle data. The platform includes: a data acquisition module for embedding feature - sensing nodes in the full life cycle of data, obtaining feature - sensing data in the data generation stage, data transfer stage, data storage stage, and data destruction stage, and determining time - sensitive metadata; a correlation analysis module for performing cross - stage correlation analysis on the time - sensitive metadata, setting classification label mapping rules under the time - effect constraint layer, and performing label conflict resolution analysis by combining data semantic similarity and context - dependence relationship to generate an adaptive classification label set; an abnormal identification module for identifying abnormal classification nodes and confidence deviation values in the full life cycle of data according to the classification label mapping rules and the adaptive classification label set; an instruction configuration module for configuring a label management instruction set including label weight adjustment parameters, classification rule iteration strategies, and life - cycle stage remapping schemes through the abnormal classification nodes and confidence deviation values; and a label classification management module for using the label management instruction set to perform multi - modal label classification management in the full life cycle of data.

[0008] In summary, one or more technical solutions provided in the present application achieve the technical effects of embedding feature - sensing nodes in the full life cycle of data, ensuring the consistency and comparability of data among different stages, performing refined data management through multi - modal label classification, identifying and correcting errors and inconsistencies in data, improving data quality, realizing rapid retrieval and efficient utilization of data, and improving data management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a schematic flow chart of a method for classifying and labeling the management of full - life - cycle data provided by the present application;

[0010] Figure 2 This application provides a structural schematic diagram of a classification and tagging management platform for full - life - cycle data.

[0011] Explanation of the reference numerals in the drawings: data acquisition module M100, correlation analysis module M200, anomaly recognition module M300, instruction configuration module M400, tag classification management module M500. Detailed implementation manners

[0012] Embodiment 1

[0013] The following describes the present application in detail with reference to the drawings. As Figure 1 shown, the present application provides a classification and tagging management method for full - life - cycle data. Among them, the method includes:

[0014] S1: Embed feature - sensing nodes in the full life cycle of data, obtain feature - sensing data in the data generation stage, data transfer stage, data storage stage, and data destruction stage, and determine time - sensitive metadata; S2: Perform cross - stage correlation analysis on the time - sensitive metadata, set classification tag mapping rules under the time - effect constraint layer, and perform tag conflict resolution analysis by combining data semantic similarity and context - dependence relationship to generate an adaptive classification tag set.

[0015] Specifically, embedding feature - sensing nodes in the full life cycle of data means setting points that can sense and obtain data features at different stages of data. These points can monitor and collect feature information of data at each stage in real time; feature - sensing data refers to the information obtained by these nodes that can reflect data features, such as data type, format, content, etc.; time - sensitive metadata refers to metadata that may affect data management and utilization over time, such as data creation time, modification time, validity period, etc.

[0016] Perform cross-phase correlation analysis on the time-sensitive metadata. Cross-phase correlation analysis refers to correlating and analyzing the time-sensitive metadata in different lifecycle phases to discover the internal connections and impacts between them; the classification label mapping rules under the timeliness constraint layer refer to the rules for establishing the mapping relationship between classification labels and data characteristics under the constraint conditions considering data timeliness; data semantic similarity refers to the degree of similarity in semantics between different data, and the relevance between data is judged by analyzing the data semantics; context dependency refers to the dependency relationship of data in different context environments, that is, the mutual association and impact of data in different scenarios and backgrounds; label conflict resolution analysis refers to when there are conflicts or overlaps between different labels, using certain analysis methods to resolve these conflicts to ensure the accuracy and consistency of labels; the adaptive classification label set refers to a set of classification labels that can automatically adjust and adapt according to data changes and lifecycle evolution.

[0017] Embed feature perception nodes in the entire data lifecycle, obtain feature perception data in the stages of data generation, transfer, storage, and destruction, determine the time-sensitive metadata, and further deploy feature perception nodes in each stage of the entire data lifecycle to comprehensively collect feature information of data in different stages, so as to accurately determine the time-sensitive metadata and provide a basis for subsequent classification label management. For example, in the data generation stage, through the feature perception node, feature information such as the type, format, and content of the data can be obtained, and then time-sensitive metadata such as the creation time of the data can be determined; in the data transfer stage, the feature perception node can monitor the changes in the data during the transmission process, such as the modification time and transfer path of the data, and further enrich the time-sensitive metadata.

[0018] Perform cross - stage correlation analysis on time - sensitive metadata, set the classification label mapping rules under the time - limit constraint layer, and conduct label conflict resolution analysis by combining data semantic similarity and context - dependence relationship to generate an adaptive classification label set. Further, through cross - stage correlation analysis, integrate and analyze time - sensitive metadata at different stages to construct a timeliness change model of data throughout its life cycle; set classification label mapping rules under the time - limit constraint layer to ensure that the classification labels can accurately reflect the timeliness characteristics of the data. For example, for a type of data with a short validity period, set the classification label mapping rules so that it is marked as valid data within the validity period and automatically marked as expired data after the expiration date. At the same time, conduct label conflict resolution analysis by combining data semantic similarity and context - dependence relationship to solve the conflicts and overlaps between different labels. For example, when two data are semantically highly similar but are assigned different labels, adjust the label assignment through semantic similarity analysis and context - dependence relationship judgment to generate an adaptive classification label set, improving the accuracy and consistency of the labels. By comprehensively considering time - sensitive metadata, semantic similarity, and context - dependence relationship, the dynamic generation and optimization of classification labels are achieved, improving the efficiency and quality of data classification management.

[0019] S3: Identify abnormal classification nodes and confidence deviation values in the entire life cycle of the data according to the classification label mapping rules and the adaptive classification label set; S4: Configure a label management instruction set including label weight adjustment parameters, classification rule iteration strategies, and life - cycle stage remapping schemes through the abnormal classification nodes and confidence deviation values; S5: Use the label management instruction set to perform multi - modal label classification management in the entire life cycle of the data.

[0020] Specifically, an abnormal classification node refers to a data node that does not meet the expected classification criteria judged according to the classification label mapping rules and the adaptive classification label set in the entire life cycle of the data; the confidence deviation value refers to the degree of deviation between the actual classification result and the expected classification result, which is used to measure the reliability of classification; the label weight adjustment parameter refers to a parameter used to adjust the label weight, and by adjusting these parameters, the importance of different labels in classification can be changed; the classification rule iteration strategy refers to a method of continuously optimizing and updating the classification rules to improve the accuracy and adaptability of classification; the life - cycle stage remapping scheme refers to a scheme for re - dividing and mapping different stages of the data life cycle to better adapt to data changes and management requirements; the label management instruction set refers to a set of instructions containing a series of label management operations, which is used to guide the specific implementation of multi - modal label classification management; multi - modal label classification management refers to comprehensively using multiple modes and methods to classify and manage labels to improve the efficiency and effect of management.

[0021] According to the classification label mapping rules and the adaptive classification label set, identify the abnormal classification nodes and confidence deviation values in the full life cycle of the data. Further, through the application of the classification label mapping rules and the adaptive classification label set, conduct classification judgments on each node in the full life cycle of the data, find the abnormal nodes that do not meet the expected classification criteria, and calculate their confidence deviation values to quantify the reliability of the classification. For example, in the data storage stage, according to the classification label mapping rules, a certain type of data should be marked as high priority, but the actual classification result is low priority, then this node is identified as an abnormal classification node, and its confidence deviation value can be calculated by comparing the probability distributions of the actual classification result and the expected classification result.

[0022] Based on the abnormal classification nodes and confidence deviation values, configure a label management instruction set that includes label weight adjustment parameters, classification rule iteration strategies, and life cycle stage remapping schemes. Further analyze the abnormal classification nodes and confidence deviation values to determine the aspects that need to be adjusted and optimized in label management. If it is found that the classification accuracy of certain labels is low in a specific stage, the weights of the corresponding labels can be adjusted to increase their influence in classification. At the same time, update the classification rules and adopt more advanced algorithms or models to improve the classification accuracy. In addition, remap the life cycle stages, merge or subdivide some stages to better adapt to the characteristics of the data and management requirements; integrate these adjustment and optimization measures into the label management instruction set to form a complete label management solution.

[0023] Use the label management instruction set to conduct multi-modal label classification management in the full life cycle of the data. Further apply the configured label management instruction set to each stage of the full life cycle of the data. Through the comprehensive use of various modes and methods to classify and manage the labels. In the data generation stage, according to the label management instruction set, use an automated label generation tool to quickly and accurately label the newly generated data with classification labels; in the data transfer stage, use a real-time monitoring system combined with the rules in the label management instruction set to dynamically adjust and update the labels during the data transfer process to ensure the accuracy and timeliness of the labels; in the data storage and destruction stages, also based on the label management instruction set, classify and store the data and safely destroy it to improve the efficiency and quality of data management. Through multi-modal label classification management, the refined management of the full life cycle of the data is realized, and the availability and security of the data are improved.

[0024] Furthermore, feature perception nodes are embedded in the full life cycle of the data. The method of this application includes:

[0025] Deploy a distributed feature collector in the data generation stage, and lay a protocol parsing probe in the transfer channel, and configure a semantic analysis agent at the storage node;

[0026] Meanwhile, a differential acquisition strategy is set, where the sampling frequency in the data generation stage ≥ 100 Hz, the protocol parsing depth in the data transfer stage ≥ 7 layers, and the semantic annotation accuracy in the data storage stage ≥ 98%;

[0027] Based on the differential acquisition strategy, timestamp chain anchoring is adopted for multi-stage traceability verification.

[0028] Specifically, in the data generation stage, a distributed feature collector is deployed. The distributed feature collector refers to a device or software module that can collect data features set at different locations and nodes during data generation, distributed throughout the data generation environment, and working collaboratively to comprehensively obtain the feature information of the data; arranging protocol parsing probes in the transfer channel means setting probe devices that can parse data transfer protocols in the data transmission channel. These probes can intercept and analyze the protocols followed by the data during transmission and extract useful information from them; configuring semantic analysis agents at storage nodes means setting agent programs that can perform semantic analysis at the data storage locations. These agent programs perform semantic-level analysis on the stored data and extract the semantic features and meanings of the data.

[0029] Meanwhile, a differential acquisition strategy is set. The differential acquisition strategy refers to formulating different data acquisition methods and parameters according to the data characteristics and management requirements in different stages to achieve precise data acquisition and efficient management. Preferably, the sampling frequency in the data generation stage ≥ 100 Hz means that in the data generation stage, at least 100 data samples are collected per second to ensure that the subtle changes and features of the data can be captured; the protocol parsing depth in the data transfer stage ≥ 7 layers means that in the data transfer stage, the protocol parsing probe parses the data transfer protocol to a depth of at least 7 layers, capable of comprehensively parsing each layer of the protocol and extracting rich protocol information; the semantic annotation accuracy in the data storage stage ≥ 98% means that in the data storage stage, the semantic analysis agent has an accuracy rate of at least 98% for semantic annotation of the data, ensuring high-precision extraction and annotation of semantic information.

[0030] Based on the differential acquisition strategy, timestamp chain anchoring is adopted for multi-stage traceability verification. Timestamp chain anchoring refers to using timestamps to link and fix data acquisition events in different stages in chronological order to form a continuous time chain. Each data acquisition event is anchored to the previous and subsequent events through timestamps to ensure the continuity and traceability of data in different stages; multi-stage traceability verification refers to analyzing and verifying the data acquisition events anchored by timestamp chains to trace and verify the flow and changes of data in all stages of the entire life cycle, ensuring the integrity and credibility of the data.

[0031] In the data generation stage, a distributed feature collector is deployed. The distributed feature collector is configured at each key node where data is generated, such as sensors, user terminals, business systems, etc., to collect the feature information of the data in real time, such as the type, format, content, generation time, etc. of the data. In an Internet of Things application scenario, sensors distributed at different locations act as distributed feature collectors, and collect the feature information of data such as temperature, humidity, pressure, etc. at a sampling frequency of 100Hz in real time to ensure that the subtle changes in the data can be captured in a timely manner.

[0032] In the data transfer channel, protocol parsing probes are deployed. The protocol parsing probes can intercept the protocol data packets used in the data transmission process and perform in-depth parsing. In a network data transmission scenario within an enterprise, the protocol parsing probes can parse the transport layer, session layer, presentation layer, and application layer, and extract the transmission path, transmission time, and transmission protocol type of the data, providing a basis for subsequent data transfer analysis and management; at the data storage node, a semantic analysis agent is configured. The semantic analysis agent program performs semantic analysis on the stored data and extracts the semantic features and meanings of the data. In a database storage scenario, the semantic analysis agent analyzes the stored text data, identifies semantic information such as key entities, concepts, and relationships therein, and performs semantic annotation with an accuracy of 98%, providing support for the semantic understanding and application of the data.

[0033] At the same time, a differential collection strategy is set. According to the data characteristics and management requirements in different stages, corresponding collection parameters are formulated. In the data generation stage, real-time and fine-grained collection of data features is ensured through a high sampling frequency (≥100Hz); in the data transfer stage, comprehensive protocol information is obtained through in-depth protocol parsing (≥7 layers); in the data storage stage, accurate extraction of semantic information is guaranteed through high-precision semantic annotation (≥98%). The differential collection strategy can effectively improve the efficiency and quality of data collection and meet the data management requirements in different stages.

[0034] Based on the differential collection strategy, timestamp chain anchoring is adopted to link and fix the data collection events in different stages in chronological order. Each data collection event records its accurate timestamp, and these timestamps are connected in sequence through a chain structure to form an immutable time chain. Thus, when multi-stage traceability verification is required, by analyzing the data collection events anchored by the timestamp chain, the transfer path and changes of the data in the entire life cycle can be traced. Specifically, in the data audit process, through timestamp chain anchoring, the time and content of each link of the data from generation to storage can be accurately verified, ensuring the integrity and credibility of the data and timely discovering and correcting possible problems in the data transfer process.

[0035] By deploying corresponding acquisition and analysis devices respectively in the data generation, circulation, and storage stages, and combining with differentiated acquisition strategies, it is possible to comprehensively and accurately obtain the characteristics, protocols, and semantic information of the data. At the same time, the timestamp chain anchoring provides a reliable basis for data traceability and verification, ensuring the traceability and credibility of the data throughout its life cycle.

[0036] Furthermore, for cross-stage correlation analysis of the time-sensitive metadata and setting the classification label mapping rules under the timeliness constraint layer, the method of the present application further includes:

[0037] Extract the timeliness characteristics in the time-sensitive metadata and construct a timeliness constraint matrix;

[0038] Based on the time-sensitive metadata, perform cross-stage data association path modeling, generate a weighted semantic dependency graph, and evaluate the timeliness compliance score in combination with the timeliness constraint matrix;

[0039] According to the comparison result between the timeliness compliance score and the preset score threshold, determine the priority adjustment coefficient of the classification label mapping rules, and generate an optimal cross-stage label mapping path.

[0040] Specifically, extracting the timeliness characteristics in the time-sensitive metadata, the timeliness characteristics refer to the characteristics that can reflect the timeliness of the data, such as the validity period of the data, update frequency, creation time, etc.; the timeliness constraint matrix refers to a matrix structure used to describe the timeliness constraints of the data in different stages, which contains various timeliness constraint conditions and parameters; based on the time-sensitive metadata, perform cross-stage data association path modeling, and cross-stage data association path modeling refers to modeling and analyzing the association paths between the data in different life cycle stages to reveal the flow and change rules of the data throughout its life cycle; the weighted semantic dependency graph refers to assigning weights to the edges in the semantic dependency graph to represent the semantic dependency strength between different data nodes; the timeliness compliance score refers to the score obtained by quantitatively evaluating the timeliness compliance of the data according to the timeliness constraint matrix and the data association path model, which is used to measure the degree of compliance of the data in terms of timeliness.

[0041] Determine the priority adjustment coefficient of the classification label mapping rule according to the comparison result between the timeliness compliance score and the preset score threshold, and generate an optimal label mapping path across stages; the preset score threshold refers to the standard value preset for judging whether the timeliness compliance score is qualified; the priority adjustment coefficient refers to the coefficient for adjusting the priority of the classification label mapping rule according to the comparison result between the timeliness compliance score and the threshold, which is used to determine the priority order of different classification label mapping rules when applied; the optimal label mapping path refers to the label mapping path that best meets the timeliness requirements and data association relationships by adjusting the priority of the classification label mapping rule considering timeliness compliance.

[0042] Extract the timeliness features in the time-sensitive metadata, construct a timeliness constraint matrix, and further analyze the time-sensitive metadata to extract its timeliness features, such as the validity period, update frequency, creation time, modification time, etc. of the data. In the process of financial data management, the creation time, update time, and validity period of transaction data are important timeliness features; construct a timeliness constraint matrix based on these timeliness features. The matrix contains the constraint conditions for data timeliness at different stages. For example, the validity period of data in the storage stage is 30 days, and the update frequency in the transfer stage is not less than once a day, etc. The construction of the timeliness constraint matrix provides a quantitative standard for subsequent timeliness compliance evaluation.

[0043] Based on the time-sensitive metadata, perform cross-stage data association path modeling to generate a weighted semantic dependency graph, and evaluate the timeliness compliance score in combination with the timeliness constraint matrix. Further, use the time-sensitive metadata to model the association path between data at different stages, analyze the transfer path and dependency relationship of data from generation to transfer, storage, destruction, etc. The transfer process can be clearly depicted through data association path modeling; generate a weighted semantic dependency graph according to the semantic dependency relationship between data, where the weight represents the intensity of semantic dependency; evaluate the timeliness compliance of data in combination with the timeliness constraint matrix to assess whether the data meets the timeliness requirements throughout the life cycle. In the process of financial data management, if a certain transaction data exceeds the validity period in the storage stage or the update frequency in the transfer stage is lower than the requirement, its timeliness compliance score will decrease.

[0044] According to the comparison result between the timeliness compliance score and the preset score threshold, determine the priority adjustment coefficient of the classification label mapping rule, and generate the optimal label mapping path across stages. Compare the timeliness compliance score with the preset score threshold. If the score is lower than the threshold, it indicates that there is a problem with the timeliness of the data, and the classification label mapping rule needs to be adjusted. The preset score threshold is 80 points, and the timeliness compliance score of a certain piece of data is 70 points, then the priority of the classification label mapping rule needs to be adjusted; determine the priority adjustment coefficient according to the comparison result, which is used to adjust the priority order of different classification label mapping rules, so that the rules that better meet the timeliness requirements are applied first, generate the optimal label mapping path across stages, ensure that the classification label mapping of the data in the whole life cycle not only conforms to the semantic dependency relationship, but also meets the timeliness requirements, ensure that the classification labels of the product data are accurate at different stages, and improve the efficiency and quality of data management.

[0045] Furthermore, the method of the present application further includes:

[0046] The timeliness constraint matrix includes the data survival period, the inter-stage transfer delay threshold, and the time decay factor;

[0047] For label nodes with semantic overlap and label nodes with timeliness conflicts, introduce an adaptive sliding window mechanism to correct the context dependency relationship, and update the weight assignment of the semantic dependency graph;

[0048] Perform a secondary fusion analysis on the updated semantic dependency graph and the timeliness constraint matrix, and output an adaptive classification label set after conflict resolution.

[0049] Specifically, the timeliness constraint matrix includes the data survival period, the inter-stage transfer delay threshold, and the time decay factor; the data survival period refers to the total time length from the generation to the invalidation of the data, and different types of data have different survival periods; the inter-stage transfer delay threshold refers to the maximum allowable delay time when the data is transferred between different life cycle stages, and exceeding this threshold will affect the timeliness and validity of the data; the time decay factor refers to a parameter used to measure the gradual decrease in value or validity of the data over time, reflecting the change rate of the data timeliness.

[0050] For tag nodes with semantic overlap and tag nodes with timeliness conflicts, an adaptive sliding window mechanism is introduced to correct the context dependence relationship, and the weight assignment of the semantic dependence graph is updated. Semantically overlapping tag nodes refer to tag nodes that have semantic overlap or similarity, resulting in classification ambiguity or conflict; timeliness-conflicting tag nodes refer to tag nodes that have conflicts in terms of timeliness. For example, one tag requires immediate data update, while another tag allows a certain delay in data; the adaptive sliding window mechanism is a mechanism that dynamically adjusts the window size and can automatically adjust the window range according to the real-time situation of the data and the context dependence relationship to better capture and process the association between data; context dependence relationship correction refers to adjusting and correcting the semantic dependence relationship between tag nodes according to the context environment and the dependence relationship between data to improve the accuracy and reliability of the semantic dependence graph; weight assignment refers to assigning weights to different edges in the semantic dependence graph to represent the semantic dependence strength between different data nodes.

[0051] The updated semantic dependence graph and the timeliness constraint matrix are subjected to secondary fusion analysis to output an adaptive classification tag set after conflict resolution; secondary fusion analysis refers to the re-fusion and comprehensive analysis of the updated semantic dependence graph and the timeliness constraint matrix. By combining semantic information and timeliness constraints, the generation and assignment of classification tags are further optimized; the adaptive classification tag set after conflict resolution refers to the classification tag set that has solved the tag conflict problem and can adapt to data changes after semantic dependence relationship correction and timeliness constraint fusion analysis.

[0052] The timeliness constraint matrix includes the data survival period, the inter-stage transfer delay threshold, and the time decay factor. When constructing the timeliness constraint matrix, the data survival period is determined. For example, for financial transaction data with high real-time requirements, its survival period is only a few days, while for some historical archive data, the survival period is up to several years; the inter-stage transfer delay threshold is set. For example, during the process of data transfer from the generation stage to the storage stage, the maximum allowable delay time is 5 minutes, and if it exceeds this time, it is regarded as a delay anomaly; the time decay factor is defined. Preferably, for news data, the time decay factor can be set to reduce the data weight by 10% every day, indicating that the value and validity of news data gradually decrease at a rate of 10% every day over time.

[0053] For tag nodes with semantic overlap and time - effect conflicts, an adaptive sliding window mechanism is introduced to correct the context - dependent relationship and update the weight assignment of the semantic dependency graph; in the actual data management scenario, when encountering tag nodes with semantic overlap and time - effect conflicts, in the e - commerce data management scenario, the two tag nodes of electronic products and digital products have semantic overlap, while the two tag nodes of limited - time offers and long - term promotions may have time - effect conflicts. Based on this, an adaptive sliding window mechanism is introduced. According to the real - time update situation and context - dependent relationship of the data, the size and range of the window are dynamically adjusted. When it is detected that the semantic overlap degree of electronic products and digital products within a certain time window is relatively high, by adjusting the window size, the analysis range of relevant data is expanded, and the semantic dependency relationship between these two tag nodes is corrected by combining context information (such as product descriptions, category attributes, etc.). At the same time, according to the corrected semantic dependency relationship, the weight assignment of the semantic dependency graph is updated. For example, the weight of the edge between electronic products and digital products is reduced to reduce the classification ambiguity caused by semantic overlap; for tag nodes with time - effect conflicts, such as limited - time offers and long - term promotions, the time - effect performance of tag nodes with time - effect conflicts at different stages is monitored in real - time through the adaptive sliding window mechanism, and the semantic dependency weights between tag nodes with time - effect conflicts are adjusted by combining the time decay factor in the time - effect constraint matrix, so that the semantic dependency graph can more accurately reflect the actual relationship between data.

[0054] The updated semantic dependency graph and the time - effect constraint matrix are subjected to secondary fusion analysis to output an adaptive classification tag set after conflict resolution. During the secondary fusion analysis process, the semantic information in the updated semantic dependency graph and the time - effect constraint conditions in the time - effect constraint matrix are comprehensively considered. For a set of tag nodes with both semantic overlap and time - effect conflicts, through fusion analysis, according to the weight assignment in the semantic dependency graph and parameters such as the data survival period and the transfer delay threshold between stages in the time - effect constraint matrix, the comprehensive score of each tag node is calculated; the tag nodes are sorted and screened according to the comprehensive score, and an adaptive classification tag set after conflict resolution is output. Specifically, the adaptive classification tag set will preferentially select tags that conform to both semantic dependency relationships and time - effect constraints to ensure the classification accuracy and timeliness of the data and improve the support efficiency and quality of decision - making.

[0055] Furthermore, updating the weight assignment of the semantic dependency graph, the method of this application includes:

[0056] Establish a grid classification architecture, the input end receives the dynamic correlation degree of the context - dependent relationship, and the output end generates a weight assignment tensor;

[0057] At the same time, an attention mechanism is set to determine the association strength score of data nodes at different stages, and dynamic weight correction is performed by combining the time decay factor in the time - effect constraint matrix.

[0058] Specifically, a grid classification architecture is established. The input end receives the dynamic correlation degree of the context-dependency relationship, and the output end generates a weight assignment tensor. The grid classification architecture refers to a classification model based on a grid structure that can perform classification processing on multi-dimensional data. The nodes and edges of the grid can represent the features of the data and the relationships between them; the dynamic correlation degree of the context-dependency relationship refers to the degree of change in the dependency relationship between data nodes in different context environments, reflecting the change of the semantic association strength between data nodes with the context; the weight assignment tensor refers to the weight assignment result represented in the form of a tensor. A tensor is a multi-dimensional array structure that can store and represent complex weight assignment information.

[0059] Meanwhile, an attention mechanism is set up to determine the association strength scores of data nodes at different stages and perform dynamic weight correction in combination with the time decay factor in the time effect constraint matrix. The attention mechanism refers to a mechanism that simulates attention and can automatically adjust the degree of attention to different data according to the importance and relevance of the data, making the model more focused on important data features; the association strength score refers to a score that measures the degree of association between data nodes at different stages. The higher the score, the stronger the association; the time decay factor refers to a parameter used to measure the gradual decrease in value or effectiveness of data over time, reflecting the change rate of data timeliness; dynamic weight correction refers to dynamically adjusting and optimizing the weight assignment according to the real-time situation of the data and relevant parameters to improve the adaptability and accuracy of the model.

[0060] A grid classification architecture is established. The input end receives the dynamic correlation degree of the context-dependency relationship, and the output end generates a weight assignment tensor; when establishing the grid classification architecture, the structure and parameters of the grid are designed according to the feature dimensions and classification requirements of the data; the input end receives the dynamic correlation degree of the context-dependency relationship, and these dynamic correlation degrees can be obtained by statistically calculating the occurrence frequency, co-occurrence relationship, etc. of the data in different contexts. Some words have a higher co-occurrence frequency under specific labels, and correspondingly, the dynamic correlation degree between these words is higher; the grid classification architecture maps and calculates features in the grid according to the input dynamic correlation degree and generates a weight assignment tensor at the output end; the weight assignment tensor stores the weight assignment situation between different feature dimensions in the form of a multi-dimensional array. In a three-dimensional grid, each element of the weight assignment tensor represents the weight value corresponding to the combination of feature dimensions, reflecting the relative importance of different features in classification.

[0061] Meanwhile, an attention mechanism is set up to determine the association strength scores of data nodes at different stages, and dynamic weight correction is carried out in combination with the time decay factor in the timeliness constraint matrix. During the data classification process, the attention mechanism automatically calculates the association strength scores of data nodes at different stages according to the characteristics and context information of the data nodes. In the video data classification scenario, the association strength scores between data nodes such as video frame images, audio, and subtitles at different stages (such as the beginning, middle, and end of the video) can be obtained by analyzing their semantic relevance, temporal continuity, etc.; in combination with the time decay factor in the timeliness constraint matrix, the weights are dynamically corrected. Further, for news video data with strong timeliness, the time decay factor is large, and as time passes after the video is released, its weight will gradually decrease. For some classic movie video data, the time decay factor is small, and the change in weight is relatively slow. Thus, the grid classification architecture can more accurately reflect the true association between data nodes, improving the accuracy and efficiency of classification.

[0062] Furthermore, to establish a grid classification architecture, the method of this application further includes:

[0063] Based on the weight assignment tensor, tensor splicing is performed to configure a cross-stage spatial feature pattern;

[0064] Based on the cross-stage spatial feature pattern, a bidirectional memory pointer is used to capture the temporal dependence relationship to generate a spatio-temporal fusion weight assignment coefficient;

[0065] According to the weight assignment coefficient, the edge weights in the semantic dependency graph are iteratively optimized, and the model parameters of the grid classification architecture are updated.

[0066] Specifically, based on the weight assignment tensor, tensor splicing is performed to configure a cross-stage spatial feature pattern; tensor splicing refers to connecting and combining multiple tensors according to certain rules and dimensions to form a larger or more complex tensor structure for integrating and utilizing more feature information; the cross-stage spatial feature pattern refers to the spatial feature distribution and combination method that spans different stages of the data life cycle. By integrating the spatial features of different stages, the changes and laws of data in the spatial dimension can be captured more comprehensively.

[0067] Based on the cross-stage spatial feature pattern, a bidirectional memory pointer is used to capture temporal dependencies and generate a spatio-temporal fusion weight allocation coefficient. The bidirectional memory pointer is a mechanism that can capture dependencies in a data sequence both forward and backward. When processing sequence data, it can consider not only the influence of previous data on the current data but also the feedback of subsequent data on the current data, thus more accurately modeling the temporal dependencies of the data. The spatio-temporal fusion weight allocation coefficient refers to the coefficient generated by combining the spatial feature pattern and temporal dependencies, which can reflect both spatial features and temporal changes and is used to more comprehensively describe the association strength between data nodes.

[0068] According to the weight allocation coefficient, the edge weights in the semantic dependency graph are iteratively optimized, and the model parameters of the grid classification architecture are updated. Iterative optimization refers to the process of gradually improving and optimizing the weight allocation coefficient and model parameters through multiple cycles of calculation and adjustment to achieve better performance and accuracy. Model parameters refer to various parameters used to define the behavior and performance of the grid classification architecture, such as weights and biases in a neural network. By optimizing the weight allocation coefficient to update these parameters, the model's ability to classify and understand data can be improved.

[0069] Based on the weight allocation tensor, tensor concatenation is performed to configure the cross-stage spatial feature pattern. The weight allocation tensor contains the weight information between data nodes in different stages. In a scenario including four stages of data generation, transfer, storage, and destruction, there is a corresponding weight allocation tensor for each stage, representing the association strength between data nodes within that stage. Through tensor concatenation, the weight allocation tensors of these four stages are connected in sequence according to the stage order to form a cross-stage spatial feature pattern tensor. The cross-stage spatial feature pattern tensor can integrate the spatial feature information of different stages. Through the concatenation operation, these spatial features of different stages are integrated together, providing a comprehensive spatial feature basis for subsequent capture of temporal dependencies.

[0070] Based on the cross-stage spatial feature pattern, a bidirectional memory pointer is used to capture the temporal dependence relationship, and a spatio-temporal fusion weight allocation coefficient is generated; when capturing the temporal dependence relationship, the bidirectional memory pointer can simultaneously consider the association of data nodes in the past and future stages. For a data node created in the data generation stage, its weight allocation in the transfer stage is not only affected by the generation stage, but may also be affected by the feedback of subsequent storage and destruction stages; the bidirectional memory pointer scans the sequence of data nodes forward and backward to calculate the dependence strength of each node at different times. In time series data, the value of a certain data point is not only related to the previous data points, but may also be affected by some subsequent data points. The bidirectional memory pointer can capture this bidirectional dependence relationship; according to the captured temporal dependence relationship and the original cross-stage spatial feature pattern, a spatio-temporal fusion weight allocation coefficient is generated.

[0071] According to the weight allocation coefficient, the edge weights in the semantic dependency graph are iteratively optimized, and the model parameters of the grid classification architecture are updated. During the iterative optimization process, the generated spatio-temporal fusion weight allocation coefficient is applied to the semantic dependency graph to adjust the edge weights. If the weight of an edge increases under the new weight allocation coefficient, it indicates that the association between the two data nodes connected by this edge becomes more important from the perspective of spatio-temporal fusion. Therefore, its weight is increased in the semantic dependency graph; conversely, if the weight decreases, the association strength of this edge is weakened. Through multiple iterative optimizations, the edge weights of the semantic dependency graph are gradually adjusted to make it more in line with the actual association of the data.

[0072] At the same time, according to the optimized weight allocation coefficient and the semantic dependency graph, the model parameters of the grid classification architecture are updated. In the grid classification architecture, some parameters determine the attention degree of the model to different spatial features. The weight allocation coefficient obtained through iterative optimization can guide the adjustment of these parameters, enabling the model to better classify and manage data, playing a role in deepening data understanding and improving model performance; by tensor splicing to integrate the cross-stage spatial feature pattern, using a bidirectional memory pointer to capture the temporal dependence relationship, and generating a spatio-temporal fusion weight allocation coefficient to more comprehensively depict the complex associations between data nodes; iteratively optimizing the edge weights of the semantic dependency graph and updating the model parameters of the grid classification architecture can continuously improve the accuracy and efficiency of data classification, and achieve refined management and dynamic optimization of the entire life cycle of data.

[0073] Furthermore, the method of the present application further includes:

[0074] When the standard deviation of the weight allocation coefficient exceeds the preset fluctuation threshold, the window expansion operation of the adaptive sliding window mechanism is triggered, and the dynamic association degree of the context dependency relationship is recalculated.

[0075] Specifically, when the standard deviation of the weight distribution coefficient exceeds the preset fluctuation threshold, the window expansion operation of the adaptive sliding window mechanism is triggered, and the dynamic correlation degree of the context dependence relationship is recalculated; the standard deviation of the weight distribution coefficient refers to the degree of dispersion of each value in the set of weight distribution coefficients from its average value. The larger the standard deviation, the greater the fluctuation of the weight distribution coefficient; the preset fluctuation threshold is a standard value preset for judging whether the fluctuation of the weight distribution coefficient is abnormal. When the standard deviation exceeds this threshold, it is considered that the weight distribution coefficient has an abnormal fluctuation; the adaptive sliding window mechanism is a mechanism that can automatically adjust the window size according to the real-time situation of the data, and is used to dynamically adjust the scope and focus of data processing; the window expansion operation refers to expanding the scope of the window on the basis of the current window in order to include more data samples, so as to more comprehensively analyze the context dependence relationship of the data; the dynamic correlation degree of the context dependence relationship refers to the degree of change of the dependence relationship between data nodes in different context environments, reflecting the change of the semantic association strength between data nodes with the change of context.

[0076] When the standard deviation of the weight distribution coefficient exceeds the preset fluctuation threshold, the window expansion operation of the adaptive sliding window mechanism is triggered, and the dynamic correlation degree of the context dependence relationship is recalculated. During the process of data classification and management, the fluctuation of the weight distribution coefficient can reflect the stability of the association relationship between data nodes. At this time, it is detected that the standard deviation exceeds the preset fluctuation threshold, and the window expansion operation of the adaptive sliding window mechanism is triggered. The window expansion operation will expand the window range of data processing, incorporating more historical data and related data nodes into the analysis scope; by recalculating the dynamic correlation degree of the context dependence relationship, the association changes between data nodes in a wider context can be captured more accurately. Further, some data nodes with relatively weak original associations will have enhanced associations due to the influence of common factors. By expanding the window range and recalculating the dynamic correlation degree, these changes can be discovered in a timely manner, providing a more accurate basis for subsequent classification label adjustment.

[0077] By monitoring the fluctuation of the weight distribution coefficient, the window expansion operation of the adaptive sliding window mechanism is triggered in a timely manner. In the face of sudden changes or abnormal fluctuations in the data, the analysis scope is automatically expanded, and the context dependence relationship between data nodes is re-evaluated, so as to more accurately reflect the true association situation of the data, providing reliable support for the dynamic adjustment of classification labels and improving the flexibility and accuracy of data management.

[0078] Furthermore, the method of the present application further includes:

[0079] The association strength score

[0080] wherein, S ij represents the association strength between data node i and node j, Qi and K j are the query vector of node i and the key vector of node j respectively, with a dimension of d, and Δt ij is the time interval between node i and node j, and λ is the time decay factor defined in the time effect constraint matrix, which is used to control the time decay rate.

[0081] Specifically, in the process of data classification and management, in order to accurately measure the association degree between different data nodes, a calculation method for the association strength score is introduced. Specifically, for the association strength score between two data nodes i and j where Q i and K j are the query vector of node i and the key vector of node j respectively, both with a dimension of d, and Δt ij is the time interval (unit: second) between node i and node j. In time series data, it is the difference in timestamps between two data points; λ is the time decay factor (unit: s -1 ) defined in the time effect constraint matrix, which is used to control the time decay rate, making the association strength of recent data nodes higher, meeting the requirements of data timeliness (the larger it is, the stronger the timeliness of the data, and the faster the association strength decays over time); the Softmax function normalizes the attention weights to ensure ∑ j S ij = 1.

[0082] The calculation of the association strength score provides a quantitative basis for the association analysis between data nodes. By comprehensively considering the semantic similarity and timeliness of the data, it accurately quantifies the association strength between data nodes, provides a theoretical basis for resolving classification label conflicts, and significantly improves the timeliness and accuracy of data classification under the full life cycle data management.

[0083] Furthermore, in the multi-modal label classification management of the data full life cycle, the method of the present application further includes:

[0084] Statistical distribution characteristics of the confidence deviation value to generate a stability heat map;

[0085] Establish a conflict judgment criterion driven by the stability heat map, set a label conflict database, and use association rules to locate abnormal classification nodes;

[0086] Activate a label purification instruction according to the cosine similarity corresponding to the abnormal classification node, which is used to trigger the iteration of the multi-modal label classification rule.

[0087] Specifically, the distribution characteristics of the confidence deviation value are statistically analyzed to generate a stability heat map; the confidence deviation value refers to the degree of deviation between the actual classification result and the expected classification result, which is used to measure the reliability of classification; the distribution characteristics refer to the distribution of the confidence deviation value under different data nodes, different stages or different classification labels, such as statistical indicators like mean, variance, standard deviation, etc.; the stability heat map is a visualization tool that intuitively displays the data stability distribution through the depth of color, and the darker the color, the lower the stability, that is, the larger the confidence deviation value.

[0088] The conflict judgment criteria driven by the stability heat map are established, a label conflict database is set up, and association rules are used to locate abnormal classification nodes; the conflict judgment criteria refer to the criteria for judging whether there are conflicts in data classification based on the distribution characteristics of the confidence deviation value in the stability heat map. The label conflict database is a database used to store information related to label conflicts, including the type, location, frequency, etc. of conflicts; the association rule refers to a rule for discovering the association relationships existing between different data nodes through data analysis and mining, which is used to locate abnormal classification nodes.

[0089] According to the cosine similarity corresponding to the abnormal classification node, a label purification instruction is activated, which is used to trigger the iteration of the multi-modal label classification rule; the cosine similarity refers to the cosine value of the included angle calculated between two vectors, which is used to measure the similarity between them, and the closer the value is to 1, the higher the similarity; the label purification instruction refers to an instruction used to start the label purification process, thereby eliminating or reducing label conflicts and abnormal classification nodes; the iteration of the multi-modal label classification rule refers to the process of continuously optimizing and updating the multi-modal label classification rule to improve the accuracy and reliability of classification.

[0090] The distribution characteristics of the confidence deviation value are statistically analyzed to generate a stability heat map. During the data classification process, the confidence deviation value reflects the reliability of the classification result; by collecting the confidence deviation values of a large number of data nodes, calculating their distribution characteristics, such as mean, variance, standard deviation, etc., to understand the overall stability of the classification result, and generating a stability heat map according to the distribution characteristics, representing the confidence deviation values of different image regions with the depth of color, intuitively displaying the stability of the classification result, and the darker the color, the larger the confidence deviation value and the more unstable the classification result.

[0091] Establish the conflict judgment criteria driven by the stability heat map, set up a label conflict database, and use association rules to locate abnormal classification nodes; according to the distribution of confidence deviation values in the stability heat map, formulate conflict judgment criteria. For example, when the confidence deviation value exceeds 0.2, it is considered that there is a label conflict. At the same time, establish a label conflict database to record the detailed information of each conflict, such as the conflicting image pairs, conflicting classification labels, etc.; use the association rule mining algorithm to analyze the data in the label conflict database, discover the association rules between different data nodes. For example, when a cat and a dog appear simultaneously in an image and there is a label conflict, according to the association rules, locate the abnormal classification nodes, providing a basis for subsequent label purification.

[0092] Activate the label purification instruction according to the cosine similarity corresponding to the abnormal classification node, which is used to trigger the iteration of the multi-modal label classification rule; calculate the cosine similarity between the abnormal classification node and other nodes to judge the similarity degree between the abnormal classification node and other nodes. For example, for an abnormal classification node A, calculate its cosine similarities with nodes B, C, and D to be 0.8, 0.6, and 0.4 respectively; according to the preset similarity threshold of 0.5, determine the nodes B and C with higher similarity to node A, activate the label purification instruction, and perform label purification operations on these similar nodes, such as relabeling or adjusting the classification labels. At the same time, use the purified labels as feedback to trigger the iterative update of the multi-modal label classification rule. In the next round of classification, add processing rules for the situation where a cat and a dog appear together to improve the accuracy and stability of classification.

[0093] By statistically analyzing the distribution characteristics of confidence deviation values and generating a stability heat map, the stability of data classification is demonstrated; establishing conflict judgment criteria and a label conflict database can systematically manage and analyze label conflict problems; using association rules to locate abnormal classification nodes and activating label purification instructions according to cosine similarities can achieve fine-tuning and optimization of classification results, trigger the iteration of multi-modal label classification rules, continuously improve the performance of the classification model, enhance the accuracy and reliability of data classification, and provide a high-quality basis for further analysis and application of data.

[0094] In summary, the beneficial effects of the embodiments of this application are:

[0095] By embedding feature perception nodes in the entire data life cycle, obtaining feature perception data at each stage, determining time-sensitive metadata, conducting cross-stage correlation analysis, setting classification label mapping rules under the time limit constraint layer, and performing label conflict resolution analysis by combining data semantic similarity and context dependency relationship to generate an adaptive classification label set; according to the classification label mapping rules and the adaptive classification label set, identifying abnormal classification nodes and confidence deviation values in the entire data life cycle, configuring a label management instruction set including label weight adjustment parameters, classification rule iteration strategies, and life cycle stage remapping schemes, and performing multi-modal label classification management in the entire data life cycle, the present application provides a classification label management method and platform for the entire life cycle data, realizes the embedding of feature perception nodes in the entire data life cycle, ensures the consistency and comparability of data among different stages, and conducts refined data management through multi-modal label classification, identifies and corrects errors and inconsistencies in the data, improves data quality, realizes rapid retrieval and efficient utilization of data, and improves the technical effect of data management efficiency.

[0096] Embodiment 2

[0097] Based on the same inventive concept as a classification label management method for the entire life cycle data in the foregoing embodiment, as Figure 2 shown, the embodiment of the present application provides a classification label management platform for the entire life cycle data, wherein the platform includes:

[0098] A data acquisition module M100, configured to embed feature perception nodes in the entire data life cycle, obtain feature perception data in the data generation stage, data transfer stage, data storage stage, and data destruction stage, and determine time-sensitive metadata.

[0099] A correlation analysis module M200, configured to conduct cross-stage correlation analysis on the time-sensitive metadata, set classification label mapping rules under the time limit constraint layer, and perform label conflict resolution analysis by combining data semantic similarity and context dependency relationship to generate an adaptive classification label set.

[0100] An abnormality identification module M300, configured to identify abnormal classification nodes and confidence deviation values in the entire data life cycle according to the classification label mapping rules and the adaptive classification label set.

[0101] An instruction configuration module M400, configured to configure a label management instruction set including label weight adjustment parameters, classification rule iteration strategies, and life cycle stage remapping schemes through the abnormal classification nodes and confidence deviation values.

[0102] A label classification management module M500, configured to use the label management instruction set to perform multi-modal label classification management in the entire data life cycle.

[0103] Further, the data acquisition module M100 is used to execute the following method:

[0104] Deploy a distributed feature collector during the data generation stage, and arrange protocol parsing probes in the transfer channel, and configure semantic analysis agents at the storage nodes; at the same time, set a differential acquisition strategy, wherein the sampling frequency during the data generation stage ≥ 100Hz, the protocol parsing depth during the data transfer stage ≥ 7 layers, and the semantic annotation accuracy during the data storage stage ≥ 98%; based on the differential acquisition strategy, use timestamp chain anchoring for multi-stage traceability verification.

[0105] Further, the correlation analysis module M200 is also used to execute the following method:

[0106] Extract the timeliness features in the time-sensitive metadata to construct a timeliness constraint matrix; based on the time-sensitive metadata, perform cross-stage data association path modeling to generate a weighted semantic dependency graph, and evaluate the timeliness compliance score in combination with the timeliness constraint matrix; according to the comparison result between the timeliness compliance score and the preset score threshold, determine the priority adjustment coefficient of the classification label mapping rule, and generate an optimal cross-stage label mapping path.

[0107] Further, the correlation analysis module M200 is also used to execute the following method:

[0108] The timeliness constraint matrix includes the data survival period, the transfer delay threshold between stages, and the time decay factor; for label nodes with semantic overlap and label nodes with timeliness conflicts, introduce an adaptive sliding window mechanism to correct the context dependency relationship, and update the weight assignment of the semantic dependency graph; perform secondary fusion analysis on the updated semantic dependency graph and the timeliness constraint matrix, and output an adaptive classification label set after conflict resolution.

[0109] Further, the correlation analysis module M200 is also used to execute the following method:

[0110] Establish a grid classification architecture, the input end receives the dynamic correlation degree of the context dependency relationship, and the output end generates a weight assignment tensor; at the same time, set an attention mechanism to determine the correlation strength score of data nodes in different stages, and perform dynamic weight correction in combination with the time decay factor in the timeliness constraint matrix.

[0111] Further, the correlation analysis module M200 is also used to execute the following method:

[0112] Perform tensor splicing based on the weight assignment tensor to configure the spatial feature pattern across stages; based on the spatial feature pattern across stages, use bidirectional memory pointers to capture temporal dependencies and generate spatio-temporal fusion weight assignment coefficients; according to the weight assignment coefficients, iteratively optimize the edge weights in the semantic dependency graph and update the model parameters of the grid classification architecture.

[0113] Further, the association analysis module M200 is also used to execute the following method:

[0114] When the standard deviation of the weight assignment coefficients exceeds a preset fluctuation threshold, trigger the window expansion operation of the adaptive sliding window mechanism and recalculate the dynamic association degree of the context dependency.

[0115] Further, the association analysis module M200 is also used to execute the following method:

[0116] The association strength score where S ij represents the association strength between data node i and node j, Q i and K j are the query vector of node i and the key vector of node j respectively, with a dimension of d, Δt ij is the time interval between node i and node j, and λ is the time decay factor defined in the time effect constraint matrix, which is used to control the time decay rate.

[0117] Further, the label classification management module M500 is also used to execute the following method:

[0118] Statistically analyze the distribution characteristics of the confidence deviation values to generate a stability heat map; establish a conflict judgment criterion driven by the stability heat map, set up a label conflict database, and use association rules to locate abnormal classification nodes; activate a label purification instruction according to the cosine similarity corresponding to the abnormal classification nodes to trigger the iteration of multi-modal label classification rules.

[0119] In summary, any step can be stored as a computer instruction or program in an unrestricted computer memory and can be called and recognized by an unrestricted computer processor, without further limitation here.

[0120] Furthermore, the above technical solutions only reflect the preferred technical solutions of the technical solutions of the embodiments of the present application. Some changes that those skilled in the art may make to some parts thereof all reflect the principles of the novel embodiments of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application.

Claims

1. A classification and labeling management method for full life cycle data, characterized in that: The method comprises: Embed feature perception nodes throughout the data life cycle to obtain feature perception data at the data generation stage, data flow stage, data storage stage, and data destruction stage, and determine time-sensitive metadata; Performing cross-stage association analysis on the time-sensitive metadata, setting classification label mapping rules under the time constraint layer, and performing label conflict resolution analysis in combination with data semantic similarity and context dependency to generate an adaptive classification label set; According to the classification label mapping rules and the adaptive classification label set, abnormal classification nodes and confidence deviation values ​​in the entire life cycle of the data are identified; By means of the abnormal classification nodes and confidence deviation values, a label management instruction set including label weight adjustment parameters, classification rule iteration strategy and life cycle stage remapping scheme is configured; The tag management instruction set is used to perform multimodal tag classification management throughout the data life cycle.

2. A classification and labeling management method for full life cycle data according to claim 1, characterized in that: Embedding feature perception nodes in the entire life cycle of data, the method comprising: Deploy distributed feature collectors in the data generation phase, place protocol parsing probes in the flow channel, and configure semantic analysis agents in the storage nodes; At the same time, a differentiated collection strategy is set, where the sampling frequency of the data generation stage is ≥100Hz, the protocol parsing depth of the data flow stage is ≥7 layers, and the semantic annotation accuracy of the data storage stage is ≥98%; Based on the differentiated collection strategy, timestamp chain anchoring is adopted to carry out multi-stage traceability verification.

3. A classification and labeling management method for full life cycle data according to claim 1, characterized in that: Performing cross-stage association analysis on the time-sensitive metadata and setting classification label mapping rules under the time constraint layer, the method further includes: Extracting timeliness features from the time-sensitive metadata and constructing a timeliness constraint matrix; Based on the time-sensitive metadata, cross-stage data association path modeling is performed to generate a weighted semantic dependency graph, and the timeliness compliance score is evaluated in combination with the timeliness constraint matrix; According to the comparison result of the timeliness compliance score and the preset score threshold, the priority adjustment coefficient of the classification label mapping rule is determined, and an optimal label mapping path across stages is generated.

4. A classification and labeling management method for full life cycle data according to claim 3, characterized in that: The time constraint matrix includes data survival period, inter-stage flow delay threshold and time decay factor; For label nodes with semantic overlap and label nodes with time conflicts, an adaptive sliding window mechanism is introduced to correct the context dependency and update the weight distribution of the semantic dependency graph; The updated semantic dependency graph and the time constraint matrix are subjected to secondary fusion analysis, and an adaptive classification label set after conflict resolution is output.

5. A classification and labeling management method for full life cycle data according to claim 4, characterized in that: Updating the weight distribution of the semantic dependency graph, the method comprising: Establishing a grid classification architecture, receiving the dynamic correlation of the context dependency at the input end, and generating a weight distribution tensor at the output end; At the same time, an attention mechanism is set to determine the association strength scores of data nodes at different stages, and dynamic weight correction is performed in combination with the time decay factor in the time constraint matrix.

6. A classification and labeling management method for full life cycle data according to claim 5, characterized in that: Establishing a grid classification framework, the method further comprises: Perform tensor splicing based on the weight distribution tensor to configure spatial feature patterns across stages; Based on the spatial feature patterns across stages, a bidirectional memory pointer is used to capture the temporal dependency and generate the weight distribution coefficients of spatiotemporal fusion; According to the weight distribution coefficients, the edge weights in the semantic dependency graph are iteratively optimized, and the model parameters of the grid classification architecture are updated.

7. A classification and labeling management method for full life cycle data according to claim 6, characterized in that: When the standard deviation of the weight allocation coefficient exceeds a preset fluctuation threshold, the window expansion operation of the adaptive sliding window mechanism is triggered to recalculate the dynamic correlation degree of the context dependency.

8. A classification and labeling management method for full life cycle data according to claim 5, characterized in that: The association strength score Among them, S ij represents the association strength between data node i and node j, Q i and K j are the query vector of node i and the key vector of node j, with dimensions d and Δt respectively. ij is the time interval between node i and node j, and λ is the time decay factor defined in the time constraint matrix, which is used to control the time decay rate.

9. The method for classifying and labeling management of full life cycle data according to claim 1, characterized in that: Multimodal label classification management is performed during the entire life cycle of the data, and the method further includes: Counting the distribution characteristics of the confidence deviation values ​​to generate a stability heat map; Establish conflict judgment criteria driven by the stability heat map, set up a label conflict database, and use association rules to locate abnormal classification nodes; According to the cosine similarity corresponding to the abnormal classification node, a label purification instruction is activated to trigger the iteration of the multimodal label classification rule.

10. A classification and labeling management platform for full life cycle data, characterized in that: The platform is used to implement the classification and labeling management method for full life cycle data according to any one of claims 1 to 9, comprising: The data acquisition module is used to embed feature perception nodes in the entire data life cycle, obtain feature perception data in the data generation stage, data flow stage, data storage stage, and data destruction stage, and determine time-sensitive metadata; An association analysis module is used to perform cross-stage association analysis on the time-sensitive metadata, set classification label mapping rules under the time constraint layer, and perform label conflict resolution analysis in combination with data semantic similarity and context dependency to generate an adaptive classification label set; An anomaly identification module is used to identify abnormal classification nodes and confidence deviation values ​​in the entire life cycle of data according to the classification label mapping rules and adaptive classification label sets; An instruction configuration module, used to configure a label management instruction set including label weight adjustment parameters, classification rule iteration strategy and life cycle stage remapping scheme through the abnormal classification node and confidence deviation value; The tag classification management module is used to use the tag management instruction set to perform multimodal tag classification management throughout the data life cycle.

Citation Information

Cited By

  • Big data management system and method based on hierarchical label system

    CN120596475A

  • Comprehensive energy optimization management system and comprehensive energy management method

    CN120765422A