AI-driven efficient data redundancy detection and cleaning method and system

Through the AI-driven fragment identification architecture and dynamic correction rules, the accurate identification and cleaning of redundant data in a multi-source heterogeneous data environment is solved, and efficient and secure redundant data management is achieved.

CN120408042AActive Publication Date: 2025-08-01CHENGDU BIG DATA GRP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510912868.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-01
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately identify and clean complex redundant data in a multi-source heterogeneous data environment, especially complex redundant data types such as function transfer, uncalled, and invalid call types in diverse contents such as text, images, and log streams, resulting in incomplete coverage of cleaning strategies, large errors in identification results, and even incorrect cleaning of valid data or missing potential redundancy.

Method used

Using an AI-driven fragment recognition architecture, combining structural features, semantic features and functional features, a multi-level cleaning instruction queue is built by building feature mapping relationship tables and dynamic correction rules, and a resource occupancy threshold and rollback protection mechanism are configured to achieve accurate identification and secure cleaning of duplicate data.

Benefits of technology

It realizes high adaptability and redundant identification in diverse data scenarios, avoids the risk of accidentally cleaning important data, ensures the accuracy of redundant judgments and the stability of business operations, and improves the safety and controllability of cleaning operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408042A_ABST
    Figure CN120408042A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-driven efficient data redundancy detection and cleaning method and system, and relates to the technical field of data processing. Comprising the steps that according to categories of different paragraphs in a target data chain, a fragment recognition framework driven by AI is constructed, and the fragment recognition framework is used for positioning the position of repeated data in the target data chain; repeated data is used as a positioning reference. Structured analysis is carried out on the target data chain through the fragment recognition architecture, positioning and recognition of repeated data in different contexts and different data types are achieved in combination with structural features, semantic features and functional features, the method is suitable for diversified data scenes such as text sequences, data streams and multimedia images, and the method is high in practicability. In addition, by constructing a feature mapping relation table, a fragment recognition framework can be constructed based on feature contents of different paragraphs, and AI recognition logic matching is adopted, so that the method has high adaptive capacity and recognition accuracy in the face of complex data structures and hidden redundancy problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to an AI-driven efficient data redundancy detection and cleaning method and system. Background Art

[0002] Data redundancy is a common problem in the field of data processing. It can lead to waste of storage space and degradation of system performance. Excessive redundant data not only increases the complexity of management, but also affects the accuracy and efficiency of data analysis. To solve this problem, a method capable of intelligently identifying and cleaning redundant data is needed to optimize storage resources and improve the overall operating efficiency of the system.

[0003] After retrieval, the Chinese patent application with the publication number "CN113535697A" proposed a "scaffold data cleaning method, scaffold control device and storage medium". By configuring corresponding data processing parameters for the data to be processed of different types and different data sources respectively, it ensures the cleaning of invalid data, reduces the data redundancy of the scaffold control device, increases the processing speed of the scaffold control device, and at the same time avoids the situation where valid data is wrongly cleaned, resulting in the scaffold control device being unable to accurately record and analyze the scaffold state. And by cleaning the data to be stored currently collected by the sensor, it prevents invalid data from entering the storage space and causing waste of storage space. By cleaning the data in the storage space and deleting redundant data, it prevents excessive data from increasing the data processing pressure of the scaffold control device, thereby improving the execution efficiency of the scaffold control device.

[0004] In addition, the Chinese patent application with the publication number "CN111258968A" proposed an "enterprise redundant data cleaning method, device and big data platform". By performing statistical item screening through data redundancy evaluation features and then cleaning redundant data, the success rate and accuracy of redundant data screening can be improved in the case of complex data content, especially when the data service is updated frequently during the data statistics process. In addition, by sending the cleaning process information to the enterprise data terminal, it is convenient for the enterprise data terminal to adjust the statistics process of the enterprise statistical data according to the cleaning process information to control the source of redundant data and avoid waste of unnecessary computing resources.

[0005] However, in the actual use process, the above-mentioned disclosed technical solutions and existing related technical solutions generally perform redundancy identification based on static rules or a single structural dimension, and it is difficult to handle complex scenarios with diverse redundancy types and significant changes in context semantics in a multi-source heterogeneous data environment. They have insufficient ability to identify data redundancy with dynamic behavior characteristics. Especially when processing data chains containing diverse content such as text, images, and log streams, they cannot accurately identify complex redundancy data types such as function transfer type, non-called type, and invalid call type, resulting in incomplete coverage of cleaning strategies, large errors in recognition results, and even the situation of mistakenly cleaning valid data or missing potential redundancy. Summary of the Invention

[0006] The purpose of the present invention is to provide an AI-driven efficient data redundancy detection and cleaning method and system to solve the problems raised in the above background technology.

[0007] To achieve the above purpose, the present invention provides the following technical solutions: In the first aspect, an AI-driven efficient data redundancy detection and cleaning method is proposed, including: According to the categories of different paragraphs in the target data chain, an AI-driven fragment recognition architecture is constructed, and the fragment recognition architecture is used to locate the positions of duplicate data in the target data chain; Taking the duplicate data as the positioning reference, the feature contents in the fragment recognition framework of the sequences before and after the positioning reference in the target data chain are associated to form a feature set; The feature sets of the same-type duplicate data in the target data chain are compared and analyzed, and then the primary redundant data is extracted and output; In a time-series data environment, an operation dataset is extracted by parsing operation instructions, a dynamic correction rule between the operation dataset and the primary redundant data is established, and an optimized redundant data correction set is generated; Based on the redundant data correction set, a multi-level cleaning instruction queue is constructed and a cleaning strategy is assigned.

[0008] As a further preference of this technical solution, the target data chain includes any one or a combination of more of a physical text sequence, a logical data stream, and a multimedia image sequence, and the feature contents include: a structural feature for describing the format and hierarchical relationship of the data, a semantic feature for representing the vector embedding of the data content, and a functional feature for reflecting its operation semantics or behavior meaning in the business scenario.

[0009] As a further preference of this technical solution, the construction method of the fragment recognition framework includes: According to the type characteristics of different paragraphs in the target data chain, an interaction path adapted to it is constructed; Based on the characteristic content of different paragraphs after being processed by the interaction path in the target data chain, construct multiple associated feature mapping relation tables. The feature mapping relation tables are used to establish an AI-driven segment recognition framework based on the characteristic content of each paragraph in the target data chain. The AI recognition logic matching the characteristic content is integrated in the segment recognition framework; Establish an AI-driven segment recognition framework according to the feature mapping relation table.

[0010] As a further preference of this technical solution, the method for the segment recognition framework to identify and locate the position of duplicate data in the target data chain includes: By segmentally integrating and storing the characteristic content of different paragraphs in the target data chain, construct the input feature set library of the segment recognition framework. The input feature set library is arranged in sequence according to the position of the paragraph in the target data chain, and the sequence arrangement order is either the logical order or the access order; Establish a duplicate data set according to the content between multiple mapping relation tables.

[0011] As a further preference of this technical solution, the formation method of the feature set includes: Construct a context association window based on the position of the duplicate data in the target data chain; Call the characteristic content in the corresponding segment recognition architecture and perform hierarchical extraction based on the type of the extracted characteristic content; Classify and integrate the extracted characteristic content to form the feature set within the context association window.

[0012] As a further preference of this technical solution, the dynamic correction rule includes: Extract the operation instruction sequence based on the operation log generated during the timing operation of the target data chain. The operation instruction sequence refers to the time sequence of generating timing operation instructions; Extract the operation instruction set that matches the serial number and type of the paragraph where the primary redundant data is located; Perform behavior pattern planning on the extracted operation instruction set and analyze the behavior variation of the extracted operation instruction set on the primary redundant data in different contexts; Construct a dynamic correction rule between the operation behavior and the redundant data according to the analysis result. The dynamic correction rule is implanted into the AI recognition logic to form a correction strategy template.

[0013] As a further preference of this technical solution, the behavior variation includes: The write behavior not called type redundancy, which is used to indicate that a certain data segment in the target data chain is overwritten by multiple write operations, but has not been accessed, read, or referenced in its subsequent operation sequence, indicating that a certain data segment is only used for temporary storage or intermediate calculation, Redundancy of repeated call behavior with unchanged results is used to indicate that within multiple operation time periods, the same data segment within the target data chain is repeatedly read or called, but the hash value of the data segment output by the system that generates the target data chain does not change after the call, indicating that these called data segments have no impact on the system behavior; Function behavior transfer type redundancy is used to indicate that the data segment within the target data chain undertakes a function during the execution stage of the system that generates the target data chain, but is replaced after the execution stage is completed.

[0014] As a further optimization of this technical solution, the generation of the cleaning strategy includes: Analyze the correction strategy template in the redundant data correction set to generate a priority instruction sequence based on the redundancy type; Configure resource occupancy thresholds and rollback protection mechanisms for each level of cleaning instruction. The resource occupancy thresholds include storage space recovery rates and peak computing resource consumption parameters. Allocate resource quotas for different priority instructions through the preset priority instruction sequence, and the rollback protection mechanism automatically generates a data snapshot before the execution of the cleaning instruction and establishes a two-way verification interface with the target data chain version control system to ensure rapid recovery of the original data chain in case of abnormal conditions; Plant the multi-level cleaning instruction queue into the AI-driven automated cleaning engine to start the policy verification closed loop.

[0015] In the second aspect, to improve an AI-driven efficient data redundancy detection and cleaning method, a cleaning system using an AI-driven efficient data redundancy detection and cleaning method is also proposed, and it includes: Data acquisition and preprocessing module, used to obtain raw data from the target system / platform / application scenario in real time or in batches; Feature extraction and mapping module, used to extract "structural features", "semantic features", and "functional features" using corresponding models for different data types, generate unified feature vectors, and then build a feature mapping relationship table based on these feature vectors; AI-driven segment recognition engine, used to execute data structure recognition logic, data application / behavior recognition logic, and data semantic recognition logic based on the feature mapping relationship table, identify duplicate or similar segments, and output the initial duplicate data set and its location information; Context association and difference discrimination module, used to build a context association window for the detected duplicate data based on its position in the target data chain, extract the features of these paragraphs, and perform semantic / structural / functional integration analysis to determine whether the duplicate data has uniqueness in different contexts and different positions; Timing operation association and dynamic correction module, used to analyze the behavioral impact on primary redundant data in a timing environment, combine operation logs, refine dynamic correction rules, and generate a redundant data correction set; The cleaning instruction queue generation and policy allocation module generates a multi-level cleaning instruction queue based on the redundant data correction set according to the redundant type priority and dynamic correction rules, and allocates resource thresholds and protection mechanisms for each instruction; The monitoring and feedback module is used to monitor the cleaning execution effect in real time, feedback to the AI engine and policy template, continuously iterate and optimize the model parameters and rules, and generate a system optimization index report, a visualization dashboard or a log report; The configuration and management interface is used to provide the system administrator or business personnel with an interactive operation interface for configuring the feature extraction scheme, defining the threshold policy, viewing the monitoring metrics, approving the cleaning plan, and restoring the snapshot.

[0016] Compared with the prior art, the beneficial effects of the present invention are: The AI-driven efficient data redundancy detection and cleaning method and system perform a structured analysis on the target data chain through the fragment recognition architecture, and combine structural features, semantic features and functional features to realize the positioning and recognition of duplicate data in different contexts and different data types. It is applicable to diverse data scenarios such as text sequences, data streams, and multimedia images. In addition, by constructing a feature mapping relationship table, a fragment recognition framework can be constructed based on the feature content of different paragraphs, and the AI recognition logic matching is adopted, so that it has high adaptability and recognition accuracy when facing complex data structures and implicit redundancy problems; It should be added that by introducing context correlation analysis, a windowed feature set is established for the preceding and following paragraphs where the duplicate data is located, so as to judge whether it has business function independence or difference in a specific position. This context-aware recognition method effectively avoids the risk of "mis-cleaning" important data in traditional redundancy cleaning, and ensures the accuracy of redundancy judgment and the stability of business operations; It is worth noting that in terms of dynamic adjustment, by analyzing the behavioral relationship between the operation instructions and the redundant data, various redundant types such as un-called writes, invalid calls, and function transfers can be identified, so that the redundancy recognition not only stays at the data level, but further correlates the operation behavior and system changes, reflecting the dynamic value of the data in the actual use process.

[0017] It should also be noted that during the cleaning execution process, a multi-level cleaning instruction queue is established, a hierarchical instruction sequence is generated according to different redundant types and correction rules, and resource occupancy thresholds and rollback protection mechanisms are configured for it, and storage space recovery rates and calculation resource consumption control parameters are set, and snapshot and version verification interfaces are introduced, so as to achieve data security guarantee and rapid recovery during the cleaning process, and improve the security and controllability of the cleaning operation. Description of the Drawings

[0018] Figure 1It is the flowchart of the steps disclosed in the present invention; Figure 2 It is the supplementary explanatory diagram of step S100 of the present invention; Figure 3 It is the supplementary explanatory diagram of step S103 of the present invention; Figure 4 It is the composition diagram of the system disclosed in the present invention. Specific embodiments

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0020] Before understanding the technical solutions proposed in this application, it should be clear that data redundancy in the prior art not only appears in the form of simple text, but also exists in the form of multimedia images. These redundant data come from duplicate storage, invalid backups, or the accumulation of intermediate data generated during the operation of the system. Different forms of data redundancy require different strategies and technical means during the detection and cleaning process. For example, for redundant data in text form, duplicate content is identified through semantic analysis and similarity comparison, while for redundant multimedia images, image feature extraction and feature matching are required for processing. In addition, during the actual cleaning process, considering the data security and integrity issues involved in the data redundancy cleaning process, corresponding protection mechanisms need to be designed to ensure that important information is not accidentally deleted or damaged.

[0021] For this reason, referring to Figure 1 it can be seen that the present invention provides a technical solution: an AI-driven efficient data redundancy detection and cleaning method, including: step S100 - step S500.

[0022] Step S100: According to the categories of different paragraphs in the target data chain, construct an AI-driven fragment recognition architecture, and the fragment recognition architecture is used to locate the positions of duplicate data in the target data chain.

[0023] It should be noted that in step S100, the target data chain refers to a data aggregate composed of multiple data fragments or data units in a system, platform, or application scenario, with a business logic order or access order. The target data chain includes any one or a combination of multiple of the following: a text sequence in a physical sense, a data stream in a logical sense, and a multimedia image sequence.

[0024] It should be added that during the actual detection process, multiple fragment recognition frameworks are set in the target data chain in step S100. The fragment recognition frameworks are set according to different paragraphs of the target data chain. It should be emphasized that the feature contents extracted by each fragment recognition framework are correlated with each other, and the feature contents include: data structure, data application, and data semantics.

[0025] It also needs to be added that structural features are used to describe the format and hierarchical relationship of data, semantic features are used to represent the vector embedding of data content, and functional features are used to reflect the operation semantics or behavioral meaning in the business scenario.

[0026] As a preferred implementation, since the target data chain may be composed of different heterogeneous data during the data processing, when actually constructing the fragment recognition framework, it is necessary to consider the diversity of data types and the compatibility issues between heterogeneous data.

[0027] Specifically, referring to Figure 2 it can be seen that this implementation is mainly used to supplement the construction of the AI-driven fragment recognition framework in step S100, including: step S101 - step S103.

[0028] Step S101: Construct an interaction path adapted to it according to the type characteristics of different paragraphs in the target data chain.

[0029] It should be noted that the interaction path embeds a feature processing engine adapted to the type characteristics of the target data chain, and the feature processing engine is used to integrate and unify the type characteristics of different paragraphs in the target data chain.

[0030] It is worth noting that the type characteristics of the target data chain in step S101 include: text feature type, data flow feature type, and multimedia image feature type. The discrimination method for the type characteristics of the target data chain is to perform automatic recognition through feature extraction algorithms and pattern matching models in the prior art.

[0031] Specifically, the text feature type adopts the word vector embedding technology in natural language processing in the prior art, and uses the known BERT model to extract syntactic structure features for recognition. The data flow feature type uses the known sliding window mechanism combined with the known LSTM network to capture temporal characteristics for recognition. The multimedia image feature type uses the known convolutional neural network (CNN) to extract spatial topology features for recognition.

[0032] It should be particularly noted that during the execution of step S101, feature processing engines that form matching relationships with different target data link types are located and implanted within the mutual path. These feature processing engines are all constructed based on mature solutions under the existing technical framework. Specifically, for text types, known NLP analysis engines are enabled; for data stream types, known real-time computing engines are enabled; for multimedia types, known computer vision engines are enabled. It should be particularly noted that the type feature extraction of the target data link and the construction of the mutual path are both realized based on the conventional technical means of those skilled in the art. Given that the technical means involved fall within the scope of industry-standardized operations, the applicant has not elaborated on the above technical details in detail.

[0033] Step S102: Based on the different paragraph feature contents after being processed by the interaction path in the target data link, construct multiple associated feature mapping relationship tables, which are used to establish an AI-driven segment recognition framework according to the feature contents of each paragraph in the target data link.

[0034] It is worth noting that in the content of step S102, each feature mapping relationship table generates a function label based on the feature content during establishment, and the function label is used to record the feature content in the corresponding paragraph.

[0035] Step S103: Establish an AI-driven segment recognition framework according to the feature mapping relationship table.

[0036] It is worth noting that in step S103, the AI recognition logic that matches the feature content is integrated into the segment recognition framework. Specifically, in the technical solution proposed in this application, the AI recognition logic is divided into data structure recognition logic, data application logic, and data semantic logic. The data structure recognition logic is used to identify duplicate structured information, the data application recognition logic is used to identify data with the same behavior pattern in the application scenario, and the data semantic recognition logic is used to identify unstructured semantic data with similar or equivalent contents.

[0037] As a preferred implementation, referring to Figure 3 it can be seen that this implementation is mainly used to supplement the working principle of the segment recognition framework in step S103, specifically including: steps S103A - S103C.

[0038] Step S103A: By segmenting and integrating the storage of different paragraph feature contents in the target data link, construct an input feature set library for the segment recognition framework.

[0039] It should be noted that step S103A is used to construct multi-type input interfaces for the AI recognition logic, and the input feature set library is arranged in sequence according to the position of the paragraph in the target data link. It should be supplemented that the sequence arrangement order is either in logical order or access order.

[0040] Step S103B: Establish a duplicate data set based on the content between multiple mapping relation tables.

[0041] It should be noted that the content between multiple mapping relation tables in step S103B refers to integrating and associating the target area chain paragraphs with the same feature content through the feature content corresponding to the function labels, and binding the duplicate features with the function labels as the prefix of the duplicate data set.

[0042] It should be added that the purpose of the duplicate data set established in step S103B is to enable the AI recognition logic to obtain an initial recognition concept template. The AI recognition logic retrieves in the input feature set library through the initial recognition concept template to determine whether there are duplicate data in different paragraphs within the target data chain.

[0043] Step S103C: Associate the duplicate data set with the input feature set library and output the corresponding duplicate data and serial numbers.

[0044] It should be added that in the association process between the duplicate data set and the input feature set library in step S103C, it depends on the similarity matching of the feature content recorded in the function labels. Specifically, according to the feature content in the function labels, locate the target paragraphs corresponding to the duplicate data set in the input feature set library and generate a result list containing the duplicate data content and its corresponding serial numbers. This result list marks the specific positions of the duplicate data in the target data chain.

[0045] Step S200: Using the duplicate data as the positioning reference, associate the feature content within the fragment recognition frameworks of the sequences before and after the positioning reference in the target data chain and form a feature set.

[0046] It should be specifically noted that step S200 aims to identify and locate the duplicate data segments in the target data chain, and contextually integrate the semantic interpretation and scenario function analysis of the target data chain paragraphs corresponding to the duplicate data with the semantic interpretation and scenario function analysis in the adjacent paragraphs, and then compare the semantic features and functional attributes presented by the duplicate data at different positions in the target data chain to establish a difference discrimination mechanism, so as to effectively verify whether the duplicate data has uniqueness.

[0047] As a preferred implementation, this implementation is used to supplement the method for forming the feature set in step S200.

[0048] Specifically, it includes: step S201 - step S203.

[0049] Step S201: Construct a context association window based on the position of the duplicate data in the target data chain.

[0050] It should be noted that in the context - related window in step S201, the sequence is mainly used to determine the order, and then determine the position of the context - related window in the target data chain.

[0051] It should be further noted that in the technical solution proposed in this application, taking the duplicate data set output in step S103C as the positioning reference, the serial numbers of the paragraphs at its location are extracted in the target data chain, and based on the sorting of the serial numbers, several serial number paragraphs are extended forward and backward to form a context - related window based on the duplicate data.

[0052] Step S202: Invoke the feature content in the corresponding fragment recognition architecture and perform hierarchical extraction based on the type of the extracted feature content.

[0053] Step S203: Classify and integrate the extracted feature content to form a feature set within the context - related window.

[0054] It should be emphasized that the feature set in step S203 not only includes the semantic features, structural features, and functional features of the duplicate data itself, but also covers the semantic features, structural features, and functional features in the serial number paragraphs before and after the duplicate data.

[0055] Step S300: Conduct a comparative analysis of the feature sets of the same - type duplicate data in the target data chain, and then extract and output the primary redundant data.

[0056] It should be noted that step S300 is mainly used to further screen and classify the duplicate data in the target data chain to identify the truly redundant parts. In this process, based on the feature set formed in step S203, through the comparative analysis of semantic, structural, and functional features, it is judged whether the duplicate data is irreplaceable in terms of semantics or function. If the duplicate data in a certain segment of the target data chain shows consistent semantic content and functional attributes at different positions, then the data is marked as primary redundant data.

[0057] It should be noted that in step S300, the comparative analysis of semantic, structural, and functional features between feature sets is realized through similarity - measurement algorithms in the prior art. Specifically, for semantic features, a cosine - similarity calculation model based on the known BERT is used for analysis, for structural features, a topological - matching degree analysis is carried out through the known graph neural network (GNN), and for functional features, a correlation - evaluation algorithm based on the known business - scenario knowledge graph is used for analysis.

[0058] Step S400: In the time - series data environment, extract the operation data set by parsing the operation instructions, establish a dynamic correction rule between the operation data set and the primary redundant data, and generate an optimized redundant - data correction set.

[0059] It should be noted that step S400 is mainly used to solve the problem of redundant discrimination in the dynamic data scenario. By capturing the timing operation instructions generated during the operation of the target data chain (including data writing, modification, and migration operation behaviors), a dynamic correction rule associated with the primary redundant data is constructed.

[0060] As a preferred implementation, this implementation is used to supplement and explain the dynamic correction rule in detail. Specifically, it includes: step S401 - step S404.

[0061] Step S401: Extract the operation instruction sequence based on the operation log generated during the timing operation of the target data chain.

[0062] It should be noted that the operation instruction sequence refers to the time sequence of generating the timing operation instructions. Specifically, in step S401, the behavior trajectory on the target data chain is captured in real time. The behavior trajectory is generated by the user's action on the target data chain based on the timing operation instructions and is sorted according to the timestamp to form an original instruction log with a complete operation process record. It should be particularly pointed out that the extracted operation instruction sequence not only records the operation actions but also includes the additional information of the operation target (i.e., the position acting on the target data chain) and the operation period.

[0063] Step S402: Extract the set of operation instructions that match the serial number and type of the paragraph where the primary redundant data is located.

[0064] It should be noted that step S402 is mainly used to associate the relationship between the primary redundant data and the set of operation instructions. Specifically, in this step, based on the operation instruction sequence, by using the serial number of the paragraph corresponding to the primary redundant data as an index, the operation instructions that have acted on this paragraph within the duplicate data are screened.

[0065] Step S403: Plan the behavior pattern of the extracted set of operation instructions and analyze the behavior variation of the extracted set of operation instructions on the primary redundant data in different contexts.

[0066] It should be noted that step S403 is mainly used to analyze the dynamic impact of the operation instructions on the primary redundant data in different context environments.

[0067] As a preferred implementation, this implementation is mainly used to elaborate on the dynamic impact. Specifically, it includes the following content: First, the write operation is not called redundant (temporary cache redundancy). It should be noted that the write operation not being called redundant is manifested as a data segment within the target data chain being overwritten by multiple write operations, but not being accessed, read, or referenced in its subsequent operation sequence, indicating that this data segment may be only used for temporary storage or intermediate calculations.

[0068] It should be noted that the discrimination basis of the write operation not being called redundant within step S403 is to detect whether the primary redundant data is called subsequently after the write operation, and whether the position of the primary redundant data within the target data chain changes after the write operation.

[0069] Second, the call behavior is repeated but the result does not change redundant (invalid call redundancy). It should be noted that the impact of the call behavior being repeated but the result not changing redundant is manifested as the same data segment within the target data chain being repeatedly read or called within multiple operation time periods, but the system state of the generated target data chain does not change significantly, indicating that these calls do not have an actual impact on the system behavior.

[0070] It should be noted that the discrimination basis of the call behavior being repeated but the result not changing redundant within step S403 is to obtain whether there are continuous or periodic read operations on the data segment of the primary redundant data within the target data chain, and then whether the data segment of the primary redundant data changes after each read.

[0071] In addition, it should be supplemented that the specific manifestations of the system state not changing significantly include: the data segment hash value output by the system that generates the target data chain does not change. Among them, the detection of the data segment hash value of the system that generates the target data chain not changing is achieved based on a standardized hash algorithm. Specifically, the standard SHA-256 algorithm is used to perform real-time hash calculation on the data segment corresponding to the primary redundant data in the target data chain, and the newly generated hash value after each read operation is compared with the previously recorded historical hash value to determine whether the system state shows no significant change.

[0072] Third, the functional behavior transfer redundant (phased redundancy). It should be noted that the impact of the functional behavior transfer redundant is manifested as: the data segment within the target data chain undertakes a function during the system execution phase, but is replaced after the system finishes the execution phase.

[0073] It should be noted that the discrimination basis of the functional behavior transfer redundant within step S403 is to obtain the call frequency of the primary redundant data within the target data chain, the interaction frequency between the primary redundant data and the new data segment, and the new call frequencies of the two after the interaction.

[0074] Step S404: Construct a dynamic correction rule between the operation behavior and the redundant data based on the analysis result.

[0075] It should be noted that in Step S404, the dynamic correction rule can be implanted into the AI recognition logic in Step S103 to form a correction strategy template. Specifically, the specific content of the dynamic correction rule in Step S404 includes: a redundant correction rule for the uninvoked write behavior, that is, setting a "non-reference timeout threshold" in the AI logic. If no reference operation related to this data segment occurs within the time range of the "non-reference timeout threshold", it is marked as "temporary cache redundancy"; a redundant correction rule for repeated call behavior but unchanged result, that is, comparing the system behavior differences (such as state hash and structure summary code) of the target data chain generated after each call. If the system behavior summaries corresponding to N calls are consistent and do not affect the data chain path or output result, the subsequent redundant calls are marked as "invalid operations", and the subsequent calls are set to an ignorable state in the AI control logic, and the "optimizable path suggestion" is recorded; a redundant correction rule for function behavior transfer type, that is, by recording the call or cooperation relationship between the primary redundant data segment and other data segments. If there is a data segment that replaces the behavior path of the primary redundant data segment and the call frequency of this data segment is higher than that of the primary redundant data segment, it is determined that the function corresponding to the primary redundant data segment has been transferred. At this time, a correction suggestion is output: let the primary redundant data segment enter the "frozen state" or be moved to the secondary call buffer.

[0076] Step S500: Based on the redundant data correction set, construct a multi-level cleaning instruction queue and allocate a cleaning strategy.

[0077] It should be noted that in Step S500, the construction of the multi-level cleaning instruction queue is hierarchically processed according to the redundant types marked in the redundant data correction set and their dynamic correction rules. Specifically, a cleaning instruction chain with a time sequence dependency is generated according to the redundant type priority (temporary cache redundancy > invalid call redundancy > phased redundancy), and at the same time, a differentiated execution strategy is configured for each level of cleaning instruction.

[0078] As a preferred implementation, this implementation is used to supplement the construction logic and strategy allocation mechanism of the multi-level cleaning instruction queue in Step S500, specifically including: Step S501 - Step S503.

[0079] Step S501: Analyze the correction strategy template in the redundant data correction set to generate a priority instruction sequence based on the redundant type.

[0080] It should be noted that in step S501, the generation of the priority instruction sequence follows the following principles: for the data segment marked as "temporary cache redundancy", an immediate cleaning instruction will be preferentially generated; for the data segment corresponding to "invalid call redundancy", a delayed cleaning instruction with a conditional trigger mechanism will be generated (the running environment status of the target data chain needs to be verified); for the "phased redundancy" data segment, an observation period retention instruction based on function substitution verification will be generated.

[0081] Step S502: Configure resource occupancy thresholds and rollback protection mechanisms for each level of cleaning instructions.

[0082] It is worth noting that the resource occupancy thresholds in step S502 include the storage space recovery rate and the peak parameter of computing resource consumption. Resource quotas are allocated to different priority instructions through the preset priority instruction sequence, while the rollback protection mechanism automatically generates a data snapshot before the execution of the cleaning instruction and establishes a two-way verification interface with the version control system of the target data chain to ensure that the original data chain can be quickly restored in case of an abnormal state.

[0083] Step S503: Implant the multi-level cleaning instruction queue into the AI-driven automated cleaning engine and start the policy verification closed-loop.

[0084] It should be added that the policy verification closed-loop in step S503 includes three layers of verification mechanisms: the first layer verifies data integrity by real-time monitoring the state hash value of the target data chain during the execution of the cleaning instruction; the second layer calls the fragment recognition framework in step S103 to recheck the features of the cleaned data chain to confirm the elimination effect of redundant features; the third layer combines the dynamic correction rules generated in step S400 to evaluate the long-term impact of the cleaning operation on the behavior pattern of the time-series data environment and generates an optimization index report.

[0085] As a preferred implementation, this implementation is used to supplement an AI-driven efficient data redundancy detection and cleaning method. Referring to Figure 4 it can be seen that an AI-driven efficient data redundancy detection and cleaning system is proposed, including: A data acquisition and preprocessing module, which is used to obtain raw data from the target system / platform / application scenario in real time or in batches, including logs, texts, data streams, and multimedia (images / videos), and perform preprocessing such as format unification, cleaning, and noise removal on the raw data to prepare for subsequent feature extraction.

[0086] It is worth noting that the data acquisition and preprocessing module is mainly related to the data acquisition before step S401 and is regarded as a precondition before the execution of an AI-driven efficient data redundancy detection and cleaning method.

[0087] Feature extraction and mapping module, which is used to extract "structural features", "semantic features", and "functional features" using corresponding models for different data types (text, time series / data stream, multimedia images), generate unified feature vectors, and then construct a feature mapping relationship table based on these feature vectors.

[0088] It should be noted that the functions included in the feature extraction and mapping module are text feature extraction engine, time series data feature extraction, image feature extraction, and extraction of functional features and scenario features. In addition, when the feature extraction and mapping module is specifically used, it outputs feature vectors, functional labels, and a feature mapping relationship table, corresponding to step S101 and step S102.

[0089] AI-driven fragment recognition engine, which is used to execute data structure recognition logic, data application / behavior recognition logic, and data semantic recognition logic based on the feature mapping relationship table, identify duplicate or similar fragments, and output an initial set of duplicate data and its location information.

[0090] It should be noted that the AI-driven fragment recognition engine includes: an input interface for accessing the "input feature set library" constructed in step S103A, and recognition logic, which is composed of a structure similarity module, a semantic similarity module, and a function / behavior similarity module. It should be added that the AI-driven fragment recognition engine corresponds to the content of steps S103A - S103C.

[0091] Context association and difference discrimination module, which is used to construct a context association window (several paragraphs before and after) based on the position of the detected duplicate data in the target data chain for the detected duplicate data, extract the features of these paragraphs, perform semantic / structural / functional integration analysis, and determine whether the duplicate data is unique in different contexts and positions (whether it can be regarded as truly redundant).

[0092] It should be noted that the context association and difference discrimination module corresponds to the content of steps S201 - S203 and step S300.

[0093] Time series operation association and dynamic correction module, which is used to analyze the behavioral impact on primary redundant data in a time series environment in combination with operation logs (sequences of operation instructions such as writing, reading, modifying, etc.), refine dynamic correction rules, and generate a redundant data correction set.

[0094] It should be noted that the time series operation association and dynamic correction module corresponds to steps S401 - S404.

[0095] Cleaning instruction queue generation and policy allocation module, which generates a multi-level cleaning instruction queue based on the redundant data correction set according to the redundant type priority and dynamic correction rules, and allocates resource thresholds and protection mechanisms for each instruction.

[0096] It should be noted that the cleaning instruction queue generates the content corresponding to steps S501 to S503 of the policy allocation module.

[0097] The monitoring and feedback module is used to monitor the cleaning execution effect in real time, feedback to the AI engine and the policy template, continuously iterate and optimize the model parameters and rules, and generate a system optimization index report, a visualization dashboard or a log report.

[0098] The configuration and management interface is used to provide an interactive operation interface for system administrators or business personnel to configure feature extraction schemes, define threshold policies, view monitoring metrics, approve cleaning plans, and restore snapshots.

[0099] It should be noted that the monitoring and feedback module and the configuration and management interface are linked, and the optimization index report, the visualization dashboard or the log report can be visualized through the configuration and management interface.

[0100] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended embodiments and their equivalents.

Claims

1. An AI-driven efficient data redundancy detection and cleaning method, characterized in that, Including: Construct an AI-driven segment recognition architecture according to the categories of different paragraphs in the target data chain. The segment recognition architecture is used to locate the positions of duplicate data in the target data chain; Taking the duplicate data as the positioning benchmark, associate the feature content within the segment recognition framework of the sequences before and after the positioning benchmark in the target data chain, and form a feature set; Conduct comparative analysis on the feature sets of the same type of duplicate data in the target data chain, and then extract and output primary redundant data; In a time-series data environment, extract an operation data set by parsing operation instructions, establish a dynamic correction rule between the operation data set and the primary redundant data, and generate an optimized redundant data correction set; Based on the redundant data correction set, construct a multi-level cleaning instruction queue and allocate cleaning strategies.

2. An AI-driven efficient data redundancy detection and cleaning method according to claim 1, characterized in that: The target data chain includes any one or a combination of multiple of the following: a text sequence in a physical sense, a data stream in a logical sense, and a multimedia image sequence. The feature content includes: a structural feature for describing the format and hierarchical relationship of data, a semantic feature for representing the vector embedding of data content, and a functional feature for reflecting its operation semantics or behavioral meaning in a business scenario.

3. An AI-driven efficient data redundancy detection and cleaning method according to claim 1, characterized in that: The construction method of the segment recognition framework includes: Construct an interaction path adapted to it according to the type characteristics of different paragraphs in the target data chain; Based on the feature content of different paragraphs processed by the interaction path in the target data chain, construct multiple associated feature mapping relation tables. The feature mapping relation tables are used to establish an AI-driven segment recognition framework according to the feature content of each paragraph in the target data chain. The segment recognition framework integrates an AI recognition logic matching the feature content; Establish an AI-driven segment recognition framework according to the feature mapping relation table.

4. An AI-driven efficient data redundancy detection and cleaning method according to claim 3, characterized in that: The method for the segment recognition framework to identify and locate the positions of duplicate data in the target data chain includes: By segmentally integrating and storing the feature content of different paragraphs in the target data chain, construct an input feature set library of the segment recognition framework. The input feature set library is arranged in sequence according to the position of the paragraph in the target data chain, and the sequence arrangement order can be either the logical order or the access order; Establish a duplicate data set according to the content between multiple mapping relation tables.

5. An AI-driven efficient data redundancy detection and cleaning method according to claim 1, characterized in that: The formation method of the feature set includes: Construct a context association window based on the position of the duplicate data in the target data chain; Call the feature content within the corresponding segment recognition architecture, and perform hierarchical extraction based on the type of the extracted feature content; Classify and integrate the extracted feature content to form a feature set within the context association window.

6. An AI-driven efficient data redundancy detection and cleaning method according to claim 3, characterized in that: The dynamic correction rule includes: Extract an operation instruction sequence based on the operation log generated during the time-series operation of the target data chain. The operation instruction sequence refers to the time sequence of generating time-series operation instructions; Extract a matching operation instruction set according to the serial number and type of the paragraph where the primary redundant data is located; Conduct behavior pattern planning on the extracted operation instruction set, and analyze the behavior variation of the extracted operation instruction set on the primary redundant data in different contexts; Construct a dynamic correction rule between the operation behavior and the redundant data according to the analysis result. The dynamic correction rule is implanted into the AI recognition logic to form a correction strategy template.

7. An AI-driven efficient data redundancy detection and cleaning method according to claim 6, wherein: The described behavioral variations include: Write behavior uncalled redundant, which is used to indicate that a certain data segment within the target data chain is overwritten by multiple write operations, but has not been accessed, read, or referenced in its subsequent operation sequence, indicating that a certain data segment is only used for temporary storage or intermediate calculation; Call behavior repeated but result unchanged redundant, which is used to indicate that within multiple operation time periods, the same data segment within the target data chain is repeatedly read or called, but the hash value of the data segment output by the system that generates the target data chain has not changed after the call, indicating that these called data segments have no impact on the system behavior; Functional behavior transfer redundant, which is used to indicate that a data segment within the target data chain undertakes a function during the execution stage of the system that generates the target data chain, but is replaced after the execution stage is completed.

8. An AI-driven efficient data redundancy detection and cleaning method according to claim 6, characterized in that: The generation of the cleaning strategy includes: Parsing the correction strategy template within the redundant data correction set to generate a priority instruction sequence based on the redundancy type; Configuring resource occupancy thresholds and rollback protection mechanisms for each level of cleaning instruction. The resource occupancy thresholds include storage space recovery rate and peak computing resource consumption parameters. Resource quotas are allocated to different priority instructions through the preset priority instruction sequence, while the rollback protection mechanism automatically generates a data snapshot before the execution of the cleaning instruction and establishes a two-way verification interface with the target data chain version control system to ensure rapid recovery of the original data chain in case of an abnormal state; Implanting the multi-level cleaning instruction queue into the AI-driven automated cleaning engine to start the strategy verification closed loop.

9. An AI-driven efficient data redundancy detection and cleaning system, which uses an AI-driven efficient data redundancy detection and cleaning method described in any one of claims 1-8, is characterized in that, Including: Data acquisition and preprocessing module, which is used to obtain raw data from the target system / platform / application scenario in real time or in batches; Feature extraction and mapping module, which is used to extract "structural features", "semantic features", and "functional features" using corresponding models for different data types, generate a unified feature vector, and then construct a feature mapping relationship table based on these feature vectors; AI-driven segment recognition engine, which is used to execute data structure recognition logic, data application / behavior recognition logic, and data semantic recognition logic based on the feature mapping relationship table, identify duplicate or similar segments, and output the initial duplicate data set and its location information; Context association and difference discrimination module, which is used to construct a context association window for the detected duplicate data based on its position in the target data chain, extract the features of these paragraphs, perform semantic / structural / functional integration analysis, and determine whether the duplicate data is unique in different contexts and positions; Temporal operation association and dynamic correction module, which is used to analyze the behavioral impact on primary redundant data in a temporal environment in combination with operation logs, refine dynamic correction rules, and generate a redundant data correction set; Cleaning instruction queue generation and strategy allocation module, which generates a multi-level cleaning instruction queue based on the redundant data correction set according to the redundancy type priority and dynamic correction rules, and allocates resource thresholds and protection mechanisms for each instruction; Monitoring and feedback module, which is used to monitor the cleaning execution effect in real time, feedback to the AI engine and strategy template, continuously iterate and optimize the model parameters and rules, and generate a system optimization index report, a visualization dashboard, or a log report; Configuration and management interface, used to provide an interactive operation interface for system administrators or business personnel to configure feature extraction schemes, define threshold policies, view monitoring metrics, approve cleanup plans, and restore snapshots.

Citation Information

Patent Citations

  • Enterprise redundant data cleaning method and device and big data platform

    CN111258968A

  • Climbing frame data cleaning method, climbing frame control device and storage medium

    CN113535697A

  • Data management method and system of industrial energy storage system

    CN119341706A

  • Context reconstruction method and system based on parent-child document matching

    CN120106083A