An AI-driven efficient data redundancy detection and cleaning method and system

Through AI-driven fragment recognition architecture and dynamic correction rules, the problem of accurate identification and cleaning of redundant data in multi-source heterogeneous data environments is solved, and efficient and secure redundant data management is achieved. It is suitable for diverse data scenarios such as text, images, and log streams.

CN120408042BActive Publication Date: 2025-09-30CHENGDU BIG DATA GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510912868.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-30
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Existing technologies have difficulty accurately identifying and cleaning complex redundant data in a multi-source heterogeneous data environment, especially complex redundant data types such as function transfer, uncalled, and invalid call types in diverse content such as text, images, and log streams. This leads to incomplete coverage of cleaning strategies, large errors in recognition results, and even incorrect cleaning of valid data or omission of potential redundancy.

Method used

It adopts an AI-driven fragment recognition architecture, builds feature mapping relationship tables and dynamic correction rules, combines structural, semantic, and functional features to identify and clean up redundant data, builds a multi-level cleaning instruction queue, and configures resource usage thresholds and rollback protection mechanisms to ensure recognition accuracy and data security.

Benefits of technology

It achieves efficient redundant data identification and cleaning in diverse data scenarios, avoids the mistaken cleaning of important data, ensures the stability of business operations, and improves the security and controllability of cleaning operations by dynamically adjusting and identifying multiple redundancy types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408042B_ABST
    Figure CN120408042B_ABST
Patent Text Reader

Abstract

The present invention discloses an AI-driven efficient data redundancy detection and cleaning method and system, which relates to the field of data processing technology. It includes constructing an AI-driven fragment recognition architecture based on the categories of different paragraphs in the target data chain, and the fragment recognition architecture is used to locate the location of duplicate data in the target data chain; and the duplicate data is used as the positioning reference. The present invention performs structured analysis on the target data chain through the fragment recognition architecture, combines structural features, semantic features and functional features, and realizes the positioning and recognition of duplicate data in different contexts and different data types. It is suitable for diversified data scenarios such as text sequences, data streams, multimedia images, etc. In addition, by constructing a feature mapping relationship table, it is possible to construct a fragment recognition framework based on the feature content of different paragraphs, and adopt AI recognition logic matching, so that it has higher adaptability and recognition accuracy when facing complex data structures and implicit redundancy problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and specifically to an AI-driven efficient data redundancy detection and cleaning method and system. Background Art

[0002] Data redundancy is a common problem in the field of data processing. It leads to waste of storage space and degradation of system performance. Excessive redundant data not only increases the complexity of management, but also affects the accuracy and efficiency of data analysis. To solve this problem, a method that can intelligently identify and clean up redundant data is needed to optimize storage resources and improve the overall operating efficiency of the system.

[0003] After searching, the Chinese invention patent application with publication number "CN113535697A" proposed a "climbing frame data cleaning method, climbing frame control device and storage medium". By configuring corresponding data processing parameters for data to be processed of different types and different data sources, the cleaning of invalid data is guaranteed, the data redundancy of the climbing frame control device is reduced, and the processing speed of the climbing frame control device is increased. At the same time, it avoids the situation where valid data is cleared incorrectly, which causes the climbing frame control device to be unable to accurately record and analyze the climbing frame status. By cleaning the data to be stored currently collected by the sensor, invalid data is prevented from entering the storage space, causing a waste of storage space. By cleaning the data in the storage space and deleting redundant data, excessive data is prevented from increasing the pressure of data processing of the climbing frame control device, thereby improving the execution efficiency of the climbing frame control device.

[0004] Furthermore, the Chinese invention patent application, publication number "CN111258968A," proposes a "Method, Device, and Big Data Platform for Enterprise Redundant Data Cleaning." This approach uses data redundancy evaluation features to screen statistical items before performing redundant data cleanup. This improves the success rate and accuracy of redundant data screening in complex data environments, particularly when data statistics are frequently updated. Furthermore, by distributing cleanup process information to enterprise data terminals, these terminals can adjust the statistical process of enterprise statistics based on this information, thereby controlling the source of redundant data and avoiding unnecessary waste of computing resources.

[0005] However, in actual use, the above-mentioned disclosed technical solutions and existing related technical solutions are generally based on static rules or a single structural dimension for redundancy identification, which is difficult to cope with complex scenarios with diverse redundancy types and significant contextual semantic changes in multi-source heterogeneous data environments. The redundancy identification capability of data with dynamic behavioral characteristics is insufficient, especially when processing data chains containing diverse content such as text, images, and log streams. Complex redundant data types such as function transfer type, uncalled type, and invalid call type cannot be accurately identified, resulting in incomplete coverage of the cleaning strategy, large errors in the recognition results, and even the mistaken cleaning of valid data or omission of potential redundancy. Summary of the Invention

[0006] The purpose of the present invention is to provide an AI-driven efficient data redundancy detection and cleaning method and system to solve the problems raised in the above background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] First, we propose an AI-driven, efficient data redundancy detection and cleaning method, including:

[0009] Based on the categories of different segments within the target data chain, an AI-driven segment recognition architecture is built to locate the location of duplicate data in the target data chain;

[0010] Using repeated data as the positioning reference, the feature content within the recognition framework of the fragments before and after the positioning reference in the target data chain is associated to form a feature set.

[0011] Compare and analyze the feature sets of similar duplicate data in the target data chain, and then extract and output primary redundant data;

[0012] In a time series data environment, the operation data set is extracted through operation instruction analysis, dynamic correction rules for the operation data set and primary redundant data are established, and an optimized redundant data correction set is generated;

[0013] Based on the redundant data correction set, a multi-level cleaning instruction queue is constructed and a cleaning strategy is assigned.

[0014] As a further preferred embodiment of the present technical solution, the target data link includes: any one or more combinations of text sequences in the physical sense, logical data flows, and multimedia image sequences, and the feature content includes: structural features for describing the format and hierarchical relationship of the data, semantic features for representing vector embedding of the data content, and functional features for reflecting its operational semantics or behavioral meaning in business scenarios.

[0015] As a further preferred embodiment of the present invention, the method for constructing the fragment recognition framework includes:

[0016] According to the type characteristics of different sections in the target data chain, construct an interactive path that is suitable for them;

[0017] Based on the characteristic content of different paragraphs in the target data chain after interactive path processing, multiple associated feature mapping relationship tables are constructed. The feature mapping relationship tables are used to establish an AI-driven segment recognition framework based on the characteristic content of each paragraph in the target data chain. The segment recognition framework integrates AI recognition logic that matches the characteristic content;

[0018] An AI-driven fragment recognition framework is established based on the feature mapping relationship table.

[0019] As a further preferred embodiment of the present technical solution, the method for identifying and locating the location of duplicate data in the target data chain by the fragment identification framework includes:

[0020] By integrating and storing the feature content of different paragraphs in the target data chain in segments, an input feature set library of the segment recognition framework is constructed. The input feature set library is arranged in sequence according to the position of the paragraphs in the target data chain. The sequence arrangement order is either logical order or access order.

[0021] A duplicate data set is created based on the contents between multiple mapping relationship tables.

[0022] As a further preferred embodiment of the present technical solution, the method for forming the feature set includes:

[0023] Building a context window based on the location of the duplicate data in the target data chain;

[0024] Call the feature content in the corresponding fragment recognition architecture and perform hierarchical extraction based on the type of extracted feature content;

[0025] The extracted feature contents are classified and integrated to form a feature set within the context association window.

[0026] As a further preferred embodiment of the present technical solution, the dynamic correction rules include:

[0027] Extracting an operation instruction sequence based on the operation log generated by the target data link during the timing operation, where the operation instruction sequence refers to the time sequence of the timing operation instructions;

[0028] Extracting a matching set of operation instructions based on the sequence number and type of the segment where the primary redundant data is located;

[0029] Conduct behavioral pattern planning for the extracted operation instruction set, and analyze the behavioral variation of the extracted operation instruction set for primary redundant data in different contexts;

[0030] Based on the analysis results, dynamic correction rules between operational behaviors and redundant data are constructed and embedded into the AI ​​recognition logic to form a correction strategy template.

[0031] As a further preferred embodiment of the present invention, the behavioral variation includes:

[0032] Write behavior uncalled redundancy is used to indicate that a data segment in the target data chain is overwritten by multiple write operations, but has never been accessed, read, or referenced in its subsequent operation sequence, indicating that a data segment is only used for temporary storage or intermediate calculation.

[0033] Repeated call behavior but unchanged result redundancy indicates that the same data segment in the target data chain is read or called repeatedly during multiple operation periods, but the hash value of the data segment output by the system that generates the target data chain after the call does not change, indicating that these called data segments have no impact on system behavior;

[0034] Functional behavior transfer redundancy is used to indicate that the data segment in the target data link assumes a function during the execution phase of the system that generates the target data link, but is replaced after the execution phase is completed.

[0035] As a further preferred embodiment of the present technical solution, the generation of the cleanup strategy includes:

[0036] Parsing the correction strategy template in the redundant data correction set to generate a priority instruction sequence based on the redundancy type;

[0037] Resource usage thresholds and rollback protection mechanisms are configured for each level of cleanup instructions. The resource usage thresholds include storage space recovery rates and computing resource consumption peak parameters. Resource quotas are allocated to instructions of different priorities through a preset priority instruction sequence. The rollback protection mechanism automatically generates a data snapshot before the execution of the cleanup instruction and establishes a two-way verification interface with the target data link version control system to ensure that the original data link can be quickly restored in abnormal conditions.

[0038] Embed a multi-level cleaning instruction queue into the AI-driven automated cleaning engine to initiate a closed-loop strategy verification.

[0039] Secondly, in order to improve an AI-driven efficient data redundancy detection and cleaning method, a cleaning system using an AI-driven efficient data redundancy detection and cleaning method is also proposed, and includes:

[0040] Data acquisition and preprocessing module, used to obtain raw data from the target system / platform / application scenario in real time or in batches;

[0041] The feature extraction and mapping module is used to extract "structural features", "semantic features" and "functional features" for different data types using corresponding models, generate unified feature vectors, and then construct a feature mapping relationship table based on these feature vectors;

[0042] An AI-driven fragment recognition engine, which is used to execute data structure recognition logic, data application / behavior recognition logic, and data semantic recognition logic based on the feature mapping relationship table, identify duplicate or similar fragments, and output the initial duplicate data set and its location information;

[0043] The context association and difference discrimination module is used to construct a context association window for the detected duplicate data based on its position in the target data chain, extract the features of these paragraphs, perform semantic / structural / functional integrated analysis, and determine whether the duplicate data is unique in different contexts and positions;

[0044] The time series operation association and dynamic correction module is used to analyze the behavioral impact on primary redundant data in a time series environment by combining operation logs, extract dynamic correction rules, and generate a redundant data correction set;

[0045] The cleanup instruction queue generation and policy allocation module generates a multi-level cleanup instruction queue based on the redundant data correction set, according to the redundancy type priority and dynamic correction rules, and allocates resource thresholds and protection mechanisms for each instruction;

[0046] The monitoring and feedback module is used to monitor the cleaning execution effect in real time, provide feedback to the AI ​​engine and policy template, continuously iterate and optimize model parameters and rules, and generate system optimization index reports, visual dashboards, or log reports;

[0047] The configuration and management interface is used to provide system administrators or business personnel with an interactive operation interface for configuring feature extraction plans, defining threshold policies, viewing monitoring indicators, approving cleanup plans, and restoring snapshots.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] This AI-driven, efficient data redundancy detection and cleanup method and system uses a fragment recognition architecture to perform structured analysis of target data chains. Combining structural, semantic, and functional features, it locates and identifies duplicate data in different contexts and data types. It is applicable to diverse data scenarios such as text sequences, data streams, and multimedia images. Furthermore, by constructing a feature mapping relationship table, it can build a fragment recognition framework based on the characteristic content of different paragraphs. Using AI recognition logic matching, it achieves high adaptability and recognition accuracy when faced with complex data structures and implicit redundancy issues.

[0050] It should be noted that by introducing contextual association analysis, a window-based feature set is established for the preceding and following paragraphs of duplicate data to determine whether the data has independent or different business functions in a specific location. This context-aware identification method effectively avoids the risk of "accidentally cleaning" important data in traditional redundancy cleaning, ensuring the accuracy of redundancy judgment and the stability of business operations.

[0051] It is worth noting that in terms of dynamic adjustment, by analyzing the behavioral connection between operating instructions and redundant data, various redundancy types such as write not being called, invalid call and function transfer can be identified, so that redundancy identification not only stays at the data level, but further associates operating behavior and system changes, reflecting the dynamic value of data in actual use.

[0052] It should also be noted that during the cleanup execution process, a multi-level cleanup instruction queue was established, and a hierarchical instruction sequence was generated based on different redundancy types and correction rules. Resource usage thresholds and rollback protection mechanisms were configured for them, and storage space recovery rates and computing resource consumption control parameters were set. Snapshot and version verification interfaces were introduced, so that data security and rapid recovery can be achieved during the cleanup process, improving the security and controllability of the cleanup operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A flowchart of the steps disclosed in the present invention;

[0054] Figure 2 A supplementary illustration of step S100 of the present invention;

[0055] Figure 3 A supplementary illustration of step S103 of the present invention;

[0056] Figure 4 This is a composition diagram of the system disclosed in the present invention. DETAILED DESCRIPTION

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0058] Before understanding the technical solution proposed in this application, it should be clear that data redundancy in the prior art is not only manifested in the form of simple text but also in the form of multimedia images. These redundant data come from repeated storage, invalid backups, or the accumulation of intermediate data generated during system operation. Different forms of data redundancy require different strategies and technical means in the detection and cleaning process. For example, for redundant data in text form, semantic analysis and similarity comparison are used to identify duplicate content, while for redundancy in multimedia image form, image feature extraction and feature matching are required for processing. In addition, in the actual cleaning process, considering the data security and integrity issues involved in the data redundancy cleaning process, corresponding protection mechanisms need to be designed to ensure that important information will not be accidentally deleted or damaged.

[0059] For this reference Figure 1 It can be seen that the present invention provides a technical solution: an AI-driven efficient data redundancy detection and cleaning method, including: steps S100 to S500.

[0060] Step S100: Based on the categories of different segments in the target data chain, an AI-driven segment recognition architecture is constructed. The segment recognition architecture is used to locate the location of duplicate data in the target data chain.

[0061] It is worth noting that in step S100, the target data chain refers to a data aggregate composed of multiple data fragments or data units and having a business logic order or access order in a system, platform or application scenario. The target data chain includes: any one or more combinations of a text sequence in the physical sense, a logical data flow, and a multimedia image sequence.

[0062] It should be added that, in the actual detection process, there are multiple fragment identification frameworks set in the target data chain in step S100, where the fragment identification frameworks are set according to different sections of the target data chain. It should be emphasized that the feature contents extracted by each fragment identification framework are related to each other, where the feature contents include: data structure, data application and data semantics.

[0063] It is also necessary to add that structural features are used to describe the format and hierarchical relationship of data, semantic features are used to represent the vector embedding of data content, and functional features are used to reflect its operational semantics or behavioral meaning in business scenarios.

[0064] As a preferred implementation scheme, since the target data chain may be composed of different heterogeneous data phases during the data processing process, the diversity of data types and the compatibility between heterogeneous data need to be considered when actually constructing the fragment recognition framework.

[0065] Specifically, refer to Figure 2It can be seen that this embodiment is mainly used to supplement the construction of the AI-driven fragment recognition framework in step S100, including: steps S101 to S103.

[0066] Step S101: constructing an interactive path adapted to different sections in the target data chain according to their type characteristics.

[0067] It should be noted that the interaction path is embedded with a feature processing engine that is adapted to the type characteristics of the target data link. The feature processing engine is used to integrate and unify the type characteristics of different sections within the target data link.

[0068] It is worth noting that in step S101, the type features of the target data link include: text feature type, data stream feature type and multimedia image feature type, among which the method for distinguishing the type features of the target data link is to automatically identify them through the feature extraction algorithm and pattern matching model in the existing technology.

[0069] Specifically, the text feature type adopts the word vector embedding technology in natural language processing in the existing technology, and extracts syntactic structure features through the known BERT model for identification. The data stream feature type uses the known sliding window mechanism combined with the known LSTM network to capture the timing characteristics for identification. The multimedia image feature type uses the known convolutional neural network (CNN) to extract spatial topology features for identification.

[0070] It should be pointed out in particular that during the execution of step S101, feature processing engines that form matching relationships with different target data link types are positioned and implanted in the mutual path, where these feature processing engines are all constructed based on mature solutions under the existing technical framework. Specifically, the text type enables the known NLP analysis engine, the data stream type enables the known real-time computing engine, and the multimedia type enables the known computer vision engine. It should be pointed out in particular that the type feature extraction of the target data link and the mutual path construction are both based on the conventional technical means of those skilled in the art. In view of the fact that the technical means involved belong to the scope of industry standardized operations, the applicant has not elaborated on the above technical details.

[0071] Step S102: Based on the feature contents of different paragraphs in the target data chain after interactive path processing, a plurality of associated feature mapping relationship tables are constructed. The feature mapping relationship tables are used to establish an AI-driven segment recognition framework based on the feature contents of each paragraph in the target data chain.

[0072] It is worth noting that in the content of step S102, each feature mapping relationship table generates a function tag based on the feature content when it is established, and the function tag is used to record the feature content in the corresponding paragraph.

[0073] Step S103: establishing an AI-driven segment recognition framework based on the feature mapping relationship table.

[0074] It is worth noting that in step S103, the fragment recognition framework is integrated with AI recognition logic that matches the feature content. Specifically, in the technical solution proposed in this application, the AI ​​recognition logic is divided into data structure recognition logic, data application logic and data semantic logic. The data structure recognition logic is used to identify repeated structured information, the data application recognition logic is used to identify data with the same behavior pattern in the application scenario, and the data semantic recognition logic is used to identify unstructured semantic data with similar or equivalent content.

[0075] As a preferred embodiment, refer to Figure 3 It can be seen that this embodiment is mainly used to supplement the working principle of the fragment identification framework in step S103, and specifically includes: step S103A to step S103C.

[0076] Step S103A: construct an input feature set library for the segment recognition framework by segmenting and integrating the feature contents of different paragraphs in the target data chain and storing them.

[0077] It should be noted that step S103A is used to construct a multi-type input interface for the AI ​​recognition logic, and the input feature set library will be arranged in sequence according to the position of the paragraph in the target data chain. It should be added that the sequence arrangement order is in either logical order or access order.

[0078] Step S103B: establishing a duplicate data set based on the contents between the multiple mapping relationship tables.

[0079] It is worth noting that in step S103B, the target area link segments with the same characteristic content are integrated and associated based on the contents between multiple mapping relationship tables through the characteristic content corresponding to the function tag, and repeated feature binding is performed using the function tag as the prefix of the repeated data set.

[0080] It should be added that the purpose of the duplicate data set established in step S103B is to allow the AI ​​recognition logic to obtain an initial recognition concept template. The AI ​​recognition logic searches the input feature set library through the initial recognition concept template to determine whether there is duplicate data in different paragraphs in the target data chain.

[0081] Step S103C: Associating the repeated data set with the input feature set library, and outputting the corresponding repeated data and sequence numbers.

[0082] It should be added that the association process of the repeated data set and the input feature set library in step S103C depends on the similarity matching of the feature content recorded in the function tag. Specifically, based on the feature content in the function tag, the target paragraph corresponding to the repeated data set is located in the input feature set library, and a result list containing the repeated data content and its corresponding serial number is generated. This result list marks the specific location of the repeated data in the target data chain.

[0083] Step S200: Using the repeated data as a positioning reference, the feature content within the recognition framework of the fragments in the sequence before and after the positioning reference in the target data chain is associated to form a feature set.

[0084] It should be noted that step S200 is intended to identify and locate duplicate data segments in the target data chain, and contextually integrate the semantic interpretation and scenario function analysis of the target data chain segments corresponding to the duplicate data with the semantic interpretation and scenario function analysis in adjacent segments, and then compare the semantic features and functional attributes presented by the duplicate data at different positions in the target data chain, establish a difference discrimination mechanism, and effectively verify whether the duplicate data is unique.

[0085] As a preferred implementation scheme, this implementation scheme is used to supplement the method for forming the feature set in step S200.

[0086] Specifically, it includes: step S201-step S203.

[0087] Step S201: constructing a context association window based on the position of the duplicate data in the target data chain.

[0088] It is worth noting that the context association windows in step S201 are mainly identified in sequence by sequence numbers, thereby determining the positions of the upper and lower association windows in the target data chain.

[0089] It should be further explained that in the technical solution proposed in this application, the repeated data set output in step S103C is used as the positioning reference, the serial number of the segment where it is located is extracted in the target data chain, and based on the sorting of the serial numbers, several series serial number segments are extended forward and backward to form a context association window based on the repeated data.

[0090] Step S202: calling the feature content in the corresponding segment recognition framework, and performing hierarchical extraction based on the type of the extracted feature content.

[0091] Step S203: Classify and integrate the extracted feature contents to form a feature set within the context association window.

[0092] It should be emphasized that the feature set in step S203 includes not only the semantic features, structural features, and functional features of the repeated data itself, but also the semantic features, structural features, and functional features of the sequence numbered paragraphs before and after the repeated data.

[0093] Step S300: Perform comparative analysis of feature sets on similar duplicate data in the target data chain, and then extract and output primary redundant data.

[0094] It should be noted that step S300 is mainly used to further screen and classify the duplicate data in the target data chain in order to identify the truly redundant parts. In this process, based on the feature set formed in step S203, through comparative analysis of semantic, structural and functional features, it is determined whether the duplicate data is irreplaceable in semantics or function. If the duplicate data in a certain segment within the target data chain exhibits consistent semantic content and functional attributes at different locations, the data is marked as primary redundant data.

[0095] It is worth noting that in step S300, the comparative analysis of semantic, structural and functional features between feature sets is achieved through the similarity measurement algorithm in the existing technology. Specifically, the semantic features are analyzed using the cosine similarity calculation model based on the known BERT, the structural features are analyzed for topological matching through the known graph neural network (GNN), and the functional features are analyzed using the association evaluation algorithm based on the known business scenario knowledge graph.

[0096] Step S400: in a time series data environment, extracting an operation data set by parsing an operation instruction, establishing a dynamic correction rule for the operation data set and primary redundant data, and generating an optimized redundant data correction set.

[0097] It should be noted that step S400 is mainly used to solve the problem of redundancy identification in dynamic data scenarios. By capturing the timing operation instructions (including data writing, modification, and migration operations) generated during the operation of the target data link, dynamic correction rules associated with the primary redundant data are constructed.

[0098] As a preferred implementation scheme, this implementation scheme is used to supplement and explain the dynamic modification rules in detail, and specifically includes: Step S401 to Step S404.

[0099] Step S401: extracting an operation instruction sequence based on an operation log generated by the target data link during the sequential operation.

[0100] It should be noted that the operation instruction sequence refers to the time sequence of the generation of sequential operation instructions. Specifically, in step S401, the behavior trajectory on the target data chain is captured in real time, where the behavior trajectory refers to the user's action on the target data chain based on the sequential operation instruction, and is sorted according to the timestamp to form an original instruction log with a complete operation process record. It is particularly pointed out that the extracted operation instruction sequence not only records the operation action, but also includes the operation target (i.e., the position acting on the target data chain) and the auxiliary information of the operation period.

[0101] Step S402: extracting a matching operation instruction set based on the sequence number and type of the segment where the primary redundant data is located.

[0102] It should be noted that step S402 is mainly used to associate the relationship between the primary redundant data and the operation instruction set. Specifically, based on the operation instruction sequence, this step obtains the serial number of the corresponding paragraph of the primary redundant data as an index, and filters the operation instructions that have acted on the paragraph in the duplicate data.

[0103] Step S403: performing behavior pattern planning on the extracted operation instruction set, and analyzing the behavior variation of the extracted operation instruction set on the primary redundant data in different contexts.

[0104] It should be noted that step S403 is mainly used to analyze the dynamic impact of the operation instruction on the primary redundant data in different context environments.

[0105] As a preferred embodiment, this embodiment is mainly used to illustrate the dynamic impact, specifically including the following contents:

[0106] First, there is the redundancy of uncalled write behavior (temporary cache redundancy). It should be noted that the redundancy of uncalled write behavior is manifested as a data segment in the target data chain being overwritten by multiple write operations, but has never been accessed, read or referenced in its subsequent operation sequence, indicating that the data segment may only be used for temporary storage or intermediate calculations.

[0107] It is worth noting that the basis for determining whether the write behavior is not called redundancy in step S403 is to detect whether the primary redundant data is subsequently called after the primary redundant data is written, and whether the position of the primary redundant data in the target data chain is changed after the write operation.

[0108] Second, the calling behavior is repeated but the result is not changed (invalid calling redundancy). It should be noted that the impact of the redundancy of repeated calling behavior but no change in result is that within multiple operation time periods, the same data segment in the target data chain is repeatedly read or called, but the system state of the target data chain after the call has no obvious change, indicating that these calls have no actual impact on the system behavior.

[0109] It is worth noting that the judgment basis for the repeated calling behavior but unchanged redundancy in step S403 is to obtain whether the data segment of the primary redundant data in the target data chain has a continuous or periodic reading operation, and then whether the data segment of the primary redundant data changes after each reading.

[0110] In addition, it should be noted that the specific manifestations of no obvious change in the system status include: the hash value of the data segment output by the system that generates the target data link has not changed. Among them, the detection of no change in the hash value of the system data segment that generates the target data link is achieved through a standardized hash algorithm. Specifically, the standard SHA-256 algorithm is used to perform real-time hash calculations on the data segments corresponding to the primary redundant data in the target data link, and the newly generated hash value after each read operation is compared with the previously recorded historical hash value to determine whether the system status has shown no obvious change.

[0111] Third, functional behavior transfer redundancy (phase redundancy). It should be noted that the impact of functional behavior transfer redundancy is manifested as follows: the data segment in the target data link assumes function during the system execution phase, but is replaced after the system completes the execution phase.

[0112] It is worth noting that the basis for judging functional behavior transfer redundancy in step S403 is to obtain the calling frequency of the primary redundant data in the target data chain, the interaction frequency between the primary redundant data and the new data segment, and the new calling frequency of the primary redundant data and the new data segment after the interaction.

[0113] Step S404: constructing dynamic correction rules between operation behaviors and redundant data based on the analysis results.

[0114] It should be noted that the dynamic correction rules in step S404 can be implanted into the AI ​​recognition logic in step S103 to form a correction strategy template. Specifically, the specific contents of the dynamic correction rules in step S404 include: redundant correction rules of the write behavior that are not called, that is, a "non-reference timeout threshold" is set in the AI ​​logic. If no reference operation related to the data segment occurs within the time range of the "non-reference timeout threshold", it is marked as "temporary cache redundancy". Redundant correction rules of repeated calling behavior but unchanged results, that is, the system behavior differences (such as state hash and structure summary code) of the target data chain are generated after each call. If N times If the system behavior summary corresponding to the call is consistent and does not affect the data chain path or output results, the subsequent redundant calls will be marked as "invalid operations". The subsequent calls will be set to an ignorable state in the AI ​​control logic, and the "optimizable path suggestions" will be recorded. The functional behavior transfer redundancy correction rule, that is, by recording the call or collaborative relationship between the primary redundant data segment and other data segments, if there is a data segment that replaces the behavior path of the primary redundant data segment and the call frequency of the data segment is higher than that of the primary redundant data segment, it is judged that the function corresponding to the primary redundant data segment has been transferred. At this time, the correction suggestion is output: let the primary redundant data segment enter the "frozen state" or be moved to the secondary call cache.

[0115] Step S500: Based on the redundant data correction set, a multi-level cleaning instruction queue is constructed and a cleaning strategy is assigned.

[0116] It should be noted that in step S500, the construction of the multi-level cleanup instruction queue is hierarchically processed according to the redundancy type marked in the redundant data correction set and its dynamic correction rules. Specifically, a cleanup instruction chain with a timing dependency will be generated according to the redundancy type priority (temporary cache redundancy > invalid call redundancy > phased redundancy), and differentiated execution strategies will be configured for each level of cleanup instructions.

[0117] As a preferred implementation scheme, this implementation scheme is used to supplement the construction logic and policy allocation mechanism of the multi-level cleaning instruction queue in step S500, and specifically includes: steps S501 to S503.

[0118] Step S501: parsing the correction policy template in the redundant data correction set to generate a priority instruction sequence based on the redundancy type.

[0119] It should be noted that the generation of the priority instruction sequence in step S501 follows the following principles: for data segments marked as "temporary cache redundancy", immediate cleanup instructions will be generated first; for data segments corresponding to "invalid call redundancy", delayed cleanup instructions with a conditional trigger mechanism will be generated (the target data link operating environment status needs to be verified); for "phased redundancy" data segments, observation period retention instructions based on functional replacement verification will be generated.

[0120] Step S502: configuring resource occupancy thresholds and rollback protection mechanisms for each level of cleanup instructions.

[0121] It is worth noting that the resource occupancy threshold in step S502 includes the storage space recovery rate and the computing resource consumption peak parameters. Resource quotas are allocated to instructions of different priorities through a preset priority instruction sequence, and the rollback protection mechanism automatically generates a data snapshot before the execution of the cleanup instruction and establishes a two-way verification interface with the target data link version control system to ensure that the original data link can be quickly restored in an abnormal state.

[0122] Step S503: Embed the multi-level cleaning instruction queue into the AI-driven automated cleaning engine to start the policy verification closed loop.

[0123] It should be added that the strategy verification closed loop in step S503 includes a three-layer verification mechanism: the first layer verifies data integrity by real-time monitoring of the state hash value of the target data link during the execution of the cleaning instruction; the second layer calls the fragment recognition framework in step S103 to recheck the features of the cleaned data link to confirm the effect of eliminating redundant features; the third layer combines the dynamic correction rules generated in step S400 to evaluate the long-term impact of the cleaning operation on the behavior pattern of the time series data environment and generate an optimization index report.

[0124] As a preferred implementation scheme, this implementation scheme is used to supplement an AI-driven efficient data redundancy detection and cleaning method, refer to Figure 4 It can be seen that an AI-driven efficient data redundancy detection and cleaning system is proposed, including:

[0125] The data acquisition and preprocessing module is used to obtain raw data from the target system / platform / application scenario in real time or in batches, including logs, text, data streams, and multimedia (images / videos), and to unify the format, cleanse, and remove noise from the raw data in preparation for subsequent feature extraction.

[0126] It is worth noting that the data collection and preprocessing module is mainly related to the data acquisition before step S401, and is regarded as a pre-link before the execution of an AI-driven efficient data redundancy detection and cleaning method.

[0127] The feature extraction and mapping module is used to extract "structural features", "semantic features" and "functional features" for different data types (text, time series / data streams, multimedia images) using corresponding models, and generate unified feature vectors, and then construct a feature mapping relationship table based on these feature vectors.

[0128] It is worth noting that the functions included in the feature extraction and mapping module include text feature extraction engine, time series data feature extraction, image feature extraction, and extraction of functional features and scene features. In addition, when the feature extraction and mapping module is used in specific situations, it outputs feature vectors, function labels, and feature mapping relationship tables, corresponding to steps S101 and S102.

[0129] The AI-driven fragment recognition engine is used to execute data structure recognition logic, data application / behavior recognition logic, and data semantic recognition logic based on the feature mapping relationship table, identify repeated or similar fragments, and output the initial repeated data set and its location information.

[0130] It is worth noting that the AI-driven fragment recognition engine includes: an input interface for accessing the "input feature set library" constructed in step S103A, and recognition logic, which is composed of a structural similarity module, a semantic similarity module, and a functional / behavioral similarity module. It should be added that the AI-driven fragment recognition engine corresponds to the contents of steps S103A to S103C.

[0131] The context association and difference discrimination module is used to construct a context association window (several segments before and after) for the detected duplicate data based on its position in the target data chain, extract the features of these segments, perform semantic / structural / functional integration analysis, and determine whether the duplicate data is unique in different contexts and positions (whether it can be considered truly redundant).

[0132] It is worth noting that the context association and difference identification module corresponds to steps S201 to S203 and step S300.

[0133] The time series operation association and dynamic correction module is used to analyze the behavioral impact on primary redundant data in a time series environment by combining operation logs (write, read, modify, and other operation instruction sequences), extract dynamic correction rules, and generate redundant data correction sets.

[0134] It is worth noting that the timing operation association and dynamic correction module correspond to steps S401 to S404.

[0135] The cleanup instruction queue generation and policy allocation module generates a multi-level cleanup instruction queue based on the redundant data correction set, according to the redundancy type priority and dynamic correction rules, and allocates resource thresholds and protection mechanisms for each instruction.

[0136] It is worth noting that the generation of the clearing instruction queue and the policy allocation module correspond to the contents of steps S501 to S503.

[0137] The monitoring and feedback module is used to monitor the cleaning execution effect in real time, provide feedback to the AI ​​engine and policy templates, continuously iterate and optimize model parameters and rules, and generate system optimization index reports, visual dashboards or log reports.

[0138] The configuration and management interface is used to provide system administrators or business personnel with an interactive operation interface for configuring feature extraction plans, defining threshold policies, viewing monitoring indicators, approving cleanup plans, and restoring snapshots.

[0139] It should be noted that the monitoring and feedback module and the configuration and management interface are linked, and the optimization index report, visual dashboard or log report can be visualized through the configuration and management interface.

[0140] Although embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is limited by the accompanying embodiments and their equivalents.

Claims

1. An AI-driven efficient data redundancy detection and cleaning method, characterized in that: include: Based on the categories of different segments within the target data chain, an AI-driven segment recognition architecture is built to locate the location of duplicate data in the target data chain; Using repeated data as the positioning reference, the feature content within the recognition framework of the fragments before and after the positioning reference in the target data chain is associated to form a feature set. Compare and analyze the feature sets of similar duplicate data in the target data chain, and then extract and output primary redundant data; In a time series data environment, the operation data set is extracted through operation instruction analysis, dynamic correction rules for the operation data set and primary redundant data are established, and an optimized redundant data correction set is generated; Based on the redundant data correction set, a multi-level cleaning instruction queue is constructed and a cleaning strategy is assigned; The target data chain includes: any one or more combinations of text sequences in the physical sense, logical data flows, and multimedia image sequences. The feature content includes: structural features for describing the format and hierarchical relationship of data, semantic features for representing vector embedding of data content, and functional features for reflecting its operational semantics or behavioral meaning in business scenarios.

2. The AI-driven efficient data redundancy detection and cleaning method according to claim 1, characterized in that: The method for constructing the fragment recognition framework includes: According to the type characteristics of different sections in the target data chain, construct an interactive path that is suitable for them; Based on the characteristic content of different paragraphs in the target data chain after interactive path processing, multiple associated feature mapping relationship tables are constructed. The feature mapping relationship tables are used to establish an AI-driven segment recognition framework based on the characteristic content of each paragraph in the target data chain. The segment recognition framework integrates AI recognition logic that matches the characteristic content; An AI-driven fragment recognition framework is established based on the feature mapping relationship table.

3. The AI-driven efficient data redundancy detection and cleaning method according to claim 2, characterized in that: The method of identifying and locating the location of repeated data in the target data chain of the fragment recognition framework includes: By integrating and storing the feature content of different paragraphs in the target data chain in segments, an input feature set library of the segment recognition framework is constructed. The input feature set library is arranged in sequence according to the position of the paragraphs in the target data chain. The sequence arrangement order is either logical order or access order. A duplicate data set is created based on the contents between multiple mapping relationship tables.

4. The AI-driven efficient data redundancy detection and cleaning method according to claim 1, characterized in that: Methods for forming feature sets include: Building a context window based on the location of the duplicate data in the target data chain; Call the feature content in the corresponding fragment recognition architecture and perform hierarchical extraction based on the type of extracted feature content; The extracted feature contents are classified and integrated to form a feature set within the context association window.

5. The AI-driven efficient data redundancy detection and cleaning method according to claim 2, characterized in that: Dynamic correction rules include: Extracting an operation instruction sequence based on the operation log generated by the target data link during the timing operation, where the operation instruction sequence refers to the time sequence of the timing operation instructions; Extracting a matching set of operation instructions based on the sequence number and type of the segment where the primary redundant data is located; Conduct behavioral pattern planning for the extracted operation instruction set, and analyze the behavioral variation of the extracted operation instruction set for primary redundant data in different contexts; Based on the analysis results, dynamic correction rules between operational behaviors and redundant data are constructed and embedded into the AI ​​recognition logic to form a correction strategy template.

6. The AI-driven efficient data redundancy detection and cleaning method according to claim 5, characterized in that: The behavioral variations include: Write behavior uncalled redundancy is used to indicate that a data segment in the target data chain is overwritten by multiple write operations, but has never been accessed, read, or referenced in its subsequent operation sequence, indicating that a data segment is only used for temporary storage or intermediate calculation. Repeated call behavior but unchanged result redundancy indicates that the same data segment in the target data chain is read or called repeatedly during multiple operation periods, but the hash value of the data segment output by the system that generates the target data chain after the call does not change, indicating that these called data segments have no impact on system behavior; Functional behavior transfer redundancy is used to indicate that the data segment in the target data link assumes a function during the execution phase of the system that generates the target data link, but is replaced after the execution phase is completed.

7. The AI-driven efficient data redundancy detection and cleaning method according to claim 5, characterized in that: Generation of a cleanup strategy includes: Parsing the correction strategy template in the redundant data correction set to generate a priority instruction sequence based on the redundancy type; Resource usage thresholds and rollback protection mechanisms are configured for each level of cleanup instructions. The resource usage thresholds include storage space recovery rates and computing resource consumption peak parameters. Resource quotas are allocated to instructions of different priorities through a preset priority instruction sequence. The rollback protection mechanism automatically generates a data snapshot before the execution of the cleanup instruction and establishes a two-way verification interface with the target data link version control system to ensure that the original data link can be quickly restored in abnormal conditions. Embed a multi-level cleaning instruction queue into the AI-driven automated cleaning engine to initiate a closed-loop strategy verification.

8. An AI-driven efficient data redundancy detection and cleaning system, using an AI-driven efficient data redundancy detection and cleaning method according to any one of claims 1 to 7, characterized in that: include: Data acquisition and preprocessing module, used to obtain raw data from the target system / platform / application scenario in real time or in batches; The feature extraction and mapping module is used to extract "structural features", "semantic features" and "functional features" for different data types using corresponding models, generate unified feature vectors, and then construct a feature mapping relationship table based on these feature vectors; An AI-driven fragment recognition engine, which is used to execute data structure recognition logic, data application / behavior recognition logic, and data semantic recognition logic based on the feature mapping relationship table, identify duplicate or similar fragments, and output the initial duplicate data set and its location information; The context association and difference discrimination module is used to construct a context association window for the detected duplicate data based on its position in the target data chain, extract the features of these paragraphs, perform semantic / structural / functional integrated analysis, and determine whether the duplicate data is unique in different contexts and positions; The time series operation association and dynamic correction module is used to analyze the behavioral impact on primary redundant data in a time series environment by combining operation logs, extract dynamic correction rules, and generate a redundant data correction set; The cleanup instruction queue generation and policy allocation module generates a multi-level cleanup instruction queue based on the redundant data correction set, according to the redundancy type priority and dynamic correction rules, and allocates resource thresholds and protection mechanisms for each instruction; The monitoring and feedback module is used to monitor the performance of cleanup execution in real time, provide feedback to the AI ​​engine and policy templates, continuously iterate and optimize model parameters and rules, and generate system optimization index reports, visual dashboards, or log reports; The configuration and management interface is used to provide system administrators or business personnel with an interactive operation interface for configuring feature extraction plans, defining threshold policies, viewing monitoring indicators, approving cleanup plans, and restoring snapshots.