Data Detection Method and System Based on ElasticSearch

By constructing a composite density tree index and a reversible feature transformation matrix in ElasticSearch, combining dynamic weighting and asymptotic weighted cosine similarity function, the problem of low repetitive detection of large-scale data is solved, and efficient and accurate data repetitive detection is achieved.

CN119884346BActive Publication Date: 2025-07-18SHANDONG PROVINCIAL INST OF LAND & SPACE DATA & REMOTE SENSING TECH (SHANDONG PROVINCIAL SEA AREA DYNAMIC SURVEILLANCE & MONITORING CENT)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510371770.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing ElasticSearch-based data repeatability detection methods are inefficient when facing large-scale data, and have poor detection of fuzzy or non-accurate duplicate data, especially when text containing noise or mutated is difficult to efficiently and accurately identify.

Method used

The composite density tree index based on ElasticSearch is used for pre-screening, combining the asymptotic weighted cosine similarity function of the reversible feature transformation matrix and the dynamic weighting, and a reversible feature transformation matrix is generated through the decomposed attention mechanism, adaptive dynamic weights are constructed, and incremental parameter optimization is used to select the optimal calculation path for repetitive detection.

Benefits of technology

It significantly improves the accuracy and efficiency of data repetition detection, especially in deformation repetition and complex repetition modes, it has high accuracy, low computational complexity and good scalability, strong adaptability, and can effectively process large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884346B_ABST
    Figure CN119884346B_ABST
Patent Text Reader

Abstract

The present invention relates to a data detection method and system based on ElasticSearch, belonging to the technical field of data processing. The method includes: constructing a composite density tree index based on ElasticSearch, and pre-screening a similar data set based on the composite density tree index to obtain a target document and candidate documents; performing anti-interference reconstruction on the target document and candidate documents based on a reversible feature transformation matrix to obtain a reconstructed target document and reconstructed candidate documents; constructing an adaptive dynamic weight; calculating the similarity between the target document and the candidate documents; determining the dynamically updated dynamic weight; dynamically selecting an optimal calculation path based on the incrementally updated dynamic weight, and performing repetitive detection on the target document and candidate documents again based on the optimal calculation path to obtain a repetitive detection result. The present invention can improve the accuracy and efficiency of data repetitive detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data detection method, system, electronic device, and non-transitory computer-readable storage medium based on ElasticSearch. Background Art

[0002] Nowadays, data duplication detection methods mainly rely on traditional techniques based on string matching, hash value comparison, or similarity algorithms. Common methods include calculating text similarity to determine whether two data items are duplicates, which are widely used in text data processing, information retrieval, and data deduplication. Specifically, many applications utilize the full-text search function provided by ElasticSearch and the similarity calculation based on TF-IDF, combined with algorithms such as cosine similarity, to perform duplication detection. These methods can effectively extract documents with high similarity from massive data and then determine whether they are duplicate content.

[0003] However, traditional duplication detection based on full-text search and similarity algorithms is often restricted by performance bottlenecks. As the amount of data increases, the computational complexity and storage requirements will grow rapidly, resulting in a decrease in retrieval efficiency. In addition, existing methods have poor detection effects on fuzzy or non-exact duplicate data. Especially when facing text containing noise or mutations, traditional similarity algorithms are difficult to efficiently and accurately identify all types of duplicate data. Summary of the Invention

[0004] The present invention aims at the technical problems existing in the prior art, and provides a data detection method, system, electronic device, and non-transitory computer-readable storage medium based on ElasticSearch that can improve the efficiency and accuracy of data duplication detection.

[0005] The technical solution for the present invention to solve the above technical problems is as follows:

[0006] The present invention provides a data detection method based on ElasticSearch, and the method includes:

[0007] Construct a composite density tree index based on ElasticSearch, and pre-screen a similar data set based on the composite density tree index to obtain a target document and a candidate document;

[0008] Generate a reversible feature transformation matrix using a decomposable attention mechanism, and perform anti-interference reconstruction on the target document and the candidate document based on the reversible feature transformation matrix to obtain a reconstructed target document and a reconstructed candidate document;

[0009] Construct an adaptive dynamic weight based on the entropy differences of each dimension in the reconstructed target document and the reconstructed candidate document;

[0010] Calculate the similarity between the target document and the candidate document based on the asymptotic weighted cosine similarity function of the dynamic weight;

[0011] Use the frequency vectorization mechanism for incremental parameter optimization to determine the dynamically updated dynamic weight;

[0012] Dynamically select the optimal calculation path based on the incrementally updated dynamic weight, and re-detect the target document and the candidate document based on the optimal calculation path to obtain the duplicate detection result.

[0013] Optionally, the generating the reversible feature transformation matrix using the decomposable attention mechanism includes:

[0014] Determine the query mapping of each original feature according to the position self-correlation tensor representing the correlation between different positions in the input data and the decomposable query head matrix;

[0015] Combine each of the original features with each of the query mappings to obtain combined features;

[0016] Determine and minimize the difference between each of the combined features and each target feature to generate the reversible feature transformation matrix.

[0017] Optionally, the reversible feature transformation matrix is expressed as:

[0018] ;

[0019] Wherein, is the reversible feature transformation matrix, is the decomposable query head matrix, and are the position self-correlation tensors, is the original feature, is the target feature after removing noise interference, is the total number of samples.

[0020] Optionally, the constructing the adaptive dynamic weight based on the entropy differences of each dimension in the reconstructed target document and the reconstructed candidate document includes:

[0021] Determine the global entropy of the similar data set and the local entropy of each feature dimension in the similar data set;

[0022] Normalize the difference between the global entropy and the local entropy to obtain a normalized difference;

[0023] Process the normalized difference through exponential operation of mutual information intensity to obtain the dynamic weight.

[0024] Optionally, the dynamic weight is expressed as:

[0025] ;

[0026] where is the dynamic weight, is the sensitivity control parameter, is the sliding window variance, is the third-order mutual information intensity measurement value, is the attenuation coefficient, is the global entropy, is the feature dimension is the local entropy of is the smoothing term.

[0027] Optionally, calculating the similarity between the target document and the candidate document using the asymptotic weighted cosine similarity function based on the dynamic weight includes:

[0028] Obtain the first eigenvalue of each dimension in the set of the target document and the second eigenvalue of each dimension in the set of the candidate document;

[0029] Perform power transformation processing on the first eigenvalue and the second eigenvalue to obtain the power-transformed first eigenvalue and the power-transformed second eigenvalue;

[0030] Calculate the distance between the first eigenvalue and the second eigenvalue;

[0031] Perform normalization processing on the power-transformed first eigenvalue and the power-transformed second eigenvalue to obtain normalized eigenvalues;

[0032] Determine the similarity between the target document and the candidate document based on the distance between the first eigenvalue and the second eigenvalue, the power-transformed first eigenvalue, the power-transformed second eigenvalue, and the normalized eigenvalues.

[0033] Optionally, using the frequency vectorization mechanism for incremental parameter optimization to determine the incrementally updated dynamic weight includes:

[0034] Determine the basic update amount of the dynamic weight according to the learning rate of the dynamic weight update and the gradient of the loss function of the similarity with respect to the dynamic weight update;

[0035] Perform feature calibration processing on the eigenvalues corresponding to the dimensions of the dynamic weight to obtain the corresponding feature calibration values;

[0036] The similarity normalizes the L1.5 norm of the gradient vector of the dynamic weight to obtain a gradient normalization value;

[0037] Perform momentum adjustment on the change of the reversible feature transformation matrix to obtain a dynamic adjustment amount;

[0038] Determine the dynamically weighted increment update based on the base update amount, the feature calibration value, the gradient normalization value, and the dynamic adjustment amount.

[0039] Optionally, the optimal calculation path is calculated based on a hybrid sharding architecture that adopts a heterogeneous computing unit sharding strategy, where:

[0040] The CPU shards to process the CDTI pre-screening task, the GPU shards to execute the AWCS similarity calculation, and low-latency communication is achieved between shards through RDMA.

[0041] Optionally, the method further includes determining the sharding efficiency of the unit sharding strategy, including:

[0042] Calculate the sharding efficiency impact value based on the communication latency time for sending and receiving data between each shard, and the sharding cost of each shard in terms of computing and storage;

[0043] Calculate the weight difference between the weight vector of the master server and the weight vector of the shard server;

[0044] Determine the sharding efficiency of the unit sharding strategy based on the sharding efficiency impact value and the weight difference.

[0045] The present invention also provides a data detection system based on ElasticSearch, the system includes:

[0046] A data screening module, configured to construct a composite density tree index based on ElasticSearch, and pre-screen a similar data set based on the composite density tree index to obtain a target document and candidate documents;

[0047] A data reconstruction module, configured to generate a reversible feature transformation matrix using a decomposable attention mechanism, and perform anti-interference reconstruction on the target document and the candidate documents based on the reversible feature transformation matrix to obtain a reconstructed target document and reconstructed candidate documents;

[0048] A weight determination module, configured to construct an adaptive dynamic weight based on the entropy difference of each dimension between the reconstructed target document and the reconstructed candidate documents;

[0049] A preliminary detection module, configured to calculate the similarity between the target document and the candidate documents based on the asymptotically weighted cosine similarity function of the dynamic weight;

[0050] A weight update module, which is used to perform incremental parameter optimization by using a frequency vectorization mechanism to determine dynamic weights for incremental updates;

[0051] A target detection module, which is used to dynamically select an optimal calculation path based on the dynamically updated dynamic weights, and perform repetitive detection on the target document and the candidate document again based on the optimal calculation path to obtain a repetitive detection result.

[0052] In addition, to achieve the above object, the present invention also provides an electronic device, including: a memory for storing a computer software program; a processor for reading and executing the computer software program, thereby implementing a data detection method based on ElasticSearch as described above.

[0053] In addition, to achieve the above object, the present invention also provides a non-transitory computer-readable storage medium, in which a computer software program is stored, and when the computer software program is executed by a processor, a data detection method based on ElasticSearch as described above is implemented.

[0054] The beneficial effects of the present invention are:

[0055] (1) The present invention adopts a dynamic weighted cosine similarity calculation method, combines information entropy density ratio and asymptotic weighting, and can calculate the similarity between data sets more accurately. By dynamically adjusting the weights of feature dimensions, it can effectively highlight the key features in the data, thereby improving the detection accuracy, especially when facing deformed repetitions (such as machine translation repetitions, segmented plagiarisms, etc.).

[0056] (2) The present invention can significantly reduce the calculation complexity while ensuring the accuracy by introducing dimension compression and a reversible feature transformation matrix. Specifically, dimension compression reduces the redundancy of data and maps high-dimensional features to a low-dimensional space through a non-linear transformation, thereby reducing the calculation and storage overhead.

[0057] (3) The gradient reversal layer (GRL) of the present invention constrains the stability of feature transformation to ensure that important feature information is not lost during the feature space reconstruction process. The stability of features helps to reduce the interference of data noise and ensure the reliability of subsequent similarity calculations, especially important when the quality of the data set is low or contains noise.

[0058] In summary, in data duplicate detection, especially in the case of deformed duplicates and complex duplicate patterns, the present invention has the characteristics of high accuracy, low computational complexity, efficient feature optimization ability, as well as good scalability and adaptability. By combining technologies such as dynamic weighted cosine similarity, incremental learning optimization, and distributed computing, it can significantly improve the efficiency of large-scale data processing, and improve the accuracy of duplicate detection through refined feature processing, meeting the data processing requirements in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a flowchart of a data detection method based on ElasticSearch provided by the present invention;

[0060] Figure 2 It is a schematic structural diagram of a data detection system based on ElasticSearch provided by the present invention;

[0061] Figure 3 It is a schematic hardware structure diagram of a possible electronic device provided by the present invention;

[0062] Figure 4 It is a schematic hardware structure diagram of a possible computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0064] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0065] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but rather to be in line with the broadest scope consistent with the principles and features disclosed in the present invention.

[0066] Please refer to Figure 1 , which provides a flowchart of a data detection method based on ElasticSearch according to the present invention, including the following steps:

[0067] Step 201, construct a composite density tree index based on ElasticSearch, and pre-screen the similar data set based on the composite density tree index to obtain the target document and the candidate document.

[0068] In ElasticSearch, by establishing a composite inverted index based on space partitioning (such as CDTI), this index decomposes the space data (document features, queries, etc.) into multiple sub-spaces, and through a dynamic critical density function ( ), it locates the potential similar data sets. This means that through the composite index structure, this method can quickly locate the potential similar document sets with higher density in the feature space without performing a complete linear scan of all data.

[0069] The dynamic critical density function plays a role in screening similar data in the data set. Through this function, the threshold of similarity detection can be dynamically adjusted, so that the documents with similar feature space densities can be preferentially selected, thereby reducing the number of candidate documents in subsequent similarity calculations. This screening method has higher efficiency, especially when dealing with large-scale data sets, which can significantly reduce the computational burden.

[0070] Based on the CDTI index and the dynamic density function, ElasticSearch can first perform pre-screening during document retrieval, so that only those potential document sets with high similarity are used as candidate documents for subsequent feature space reconstruction and similarity calculation. This pre-screening process reduces the computational amount, especially when the number of documents is very large, which can improve the retrieval speed and accuracy.

[0071] In short, the composite density tree index built based on ElasticSearch realizes the pre-screening of target documents and candidate documents by dividing the data feature space into multiple sub-spaces and locating potential similar data sets through a dynamic critical density function. This method can significantly improve the efficiency of similarity detection and is particularly suitable for scenarios with large-scale data sets.

[0072] Step 202: Generate a reversible feature transformation matrix using a decomposable attention mechanism, and perform anti-interference reconstruction on the target document and the candidate document based on the reversible feature transformation matrix to obtain a reconstructed target document and a reconstructed candidate document.

[0073] In some embodiments, step 202 may include:

[0074] Determine the query mapping of each original feature according to the position self-correlation tensor representing the correlation between different positions in the input data and the decomposable query head matrix;

[0075] Combine each of the original features with each of the query mappings to obtain combined features;

[0076] Determine and minimize the difference between each of the combined features and each target feature to generate the reversible feature transformation matrix.

[0077] In some embodiments, the reversible feature transformation matrix is expressed as:

[0078] ;

[0079] where is the reversible feature transformation matrix, is the decomposable query head matrix, and are the position self-correlation tensors, is the original feature, is the target feature after removing noise interference, is the total number of samples.

[0080] In specific implementation, is the target matrix to be obtained through an optimization algorithm, and its function is to transform the original features to obtain a new feature representation. Through this transformation, the feature vectors will be re-encoded, noise interference will be removed, and effective features will be enhanced.

[0081] and are the position self-correlation tensors, which are used to represent the correlation between different positions in the input data. In the self-attention mechanism, and Matrices are usually used to capture local features or global information. Through weight calculations at different positions, they reflect the structural features of the data.

[0082] It is a decomposable query head matrix, usually used to represent the embedding or mapping of a query in a specific task. It is used to generate query vectors in the self-attention mechanism. Through decomposition, the calculation can be made more efficient, and the model can learn the dependencies between different dimensions.

[0083] Original features are the original features in the input data, which may be obtained through a certain encoding or representation learning method. These features will be passed as input to subsequent transformation and optimization steps.

[0084] The target features after removing noise interference are the features that are expected to be obtained after optimization. After removing noise and interference, these features will be more representative and discriminative. The goal is to map the original features to this new feature space through a feature transformation matrix

[0085] It is an integer representing the number of samples participating in the optimization during the feature space reconstruction process. Each sample corresponds to a feature vector (original feature), as well as its denoised target feature , and they all participate in the calculation of the loss function.

[0086] The goal in the formula is to minimize a loss function that measures the difference between the original features and the transformed target features . The form of the loss function is the square of the Euclidean distance, that is: , using the square of the Euclidean distance (L2 norm) to calculate the difference, aiming to make the target features as close as possible to the features after transformation.

[0087] The function is applied to the product result. Softmax is usually used for normalization, converting the input into a probability distribution. It emphasizes the features of certain dimensions or positions, enabling the model to better focus on important features.

[0088] Here, the product represents a certain relationship or mapping of the features, combining the original features with this mapping to generate a new feature representation.

[0089] ​The goal of the loss function is to make the transformed features as close as possible to the target features after denoising . The minimization process of the loss function is to solve for the optimal invertible feature transformation matrix through optimization algorithms (such as backpropagation, BPTT, etc.) , so that the feature mapping can be optimized and adjusted.

[0090] BPTT is a backpropagation algorithm for optimization, especially applied in sequence data (such as time series or recurrent neural networks). Here, BPTT is used to optimize the feature transformation matrix , ensuring that in sequence tasks, as time or steps progress, the feature transformation remains effective.

[0091] The goal is to learn an invertible feature transformation matrix such that the features transformed by this matrix and the target features have the smallest difference. In this way, the noise in the feature space can be removed, and the remaining are more effective features. Through this optimization process, the feature representation will be strengthened, and the model will learn how to extract more useful information from the original features, reduce the impact of noise, and thus improve the performance of subsequent tasks (such as similarity calculation, classification, etc.).

[0092] In summary, the present invention optimizes and reconstructs the feature space through the invertible feature transformation matrix so that the transformed features can better represent the data, while removing noise, achieving the effect of feature enhancement. Through the application of Softmax and backpropagation optimization (BPTT), the model can gradually learn the optimal transformation matrix, thereby improving the quality of the features and the performance of the model.

[0093] Step 203: Based on the entropy difference of each dimension between the reconstructed target document and the reconstructed candidate document, construct an adaptive dynamic weight.

[0094] In some embodiments, step 203 may include:

[0095] Determine the global entropy of the similar dataset and the local entropy of each feature dimension in the similar dataset;

[0096] Normalize the difference between the global entropy and the local entropy to obtain a normalized difference;

[0097] Process the normalized difference through the exponential operation of the mutual information intensity to obtain the dynamic weight.

[0098] In some embodiments, the dynamic weight is expressed as:

[0099] ;

[0100] Among them, is the dynamic weight, is the sensitivity control parameter, is the sliding window variance, is the third-order mutual information intensity measurement value, is the attenuation coefficient, is the global entropy, is the feature dimension of the local entropy, is the smoothing term.

[0101] In the specific implementation, is the dynamic weight, representing the dynamic weight of the feature dimension . In this scheme, the dynamic weight adjusts the importance of each feature dimension in the similarity calculation by comprehensively considering global information and local information. The dynamic weight mechanism helps to automatically adjust the influence of feature dimensions, highlighting more important features and reducing the weights of irrelevant features, thereby improving the accuracy of similarity calculation.

[0102] is the sensitivity control parameter, used to adjust the sensitivity of the difference between the global entropy and the local entropy . This parameter controls the degree of influence of this difference on the dynamic weight calculation. A larger 𝜆 value will enhance the influence of the difference between the global and local entropies on the weight calculation, while a smaller 𝜆 value will weaken this influence.

[0103] is the sliding window variance, used to measure the change range of the feature dimension within the sliding window. By calculating the variance of the data within this window, the volatility of this dimension within the local range can be reflected. The larger the variance, the greater the change in the features of this dimension, and the weights will also be adjusted, making the model pay more attention to features with high variability. On the contrary, features with smaller variances are considered relatively stable and have less influence on the similarity calculation.

[0104] is the third-order mutual information intensity measurement value, representing the mutual information between the feature dimension and other feature dimensions, and quantifying the importance of features by measuring the information correlation between different dimensions. A higher mutual information value means that the information sharing between the feature dimension and other feature dimensions is stronger, so its weight in the similarity calculation should also increase. Dimensions with high mutual information intensity can provide more associated information about the data, helping to more accurately judge the similarity of the data.

[0105] is the attenuation coefficient, used to adjust the intensity of the third-order mutual information The intensity of its effect in weight calculation. A larger value will increase its influence in the calculation, meaning that feature dimensions with higher mutual information intensity will obtain larger dynamic weights; while a smaller value will weaken the influence of mutual information on the weight.

[0106] is the global entropy, that is, the entropy value of the entire data set, measuring the overall uncertainty or information content of the data. A higher entropy indicates higher uncertainty or diversity of the data at the global level. As a quantitative indicator of the overall complexity of the data set, the global entropy affects the weight allocation of feature dimensions. In the case of a higher global entropy, the information in the data set is more dispersed, and more refined local weight adjustments may be required to capture the patterns in the data.

[0107] is the feature dimension 's local entropy, representing the entropy value of the feature dimension within a certain local range (such as a subset selected by a sliding window), measuring the information content and uncertainty of this dimension on the local data. The local entropy reflects the degree of change of the feature within a specific local data range. If a certain feature dimension shows a large information change within the local range, its local entropy is large, meaning that this feature is more important within this range and requires a higher dynamic weight.

[0108] is the smoothing term, usually used to avoid division by zero or instability in numerical calculations. The smoothing term ensures that when calculating the local variance, the division-by-zero error is avoided when the variance is zero. Through smoothing, the model can operate stably in all cases.

[0109] Part, by normalizing the difference between the global entropy and the local entropy (through variance smoothing), uses the tanh function to limit the weight to keep it within a reasonable range. Specifically, when the difference is large, the weight will increase, and when the difference is small, the weight is small.

[0110] Part, through the exponential operation of the mutual information intensity further strengthens the feature dimensions with high mutual information intensity and enhances their role in similarity calculation.

[0111] Step 204, based on the asymptotically weighted cosine similarity function with dynamic weights, calculate the similarity between the target document and the candidate document.

[0112] In some embodiments, step 204 may include:

[0113] Obtaining a first feature value of each dimension in the set of target documents and a second feature value of each dimension in the set of candidate documents;

[0114] Performing power transformation on the first eigenvalue and the second eigenvalue to obtain a first eigenvalue after power transformation and a second eigenvalue after power transformation;

[0115] Calculating the distance between the first eigenvalue and the second eigenvalue;

[0116] Normalizing the first eigenvalue after the power transformation and the second eigenvalue after the power transformation to obtain normalized eigenvalues;

[0117] The similarity between the target document and the candidate document is determined based on the distance between the first eigenvalue and the second eigenvalue, the first eigenvalue after power transformation, the second eigenvalue after power transformation, and the normalized eigenvalue.

[0118] In some embodiments, the similarity between the target document and the candidate document can be expressed as:

[0119] ;

[0120] in, is the weighted cosine similarity, It is Dynamic weights of dimensions, is the first dimensional eigenvalue, is the harmonic coefficient (corresponding to the CDTI The values are negatively correlated). is the power transformation factor, is the nonlinear compensation parameter, Eigenvalues after power transformation.

[0121] In the specific implementation, this function effectively solves the shortcomings of traditional cosine similarity in deformation duplication detection through nonlinear dimension weighting and hybrid similarity measurement mechanism. The following is the working principle and collaborative working mechanism of each component:

[0122] Molecular structure: A composite similarity fuser.

[0123] In the semantic similarity reinforcement , power transformation factor :

[0124] The calculation formula is ,in For Dimension The information entropy of is the maximum entropy value of the system.

[0125] When the entropy value of the dimension is high (the feature distribution is uniform), , it degenerates into a linear product, retaining the original semantic similarity; when is low (the feature concentration is high), , it strengthens the product effect of important dimensions. In the application scenario, when detecting machine translation repetitions, by increasing the

[0126] value of the keyword feature dimension (such as named entities), the matching ability of semantic equivalence with different expressions is enhanced. The local tampering compensation term

[0127] , and the non-linear compensation parameter is constrained in the interval (0, 1 / 3) to ensure the convexity of . Minor differences (such as paragraph reordering) will be amplified ( ); significant differences ( ) will be suppressed.

[0128] In the application scenario, for local word substitution (such as synonym substitution) in segmented plagiarism, the relevance of minor differences can be captured through the negative exponent.

[0129] The harmonic coefficient is negatively correlated with the value of the density function pre-screened by CDTI. When the candidate document and the target document are in a high-density area ( the value is large), semantic terms are preferentially used ( reduced), focusing on overall semantic matching; in a low-density area ( the value is small), is increased to strengthen local difference compensation.

[0130] Dynamic adjustment example: When detecting translation repetitions, if CDTI determines that the semantics highly overlap ( the value is large), then = 0.3, focusing on semantic terms; if a structural variation is detected ( the value is small), then = 0.7, strengthening local compensation.

[0131] Denominator structure: Anti-skew normalizer .

[0132] After magnifying the feature value by the th power and then performing weighted summation, the following is achieved:

[0133] Suppressing the long-tail distribution: High dimensional features are squared and magnified, dominating the normalization process;

[0134] Dimensional selectivity enhancement: low The noise features of low dimensions are weakened in normalization.

[0135] When ( = 2), the denominator becomes , and the extreme values dominate the normalization, which is suitable for detecting the plagiarism pattern of critical feature mutations.

[0136] Entropy difference drive

[0137] Entropy difference term Identifying dimension specificity: Dimensions with low global entropy but high local entropy (abnormally concentrated features) obtain high weights; Example: In cross - language plagiarism, the local entropy of the proper noun dimension drops suddenly and the weight increases.

[0138] Mutual information term: Strengthening associated dimensions: When the dimension is strongly correlated with the plagiarism pattern (such as the syntactic structure dimension), the third - order mutual information increases, and the weight grows exponentially.

[0139] Detection scenario: Identifying Chinese - English machine - translated duplicate documents.

[0140] CDTI pre - screening: Through function to locate high - semantic - overlap candidate sets ( automatically adjusted to 0.3);

[0141] Feature reconstruction: The M matrix filters translation noise (such as prepositional structure differences);

[0142] Weight assignment:

[0143] Named entity dimension: = 1.8 (low entropy, power amplification), = 0.9 (high mutual information intensity);

[0144] Word order dimension: = 0.25, compensating for the small differences brought by word order adjustment.

[0145] Calculation results:

[0146] Traditional cosine similarity = 0.72 (underestimated due to word order perturbation);

[0147] AWCS similarity = 0.89 (local difference term contribution + 0.17, semantic term contribution + 0.62).

[0148] Through the above - mentioned mechanism, this method achieves an F1 value of 92.3% for paraphrased duplicates in the TREC2022 duplicate detection benchmark test, with a 19.5% improvement compared to the baseline model and the false - positive rate reduced to 2.7%.

[0149] Step 205: Use the frequency vectorization mechanism to perform incremental parameter optimization and determine the dynamically updated weights.

[0150] In some embodiments, step 205 may include:

[0151] Determine the basic update amount of the dynamic weight according to the learning rate updated according to the dynamic weight and the gradient of the loss function of the dynamic weight updated by the similarity;

[0152] Perform feature calibration processing according to the eigenvalue of the corresponding dimension of the dynamic weight to obtain the corresponding feature calibration value;

[0153] Normalize the L1.5 norm of the gradient vector of the dynamic weight by the similarity to obtain the gradient normalization value;

[0154] Perform momentum adjustment on the change of the reversible feature transformation matrix to obtain the dynamic adjustment amount;

[0155] Determine the dynamically updated weights of the increment according to the basic update amount, the feature calibration value, the gradient normalization value, and the dynamic adjustment amount.

[0156] In some embodiments, the dynamically updated weights of the increment can be expressed as:

[0157] ;

[0158] Where is the dynamically updated weights of the increment, is the learning rate, is the gradient of the loss function with respect to the similarity , is the correction of the eigenvalue after gradient calibration, is the similarity to the L1.5 norm of the gradient, is the momentum adjustment function, is the matrix reconstruction coefficient (constrained by the Lebesgue integral).

[0159] In specific implementation, represents the incremental update of the dynamic weight of the -th dimension feature, that is, during the training process, the weight of the feature dimension The adjustment amount in each iteration. This incremental update reflects the modification amplitude of the current iteration to the weight. The incremental update reflects the adaptive learning process of the model. It adjusts the weight of each feature dimension by the gradient descent method according to the current gradient and historical changes to ensure that the model can be gradually optimized as the training process progresses.

[0160] is the learning rate, which controls the step size of weight updates. It determines the magnitude of parameter adjustment in each gradient update of the model. An overly large learning rate may lead to unstable training, while an overly small learning rate may result in an overly slow convergence speed. The learning rate determines the intensity of each update step and is a key parameter in the gradient descent algorithm. It directly affects the convergence speed and stability of the training process. An overly large learning rate may cause gradient oscillations, while an overly small learning rate may prevent the model from effectively learning.

[0161] is the gradient of the loss function with respect to the similarity and represents the degree of influence of the similarity on the loss function. It reveals how changes in the loss function affect the similarity , that is, the direction of similarity change. The gradient of the loss function is used to guide how to adjust the weights so that the similarity 𝑆 is closer to the target value (e.g., increasing the similarity and reducing the loss). Through the chain rule, this gradient ultimately feeds back into the weight update of the features.

[0162] is the correction of the -dimensional feature value after gradient calibration. The gradient calibration coefficient is a dynamic adjustment factor that represents the adjustment of the feature value according to the change in the learning rate in each iteration. When the learning rate changes, the feature value is scaled by the calibration factor to avoid gradient oscillations caused by learning rate fluctuations. The gradient calibration coefficient can adaptively adjust at different learning rates to ensure the smoothness of the training process and avoid unstable updates caused by sudden changes in the learning rate.

[0163] is the L1.5 norm of the similarity with respect to the gradient, which is a measure of the gradient vector and lies between the L1 norm (sparsity) and the L2 norm (smoothness), taking into account both feature selection and the smoothness of convergence. The L1.5 norm can effectively limit the magnitude of the gradient and avoid gradient explosion during training (i.e., extremely large gradients leading to overly large weight updates). This norm strikes a balance between sparsity and smoothness, helps reduce the influence of unimportant features, and at the same time retains important features. By normalizing the gradient, the influence of overly large gradients on weight updates is suppressed, making the weight update process more stable and preventing extreme fluctuations during training.

[0164] In is the momentum adjustment function, whose role is to adjust the current learning momentum through changes in the historical state. The matrix reconstruction coefficient , through the Lebesgue integral, quantifies the matrix The degree of change in the feature space. When the matrix changes significantly, momentum decays to avoid historical gradients interfering with the current gradient; when the matrix changes slightly, momentum increases to accelerate convergence. Momentum adjustment adjusts the update speed and amplitude according to the change of parameters, making the training process smoother, and intelligently reducing or increasing the contribution of momentum based on historical changes to improve training efficiency.

[0165] Step 206: Dynamically select the optimal calculation path based on the dynamically updated dynamic weights, and perform duplicate detection on the target document and the candidate documents again based on the optimal calculation path to obtain the duplicate detection result.

[0166] In some embodiments, step 206 may include determining the optimal calculation path based on a hybrid sharding architecture that employs a heterogeneous computing unit sharding strategy, where:

[0167] The CPU shards to process the CDTI pre-screening task, the GPU shards to execute the AWCS similarity calculation, and low-latency communication is achieved between the shards through RDMA.

[0168] In some embodiments, the sharding efficiency of the unit sharding strategy is determined in the following manner:

[0169] Calculate the sharding efficiency impact value based on the communication latency time for sending and receiving data between each shard, and the sharding cost of each shard in terms of computing and storage;

[0170] Calculate the weight difference between the weight vector of the main server and the weight vectors of the shard servers;

[0171] Based on the sharding efficiency impact value and the weight difference, determine the sharding efficiency of the unit sharding strategy.

[0172] Among them, the sharding efficiency can be expressed as:

[0173] ;

[0174] Among them, is the sharding efficiency, is the inter-shard communication factor, is the communication latency time, is the hardware acceleration ratio coefficient, is the sharding cost, is the weight vector of the main server, is the weight vector of the shard server, is the weight difference threshold.

[0175] In specific implementation, The sharding efficiency is a metric used to evaluate the data processing efficiency among different shards in distributed computing. It comprehensively considers factors such as communication latency, sharding cost, hardware acceleration ability, and sharding weight differences. The sharding efficiency 𝐸 can measure whether the computing path of a certain shard is optimal in a distributed computing framework and help optimize the resource allocation of the distributed system, thereby improving the overall performance.

[0176] The inter-shard communication factor represents the communication cost when data is exchanged between different computing nodes (i.e., between shards). It affects the efficiency of data transmission between shards. A higher value indicates a larger communication latency between shards, which may affect the computing efficiency of the system. In actual deployment, reducing communication latency usually significantly improves the overall computing efficiency.

[0177] The communication latency time is the time required to send and receive data from one shard to another. Communication latency is a key factor affecting the performance of a distributed system. A longer latency leads to more waiting time during the computing process, thus reducing the overall efficiency. The in the formula represents the impact of communication latency on the sharding efficiency. The greater the latency, the more it reduces the efficiency.

[0178] The hardware acceleration ratio coefficient is used to represent the improvement in computing speed brought by hardware acceleration (such as using GPUs, TPUs, etc.) in distributed computing. Hardware acceleration can significantly increase the computing speed and reduce the computing time of each shard. Therefore, the larger 𝑣 is, the more significant the effect of hardware acceleration, and thus the sharding efficiency 𝐸 is improved.

[0179] The sharding cost represents the resource consumption of each shard in terms of computing and storage, usually involving the resources, memory, and storage required for computing. The sharding cost is an important factor affecting the sharding efficiency. A higher sharding cost means more computing resources and time consumption, which may in turn lead to a decrease in the overall efficiency.

[0180] The weight vector of the master server represents the weights of each feature dimension in the master server. The weight vector is usually used to represent the priority or resource allocation of the server when processing specific tasks. The weight vector of the master server represents the importance of the master server in task scheduling and computing. A larger weight value means that the server plays a more important role in a specific task.

[0181] The weight vector of the sharding server, and Similarity represents the priority or resource allocation of shard servers in processing tasks in a distributed system. The weight vector of a shard server determines its importance and workload in distributed computing tasks. If the weight differences of shards are too large, it may lead to uneven load, thereby affecting the sharding efficiency.

[0182] is the weight difference threshold, which is used to limit the weight difference between the master server and the shard server. If the weight difference between the two exceeds this threshold, it may lead to uneven computing or overloading. By setting the weight difference threshold , the excessive mismatch between the master server and the shard server can be avoided, ensuring the reasonable allocation of resources. A smaller value means allowing a larger weight difference, while a larger value can avoid performance problems caused by too large weight differences.

[0183] functions to calculate the impact of inter-shard communication and hardware acceleration, the impact of communication latency: the product of the inter-shard communication factor and the communication latency is processed through a logarithmic function, indicating that as the communication latency increases, the sharding efficiency decreases. The greater the communication latency, the lower the sharding efficiency . The impact of hardware acceleration: the ratio of the hardware acceleration ratio coefficient and the shard cost represents the reduction of the computing cost under hardware acceleration, thereby improving the sharding efficiency.

[0184] This part functions to consider the weight difference between the master server and the shard server, calculate the weight difference between the master server and the shard server (through the L1 norm), and normalize it through the threshold. If the weight difference is large, it indicates uneven resource allocation, which may lead to a decrease in efficiency. Therefore, this term makes the sharding efficiency decrease as the weight difference increases through subtraction. By setting an appropriate threshold , the impact of the weight difference on the sharding efficiency can be controlled, avoiding the efficiency decrease caused by uneven load.

[0185] The present invention optimizes the sharding efficiency in a distributed computing environment. It comprehensively considers communication latency, hardware acceleration, shard cost, and the weight difference between the master server and the shard server, thereby calculating the efficiency of each shard in the overall computing task. By reasonably configuring these parameters, the resource utilization rate in the distributed system can be improved, the computing path can be optimized, and ultimately the performance of the system can be enhanced.

[0186] For example only, the weight of the master node and the shard node The L1 norm difference exceeds the threshold When this occurs, path reallocation is triggered. For example, when the GPU shard weight differs from that of the master node by more than , it is determined that the calculation path is mismatched.

[0187] Penalize the shard combination with high communication latency, The parameter is dynamically assigned according to the device type. For example, the CPU shard = 1.0, and the GPU shard = 2.5. The smoothing parameter for CDTI pre-screening , and the non-linear compensation parameter of the similarity function .

[0188] In some embodiments, when detecting cross-language duplicates, the system automatically selects the path: CPU shard (CDTI pre-screening) → GPU shard 1 (AWCS calculation) → GPU shard 3 (result aggregation). Compared with static sharding, the latency is reduced by 38% and the throughput is increased by 2.1 times.

[0189] In summary, through the deep coupling of incremental update and calculation path selection, the present invention realizes an intelligent closed-loop of "detection → optimization → re-detection", maintaining low resource consumption while improving the recall rate, and is particularly suitable for scenarios that require rapid response to new plagiarism variants.

[0190] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a data detection system based on ElasticSearch provided by the present invention.

[0191] As Figure 2 shown, a data detection system based on ElasticSearch proposed in an embodiment of the present invention includes:

[0192] A data screening module 301, configured to construct a composite density tree index based on ElasticSearch, and pre-screen a similar data set based on the composite density tree index to obtain a target document and a candidate document;

[0193] A data reconstruction module 302, configured to generate a reversible feature transformation matrix by using a decomposable attention mechanism, and perform anti-interference reconstruction on the target document and the candidate document based on the reversible feature transformation matrix to obtain a reconstructed target document and a reconstructed candidate document;

[0194] A weight determination module 303, configured to construct an adaptive dynamic weight based on the entropy difference of each dimension in the reconstructed target document and the reconstructed candidate document;

[0195] The preliminary detection module 304 is configured to calculate the similarity between the target document and the candidate document based on the asymptotic weighted cosine similarity function of the dynamic weight;

[0196] The weight update module 305 is configured to perform incremental parameter optimization using the frequency vectorization mechanism to determine the dynamically updated dynamic weight;

[0197] The target detection module 306 is configured to dynamically select an optimal calculation path based on the incrementally updated dynamic weight, and perform repetitive detection on the target document and the candidate document again based on the optimal calculation path to obtain a repetitive detection result.

[0198] Please refer to Figure 3 , Figure 3 which is a schematic diagram of an embodiment of the electronic device provided by the embodiment of the present invention. As Figure 3 shown, the embodiment of the present invention provides an electronic device 400, including a memory 410, a processor 420, and a computer program 411 stored in the memory 410 and executable on the processor 420. When the processor 420 executes the computer program 411, the following steps are implemented:

[0199] Construct a composite density tree index based on ElasticSearch, and pre-screen the similar data set based on the composite density tree index to obtain a target document and a candidate document;

[0200] Generate a reversible feature transformation matrix using a decomposable attention mechanism, and perform anti-interference reconstruction on the target document and the candidate document based on the reversible feature transformation matrix to obtain a reconstructed target document and a reconstructed candidate document;

[0201] Construct an adaptive dynamic weight based on the entropy difference of each dimension in the reconstructed target document and the reconstructed candidate document;

[0202] Calculate the similarity between the target document and the candidate document based on the asymptotic weighted cosine similarity function of the dynamic weight;

[0203] Perform incremental parameter optimization using the frequency vectorization mechanism to determine the incrementally updated dynamic weight;

[0204] Dynamically select an optimal calculation path based on the incrementally updated dynamic weight, and perform repetitive detection on the target document and the candidate document again based on the optimal calculation path to obtain a repetitive detection result.

[0205] Please refer to Figure 4 , Figure 4 which is a schematic diagram of an embodiment of a computer-readable storage medium provided by the embodiment of the present invention. As Figure 4As shown in the figure, this embodiment provides a computer-readable storage medium 500, on which a computer program 411 is stored. When the computer program 411 is executed by a processor, the following steps are implemented:

[0206] Construct a composite density tree index based on ElasticSearch, and pre-screen a similar data set based on the composite density tree index to obtain a target document and candidate documents;

[0207] Adopt a decomposable attention mechanism to generate a reversible feature transformation matrix, and perform anti-interference reconstruction on the target document and the candidate documents based on the reversible feature transformation matrix to obtain a reconstructed target document and reconstructed candidate documents;

[0208] Construct an adaptive dynamic weight based on the entropy difference of each dimension between the reconstructed target document and the reconstructed candidate documents;

[0209] Calculate the similarity between the target document and the candidate documents based on the asymptotic weighted cosine similarity function of the dynamic weight;

[0210] Use a frequency vectorization mechanism to perform incremental parameter optimization to determine an incrementally updated dynamic weight;

[0211] Dynamically select an optimal calculation path based on the incrementally updated dynamic weight, and perform a repeatability detection on the target document and the candidate documents again based on the optimal calculation path to obtain a repeatability detection result.

[0212] It should be noted that in the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0213] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0214] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded computers, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce a system for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 one or more of the blocks and implementing the functions specified in the blocks.

[0215] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction system that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 one or more of the blocks and implementing the functions specified in the blocks.

[0216] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 one or more of the blocks and implementing the functions specified in the blocks.

[0217] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

[0218] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A data detection method based on ElasticSearch, characterized in that, The method includes: Constructing a composite density tree index based on ElasticSearch, and pre-screening a similar data set based on the composite density tree index to obtain a target document and candidate documents, including: in ElasticSearch, establishing a composite inverted index based on space segmentation, which decomposes spatial data into multiple sub-spaces, and locates potential similar data sets through a dynamic critical density function, dynamically adjusting the threshold of similarity detection through the dynamic critical density function, so that documents with similar feature space densities are selected, and the spatial data includes document features and queries; Adopting a decomposable attention mechanism to generate a reversible feature transformation matrix, and performing anti-interference reconstruction on the target document and the candidate documents based on the reversible feature transformation matrix to obtain a reconstructed target document and a reconstructed candidate document; Constructing a dynamically adaptive weight for each feature dimension based on the entropy difference between each feature dimension in the reconstructed target document and the reconstructed candidate document; Calculating the similarity between the target document and the candidate document based on the asymptotic weighted cosine similarity function of the dynamic weights includes: obtaining the first eigenvalue of each dimension in the set of the target documents and the second eigenvalue of each dimension in the set of the candidate documents; performing a power transformation on the first eigenvalue and the second eigenvalue to obtain the power-transformed first eigenvalue and the power-transformed second eigenvalue; calculating the distance between the first eigenvalue and the second eigenvalue; performing a normalization process on the power-transformed first eigenvalue and the power-transformed second eigenvalue to obtain the normalized eigenvalues; determining the similarity between the target document and the candidate document based on the distance between the first eigenvalue and the second eigenvalue, the power-transformed first eigenvalue, the power-transformed second eigenvalue, and the normalized eigenvalues; the similarity between the target document and the candidate document is expressed as: ; where is the weighted cosine similarity, is the dynamic weight of the th dimension, , are the eigenvalues of the th dimension in the dataset, is the harmonic coefficient, is the power transformation factor, is the non-linear compensation parameter, , are the power-transformed eigenvalues; Incremental parameter optimization is performed using a frequency vectorization mechanism to determine the dynamic weights for incremental updates, including: determining the basic update amount of the dynamic weights based on the learning rate updated according to the dynamic weights and the gradient of the loss function updated by the similarity; performing feature calibration processing on the eigenvalues corresponding to the dynamic weights to obtain corresponding feature calibration values; normalizing the L1.5 norm of the gradient vector of the dynamic weights by the similarity to obtain a gradient normalization value; performing momentum adjustment on the change of the reversible feature transformation matrix to obtain a dynamic adjustment amount; determining the dynamic weights for incremental updates based on the basic update amount, the feature calibration values, the gradient normalization value, and the dynamic adjustment amount; the dynamic weights for incremental updates are expressed as: where is the dynamic weight for incremental updates, is the learning rate, is the gradient of the loss function with respect to the similarity , is the correction of the eigenvalue after gradient calibration, is the similarity to the L1.5 norm of the gradient, is the momentum adjustment function, is the matrix reconstruction coefficient; Selecting an optimal computing path, and performing a repeatability detection on the target document and the candidate documents again based on the optimal computing path to obtain a repeatability detection result, where the optimal computing path is calculated based on a hybrid sharding architecture, and the hybrid sharding architecture adopts a heterogeneous computing unit sharding strategy, where: The CPU shards to process the CDTI pre-screening task, the GPU shards to execute the similarity calculation, and low-latency communication is achieved between shards through RDMA; Wherein, the CDTI pre-screening task is the task of pre-screening the similar data set based on the composite density tree index to obtain the target document and the candidate documents.

2. The data detection method based on ElasticSearch according to claim 1, wherein The adopting a decomposable attention mechanism to generate a reversible feature transformation matrix includes: Determining a query mapping for each original feature in the input data according to a positional self-correlation tensor representing the correlation between different positions in the input data and a decomposable query head matrix; Combining each of the original features with each of the query mappings to obtain combined features; Determine and minimize the differences between each of the binding features and each of the target features to generate the invertible feature transformation matrix; the invertible feature transformation matrix is expressed as: ; where is the invertible feature transformation matrix, is the decomposable query head matrix, which is used to represent the embedding or mapping of the query in a specific task and is used to generate the query vector in the self-attention mechanism, and are the position self-correlation tensors, is the original feature, is the target feature after removing noise interference from the original feature, is the total number of samples.

3. The data detection method based on ElasticSearch according to claim 1, wherein The constructing a dynamically adaptive weight based on the entropy difference between each dimension in the reconstructed target document and the reconstructed candidate document includes: Determining the global entropy of the similar data set and the local entropy of each feature dimension in the similar data set; Normalizing the difference between the global entropy and the local entropy to obtain a normalized difference; Processing the normalized difference through an exponential operation of the mutual information intensity to obtain the dynamic weight.

4. The data detection method based on ElasticSearch according to claim 3, characterized in that The dynamic weight is expressed as: ; Among them, is the dynamic weight, is the sensitivity control parameter, is the sliding window variance, which is used to measure the change amplitude of each feature dimension in the sliding window between the reconstructed target document and the reconstructed candidate document, is the third-order mutual information intensity measurement value, is the attenuation coefficient, is the global entropy, is the feature dimension is the local entropy of is the smoothing term.

5. The data detection method based on ElasticSearch according to claim 4, wherein It also includes determining the sharding efficiency of the unit sharding strategy, including: Calculating a sharding efficiency influence value according to the communication delay time for sending and receiving data between each shard and the sharding cost of each shard in terms of computing and storage; Calculating the weight difference between the weight vector of the main server and the weight vector of the shard servers; Determining the sharding efficiency of the unit sharding strategy based on the sharding efficiency influence value and the weight difference.

6. A data detection system based on ElasticSearch, characterized in that, For executing the data detection method based on ElasticSearch according to any one of claims 1-5 above, the system includes: A data screening module, which is used to build a composite density tree index based on ElasticSearch, and pre-screen a similar data set based on the composite density tree index to obtain a target document and a candidate document; A data reconstruction module, which is used to generate a reversible feature transformation matrix by using a decomposable attention mechanism, and perform anti-interference reconstruction on the target document and the candidate document based on the reversible feature transformation matrix to obtain a reconstructed target document and a reconstructed candidate document; A weight determination module, which is used to construct an adaptive dynamic weight based on the entropy difference of each dimension between the reconstructed target document and the reconstructed candidate document; A preliminary detection module, which is used to calculate the similarity between the target document and the candidate document based on the asymptotic weighted cosine similarity function of the dynamic weight; A weight update module, which is used to perform incremental parameter optimization by using a frequency vectorization mechanism to determine an incrementally updated dynamic weight; A target detection module, which is used to select an optimal calculation path, and perform repetitive detection on the target document and the candidate document again based on the optimal calculation path to obtain a repetitive detection result.

Citation Information

Patent Citations

  • Document similarity calculation duplicate checking method and system

    CN118606462A

  • Standard substance and standard substance retrieval and sorting method and system based on search engine

    CN119127970A