On-chain k line data cleaning method for sandwich attack
Patent Information
- Application Number
- HK22026125420
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-28
Smart Images

Figure 00000018_0000 
Figure 00000019_0000 
Figure 00000020_0000
Abstract
Description
This invention relates to the field of on-chain data processing, and more particularly to an on-chain candlestick chart cleaning method for sandwich attacks. Background Technology: Decentralized exchanges (DEXs) execute transactions based on on-chain smart contracts. All transaction records are publicly verifiable and serve as the original data source for candlestick charts generated by market data platforms. However, on-chain transaction data suffers from systemic price pollution caused by sandwich attacks, leading to distortion in candlestick charts generated from the original on-chain data, failing to accurately reflect market supply and demand. Maximum Extractable Value (MEV) on the Blockchain: A robot monitors pending transactions on-chain. Upon discovering a target transaction, it inserts a large operation before and after it within the same block: first submitting a large buy order in the same direction as the target transaction (front run), then submitting a large sell order in the opposite direction (back run), thereby profiting from the price difference by driving up the target transaction's price. Neither the front run nor the back run is intended for actual execution, creating abnormal price deviations in the candlestick data and affecting the accuracy of the highest, lowest, and average transaction prices. Existing on-chain data analysis platforms and candlestick chart generation services typically use all on-chain transaction records to generate candlestick charts, indiscriminately including pre-attack buys and post-attack sells in the calculations. This results in abnormal transaction records of attackers being mixed into the candlestick data. Furthermore, the lack of cross-transaction correlation verification to identify attackers makes it impossible to identify and remove pre-attack pairing operations in sandwich attacks, leading to persistent candlestick price distortion. Summary of the Invention 1 HK 20138186 A Specification The purpose of this invention is to provide an on-chain candlestick data cleaning method for sandwich attacks, addressing the technical problem in the background art where sandwich attacks affect candlestick charts, causing inaccuracies. To achieve the above objectives, this invention provides the following technical solution: an on-chain candlestick data cleaning method for sandwich attacks, comprising: S1: constructing a four-key composite sorting mechanism based on liquidity pools; S2: establishing a unique time series and reconstructing the pre-transaction reserve data transaction by transaction; S3: identifying large buy orders from the previous run in the pre-transaction reserve data and pushing them into the previous run temporary stack, performing multi-factor stack-style pairing on subsequent sell orders to identify the attacker's previous buy orders and subsequent sell orders, while covering nested and interleaved sandwiches and retaining unpaired legitimate large transactions; S4: based on the pairing results, precisely deleting the attacker's two endpoint operations, retaining the victim's transactions, outputting the cleaned on-chain transaction data and providing a candlestick pollution metric. By adopting the above technical solution, the attacker's two endpoint operations can be precisely deleted, while the victim's transactions are retained, making the on-chain transaction data cleaner and reducing interference.Furthermore, in step S1, each time a time window is triggered, all on-chain transaction records of the same token address within that time window are obtained. Transactions within the window are first grouped by liquidity pool identifier, and then sorted and subjected to subsequent sandwich testing for each pool. Each pool operates independently, preventing crosstalk between transactions from different pools in the same sequence. This technical solution ensures that each pool independently models its price formation mechanism, allowing the cleaning results to accurately reflect the true liquidity constraints of a single pool; and provides an unambiguous time-series basis for subsequent reserve reconstruction. The four-key total order ensures the unique order of transactions within a block under the same timestamp. Further, in step S1, the priority of the four-key composite sorting is: block timestamp > block height > transaction position number within the block > instruction position number within the transaction. Through this technical solution, the sorting mechanism strictly follows the underlying execution logic of the blockchain. Furthermore, in step S2, the method for reconstructing the pre-trade reserve volume is as follows: A percentage indicator with a value consistently between 0 and 1 is used to replace the unbounded ratio of purchase volume to reserve volume, B / V. This ensures that the meaning of the threshold value does not change with the absolute size of the pool liquidity depth and does not require repeated adjustments when the pool depth changes. Under the constant product market-making mechanism, it is proven that P is monotonically positively correlated with price shocks, making the threshold value equivalent to the upper limit constraint of price shocks, and also serving as a wash strength switch; when increased, only high-impact sandwiches are washed. The reserve volume V before the trade execution is taken. By adopting the above scheme, it prevents the underestimation of the proportion of large-volume purchases after price increases, ensuring that the identification always describes the true shock starting point, thereby maintaining the stability and interpretability of the threshold semantics in a dynamic liquidity environment. Further, in step S3, the method for identifying the preceding and following runs is to traverse the sorted trade list once to complete the identification; S3.1 Identification of large-volume purchases in the preceding run: Traverse the sorted trade list in full order using four keys. For each buy transaction, calculate its liquidity pool percentage P = BI (B + V), where B is the base token transaction volume and V is the base token reserve in the pool before the transaction is executed. If P is greater than the threshold Large Ratio, the buy is considered a large buy by the front runner, and its position (buy Index), signer (S), base token purchase volume (Bbuy), pricing token payment volume (Qbuy), and block height are added to the front runner's temporary stack, waiting to be paired with subsequent sell orders.The Large Ratio also acts as a switch for cleaning intensity: increasing it only cleans the sandwiches with the greatest impact, while decreasing it expands the cleaning range; 3 HK 20138186 A Manual S3.2: Multi-factor pairing judgment for subsequent sell transactions: When traversing to a transaction with a sell direction, search from the top of the previous run temporary stack downwards for the nearest previous run buy that satisfies all pairing conditions with the sell; S3.3: If a previous run that satisfies all pairing conditions is found, a sandwich pairing is confirmed, its buy index and the current sell position sell index are added to the set of indexes to be deleted, and the previous run is removed from the stack. Paired items will not participate in subsequent pairings; In the above technical solution, subsequent sell transactions no longer require P to be greater than the Large Ratio: its scale is indirectly guaranteed by the near conservation of buy and sell volume, ensuring that a sell volume is equivalent to a large previous run buy volume, so the sell itself must also be large; and the pool depth changes after the previous run buy raises the pool price. If a separate percentage threshold is set for subsequent sell transactions, it would be a repetition of condition four. Furthermore, changes in pool depth may cause sell orders that should have been matched to be mistakenly excluded and missed during the detection process. The aforementioned multi-factor evidence is sufficient to confirm a match, and no additional large threshold is required. Furthermore, the pairing conditions in step S3.2 include: Condition 1, Consistent Signer: The signer of the sell transaction is the same as the signer S of the previous run, confirming it as a reverse operation by the same attacker; Condition 2, Block Proximity: The block height of the sell transaction is not earlier than that of the previous run, and the difference between the block heights of the two transactions does not exceed the configurable upper limit Max Block Gap. The previous and subsequent runs of a true sandwich are usually bound to the same block, so this upper limit is set to a smaller value to exclude accidental reverse transactions that are far apart in time; Condition 3, Positive Profit: The effective selling price Esell = Qsell / Bsell of the subsequent run is greater than the effective buying price Ebuy = Qbuy / Bbuy of the previous run, where Qsell and Bsell are the proceeds from the sale of the pricing tokens and the amount of the base tokens sold, respectively. Attackers only have arbitrage motives when selling above the entry price. This condition excludes normal reverse transactions that are not profitable; Condition 4, Approximate Conservation of Buying and Sell Volume: The difference in base token volume between the previous and subsequent runs is IBsell. Bbuyl / Bbuy does not exceed the configurable tolerance Vol Deviation - the attacker will sell the base tokens bought by the previous run as is, while normal transactions do not have this equal amount of reselling feature; Condition 5, sandwich structure: there is at least one transaction between the previous run and the current sell where the signer is not equal to S, that is, there is indeed a victim transaction sandwiched in the middle, excluding self-loop operations with the same signer but no victim.Furthermore, the method for deleting the attacker's endpoint operation in step S4 is as follows: when traversing to a sell transaction, search from the top of the previous run temporary stack downwards for the nearest previous run buy that satisfies all the conditions of consistent signer, block proximity, positive profit, approximately conserved buying and selling volume, and sandwich structure. After confirming the pairing, add it to the set of indexes to be deleted and remove it from the stack. The multi-factor pairing determination is formalized as the predicate sandwich(i, k), which is true when certain conditions are met. The protection of legitimate large transactions and the retention of large buys that have not been paired at the end of the traversal are implemented. The transactions corresponding to the indexes to be deleted are precisely deleted, while the victim and other transaction outputs are retained. The K-line contamination metric calculates the highest and lowest prices before and after cleaning, resulting in a high-price contamination degree of 8H = (H-H') / H' and a low-price contamination degree of (L'-L) / L'. Block proximity refers to: the block height difference not exceeding the Max Block Gap; positive profit means: the effective selling price of the later run is greater than the effective buying price of the earlier run; approximate conservation of buying and selling volume means: the volume deviation does not exceed Vol Deviation; the sandwich structure means: there is at least one transaction with different signers between the earlier and later runs. In summary, this invention has the following beneficial effects compared to existing technologies: 1. It establishes the minimum complete sequence of on-chain transaction records using a four-key composite sorting method, based on liquidity pools, eliminating temporal ambiguities in scenarios such as multiple blocks sharing timestamps and the same transaction containing multiple instructions. Based on this, it reconstructs the liquidity reserve before each transaction, providing a correct basis for accurate deletion and proportion calculation based on position indexes. 2. A method for identifying large-scale transactions before the transaction is executed is proposed, based on the liquidity pool ratio P = B / (B+V). This method replaces the unbounded ratio of buy volume to reserve volume B / V with a ratio indicator that remains constant between 0 and 1. This ensures that the threshold value does not change with the absolute size of the pool's liquidity depth and does not require repeated adjustments when the pool depth changes. Under a constant product market-making mechanism, it is proven that P is monotonically positively correlated with price shocks, making the threshold value equivalent to a price shock upper limit constraint. It also acts as a washout strength switch, only washing out high-impact sandwiches when the threshold value is increased. The reserve volume V before the transaction is executed is used to prevent the proportion of large-scale buys from being underestimated and missed after the price is raised. 3. A pairing and discrimination mechanism combining five types of evidence—signer consistency, block proximity, positive profit, near-constant buy / sell volume, and sandwich structure—is proposed. This mechanism replaces single-condition matching with multiple pieces of evidence, eliminating the possibility of misjudging normal buy-sell transactions by the same signer and accurately identifying genuine attack pairs, significantly reducing the false deletion rate. 4. A stack-based detection method is adopted, where the front runs into the stack and the back runs from the top down to the nearest pair. A single traversal can cover multiple attacks within the same window, as well as nested attacks and interleaved sandwiches. Large buy orders that fail to pair are not terminated and are retained to prevent legitimate whales from mistakenly deleting large one-way positions.5. Only the attacker's buy-in and sell-out endpoints are deleted, while the legitimate transaction records of the victim in the middle are retained. The purified transaction data after eliminating price deviations is output. At the same time, high and low price contamination levels 8H and 8L are given to quantify the degree of distortion of the K-line OHLC by the sandwich attack, which can be used as an auxiliary basis for assessing the severity of the attack and whether to trigger cleaning. 6 HK 20138186 A Specification Drawings Figure 1 is the overall architecture diagram of the on-chain K-line data cleaning system for sandwich attacks; Figure 2 is the flowchart of multi-factor stacked pairing identification for sandwich attacks; Figure 3 is a schematic diagram of stacked pairing and nested interlaced sandwich processing. Detailed Implementation The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative labor are within the scope of protection of the present invention. Example 1, referring to Figures 1-3, describes an on-chain candlestick data cleaning method for sandwich attacks. This method includes two steps: multi-key sorting of transaction data and multi-factor pairing identification and precise deletion of transactions before and after the sandwich attack. Specifically, it comprises four steps, S1-S4: S1: Constructing a four-key composite sorting mechanism based on liquidity pools; S2: Establishing a unique time series and reconstructing the pre-transaction reserve data transaction by transaction; S3: Identifying large-scale buy orders from the pre-transaction reserve data and pushing them into the pre-transaction temporary stack, performing multi-factor stack pairing on subsequent sell orders to identify the attacker's pre-transaction buy and subsequent sell orders, while covering nested and interleaved sandwiches and retaining unpaired legitimate large-scale transactions; S4: Based on the pairing results, precisely deleting the attacker's two endpoint operations, retaining the victim's transactions, outputting the cleaned on-chain transaction data, and providing a candlestick pollution metric. Each time a time window is triggered, all on-chain transaction records of the same token address within that time window are retrieved. Since the same token may be traded in multiple liquidity pools (trading pairs) at the same time, and the candlestick chart is generated by trading pair, this method processes transactions independently by liquidity pool: First, the transactions in the window are grouped by liquidity pool identifier, and then the transactions in each pool are sorted and then subjected to sandwich detection. The transactions in each pool do not interfere with each other, thus avoiding crosstalk between transactions from different pools in the same sequence.A four-key composite sort is performed on the transaction records in each pool, with the following priority: (1) Block timestamp (ascending order): ensuring global consistency of the overall time sequence; (2) Block height (ascending order): arranged according to the on-chain block production order within the same second; (3) Transaction position number within the block (ascending order): arranged according to the original transaction order within the same block; (4) Instruction position number within the transaction (ascending order): arranged according to the instruction order within the same transaction. The four-key composite sort constitutes the minimum total order definition of blockchain transaction records: In cases where multiple blocks on the blockchain share the same timestamp, sorting by timestamp alone cannot distinguish them; The same transaction may contain multiple instructions that affect the liquidity pool separately, and sorting by transaction position alone will misjudge different instructions of the same transaction as independent transactions; The four-key combination is the prerequisite for the correctness of all subsequent precise deletion operations based on position index. Based on the sorting, the underlying token reserve V of the liquidity pool before each transaction is determined, which is used in step two to calculate the liquidity pool ratio. V can be directly taken from the snapshot of the pool reserve before each transaction record; if this snapshot is missing, it can also be reconstructed by starting from the pool reserve baseline at the beginning of the window and accumulating the net change of the underlying token reserve of each previous transaction in a four-key full order. Transaction-by-transaction reconstruction relies on a strict full order of transactions, which is one of the necessities of the four-key composite sorting. In steps S3 and S4, the multi-factor pairing identification and precise deletion of pre- and post-sandwich attack transactions are specifically as follows: Sandwich attack detection is performed on the transaction list of each liquidity pool after sorting within the time window. In a sandwich attack, the attacker submits a large buy order before the victim's transaction (pre-run) and a large sell order after the victim's transaction (post-run). Neither operation is intended for actual execution, but rather to extract profit by sandwiching the victim's transaction, introducing abnormal price deviations into the K-line data. This step only requires one traversal of the sorted transaction list to complete the identification: When a large purchase is encountered during the traversal, it is first recorded in a forward temporary storage stack. When a sale is encountered, it is paired with the purchase in the stack. For each pair of forward and backward transactions, multiple types of evidence are applied for joint judgment based on liquidity ratio, consistent signer, block proximity, positive profit, approximate conservation of buying and selling volume, and sandwich structure. Only the two ends of the attacker confirmed by the pair are deleted, and the intermediate victim transactions are retained. The forward temporary storage stack is only a temporary storage structure during the traversal process. It stores transactions that have been identified as large purchases but have not yet encountered a pairing with a subsequent sale. Each purchase in the stack has only two possible results. If a pairing is found in the traversal, the subsequent transaction is deleted together. If no pairing is found after the traversal, it is considered a legitimate transaction and is retained. There is no third intermediate state. The specific process is as follows: (1) Identification of large purchases in the forward process: Traverse the sorted transaction list by four keys in full order.For each transaction in the buy direction, calculate its liquidity pool ratio P = B / (B + V), where B is the trading volume of the base token in this transaction, and V is the reserve amount of the base token in the pool before the transaction is executed. If P is greater than the threshold Large Ratio, the buy is determined as a front-run large buy, and its position buy Index, signatory S, base token buy amount Bbuy, pricing token payment amount Qbuy, and the block height it is located in are all recorded in the front-run temporary storage stack, waiting to be paired with the subsequent back-run sell. The threshold Large Ratio also acts as a cleaning intensity switch: when increased, only the most impactful sandwich transactions are cleaned; when decreased, the cleaning scope is expanded. Description of the design of the ratio formula: this method adopts the ratio P = B / (B+V), instead of directly using the ratio of the buy amount to the pool reserve B / V. The reason is that the value of P is always between 0 and 1, and the meaning of the threshold Large Ratio is fixed as "the proportion of the current buy amount in the total amount of the new pool after the transaction", which does not change with the absolute size of the pool liquidity depth; while B / V has no upper bound, the same value of B / V has different meanings in deep pools and shallow pools, and the appropriate threshold will drift as the pool depth changes, requiring repeated adjustments. The reserve amount takes V before transaction execution, which is to measure the real impact of the current transaction on the original pool depth, avoiding missed detection due to underestimated self-proportion caused by calculating the reserve after the large buy first pushes up the pool price. Further, under the constant product market maker mechanism, the pool state is (V, Q) (base token reserve and pricing token reserve), the pricing token to be paid for buying B base tokens is Q·B / (V−B), the corresponding effective transaction price is E = Q / (V−B), the spot price before transaction is E₀ = Q / V, and the price impact is ΔE / E₀ = (E−E₀) / E₀ = B / (V−B); it can be seen that both the price impact B / (V−B) and the ratio P = B / (B+V) increase monotonically with the buy amount B, and increase or decrease together with each other. Therefore, setting a threshold for P is equivalent to setting an upper limit for price impact, and this corresponding relationship does not change with the pool depth, so that the threshold still maintains a stable meaning when the pool depth fluctuates. (2) Multi-factor pairing discrimination of back-run sells: when traversing to a transaction with the exchange direction of sell, search from the top of the front-run temporary storage stack downward for the most recent front-run buy that satisfies all pairing conditions with this sell.The pairing conditions are as follows: - Condition 1 (Same Signer): The signer of the sell transaction is the same as the signer S of the previous transaction, confirming it as a reverse operation by the same attacker; - Condition 2 (Block Proximity): The block height of the sell transaction is not earlier than that of the previous transaction, and the difference between the block heights of the two transactions does not exceed the configurable upper limit Max Block Gap. For a true sandwich transaction, the previous and subsequent transactions are usually tied to the same block, so this upper limit is set to a smaller value to exclude accidental reverse transactions that are far apart in time; - Condition 3 (Positive Profit): The effective selling price of the subsequent transaction, Esell = Qsell / Bsell, is greater than the effective buying price of the previous transaction, Ebuy = Qbuy / Bbuy, where Qsell and Bsell are the proceeds from the sale of the denominated tokens and the amount of base tokens sold, respectively. An attacker only has an arbitrage motive when selling above the initial purchase price; this condition excludes normal reverse transactions that are not profitable; - Condition 4 (Approximate Conservation of Buying and Sell Volume): The deviation in the base token volume between the previous and subsequent transactions, IBsell — Bbuy / Bbuy, does not exceed the configurable tolerance Vol. Deviation - Attackers will essentially dump the base tokens bought by the previous runner as is, while normal transactions do not exhibit this equal-volume dumping characteristic; - Condition 5 (Sandwich Structure): There is at least one transaction between the previous runner and the current sellr where the signer is not equal to S, meaning there is indeed a victim transaction sandwiched in the middle, excluding self-looping operations with the same signer but no victim. If a previous runner that meets all conditions is found, a sandwich pairing is confirmed, its buy index and the current sell position sell index are added to the set of indexes to be deleted, and the previous runner is removed from the stack (already paired, no longer participating in subsequent pairings). The subsequent runner's sell itself no longer requires P to be greater than the Large Ratio: its scale has been indirectly guaranteed by the near-conservation condition of buy and sell volume, ensuring that a sell volume is equivalent to a large amount of previous runner's buy volume, so the sell itself must also be large; and after the previous runner's buy raises the pool price, the pool depth changes. If a separate percentage threshold is set for the subsequent runner, it will not only overlap with condition 4, but may also cause the sell that should have been paired to be mistakenly excluded due to changes in pool depth, thus missing detection. The aforementioned multi-factor evidence is sufficient to confirm the pairing, without the need for additional large thresholds. The stack-based strategy of pairing from the top down to the nearest pairing allows the method to cover nested and interleaved sandwiches: when there are attackers nested within attackers (one group of attackers running forward and backward, with another group of attackers running forward and backward inside) or multiple attackers running forward and backward interleaved, the inner sell is first paired with the inner forward run at the top of the stack and popped off the stack, and the outer sell is then paired with its forward run. Each pairing is correctly identified, and the inner layer is not skipped because the outer layer is paired first.(3) Formalization of multi-factor pairing determination: The above determination can be formalized as the predicate sandwich(i, k), which is true if and only if the following conditions are all satisfied simultaneously: P(Ti) is greater than Large Ratio, the direction of Ti is buy, and the direction of Tk is sell; signer(L) = signer(T k) = S; the difference between blockOf(Tk) and blockOf(L) falls within the interval [0, Max Block Gap]; the effective selling price of Tk is greater than the effective buying price of Ti; the deviation of the base token amount of the two does not exceed Vol Deviation; and there exists Tm such that i<m<k and signer(Tm) ≠ S. For stack-based pairing, L that satisfies the predicate and has the closest distance to Ti in the stack is selected to form a front-running and back-running pairing, and the transactions at indexes i and k are precisely deleted. (4) Legitimate large transaction protection and one-pass completion: After the one-pass traversal is completed, all unpaired large buy orders remaining in the front-running temporary storage stack (such as one-way large position building by legitimate whales, and large buy orders that are not sold back in equal volume within the window) are retained and not added to the set to be deleted. Since the failure of front-running pairing will not terminate the detection, and the unpaired transaction remains in the stack to wait for possible subsequent back-running, one pass stack-based traversal can completely cover multiple sandwich attacks within the same window; this mechanism not only avoids mistakenly deleting legitimate large transactions, but also eliminates the need for multiple rounds of rescanning. (5) Precise deletion: Delete all transactions corresponding to the indexes to be deleted from the transaction list (that is, the front-running buy and back-running sell of each sandwich pairing), retain the victim transactions and all other transactions within the interval as the output. Precise deletion only removes the two endpoint operations performed by the attacker not for the purpose of transaction completion, and does not affect the legitimate transaction records of the victim and other participants, and the output is a purified transaction list that can be directly used for candlestick data generation. (6) Candlestick pollution measurement: Let the transaction price sequence of n transactions within the time window be C1, C2, …, Cn, and the set of attacker indexes to be deleted be D. The highest candlestick price before cleaning is H= max{Ci |1 ≤ i ≤ n}, the highest candlestick price after cleaning is H'= max{Ci | i ∉ D}, the high price pollution degree δH=(H — H') / H'; the lowest candlestick price before cleaning is L=min{Ci |1 ≤ i ≤ n}, the lowest candlestick price after cleaning is L'= min { Ci | i ∉ D} , the low price pollution degree δL = (L' —L) / L' . δH and δL quantify the actual price distortion degree of sandwich attacks on candlestick OHLC, which can be used as an evaluation indicator of attack severity, and can also be used as an auxiliary judgment basis for setting whether to trigger the cleaning process.In another embodiment, the identification of large-scale pre-running purchases can be achieved not only by using the liquidity pool ratio P = B / (B+V), but also by using the ratio of the purchase amount to the pre-trade reserve amount, combined with a reading value that adapts to the pool depth. That is, while maintaining the multi-factor stack-based pairing main chain unchanged, equivalent scale measures are used to screen pre-running purchases, achieving the same effect of identifying and accurately deleting pre- and post-running endpoints of sandwich attacks while preserving the victim's transactions. The terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms "a," "described," and "the" used in this invention and the appended claims are also intended to include the majority forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more associated listed items. It should be understood that although the invention may use terms such as first, second, third, etc., to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of the invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination." Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents. 14 HK 20138186 A Claim 1: A method for cleaning on-chain candlestick data against sandwich attacks, comprising: S1: Constructing a four-key composite sorting mechanism based on liquidity pools; S2: Establishing a unique time series and reconstructing the pre-transaction reserve data one by one; S3: Identifying large buy orders from the pre-transaction reserve data and pushing them into the previous run temporary stack, performing multi-factor stack-style pairing on subsequent sell orders, identifying the attacker's previous run buy orders and subsequent run sell orders, while covering nested and interleaved sandwiches and retaining unpaired legitimate large transactions; S4: Based on the pairing results, finally precisely deleting the attacker's two endpoint operations, retaining the victim's transactions, outputting the cleaned on-chain transaction data and providing a candlestick pollution metric.2. The on-chain candlestick data cleaning method for sandwich attacks according to claim 1, characterized in that, in step S1, each time a time window is triggered, all on-chain transaction records of the same token address within that time window are obtained. Transactions within the window are first grouped according to liquidity pool identifiers, and then sorting and subsequent sandwich detection are performed on the transactions within each pool separately. The pools do not interfere with each other, avoiding crosstalk between transactions from different pools in the same sequence. 3. The on-chain candlestick data cleaning method for sandwich attacks according to claim 1, characterized in that, in step S1, the priority of the four-key composite sorting is: block timestamp > block height > transaction position number within the block > instruction position number within the transaction. 4. The on-chain K-line data cleaning method for sandwich attacks according to claim 1, characterized in that, in step S2, the method for reconstructing the reserve volume before the transaction is as follows: replacing the unbounded ratio of purchase volume to reserve volume B / V with a percentage indicator whose value is always between 0 and 1, so that the meaning of the threshold does not change with the absolute size of the pool liquidity depth and does not need to be repeatedly adjusted when the pool depth changes; proving that P is monotonically positively correlated with price shock under the constant product market-making mechanism, so that the threshold is equivalent to the upper limit constraint of price shock, and also serves as a cleaning intensity switch, when it is increased, only high-impact sandwiches are cleaned; taking the reserve volume V before the transaction is executed. 5. The on-chain K-line data cleaning method for sandwich attacks according to claim 1, characterized in that, in step S3, the method for identifying the previous run and the next run is to traverse the sorted transaction list once to complete the identification; S3.1 Identification of large purchases in the previous run: traversing the sorted transaction list in full order by four keys. For each buy transaction, calculate its liquidity pool percentage P = BI (B + V), where B is the base token transaction volume and V is the base token reserve in the pool before the transaction is executed. If P is greater than the threshold Large Ratio, the buy is considered a large buy by the front runner, and its position buy index, signer S, base token buy volume Bbuy, pricing token payment volume Qbuy, and block height are recorded in the front runner's temporary stack, waiting to be paired with subsequent sell transactions.The Large Ratio value also acts as a switch for cleaning intensity: increasing it only cleans the sandwiches with the greatest impact, while decreasing it expands the cleaning range; S3.2: Multi-factor pairing judgment for subsequent sell transactions: when traversing to a transaction with a sell direction, search from the top of the previous transaction's temporary stack downwards for the nearest previous transaction's buy that satisfies all pairing conditions with the sell; S3.3: if a previous transaction that satisfies all pairing conditions is found, a sandwich pairing is confirmed, its buy index and the current sell position sell index are added to the set of indices to be deleted, and the previous transaction is removed from the stack. Paired transactions will no longer participate in subsequent pairings. 6. A method for cleaning on-chain candlestick data against sandwich attacks according to claim 5, characterized in that the pairing conditions in step S3.2 include: Condition 1, consistent signer: the signer of the sell transaction is the same as the signer S of the previous transaction, confirming it as a reverse operation by the same attacker; Condition 2, block proximity: the block height of the sell transaction is not earlier than the previous transaction, and the difference between the block height of the sell transaction and the previous transaction does not exceed the configurable upper limit Max Block Gap -- the previous and subsequent transactions of a real sandwich are usually bound to the same block, so this upper limit is taken as a smaller value to exclude accidental reverse transactions that are far apart in time; Condition 3, positive profit: the effective selling price Esell = Qsell / Bsell of the subsequent transaction is greater than the effective buying price Ebuy = Qbuy I Bbuy of the previous transaction, where Qsell and Bsell The conditions are: 1) Proceeds from the sale of the denominated tokens and 2) Sales volume of the base tokens – Attackers only have an arbitrage motive when selling at a price higher than the initial purchase price; this condition excludes unprofitable normal reverse transactions. 2) Approximate conservation of buying and selling volume: The difference in base token volume between the previous and current runs (IBsell — Bbuyl / Bbuy) does not exceed the configurable tolerance (Vol Deviation) – Attackers will sell the base tokens bought by the previous run almost exactly as they were; normal transactions do not exhibit this equal-volume sell-back characteristic. 3) Sandwich structure: There is at least one transaction between the previous run and the current sell where the signer is not equal to S, meaning there is indeed a victim transaction sandwiched in the middle, excluding self-circulating operations with the same signer but no victims.7. A method for cleaning on-chain K-line data against sandwich attacks according to claim 6, characterized in that the method for deleting the attacker's endpoint operation in step S4 is as follows: when traversing to a sell transaction, search from the top of the previous run temporary stack downwards for the nearest previous run buy that satisfies the conditions of consistent signer, block proximity, positive profit, approximately conserved buying and selling volume, and all conditions of sandwich structure. After confirming the pairing, add it to the set of indexes to be deleted and remove it from the stack; the multi-factor pairing determination of claim 3 HK 20138186 A is formalized as the predicate sandwich(i, k), which is true when certain conditions are met. The protection of legal large transactions and the retention of large buys that have not been paired at the end of the traversal are completed in a single pass. The transactions corresponding to the indexes to be deleted are accurately deleted, and the outputs of victims and other transactions are retained; the K-line pollution metric calculates the highest and lowest prices before and after cleaning, and obtains the high price pollution degree 8H = (H-H') / H' and the low price pollution degree = (L'-L) / L'. 4 HK 20138186 A HK 20138186 A 1 HK 20138186 A 2 HK 20138186 A 3.