Data filtering method and system, and electronic device and storage medium
Through the two-stage filter structure, the data is judged in label value, and efficient duplicate data filtering is achieved, which solves the problem that duplicate data filtering affects data forwarding performance in the prior art, and improves data processing efficiency and accuracy.
Patent Information
- Application Number
- PCT/CN2025/073630
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2025-01-21
- Publication Date
- 2025-07-03
AI Technical Summary
The duplicate data filtering method in the prior art affects the data forwarding performance, resulting in a reduced processing efficiency.
A two-stage filter structure is adopted, by labeling data, using a primary filter to filter data of continuous label value, and a secondary filter to filter data of non-continuous label value to achieve efficient data filtering.
Improve data processing efficiency, avoid additional deployment and damage to original network data packets, and improve data processing accuracy and efficiency.
Smart Images

Figure CN2025073630_03072025_PF_FP_ABST
Abstract
Description
Data filtering method and system, electronic device and storage medium Technical Field
[0001] The present application relates to the field of network communication technology, and in particular to a data filtering method and system, an electronic device, and a storage medium. Background Art
[0002] A deduplication application is a tool used to identify and remove duplicate records from a dataset. It can be applied to a variety of data types, including text, numbers, images, and more. Duplicate data refers to the presence of multiple identical or similar records in a dataset, which may be caused by data entry errors, system failures, or other reasons. The main goal of a deduplication application is to improve data quality and accuracy. By detecting and removing duplicate records, data redundancy can be reduced and the efficiency of data analysis and processing can be improved. In related art, deduplication applications are typically based on algorithms and techniques to identify duplicate records. These algorithms can compare different records in a dataset and determine the similarity between them based on predefined similarity metrics. Once duplicate records are identified, the application can take appropriate actions, such as deleting, merging, or marking the duplicate records. However, filtering methods in related art often reduce data forwarding performance and affect data processing messages. Summary of the Invention
[0003] The main purpose of the embodiments of the present application is to provide an efficient data filtering method and system, electronic device and storage medium.
[0004] To achieve the above-mentioned purpose, one aspect of an embodiment of the present application proposes a data filtering method, the method comprising: receiving data to be filtered and determining the label value of the data; inputting the data into a first-level filter to determine an expected value and a tail value; the expected value is used to characterize the label value determined based on the continuity of the data, and the tail value is used to characterize the earliest valid label value in the second-level filter; if the label value is equal to the expected value, filtering the data through the first-level filter to determine the filtered target data; the first-level filter is used to filter data with continuous label values; or, if the label value is not equal to the expected value, and the label value is greater than or equal to the tail value, inputting the data into the second-level filter, filtering the data through the second-level filter to determine the filtered target data; the second-level filter is used to filter data corresponding to non-continuous label values. The embodiment of the present application determines the label value of the data by marking the data, and then filters the data through the first-level filter and the second-level filter, without the need for additional deployment and without destroying the data packets of the original network. Therefore, the embodiment of the present application is conducive to improving data processing efficiency.
[0005] In some embodiments, the method provided by the embodiments of the present application, the secondary filter includes a plurality of filter nodes, each of which includes a boundary record and a marking unit; and filtering the data through the secondary filter includes:
[0006] determining, based on the boundary record and the tag value, whether the tag value is registered in the secondary filter;
[0007] If the tag value is registered in the secondary filter, determining a first boundary corresponding to the tag value; and filtering the data according to the first marking unit to determine target data; wherein the first filtering node includes a first boundary record and a first marking unit;
[0008] Alternatively, if the label value is not registered in the secondary filter, a second filter node is registered in the secondary filter according to the label value.
[0009] In some embodiments, in the method provided by embodiments of the present application, the second filter node includes a second boundary record and a second marking unit; and registering the second filter node in the secondary filter according to the label value includes:
[0010] Determine a first boundary value based on the current expected value;
[0011] Determine a second boundary value according to the tag value; the first boundary value and the second boundary value are two endpoints of the interval recorded by the second boundary value;
[0012] A second marking unit is determined according to the first boundary value and the second boundary value.
[0013] In some embodiments, the method provided in the embodiments of the present application, filtering the data according to the first marking unit includes:
[0014] If the tag value has been registered in the first marking unit, discard the data, and determine that the target data does not include the data;
[0015] Alternatively, if the tag value has not been registered in the first marking unit, the tag value is registered in the first marking unit according to the first boundary value and the tag value, and it is determined that the target data includes the data.
[0016] In some embodiments, in the method provided by the embodiments of the present application, the filtering node further includes a timer, and the timer starts counting when the filtering node is established; the method further includes:
[0017] If the duration of the timer is greater than a preset duration, the filtering node is deleted; the preset duration is used to represent the aging time of the filtering node.
[0018] In some embodiments, in the method provided by the embodiments of the present application, the primary filter includes a historical value, where the historical value is a valid tag value received last time and not filtered; the method further includes:
[0019] If the tag value is equal to the historical value, discard the data, and determine that the target data does not include the data;
[0020] Alternatively, if the tag value is smaller than the tail value, the data is discarded, and it is determined that the target data does not include the data.
[0021] In some embodiments, the method provided by the embodiments of the present application further includes:
[0022] If the secondary filter is not empty, determining the tail value to be the earliest valid label value in the secondary filter;
[0023] Alternatively, if the secondary filter is empty, the tail value is determined to be an expected value.
[0024] To achieve the above objectives, another aspect of the present application provides a data filtering system, comprising:
[0025] The first module is used to receive the data to be filtered and determine the label value of the data;
[0026] The second module is configured to input the data into a primary filter and determine an expected value and a tail value; the expected value is used to represent a label value determined based on data continuity, and the tail value is used to represent the earliest valid label value in the secondary filter;
[0027] A third module is configured to filter the data through the first-level filter to determine filtered target data if the tag value is equal to the expected value; the first-level filter is configured to filter data with continuous tag values;
[0028] or,
[0029] The fourth module is used to input the data into the secondary filter if the label value is not equal to the expected value and the label value is greater than or equal to the tail value, filter the data through the secondary filter, and determine the filtered target data; the secondary filter is used to filter the data corresponding to non-continuous label values.
[0030] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned method when executing the computer program.
[0031] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above-mentioned method when executed by a processor.
[0032] The embodiments of the present application include at least the following beneficial effects: the data filtering method provided by the embodiments of the present application includes: receiving data to be filtered and determining the label value of the data; inputting the data into a first-level filter to determine an expected value and a tail value; the expected value is used to characterize the label value determined based on the continuity of the data, and the tail value is used to characterize the earliest valid label value in the second-level filter; if the label value is equal to the expected value, filtering the data through the first-level filter to determine the filtered target data; the first-level filter is used to filter data with continuous label values; or, if the label value is not equal to the expected value, and the label value is greater than or equal to the tail value, inputting the data into the second-level filter, filtering the data through the second-level filter to determine the filtered target data; the second-level filter is used to filter data corresponding to non-continuous label values. The embodiments of the present application determine the label value of the data by marking the data, and then filter the data through the first-level filter and the second-level filter, without the need for additional deployment and without destroying the data packets of the original network. Therefore, the embodiments of the present application are conducive to improving data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] FIG1 is a flow chart of an embodiment of a data filtering method provided by the present application;
[0034] FIG2 is a schematic diagram of the architecture of an embodiment of a filter provided by the present application;
[0035] FIG3 is a schematic diagram of the architecture of an embodiment of a secondary filter provided by the present application;
[0036] FIG4 is a flow chart of an embodiment of a filtering process of a secondary filter provided by the present application;
[0037] FIG5 is a flow chart of another embodiment of the filtration process of the secondary filter provided by the present application;
[0038] FIG6 is a flow chart of an embodiment of a filtering process of a primary filter provided by the present application;
[0039] FIG7 is a flow chart of an embodiment of a data filtering process provided by the present application;
[0040] FIG8 is a filtering effect diagram of an embodiment of the data filtering process provided by the present application;
[0041] FIG9 is a schematic diagram of the structure of a data filtering system provided in an embodiment of the present application;
[0042] FIG10 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0044] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0045] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0047] Before explaining the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0048] Ring queue: a circular data storage structure.
[0049] Bitmap: A data structure based on bit storage.
[0050] In related technologies, deduplication applications are tools used to identify and remove duplicate records from datasets. They can be applied to a variety of data types, including text, numbers, and images. Duplicate data refers to the presence of multiple identical or similar records in a dataset, which may be caused by data entry errors, system failures, or other reasons. The primary goal of deduplication applications is to improve data quality and accuracy. By detecting and removing duplicate records, data redundancy can be reduced, improving the efficiency of data analysis and processing. Furthermore, deduplication applications can help protect data consistency and integrity, ensuring that the information in a dataset is accurate and reliable. Deduplication applications typically use algorithms and techniques to identify duplicate records. These algorithms compare different records in a dataset and determine their similarity based on predefined similarity metrics. However, these algorithms can affect data forwarding performance, resulting in reduced data processing efficiency.
[0051] As you can see, once duplicate records are identified, the application can take appropriate actions, such as deleting, merging, or marking them. Duplicate data filtering applications are widely used in various fields, including database management, data cleaning, data mining, and data analysis. They can help organizations and individuals better manage and utilize data resources, improving data quality and credibility. In many application scenarios, duplicate data filtering is crucial to ensure that duplicate data is not received at the end point.
[0052] In view of this, in order to filter duplicate data without losing data forwarding performance, a data filtering method is provided in an embodiment of the present application to achieve functions such as filtering duplicate data in network data packets and removing timed data.
[0053] The data filtering method provided in the embodiment of the present application relates to the field of network communication technology. The data filtering method provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the data filtering method, etc., but is not limited to the above forms.
[0054] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0055] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0056] FIG1 is an optional flowchart of a data filtering method provided in an embodiment of the present application; the method in FIG2 may include but is not limited to steps S100 to S400.
[0057] Step S100, receiving data to be filtered and determining the label value of the data;
[0058] Step S200: Input data into the primary filter to determine the expected value and tail value; the expected value is used to represent the label value determined based on data continuity, and the tail value is used to represent the earliest valid label value in the secondary filter;
[0059] Step S300: If the tag value is equal to the expected value, the data is filtered through a first-level filter to determine the filtered target data; the first-level filter is used to filter data with continuous tag values;
[0060] Alternatively, in step S400, if the tag value is not equal to the expected value and the tag value is greater than or equal to the tail value, the data is input into the secondary filter, and the data is filtered by the secondary filter to determine the filtered target data; the secondary filter is used to filter the data corresponding to the non-continuous tag values.
[0061] In the embodiment of the present application, the data tag value is used to distinguish different data. In some embodiments, the method of marking network data in the embodiment of the present application can be: the tag data is marked after the network layer 2 data frame; that is, the determination of the flag value in the embodiment of the present application can be after the network layer 2 data frame; it does not destroy the original network data packet form, and improves the efficiency of data filtering processing. It can be understood that the premise of data filtering is to mark the data packet, and then the multiple data channels are aggregated into one channel of data at the receiving end, duplicate data is removed, and timed data is removed. In order to mark the uniqueness of the data packet, for example, the present invention tags the data, and the same data will be marked with the same tag (same ID) to mark the uniqueness of the data. The data format of the tag: message ID (4Bytes). The message ID uses an unsigned 32-bit integer data type. It should be noted that the first-level filter in the embodiment of the present application is a ring data structure, and the first-level filter includes an expected value and a tail value to realize the recording and screening of the tag value.
[0062] It can be understood that the data filtering method in the embodiment of the present invention, with reference to the structure of the filter provided in the embodiment of the present application shown in Figure 2, is composed of two parts: a primary filter and a secondary filter. Among them, the primary filter is used to filter continuous data, that is, data with continuous data tag values, such as the current data value is 10, and the next tag value is 11. Such data will first be filtered by the primary filter. The secondary filter in the embodiment of the present application is used to filter non-continuous data, that is, data with non-continuous data tags. For example, the current data tag value is 10, and the next one is 20, indicating that there are 10 data between the two data tags that have not been filtered. In this case, the embodiment of the present application filters the data through the secondary filter. The embodiment of the present application implements data filtering through a two-stage filter structure. The filtering method provided in the embodiment of the present application is based on user state implementation, does not require additional deployment and maintenance, and improves data processing efficiency.
[0063] In the embodiment of the present application, after the tagged data passes through the data filtering module, it will first enter the first-level filter, which is composed of an expected value counter (expert, i.e., the expected value in the embodiment of the present application), a historical value counter (last, i.e., the historical value in the embodiment of the present application), and a tail value counter (tail, i.e., the tail value in the embodiment of the present application). The first-level filter will compare the data tag value and the expected value. If they are the same, no filtering is required, and the expected value is increased to filter the next data. That is, if the tag value is equal to the expected value, it is determined that the target data includes the data. The data in the embodiment of the present application is used to represent the data currently being filtered.
[0064] In some embodiments, referring to the architecture of the secondary filter shown in FIG3 and the filtering process of the secondary filter shown in FIG4 , the secondary filter includes a plurality of filtering nodes, each of which includes a boundary record and a marking unit. Filtering data through the secondary filter includes:
[0065] Step S410, judging whether the tag value is registered in the secondary filter based on the boundary record and the tag value;
[0066] Step S420: If the tag value is registered in the secondary filter, determine the first boundary corresponding to the tag value; and filter the data according to the first marking unit to determine the target data; wherein the first filtering node includes the first boundary record and the first marking unit;
[0067] Alternatively, in step S430, if the tag value is not registered in the secondary filter, a second filter node is registered in the secondary filter according to the tag value.
[0068] In some possible implementations, the boundary record in the embodiments of the present application is used to represent the boundary of the label value recorded by the filter node; the marking unit is used to mark the label value recorded by the filter node. The second filter in the present application can include multiple filter nodes. As shown in Figure 3, the embodiment of the present application can logically control X filter nodes through a circular queue in the secondary filter.
[0069] The secondary filter in the embodiment of the present application is used to filter data that cannot be filtered by the primary filter. As can be seen from the above description, when the received data tag value is not the expected value and is greater than the tail value, the data will enter the secondary filter. Referring to Figure 2, the secondary filter can be composed of a filter node and a circular queue, wherein the filter node consists of a boundary record + timer + bitmap marking unit. The circular queue SlaveRring is used to manage the filter node. Referring to an embodiment shown in Figure 5, the secondary filter in the embodiment of the present application includes: Boundary record: It consists of a maximum tag value (min, i.e., the first boundary value in the embodiment of the present application) and a minimum tag value (max, i.e., the second boundary value in the embodiment of the present application), which is used to record the boundary value of the filter node. Timer: Records the aging time of the node. When it exceeds a certain time, the filter node is eliminated. Bitmap marking unit: A 128-bit bitmap that marks the data tag value recorded by the current filter node. For example, 1 can be used to indicate filtered, and 0 can be used to indicate unfiltered. In a specific embodiment, after the data enters the secondary filter, it is first checked whether it is registered in the filter node. If not, the node information to be filtered needs to be registered first. Otherwise, the data tag value is used to filter in the filter node. Data is filtered through the primary and secondary filters, thereby filtering duplicate data within a certain time range. It is understood that if the tag value has been registered in the secondary filter, the first boundary to which the tag value belongs is determined, and the data is filtered based on the first marking unit corresponding to the first boundary. If the tag value has not been registered in the secondary filter, a new second filter node is registered. In the embodiment of the present application, the second filter node is used to represent an unregistered filter node.
[0070] In some embodiments, the method provided by the embodiments of the present application, the second filter node includes a second boundary record and a second marking unit; and registering the second filter node in the secondary filter according to the label value includes:
[0071] Determine a first boundary value based on the current expected value;
[0072] Determine the second boundary value according to the tag value; the first boundary value and the second boundary value are two endpoints of the interval recorded in the second boundary;
[0073] A second marking unit is determined according to the first boundary value and the second boundary value.
[0074] In some possible implementations, the functions of the secondary filter include node registration, data filtering, and timeout aging. When data enters the secondary filter, it is first necessary to determine whether the registration of the filter node is required and complete the registration of relevant information before subsequent data tag filtering can be completed. It is understandable that if the current tag value has not been registered in the secondary filter, it is first necessary to apply for an idle filter node from the circular queue (that is, the second filter node in the embodiment of the present application is an idle filter node) and then complete the data registration. The data registration requires the registration of the following information:
[0075] Boundary minimum value (i.e., the first boundary value in the embodiment of the present application): the current expected value (not received). Boundary maximum value (i.e., the second boundary value in the embodiment of the present application): the received data tag value. Timer value: the reception time of the current data. Bitmap tag (i.e., the second tag unit in the embodiment of the present application): completes the data tag registration in this table, and the registration value = setBitmap (tag value - boundary minimum value).
[0076] After entering the secondary filter, check whether the data tag has been registered. If it has been registered, it is necessary to use the data tag to perform a one-to-one node match on the registered node boundary value to determine the first boundary corresponding to the tag value, and then perform subsequent filtering.
[0077] In some embodiments, the method provided in the embodiments of the present application filters the data according to the first marking unit, including:
[0078] If the tag value has been registered in the first tag unit, the data is discarded and it is determined that the target data does not include the data;
[0079] Alternatively, if the tag value has not been registered in the first marking unit, the tag value is registered in the first marking unit according to the first boundary value and the tag value, and it is determined that the target data includes the data.
[0080] In some possible implementations, after matching a suitable filter node, the data is first checked to see if it is registered in the bitmap registration table. Specifically, if the first filter node is matched, the first tag unit of the first filter node is used to determine whether the current tag value has been registered. If not, the value is registered; otherwise, the data is discarded, completing the secondary data filtering.
[0081] In some embodiments, the method provided by the embodiments of the present application, the filtering node further includes a timer, the timer starts timing when the filtering node is established, and the method further includes:
[0082] If the timer duration is longer than the preset duration, the filter node is deleted; the preset duration is used to represent the aging time of the filter node.
[0083] In some possible implementations, since the data to be filtered is inherently time-sensitive, the aging of filtering nodes needs to be considered. When the timer expires, the filtering node is removed from the SlaveRing. First, when registering the data tag, the data's reception time is registered, so that each filtering node has its maximum timeout. The secondary filter removes the timed-out nodes from the ring queue at regular intervals (e.g., 500ms), completing data aging.
[0084] It can be seen from this that the secondary filtering of data in the embodiment of the present application is slightly complicated, and its core idea is: first register the information, then match the information, and then periodically check the data aging.
[0085] In some embodiments, the method provided by the embodiments of the present application, the first-level filter includes a historical value, where the historical value is a valid tag value received last time and not filtered, and the method further includes:
[0086] If the tag value is equal to the historical value, the data is discarded and the target data is determined to not include the data;
[0087] Alternatively, if the tag value is less than the tail value, the data is discarded and it is determined that the target data does not include the data.
[0088] In some possible implementations, if the data tag and the historical value are the same, the data is discarded, indicating that the current data duplicates the historical data. If the data tag is greater than the expected value, the data enters the second-level filtering. If the data tag is less than the historical value and greater than or equal to the tail value, the data enters the second-level filtering. If the data tag is less than the tail value, the data is discarded (indicating that the current data has timed out). As can be seen, the first-level filter can filter out most duplicate data. Only a small amount of data enters the second-level filter due to link delays.
[0089] The first-level filtering in the embodiment of the present application can be implemented in programming by using only three 32-bit unsigned integer data (expert, last, tail). Referring to an embodiment shown in Figure 6, the expected value counter (expert) in the embodiment of the present application: the expected value is the data label value that the first-level filter wants to receive. When the received data label is the same as the expected value, the data is discarded and the value is automatically incremented. Historical value counter (last): The historical value refers to the valid data label value that was received last time and was not filtered. Tail counter (tail): The tail value refers to the last data label value to be filtered. If the secondary filter is empty, the value is the expected value, otherwise it is the last valid data label value of the secondary filter node.
[0090] The reason the primary filter in this embodiment uses 32-bit unsigned integer data is that it is a ring data structure. When the maximum value is reached, the count starts again from 0. The difference between the previous value and the next value is always the difference between the two numbers. This is equivalent to a ring counter.
[0091] In some embodiments, the method provided by the embodiments of the present application further includes:
[0092] If the secondary filter is not empty, determine the tail value as the earliest valid label value in the secondary filter;
[0093] Alternatively, if the secondary filter is empty, the tail value is determined to be the expected value.
[0094] 7 , the solution of the embodiment of the present invention will be described in detail and in combination with a specific application example.
[0095] The first process shows that the first-level filter has started to filter data, and the second-level filter has registered two filter nodes (i.e., node0 and node1). At this time, the tail value is the first boundary corresponding to the earliest registered node0 filter node. The second process in Figure 7 demonstrates that a filter node (node0 filter node) has failed and been removed. At this time, the tail value is the first boundary corresponding to the earliest registered node1 filter node; and the expected value of the first-level filter has overflowed. According to the characteristics of the circular data structure, the expected value redefines the label value from label 0. The third process in Figure 7 demonstrates that the node1 filter node of the second-level filter has failed and been removed. There is no valid filter node in the second-level filter. At this time, the second-level filter does not work, so the tail value is the expected value. At this time, the first-level filter filters the data normally. Referring to Figure 8, the data is filtered by the data filtering method provided by the embodiment of the present application. In Figure 8, different data are represented by different symbols. The same data is discarded by the first-level filter and the second-level filter to obtain non-repeated data. In Figure 7, different data are represented by different symbols.
[0096] It can be understood that on the data sender, data tagging application software is integrated to complete data tagging. On the data receiver, data filter software is integrated to filter out duplicate data and timed-out invalid data. Once both are integrated, the data filtering function is complete.
[0097] The embodiment of the present application can implement the above method based on the Linux user state. Specifically, the network data is labeled in the following way: the label data is printed behind the network layer 2 data frame; the specific format of the data label is: device identifier + status identifier + frame ID + magic number.
[0098] Compared to existing kernel-mode implementations, the embodiments of this application improve the portability and deployability of the multiple-transmit selective reception function, enhancing product value and reducing maintenance and deployment costs. In the event of a software failure, only the software needs to be restarted, without restarting the entire system, allowing for rapid network recovery. Data is labeled without destroying the original Layer 2 network data packets.
[0099] The data filtering method provided by the embodiment of the present application includes: receiving data to be filtered and determining the label value of the data; inputting the data into a first-level filter to determine the expected value and the tail value; the expected value is used to characterize the label value determined based on the continuity of the data, and the tail value is used to characterize the earliest valid label value in the second-level filter; if the label value is equal to the expected value, the data is filtered through the first-level filter to determine the target data after filtering; the first-level filter is used to filter data with continuous label values; or, if the label value is not equal to the expected value, and the label value is greater than or equal to the tail value, the data is input into the second-level filter, and the data is filtered through the second-level filter to determine the target data after filtering; the second-level filter is used to filter data corresponding to non-continuous label values. The embodiment of the present application determines the label value of the data by marking the data, and then filters the data through the first-level filter and the second-level filter, without the need for additional deployment and without destroying the data packets of the original network. Therefore, the embodiment of the present application is conducive to improving data processing efficiency.
[0100] Referring to FIG. 9 , an embodiment of the present application further provides a data filtering system that can implement the above-mentioned data filtering method. The system includes:
[0101] The first module 810 is configured to receive data to be filtered and determine a tag value of the data;
[0102] The second module 820 is used to input data into the first-level filter and determine the expected value and the tail value; the expected value is used to represent the label value determined based on the continuity of the data, and the tail value is used to represent the earliest valid label value in the second-level filter;
[0103] The third module 830 is configured to filter the data through a primary filter to determine filtered target data if the tag value is equal to the expected value; the primary filter is configured to filter data with continuous tag values;
[0104] or,
[0105] The fourth module 840 is used to input the data into the secondary filter if the label value is not equal to the expected value and the label value is greater than or equal to the tail value, filter the data through the secondary filter, and determine the filtered target data; the secondary filter is used to filter the data corresponding to non-continuous label values.
[0106] In some embodiments, the system provided by the embodiments of the present application, the filtering node also includes a timer, and the timer starts timing when the filtering node is established; the system also includes a sixth module, which is used to delete the filtering node if the timer duration is greater than a preset duration; the preset duration is used to represent the aging time of the filtering node.
[0107] In some embodiments, the system provided by the embodiments of the present application includes a primary filter including a historical value, where the historical value is a valid tag value that was received last time and has not been filtered; the system further includes a seventh module for discarding data if the tag value is equal to the historical value, and determining that the target data does not include the data;
[0108] Alternatively, if the tag value is less than the tail value, the data is discarded and it is determined that the target data does not include the data.
[0109] In some embodiments, the system provided by the embodiments of the present application further includes an eighth module for determining, if the secondary filter is not empty, that the tail value is the earliest valid label value in the secondary filter;
[0110] Alternatively, if the secondary filter is empty, the tail value is determined to be the expected value.
[0111] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0112] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned data filtering method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.
[0113] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0114] Please refer to FIG10 , which illustrates a hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0115] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0116] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the data filtering method of the embodiments of this application;
[0117] Input / output interface 903, used to implement information input and output;
[0118] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0119] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0120] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0121] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned data filtering method is implemented.
[0122] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0123] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0124] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0125] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0126] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0127] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0128] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0129] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0130] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0131] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0132] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0134] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A data filtering method, characterized in that, The method includes: Receiving the data to be filtered and determining the tag value of the data; Inputting the data into a first-level filter to determine an expected value and a tail value; the expected value is used to represent the tag value determined based on data continuity, and the tail value is used to represent the earliest valid tag value in the second-level filter; If the tag value is equal to the expected value, filtering the data through the first-level filter to determine the filtered target data; the first-level filter is used to filter data with continuous tag values; Or, if the tag value is not equal to the expected value and the tag value is greater than or equal to the tail value, inputting the data into the second-level filter and filtering the data through the second-level filter to determine the filtered target data; the second-level filter is used to filter data corresponding to non-continuous tag values.
2. The method according to claim 1, wherein The second-level filter includes a plurality of filtering nodes, and each filtering node includes a boundary record and a marking unit; filtering the data through the second-level filter includes: Judging whether the tag value is registered in the second-level filter according to the boundary record and the tag value; If the tag value is registered in the second-level filter, determining a first boundary corresponding to the tag value; and filtering the data according to the first marking unit to determine the target data; wherein the first filtering node includes a first boundary record and a first marking unit; Or, if the tag value is not registered in the second-level filter, registering a second filtering node in the second-level filter according to the tag value.
3. The method according to claim 2, characterized in that, The second filtering node includes a second boundary record and a second marking unit; registering the second filtering node in the second-level filter according to the tag value includes: Determining a first boundary value according to the current expected value; Determining a second boundary value according to the tag value; the first boundary value and the second boundary value are two interval endpoints of the second boundary record; Determining the second marking unit according to the first boundary value and the second boundary value.
4. The method according to claim 2, characterized in that Filtering the data according to the first marking unit includes: If the tag value has been registered in the first marking unit, discarding the data and determining that the target data does not include the data; Or, if the tag value has not been registered in the first marking unit, registering the tag value in the first marking unit according to the first boundary value and the tag value and determining that the target data includes the data.
5. The method according to claim 2, wherein The filtering node further includes a timer, and the timer starts timing from the establishment of the filtering node; the method further includes: If the duration of the timer is greater than a preset duration, deleting the filtering node; the preset duration is used to represent the aging time of the filtering node.
6. The method according to any one of claims 1 to 5, characterized in that, The first-level filter includes a historical value, and the historical value is the valid tag value received last time and not filtered; the method further includes: If the tag value is equal to the historical value, discarding the data and determining that the target data does not include the data; Or, if the tag value is less than the tail value, discarding the data and determining that the target data does not include the data.
7. The method according to any one of claims 1 to 5, characterized in that The method further includes: If the secondary filter is not empty, determine that the tail value is the earliest valid tag value in the secondary filter; Alternatively, if the secondary filter is empty, determine that the tail value is the expected value.
8. A data filtering system, characterized in that, The system includes: A first module, configured to receive data to be filtered and determine the tag value of the data; A second module, configured to input the data into a primary filter and determine an expected value and a tail value; the expected value is used to represent the tag value determined based on data continuity, and the tail value is used to represent the earliest valid tag value in the secondary filter; A third module, configured to, if the tag value is equal to the expected value, filter the data through the primary filter to determine the filtered target data; the primary filter is used to filter data with continuous tag values; Or, A fourth module, configured to, if the tag value is not equal to the expected value and the tag value is greater than or equal to the tail value, input the data into the secondary filter and filter the data through the secondary filter to determine the filtered target data; the secondary filter is used to filter data corresponding to non - continuous tag values.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
LPWAN technology-based redundant data filtering method for network communication management platform
CN108334424A
Data deletion method and device, computer equipment and storage medium
CN109828721A
Data filtering method and system, electronic equipment and storage medium
CN117851391A
Deduplication device and deduplication method
US20140059016A1