Multi-domain collaborative industrial storage full-link redundancy correction system
By using a multi-domain collaborative industrial storage end-to-end redundancy correction system, which utilizes five-tuple information and device model clustering, the system solves the problem of low redundancy identification accuracy in industrial data storage. It achieves accurate quantification and efficient correction of cross-device data redundancy, thereby improving the accuracy of data processing and storage utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN LARIX TECH CO LTD
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-30
AI Technical Summary
Existing technologies have low redundancy identification accuracy and poor adaptability in industrial data storage, making it difficult to adapt to multi-domain collaborative scenarios. They lack effective cross-domain data integration and processing mechanisms, and the data processing flow is not standardized enough, resulting in limited efficiency and reliability.
A multi-domain collaborative industrial storage end-link redundancy correction system is adopted. Through preliminary processing of industrial data, acquisition of redundancy features, classification of redundancy features and redundancy filtering and elimination, the system uses five-tuple information to perform fine identification and structured encapsulation of data packets. Combined with cluster grouping of equipment models and geometric redundancy quantification methods, it realizes cross-device data horizontal comparison and intelligent elimination mechanism based on the reception time of new and old data packets.
It improves the accuracy and adaptability of redundancy identification, ensures priority processing of data from critical equipment, eliminates data from illegal equipment sources or abnormal data, provides high-quality data input with unified standards, improves the accuracy and efficiency of end-to-end correction, reduces storage redundancy, and enhances the timeliness and value of data.
Smart Images

Figure REF-OBJ-1773387762081-000002 
Figure REF-OBJ-1773387762081-000003 
Figure REF-OBJ-1773387762081-000004
Abstract
Description
Technical Field
[0001] This invention belongs to the field of telecommunications technology, specifically involving a multi-domain collaborative industrial storage end-link redundancy correction system. Background Technology
[0002] With the booming development of the economy and the continuous acceleration of industrialization, the scale and complexity of industrial equipment have grown exponentially. In the field of industrial production, various sensors are widely deployed to monitor the operating status of key equipment in real time, and these sensors continuously generate massive amounts of data packets. In order to ensure efficient data processing and storage, industrial data storage technology has also developed, and a series of processes from data acquisition to storage have been initially formed.
[0003] Traditional methods for handling redundant data often rely on simple timestamps or hash matching, which struggles to accurately identify redundant features in complex time-series data, limiting the accuracy and flexibility of redundancy identification. Furthermore, existing technologies are rigid in their handling of data priorities, making it difficult to dynamically adjust data retention strategies based on device importance. In addition, traditional methods suffer from limited application scenarios, failing to adapt to multi-domain collaborative scenarios and lacking effective cross-domain data integration and processing mechanisms. Moreover, the lack of standardized data processing workflows limits efficiency and reliability.
[0004] To address the aforementioned issues, this invention proposes a multi-domain collaborative industrial storage end-to-end redundancy correction system. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a multi-domain collaborative industrial storage end-to-end redundancy correction system, which solves the problems of low accuracy and poor adaptability in redundancy identification in industrial data storage.
[0006] The objective of this invention can be achieved through the following technical solutions: A multi-domain collaborative industrial storage end-to-end redundancy correction system, comprising the following components: The industrial data preliminary processing end receives data packets from various industrial sensors monitoring target industrial equipment and transmitting them in a periodic form within a preset monitoring period. It extracts the five-tuple information associated with the corresponding data packets, performs dirty data cleaning on the data packets based on the five-tuple information, outputs standard data blocks and stores them in the cache pool. The industrial data redundancy feature acquisition end extracts the standard data block associated with the target industrial equipment from the cache pool, determines the industrial equipment cluster associated with the target industrial equipment based on the five-tuple information associated with this standard data block, extracts the standard data block in the same monitoring period within the industrial equipment cluster from the cache pool for analysis, outputs the redundancy feature vector associated with each standard data block, forms a set of redundancy feature vectors, and stores them in the cache pool. The industrial data redundancy feature classification terminal retrieves a set of redundancy feature vectors from the cache pool according to priority, and evaluates redundant data packets and non-redundant data packets. The industrial data redundancy filtering and elimination end extracts all redundant data packets associated with the same set of redundant feature vectors, and based on the reception time of the continuous data streams associated with the redundant data packets, removes old data packets and retains new data packets.
[0007] As a further aspect of the present invention, the specific method for extracting the five-tuple information associated with the corresponding data packet in the preliminary industrial data processing terminal is as follows: Obtain the monitoring cycle preset by the operator. ; The monitoring period is based on the current time. The start time is denoted as ; exist Based on the monitoring cycle The end time is obtained and recorded as . ; Obtain the time window slice preset by the operator. ; use right Slicing is performed to obtain several time windows. The total number of time windows is counted and denoted as . ; Will The time windows are sorted in chronological order to obtain the time window sequence. ; Identify the target industrial equipment and any associated industrial sensor; Industrial sensors Any time window The continuous data stream obtained from the internal monitoring target industrial equipment is regarded as a single data packet and marked as... ,in, For counting index, The first time point at which the industrial sensor uploads a continuous data stream to the local end is recorded as the receiving time associated with the continuous data stream. Extract the industrial equipment number of the target industrial equipment and denote it as... ; Determine the data type of a continuous data stream, denoted as . ; The amount of data extracted from the continuous data stream is denoted as . ; Priority for acquiring target industrial equipment, denoted as The priority of the target industrial equipment is directly proportional to its importance, which is determined by the operator. combination , , , as well as constituting a data packet The associated quintuple information ; By analogy, the quintuple information associated with other data packets can be determined.
[0008] As a further aspect of the present invention, the specific method for cleaning dirty data packets based on five-tuple information in the preliminary industrial data processing terminal is as follows: S31. Obtain data packet quintuple information ; S32. Preliminary Verification: Extraction In Verify the equipment against the whitelist of industrial equipment pre-registered by the operators; Industrial equipment whitelist, discard data packets and quintuple information Conversely, proceed to step S33; S33, Secondary Verification: Data Acquisition Volume and data types The historical average data range associated within the time window ; Discard data packets and quintuple information Conversely, proceed to step S34; S34, Transfer data packet and quintuple information Mark as valid data; S35. Similarly, mark or discard other data packets and their associated 5-tuple information.
[0009] As a further aspect of the present invention, in the industrial data preliminary processing terminal, if the data packet is determined... and quintuple information If the data is valid, then... and Combine, and as The associated standard data blocks are stored in a cache pool pre-built by the operator.
[0010] As a further aspect of the present invention, the industrial data redundancy feature acquisition terminal outputs the redundancy feature vectors associated with each standard data block, and the specific method for forming the redundancy feature vector set is as follows: Retrieve the standard data block associated with the target industrial equipment from the cache pool; Separate data packets from standard data blocks and quintuple information ; from Extract industrial equipment numbers ; based on Acquire other industrial equipment of the same type and model as the target industrial equipment to form an industrial equipment cluster. Among them, target industrial equipment ; statistics The total number of industrial equipment in China is denoted as ; Will The industrial equipment is represented in the order of acquisition as follows: Among them, the target industrial equipment is , For counting index, ; Get outside Industrial equipment within the time window The data types associated with it and Standard data blocks of the same data type as those in the dataset, along with... The associated standard data blocks are Sort the data in the specified order to obtain the standard data block sequence. ,in, The standard data block is denoted as ; right The analysis outputs the redundant feature vectors associated with any two standard data blocks, forming a set of redundant feature vectors.
[0011] As a further aspect of the present invention, the specific method by which the redundant feature vector associated between any two standard data blocks is output in the industrial data redundancy feature acquisition terminal is as follows: Obtain the standard data block sequence ; Get any standard data block ; extract quintuple information and data packets ; Sure Time window in ; With time window Construct a two-dimensional coordinate system with the horizontal axis representing the data packets and the vertical axis representing the values of the continuous data stream. Within the time window The continuous data stream within is fitted sequentially over time into the constructed two-dimensional coordinate system to obtain the continuous data stream variation curve, denoted as... ; Similarly, to obtain The remaining -1 standard data block associated with the continuous data stream change curve, along with according to Sort in order to get ; from Extract any one of the divisions External continuous data stream change curve ,Will Fit to In the two-dimensional coordinate system, where, For counting index, ,and ; Build started time And a straight line perpendicular to the horizontal axis and parallel to the vertical axis ; Build completion time And a straight line perpendicular to the horizontal axis and parallel to the vertical axis ; Summary , , and The total area of one or more closed intervals, and the value of the total area is used as the standard data block. With standard data blocks The redundant feature vectors between them are denoted as ,in, Equivalent to ; Similarly, determine Redundant feature vectors between any two standard data blocks.
[0012] As a further aspect of the present invention, the specific method for composing the set of redundant feature vectors in the industrial data redundancy feature acquisition terminal is as follows: Extracting standard data block sequences The redundant feature vectors between any two standard data blocks are sorted according to the order in which the redundant feature vectors were determined, resulting in... The associated set of redundant feature vectors is stored in a cache pool.
[0013] As a further aspect of the present invention, the specific method for evaluating redundant data packets and non-redundant data packets in the industrial data redundancy feature classification terminal is as follows: Extract the quintuple information and data packets associated with standard data blocks from all existing redundant feature vector sets in the cache pool; The set of redundant feature vectors with the highest priority is determined as the set of redundant feature vectors to be processed. Determine the total number of redundant feature vectors to be processed in the set of redundant feature vectors to be processed, denoted as . ; extract Each redundant feature vector to be processed is sequentially compared with a value preset by the operator. Perform a comparison; If the value of any redundant feature vector to be processed is less than If it is, then it is marked as redundant; Conversely, it is marked as non-redundant; Two data packets associated with any redundant feature vector to be processed are considered redundant data packets. Two data packets associated with a non-redundant feature vector to be processed are considered non-redundant data packets. As a further aspect of the present invention, the specific method for clearing outdated data packets and retaining new data packets in the industrial data redundancy filtering and elimination terminal is as follows: Extract any redundant feature vector marked as redundant, and obtain the two redundant data packets associated with it; Determine the reception time of the consecutive data streams corresponding to the two redundant data packets, and treat the redundant data packets corresponding to the consecutive data stream with the latest reception time as stale data packets and perform a clearing operation. The redundant data packets corresponding to the earliest received continuous data stream are retained as new data packets.
[0014] The beneficial effects of this invention are: This invention achieves refined identification and structured encapsulation of data sources through five-tuple information. Firstly, the priority field ensures that critical equipment data is processed first. Secondly, a dual verification mechanism (device whitelist filtering + historical data volume range verification) eliminates dirty data from illegal equipment sources or with abnormal data volumes, significantly improving data quality and preventing invalid data from occupying storage resources and interfering with analysis. Finally, binding valid data with five-tuples into standard data blocks and caching them not only ensures data integrity and facilitates subsequent traceability and collaborative processing, but also provides unified, standardized, and high-quality data input for end-to-end redundancy correction. This invention achieves cross-device data horizontal comparison based on a cluster grouping mechanism of equipment model, breaking through the limitations of single-device analysis. Secondly, it introduces an innovative geometric redundancy quantification method, which transforms abstract data similarity into calculable numerical features, making it intuitive and efficient, and avoiding the high complexity problem of traditional pattern matching. Finally, by constructing an ordered set of redundant feature vectors, it preserves the topological relationship between industrial equipment, achieving accurate quantification of multi-device data redundancy with low computational cost, significantly improving the accuracy and efficiency of end-to-end correction. This invention, based on a high-priority dynamic focusing mechanism (automatically selecting the highest-priority set of redundant feature vectors), ensures that the system always prioritizes processing data from critical equipment, significantly improving correction timeliness. Secondly, the intelligent binary search method for redundant feature vector groups achieves accurate quantitative classification between redundant and non-feature vector groups, where numerical values... Its configurability gives operators the ability to flexibly adapt to different scenarios; finally, the efficient aggregation of the group structure (pairwise combination to eliminate duplication) transforms massive vectors into group units that can be processed in batches, compressing the scale of data processing while retaining the correlation characteristics between industrial equipment, effectively avoiding resource waste in the redundancy correction process. This invention introduces an intelligent data packet replacement mechanism based on reception time, which accurately identifies and removes outdated data packets while retaining data with higher timeliness. It significantly reduces storage redundancy by eliminating duplicate data through the time dimension, directly freeing up storage space. Secondly, it enhances the timeliness value of data, ensuring that subsequent analysis and processing are always based on the latest data samples, thus improving the accuracy of correction decisions. Finally, it achieves lightweight dynamic maintenance, performing cleanup operations with a computational complexity of O(1), efficiently filtering low-value redundant data, and providing sustainable optimization power for the entire storage system. Attached Figure Description
[0015] The invention will now be further described with reference to the accompanying drawings.
[0016] Figure 1 This is a schematic diagram of the system described in this invention; Figure 2 This is a flowchart illustrating the method described in Embodiment 2 of the present invention; Figure 3 This is a flowchart illustrating the method described in Embodiment 3 of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1 Multi-domain collaborative industrial storage end-to-end redundancy correction system, such as Figure 1 As shown, this system includes the following: This system mainly includes an industrial data preliminary processing end, an industrial data redundancy feature acquisition end, an industrial data redundancy feature classification end, and an industrial data redundancy filtering and elimination end. The industrial data preliminary processing terminal is mainly used to interact with various industrial sensors, including but not limited to: ammeters, voltmeters, pressure sensors, temperature sensors, photoelectric sensors, etc.
[0019] While interacting with various industrial sensors, the industrial data preliminary processing terminal periodically receives data packets from industrial sensors monitoring industrial equipment. The periodicity is based on the monitoring cycle preset by the operator and also depends on the time window slice preset by the operator. The monitoring cycle is sliced using the time window slice to obtain several time windows within a monitoring cycle. Then, the continuous data stream transmitted by the industrial sensors within each time window is regarded as a data packet.
[0020] Further detailed analysis is then performed on any one of the determined data packets to extract the quintuple information associated with the corresponding data packet. The quintuple information represents the significant characteristics of the corresponding data packet, including: industrial equipment number, time window, data type, data volume, and priority.
[0021] Next, based on the five-tuple information of the corresponding data packets, preliminary dirty data cleaning is performed to remove dirty data that obviously causes monitoring errors, so as to prevent the dirty data from flowing into the cache pool (the cache pool is a cache space pre-built by the operator) and the database for subsequent persistent storage of normal data (the database is a storage space pre-built by the operator for final storage of data and completed persistence).
[0022] If the data packet is cleaned of dirty data and is not cleared, it is considered a normal data packet. The five-tuple information of the corresponding data packet is combined with the corresponding data packet, and the combined result is recorded as a standard data block and stored in the cache pool.
[0023] At this point, the initial processing of industrial data is complete. The above content is an example of processing any one data packet; the remaining data packets should be processed in the same way.
[0024] The industrial data redundancy feature acquisition terminal is mainly used to further process the standard data blocks in the cache pool. The ultimate goal is to output the redundancy feature vectors associated with each standard data block and form a set of redundancy feature vectors, which are then stored in the cache pool (the cache pool mentioned here is the same as the cache pool mentioned above, but the storage partitions are different to avoid data interference).
[0025] After the local end retrieves the standard data block associated with the target industrial equipment from the cache pool, it determines the industrial equipment cluster associated with this target industrial equipment through the five-tuple information associated with this standard data block (determining the industrial equipment cluster requires the industrial equipment number in the five-tuple information, which contains information including the equipment type, equipment model, and equipment details associated with the industrial equipment). ).
[0026] Then, standard data blocks within the same monitoring period (also based on 5-tuple information) of this industrial equipment cluster are extracted from the cache pool for analysis. Finally, the redundancy feature vectors associated with each standard data block are output. At this time, the redundancy feature vectors cannot indicate which of the two data packets associated with the two standard data blocks is a redundant data packet and which is a non-redundant data packet. Further determination is required.
[0027] The main function of the industrial data redundancy feature classification terminal is to extract a set of redundant feature vectors from the cache pool according to priority, and to further analyze any redundant feature vector in the set to evaluate redundant data packets and non-redundant data packets.
[0028] The main function of the industrial data redundancy filtering and elimination terminal is to perform a clearing operation after the industrial data redundancy feature classification terminal evaluates redundant data packets and non-redundant data packets, clearing out old data packets and retaining new data packets.
[0029] The main purpose of this embodiment is to effectively solve the redundancy problem in industrial data. The method mainly involves comparing the same type of data from different devices within the same time window to identify redundancy between industrial equipment clusters (redundant data from different industrial devices at the same time point). The main problems it solves are the data duplication and data disorder problems caused by the simultaneous writing of data during the real-time transmission of industrial sensors. This improves data processing efficiency and storage utilization, ensures the accuracy and timeliness of industrial data, and thus optimizes industrial production monitoring and analysis.
[0030] Example 2 This embodiment, based on Embodiment 1, discloses a method for extracting five-tuple information from data packets and performing dirty data cleaning on the data packets, such as... Figure 2 As shown, the specific steps include the following: First, as described in Example 1, it is necessary to obtain the monitoring cycle preset by the operator. The duration of the monitoring cycle is also determined by the operator, and in this solution, it is marked as... For ease of representation, and all subsequent data will be in the form of monitoring cycles. A monitoring period is represented in a way that is equal to or equal to the duration of the monitoring period. .
[0031] The following method is an example of the method described in this embodiment.
[0032] Obtain the current time (the current time is a point in time) and use the current time as a monitoring period. The start time is denoted as ; Next, at the start time Add a monitoring cycle on top of that. The duration, and will be superimposed with the monitoring cycle. The end time, denoted as the end time, is... (The start time and end time mentioned refer to a monitoring cycle.) (start time and end time).
[0033] Next, the operator's monitoring cycle is obtained. Preset time window slices (The operator shall determine this based on the actual situation), among which, The duration of the time window slice is also determined by the operator.
[0034] Next, the time window slice is used. The monitoring period determined in the above content After slicing, several time windows are obtained. The total number of all time windows is then counted and marked as follows. .
[0035] Thus, the monitoring period within the determined timeframe is obtained. A time window, then The time windows are sorted according to the chronological order of the timeline to obtain the monitoring period. The associated time window series, and represent it as: ,in, to Indicates the monitoring period The first time window within the period to the A time window.
[0036] Next, select any industrial sensor (the industrial sensor referred to here is any industrial sensor that conforms to the content described in this solution, and as mentioned above, the following method is also an example processing, and the data processing methods for other industrial sensors and their related data are the same as the following content).
[0037] The identified industrial sensors will be monitored during the cycle. Associated time window series any time window The continuous data stream obtained from the internal monitoring target industrial equipment is recorded as a single data packet, and this data packet is marked as... ,in, This is a counting index, with values ranging from 1 to... .
[0038] It should be added here that the steps and methods described in this embodiment are all performed in the industrial data preliminary processing terminal. The time point at which the industrial data preliminary processing terminal receives the continuous data stream (data packet) obtained by the industrial sensor monitoring the target industrial equipment is recorded as the reception time associated with the continuous data stream (data packet).
[0039] Next, the industrial equipment number associated with the target industrial equipment monitored by the industrial sensor is extracted and recorded as follows: The industrial equipment number includes the equipment type, equipment model, and equipment information associated with the target industrial equipment. .
[0040] Next, determine the data type of the resulting continuous data stream and denote it as... The data types mentioned include current, voltage, pressure, temperature, photoelectric data, etc.
[0041] Next, determine the amount of data in the resulting continuous data stream and denot it as... The data volume refers to the size of the continuous data stream in bytes.
[0042] Next, obtain the priority of the target industrial equipment monitored by the industrial sensor and record it as... The priority of industrial equipment is directly proportional to its importance. The importance of industrial equipment is determined by the operator based on the actual situation. The higher the importance of industrial equipment, the higher its priority, and vice versa.
[0043] Thus, the industrial equipment number was obtained. Time window Data types Data volume and priority Based on the five types of information obtained above, a corresponding data packet is constructed. The associated quintuple information , is represented as: .
[0044] Similarly, when data packets are obtained from other industrial sensors, the quintuple information associated with the corresponding data packet is determined using the method described above.
[0045] Next, extract the data packet. The associated quintuple information ; Information on quintuples Perform preliminary verification to obtain quintuple information. Industrial equipment number in and number the industrial equipment Verify against the whitelist of industrial equipment pre-registered by operators (the whitelist refers to the industrial equipment number associated with an industrial equipment that is put into use when it is put into use, and is removed from the whitelist if it is taken out of use). If the target industrial equipment's industrial equipment number If the industrial equipment is whitelisted, the data packet will be discarded. and its associated quintuple information If the target industrial equipment's industrial equipment number For industrial equipment whitelists, a second verification is performed; Get the quintuple information again The amount of data acquired and data types The range of historical average data volume associated within a time window , the amount of data Compared with the historical average data volume range Compare the data, if the amount of data is... If so, the data packet is discarded. and its associated quintuple information ,if This indicates that both the preliminary and secondary checks have passed, and the data packet will be sent to the next location. Mark as valid data.
[0046] Then send the data packet With data packets The associated quintuple information The data is combined, and the result is denoted as a data packet. The associated standard data blocks are stored in a cache pool pre-built by the operator.
[0047] The core of this embodiment lies in processing the data packets collected by industrial sensors during the monitoring period to generate standard data blocks for subsequent processing, with the ultimate goal of ensuring the validity and accuracy of the data.
[0048] Example 3 This embodiment discloses a method for outputting redundant feature vectors associated with each standard data block and forming a set of redundant feature vectors, such as... Figure 3 As shown, it specifically includes the following: Based on the content described in Example 2, data packets can be obtained from the cache pool. The associated standard data blocks are then extracted and separated to obtain data packets. and quintuple information .
[0049] Then, from the information obtained by separation of the quintuples Extracting data packets The associated industrial equipment number As described in Examples 1 and 2, based on the industrial equipment number Obtain the equipment type and model associated with the target industrial equipment.
[0050] Based on the equipment type and model associated with the target industrial equipment, other industrial equipment of the same type and model as the target industrial equipment are obtained, and together with the target industrial equipment, an industrial equipment cluster is formed (the target industrial equipment belongs to the industrial equipment cluster), and the industrial equipment cluster is marked as: .
[0051] Next, acquire the industrial equipment cluster. The total number of industrial equipment in the country, and record the total number as At this point, industrial equipment clusters can be... In Individual industrial equipment can be differentiated and represented as follows: The sorting order is the order in which the industrial equipment was acquired, and the industrial equipment number is... The corresponding target industrial equipment is , This is a counting index, with values ranging from 1 to... .
[0052] Next, acquire the industrial equipment cluster. In Industrial equipment within the time window The data types associated with the target industrial equipment Standard data blocks with the same data type in the standard data blocks (that is, industrial equipment clusters) All industrial sensors associated with the industrial equipment in this example are the same as the target industrial equipment described in Example 2. Since they are the same industrial sensors, the data types obtained from industrial sensor monitoring are also the same.
[0053] Due to industrial equipment clusters Including Each industrial device will be available within a time window. Internally generate a data type and target industrial equipment The standard data blocks are of the same data type as the standard data blocks, so in the time window Nekode Industrial Equipment Cluster In Associated with individual industrial equipment A standard data block, the resulting Each standard data block is arranged according to industrial equipment clusters. The data blocks are sorted in the following order, and the sorted result is denoted as the standard data block sequence, represented as: Among them, the target industrial equipment is The standard data block is represented as: .
[0054] Continue to determine standard data blocks Five-tuple information and data packets Further determine the quintuple information Time window in .
[0055] Then, using the timeline as the horizontal axis of a two-dimensional coordinate system, the data packet... The corresponding value of the continuous data stream (for example, if the data type of the continuous data stream is voltage, then the time window) The voltage value within the continuous data stream is used as the horizontal axis and vertical axis of a two-dimensional coordinate system. The usable range of the horizontal axis of this two-dimensional coordinate system is a time window. Duration.
[0056] data packet The values of the corresponding continuous data stream are fitted onto the constructed two-dimensional coordinate system according to the timeline to obtain the data packets. The associated continuous data stream change curve, and denoted as .
[0057] Repeat the above steps to obtain the standard data block sequence. Excluding standard data blocks Other The continuous data stream associated with -1 standard data block is fitted into the constructed two-dimensional coordinate system (each standard data block corresponds to a new two-dimensional coordinate system), resulting in... The resulting continuous data stream variation curve associated with -1 standard data block will be... -1 continuous data stream variation curve along with the continuous data stream variation curve According to the standard data block sequence Sort them in the following order, and the sorted result is represented as: .
[0058] Then from Extract any one of the divisions Other than the continuous data stream change curve, and labeled as: ,in This is a counting index, with values ranging from 1 to... ,and Not equal to ; The extracted continuous data stream change curve Fit to continuous data stream change curve In the two-dimensional coordinate system in which it exists, and after passing through the time window The associated start time Construct a straight line parallel to both the horizontal and vertical axes, denoted as . ; Then through the time window The associated end time Construct a straight line parallel to both the horizontal and vertical axes, denoted as . ; At this point, the straight line ,straight line Continuous data stream change curve and continuous data stream change curves Form one or more closed intervals, calculate the total area of the one or more closed intervals using the principles of calculus, and use the calculated total area as the standard data block. With standard data blocks Redundant feature vectors between them, that is, standard data blocks With standard data blocks The redundant feature vector between the two corresponding data packets is denoted as . What needs to be explained here is: Equivalent to .
[0059] Repeat the above steps to determine the standard data block sequence. The redundant feature vectors between any two standard data blocks are identified, and these redundant feature vectors are sorted according to the order in which they were determined. The sorted result is then used as the standard data block sequence. The associated set of redundant feature vectors is stored in the cache pool.
[0060] The main purpose of this embodiment is to extract standard data blocks from the cache pool, separate data packets and quintuple information, obtain the type and model of the target device through the industrial equipment number in the quintuple information, and then identify other industrial devices of the same type and model to form an industrial equipment cluster. Next, standard data blocks of the same data type for each industrial device in the industrial equipment cluster within the time window are obtained, a standard data block sequence is constructed, and then the continuous data flow change curve of the data packet is fitted with the time line as the horizontal axis and the data flow value as the vertical axis. By calculating the area of the closed interval between different continuous data flow change curves, a set of redundant feature vectors is formed and stored in the cache pool.
[0061] All data in the formulas described above are numerical calculations performed after removing their dimensions. Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.
[0062] The above description is merely an example and illustration of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.
[0063] It should be stated that all user data collected in this application was collected with the user's consent and authorization. Furthermore, the uses of user data are legal and compliant, and the use and processing of user data comply with the relevant laws, regulations, and standards of the relevant regions.
Claims
1. A multi-domain collaborative industrial storage end-to-end redundancy correction system, characterized in that, This system includes the following: The industrial data preliminary processing end receives data packets from various industrial sensors monitoring target industrial equipment and transmitting them in a periodic form within a preset monitoring period. It extracts the five-tuple information associated with the corresponding data packets, performs dirty data cleaning on the data packets based on the five-tuple information, outputs standard data blocks and stores them in the cache pool. The industrial data redundancy feature acquisition end extracts the standard data block associated with the target industrial equipment from the cache pool, determines the industrial equipment cluster associated with the target industrial equipment based on the five-tuple information associated with this standard data block, extracts the standard data block in the same monitoring period within the industrial equipment cluster from the cache pool for analysis, outputs the redundancy feature vector associated with each standard data block, forms a set of redundancy feature vectors, and stores them in the cache pool. The industrial data redundancy feature classification terminal retrieves a set of redundancy feature vectors from the cache pool according to priority, and evaluates redundant data packets and non-redundant data packets. The industrial data redundancy filtering and elimination end extracts all redundant data packets associated with the same set of redundant feature vectors, and based on the reception time of the continuous data streams associated with the redundant data packets, removes old data packets and retains new data packets.
2. The multi-domain collaborative industrial storage end-to-end redundancy correction system according to claim 1, characterized in that, In the preliminary industrial data processing terminal, the specific method for extracting the five-tuple information associated with the corresponding data packet is as follows: Obtain the monitoring cycle preset by the operator. ; The monitoring period is based on the current time. The start time is denoted as ; exist Based on the monitoring cycle The end time is obtained and recorded as . ; Obtain the time window slice preset by the operator. ; use right Slicing is performed to obtain several time windows. The total number of time windows is counted and denoted as . ; Will The time windows are sorted in chronological order to obtain the time window sequence. ; Identify the target industrial equipment and any associated industrial sensor; Industrial sensors Any time window The continuous data stream obtained from the internal monitoring target industrial equipment is regarded as a single data packet and marked as... ,in, For counting index, The first time point at which the industrial sensor uploads a continuous data stream to the local end is recorded as the receiving time associated with the continuous data stream. Extract the industrial equipment number of the target industrial equipment and denote it as... ; Determine the data type of a continuous data stream, denoted as . ; The amount of data extracted from the continuous data stream is denoted as . ; Priority for acquiring target industrial equipment, denoted as The priority of the target industrial equipment is directly proportional to its importance, which is determined by the operator. combination , , , as well as constituting a data packet The associated quintuple information ; By analogy, the quintuple information associated with other data packets can be determined.
3. The multi-domain collaborative industrial storage end-to-end redundancy correction system according to claim 2, characterized in that, In the aforementioned industrial data preliminary processing terminal, the specific method for dirty data cleaning of data packets based on five-tuple information is as follows: S31. Obtain data packet quintuple information ; S32. Preliminary Verification: Extraction In Verify the equipment against the whitelist of industrial equipment pre-registered by the operators; Industrial equipment whitelist, discard data packets and quintuple information Conversely, proceed to step S33; S33, Secondary Verification: Data Acquisition Volume and data types The historical average data range associated within the time window ; Discard data packets and quintuple information Conversely, proceed to step S34; S34, Transfer data packet and quintuple information Mark as valid data; S35. Similarly, mark or discard other data packets and their associated 5-tuple information.
4. The multi-domain collaborative industrial storage end-to-end redundancy correction system according to claim 3, characterized in that, In the aforementioned industrial data preliminary processing terminal, if the data packet is determined... and quintuple information If the data is valid, then... and Combine, and as The associated standard data blocks are stored in a cache pool pre-built by the operator.
5. The multi-domain collaborative industrial storage end-to-end redundancy correction system according to claim 1, characterized in that, The industrial data redundancy feature acquisition terminal outputs the redundancy feature vectors associated with each standard data block, and the specific method for forming the redundancy feature vector set is as follows: Retrieve the standard data block associated with the target industrial equipment from the cache pool; Separate data packets from standard data blocks and quintuple information ; from Extract industrial equipment numbers ; based on Acquire other industrial equipment of the same type and model as the target industrial equipment to form an industrial equipment cluster. Among them, target industrial equipment ; statistics The total number of industrial equipment in China is denoted as ; Will The industrial equipment is represented in the order of acquisition as follows: Among them, the target industrial equipment is , For counting index, ; Get outside Industrial equipment within the time window The data types associated with it and Standard data blocks of the same data type as those in the dataset, along with... The associated standard data blocks are Sort the data in the specified order to obtain the standard data block sequence. ,in, The standard data block is denoted as ; right The analysis outputs the redundant feature vectors associated with any two standard data blocks, forming a set of redundant feature vectors.
6. The multi-domain collaborative industrial storage end-to-end redundancy correction system according to claim 5, characterized in that, The specific method for outputting the redundancy feature vector associated between any two standard data blocks in the industrial data redundancy feature acquisition terminal is as follows: Obtain the standard data block sequence ; Get any standard data block ; extract quintuple information and data packets ; Sure Time window in ; With time window Construct a two-dimensional coordinate system with the horizontal axis representing the data packets and the vertical axis representing the values of the continuous data stream. Within the time window The continuous data stream within is fitted sequentially over time into the constructed two-dimensional coordinate system to obtain the continuous data stream variation curve, denoted as... ; Similarly, to obtain The remaining -1 standard data block associated with the continuous data stream change curve, along with according to Sort in order to get ; from Extract any one of the divisions External continuous data stream change curve ,Will Fit to In the two-dimensional coordinate system, where, For counting index, ,and ; Build started time And a straight line perpendicular to the horizontal axis and parallel to the vertical axis ; Build by end time And a straight line perpendicular to the horizontal axis and parallel to the vertical axis ; Summary , , and The total area of one or more closed intervals, and the value of the total area is used as the standard data block. With standard data blocks The redundant feature vectors between them are denoted as ,in, Equivalent to ; Similarly, determine Redundant feature vectors between any two standard data blocks.
7. The multi-domain collaborative industrial storage end-to-end redundancy correction system according to claim 6, characterized in that, In the industrial data redundancy feature acquisition terminal, the specific method for composing the set of redundancy feature vectors is as follows: Extracting standard data block sequences The redundant feature vectors between any two standard data blocks are sorted according to the order in which the redundant feature vectors were determined, resulting in... The associated set of redundant feature vectors is stored in a cache pool.
8. The multi-domain collaborative industrial storage end-to-end redundancy correction system according to claim 6, characterized in that, In the industrial data redundancy feature classification terminal, the specific method for evaluating redundant data packets and non-redundant data packets is as follows: Extract the quintuple information and data packets associated with standard data blocks from all existing redundant feature vector sets in the cache pool; The set of redundant feature vectors with the highest priority is determined as the set of redundant feature vectors to be processed. Determine the total number of redundant feature vectors to be processed in the set of redundant feature vectors to be processed, denoted as . ; extract Each redundant feature vector to be processed is sequentially compared with a value preset by the operator. Perform a comparison; If the value of any redundant feature vector to be processed is less than If it is, then it is marked as redundant; Conversely, it is marked as non-redundant; Two data packets associated with any redundant feature vector to be processed are considered redundant data packets. Two data packets associated with a non-redundant feature vector to be processed are considered non-redundant data packets.
9. The multi-domain collaborative industrial storage end-to-end redundancy correction system according to claim 8, characterized in that, In the aforementioned industrial data redundancy filtering and elimination terminal, the specific method for clearing outdated data packets and retaining new data packets is as follows: Extract any redundant feature vector marked as redundant, and obtain the two redundant data packets associated with it; Determine the reception time of the consecutive data streams corresponding to the two redundant data packets, and treat the redundant data packets corresponding to the consecutive data stream with the latest reception time as stale data packets and perform a clearing operation. The redundant data packets corresponding to the earliest received continuous data stream are retained as new data packets.