A method for real-time judging the given delay repeatability of big data itself

Through iterative calculation autocorrelation method, only newly added and removed data elements and their adjacent elements in the big data calculation window are processed, which solves the problem of inefficient real-time autocorrelation calculation of big data and realizes efficient and energy-saving autocorrelation calculation.

CN112035792BActive Publication Date: 2025-07-18吕纪竹
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910478187.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-03
Publication Date
2025-07-18
Estimated Expiration
2039-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently calculate autocorrelation in real time on big data to judge the repetition of a given delay, resulting in waste of computing resources and inefficient computing efficiency.

Method used

Through iterative calculation, only newly added and removed data elements and their adjacent elements in the big data calculation window are accessed and calculated, avoiding repeated calculations and access to all data elements, and achieving efficient autocorrelation calculations.

Benefits of technology

Improve computing efficiency, save computing resources and energy consumption, making it possible to judge the given delay repetition of big data in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112035792B_ABST
    Figure CN112035792B_ABST
Patent Text Reader

Abstract

The autocorrelation with a given delay can be used to determine the repeatability of big data with a given delay. The present invention discloses a method, system, and computing system program product for iteratively calculating the autocorrelation with a specified delay of a computing window of a given scale, thereby enabling real-time determination of the repeatability of big data with a given delay. Embodiments of the present invention include components for iteratively calculating the autocorrelation with a specified delay of an adjusted computing window based on more than two components of the autocorrelation with a specified delay of a pre-adjustment computing window, and then generating the autocorrelation with a specified delay of the adjusted computing window based on the more than two components obtained by iterative calculation as needed. Iteratively calculating the autocorrelation avoids accessing all data elements in the adjusted computing window and performing repeated calculations, thereby improving the computing efficiency, saving computing resources, and reducing the energy consumption of the computing system, making it possible to efficiently and with low power consumption perform real-time determination of the repeatability of big data with a given delay and making some scenarios of real-time determination of the repeatability of big data with a given delay possible that were previously impossible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Big data or stream data analysis. Background Art

[0002] The Internet, mobile communication, navigation, online games, sensing technology, and large-scale computing infrastructure generate a vast amount of data every day. Big data is data that exceeds the processing capacity of traditional database systems and the analysis capabilities of traditional analysis methods due to its huge scale, rapid change, and growth rate. Current big data analysis methods involve applying a large amount of computing resources, which are very expensive but still cannot meet the need to make real-time decisions using the latest data information, especially in the Internet of Things, the financial industry, etc. How to process and analyze big data efficiently, in real time, and resource-saving is a difficult challenge for data analysts and computer scientists.

[0003] Autocorrelation, also known as lag correlation or serial correlation, is a measure of the degree of correlation between a specific time series and the time series itself with a lag of l time points. It can be obtained by dividing the cross-correlation of the observations of a time series separated by l time points by its standard variance. An autocorrelation value of 1 or close to 1 for a certain lag can be considered that big data shows its own repeating pattern after that lag. Therefore, it is obvious to judge the repeatability of big data with a given lag based on autocorrelation, while the difficulty and challenge lie in how to calculate autocorrelation on big data in real time.

[0004] Autocorrelation may need to be recalculated after some data changes in big data to reflect the latest data situation. For example, perhaps the autocorrelation needs to be calculated for a calculation window containing n data elements of a big data set newly added to the storage medium. In this way, every time two data elements are received or accessed, one data element is added to the calculation window while the other data element is removed from the calculation window, and the n data elements in the calculation window will be accessed to recalculate the autocorrelation. In this way, each data change may only change a small part of the data in the calculation window. Recalculating the autocorrelation using all the data elements in the calculation window involves repeated data access and calculation, so it is time-consuming and resource-wasting.

[0005] Depending on the need, the scale of the calculation window may be very large. For example, the data elements in the calculation window may be distributed on thousands of computing devices in the cloud platform. Recalculating the autocorrelation on big data using traditional methods after some data changes cannot achieve real-time processing and occupies and wastes a large amount of computing resources, and also makes it impossible to meet the demand to realize the repeatability of judging the given lag of big data in real time. Summary of the Invention

[0006] The present invention extends to a method, a system, and a computer system program product for calculating the autocorrelation of a given latency of big data in an iterative manner so as to be able to judge the repeatability of the given latency of big data itself in real time. Iteratively calculating the autocorrelation of a specified latency l (l>0) for an adjusted calculation window includes iteratively calculating more than two (p (p>1)) components of the autocorrelation of the specified latency for the adjusted calculation window based on more than two components of the autocorrelation of the specified latency for the pre-adjustment calculation window, and then generating the autocorrelation of the specified latency for the adjusted calculation window based on the more than two components obtained by the iterative calculation as needed. Iteratively calculating the autocorrelation only requires accessing and using the components obtained by the iterative calculation, the newly added and removed data elements, and l data elements adjacent to the newly added and removed data elements on both sides of the calculation window respectively, thereby avoiding accessing all the data elements in the adjusted calculation window and performing repeated calculations, reducing the data access latency, improving the calculation efficiency, saving the calculation resources, and reducing the energy consumption of the computer system, making it possible to judge the repeatability of the given latency of big data itself in real time with high efficiency and low consumption, and making some scenarios of judging the repeatability of the given latency of big data itself that were impossible become possible.

[0007] The computer system initializes more than two (p (p>1)) components of the autocorrelation of a pre-adjustment calculation window of a big data set stored on one or more storage media. The initialization of the more than two components includes calculating the more than two components based on the data elements in the pre-adjustment calculation window through the definition of the components, or receiving or accessing the already calculated more than two components from a computer-readable medium.

[0008] The computer system accesses a data element to be removed from the pre-adjustment calculation window and a data element to be added to the pre-adjustment calculation window.

[0009] The computer system adjusts the pre-adjustment calculation window by removing the data element to be removed from the pre-adjustment calculation window and adding the data element to be added to the pre-adjustment calculation window.

[0010] The computer system directly iteratively calculates one or more (let v (1≤v≤p)) components of the autocorrelation of the specified latency for the adjusted calculation window. Directly iteratively calculating the one or more components includes: accessing v components of the specified latency of the pre-adjustment calculation window; mathematically removing the contribution of the removed data element from each of the accessed components; and mathematically adding the contribution of the added data element to each of the accessed components.

[0011] The computing system indirectly iteratively computes, as needed, w = p - v components of the autocorrelation of a specified latency of an adjusted computation window. Indirectly iteratively computing the w components of the specified latency includes indirectly iteratively computing each of the w components one by one. Indirectly iteratively computing one component of the specified latency includes: accessing and using one or more components of the specified latency other than that component to compute that component. These one or more components may have been initialized, directly iteratively computed, indirectly iteratively computed, or computed in any other way.

[0012] The computing system generates an autocorrelation of a specified latency of an adjusted computation window based on components of the autocorrelation of the specified latency of one or more iteratively computed adjusted computation windows.

[0013] The computing system can continuously access a data element to be removed and a data element to be added, adjust the computation window, directly iteratively compute v components of the specified latency, indirectly iteratively compute w = p - v components of the specified latency as needed, and compute the autocorrelation of the specified latency. The computing system can repeat this process as needed multiple times.

[0014] This summary introduces some selected concepts in a simplified manner, which will be described in further detail below. This summary is neither intended to identify the key features or essential features of the claimed subject matter nor to be used to help confirm the scope included in the claimed subject matter.

[0015] Other features and advantages of the present invention will be apparent from the following description, will be partially apparent from the description, or will be learned from the practice of the present invention. The features and advantages of the present invention can be realized and obtained from the methods, apparatuses, and combinations thereof particularly pointed out in the appended claims. These and other features of the present invention will become more fully apparent and clear in the following description and the appended claims or from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To describe the manner in which the above and other advantages and features of the present invention can be obtained, a more specific description of the present invention briefly described above will be presented by referring to specific embodiments shown in the appended drawings. These drawings merely depict typical embodiments of the present invention and should not, therefore, be construed as limiting the scope of the present invention:

[0017] Figure 1 Illustrates a high-level overview of an example computing system that supports iterative computation of autocorrelation.

[0018] Figure 1-1 Shows an example computing system architecture that supports iterative computation of the autocorrelation of big data and where all components are computed in a direct iterative manner.

[0019] Figure 1-2 Shows an example of a computing system architecture that supports autocorrelation for iterative calculation of big data, where some components are calculated in a direct iterative manner and some components are calculated in an indirect iterative manner.

[0020] Figure 2 Shows a flowchart of an example method for iterative calculation of the autocorrelation of big data.

[0021] Figure 3-1 Shows the data removed and added when the calculation window 300A moves to the right.

[0022] Figure 3-2 Shows the data accessed for iterative calculation of autocorrelation when the calculation window 300A moves to the right.

[0023] Figure 3-3 Shows the data removed and added when the calculation window 300B moves to the left.

[0024] Figure 3-4 Shows the data accessed for iterative calculation of autocorrelation when the calculation window 300B moves to the left.

[0025] Figure 4-1 Shows the definition of autocorrelation and the traditional equation for calculating autocorrelation.

[0026] Figure 4-2 Shows the first autocorrelation iterative calculation algorithm (Iterative Algorithm 1).

[0027] Figure 4-3 Shows the second autocorrelation iterative calculation algorithm (Iterative Algorithm 2).

[0028] Figure 4-4 Shows the third autocorrelation iterative calculation algorithm (Iterative Algorithm 3).

[0029] Figure 5-1 Shows the first calculation window for a computing instance.

[0030] Figure 5-2 Shows the second calculation window for a computing instance.

[0031] Figure 5-3 Shows the third calculation window for a computing instance.

[0032] Figure 6-1 Shows the comparison of the computational amounts of traditional and iterative autocorrelation algorithms when the calculation window size is 4 and the delay is 1.

[0033] Figure 6-2 Shows the comparison of the computational amounts of traditional and iterative autocorrelation algorithms when the calculation window size is 1000000 and the delay is 1. Detailed implementation mode

[0034] Calculating autocorrelation is an effective method for judging the self-given delay repeatability of a time series or streaming big data. The present invention extends to a method, system and computing system program product for judging the self-given delay repeatability of big data in real time by iteratively calculating the autocorrelation of a specified delay l (1 ≤ l < n) of a computing window with a computing scale of n (n > 1) for big data. A computing system includes one or more processor-based computing devices. Each computing device includes one or more processors. The computing system includes one or more storage media. There is a data set on at least one of the one or more storage media. A plurality of data elements related to autocorrelation calculation from the data set form a pre-adjustment computing window. The computing window scale n (n > l) indicates the number of data elements in a computing window of the data set. The delay l indicates the delay used for autocorrelation calculation. Embodiments of the present invention include iteratively calculating the autocorrelation of more than two (p (p > 1)) components of the specified delay of the post-adjustment computing window based on the autocorrelation of the specified delay of the pre-adjustment computing window, and then generating the autocorrelation of the specified delay of the post-adjustment computing window based on more than two components of the iterative calculation as needed. Iteratively calculating autocorrelation avoids accessing all data elements in the post-adjustment computing window and performing repeated calculations, thereby improving computing efficiency, saving computing resources and reducing the energy consumption of the computing system, making it possible to efficiently and with low consumption judge the self-given delay repeatability of big data in real time and making some scenarios of judging the self-given delay repeatability of big data in real time possible that were previously impossible.

[0035] Autocorrelation, also known as lag correlation or serial correlation, is a measure of the degree of correlation between a particular time series and the time series itself delayed by l time points. It can be obtained by dividing the covariance of observations of a time series separated by l time points by its standard variance. If the autocorrelation for all different delay values of a time series is calculated, the autocorrelation function of the time series is obtained. For a time series that does not change over time, its autocorrelation values will exponentially decrease to 0. The values of autocorrelation range between -1 and +1. A value of +1 indicates a perfect positive linear relationship between past and future values of the time series, while a value of -1 indicates a perfect negative linear relationship between past and future values of the time series.

[0036] In this article, big data includes a plurality of data elements with a certain order and stored on one or more computing device-readable media. The difference between big data and streaming data is that when processing big data, all historical data can be accessed, so it may not be necessary to create another data buffer to store newly received data elements.

[0037] In this document, a computation window is a moving window on a large dataset that contains the data involved in the autocorrelation computation. The computation window can move left or right. For example, when processing data newly added to the large dataset, the computation window moves to the right. At this time, a data is added to the right side of the computation window and a data element on the left side of the computation window is removed from the computation window. When recomputing the autocorrelation of data elements previously added to the large dataset, the computation window moves to the left. At this time, a data is added to the left side of the computation window and a data element on the right side of the computation window is removed from the computation window. The goal is to iteratively compute the autocorrelation at a given delay of the data elements in the computation window whenever the computation window moves left or right by one or more data. These two cases can be handled in the same way but only the equations used for iterative computation are different. For the purpose of illustration rather than limitation, in the following description, the first case (the computation window moves to the right) will be used as an example to describe and explain the implementation of the present invention.

[0038] In this document, a component of the autocorrelation is a quantity or expression that appears in the autocorrelation definition formula or any transformation of its definition formula. The autocorrelation is its own largest component. The following are some examples of components of the autocorrelation.

[0039]

[0040] The autocorrelation can be computed based on one or more components or their combination, so multiple algorithms support iterative autocorrelation computation.

[0041] A component can be iteratively computed directly or indirectly. The difference between them is that when a component is iteratively computed directly, the component is computed based on the value of the component in the previous round of computation, while when the component is iteratively computed indirectly, the component is computed using other components outside the component.

[0042] For a given component, it may be iteratively computed directly in one algorithm but iteratively computed indirectly in another algorithm.

[0043] For any algorithm, at least two components will be iteratively computed, one component is iteratively computed directly and the other component is iteratively computed directly or indirectly. For a given algorithm, assume the total number of different components used is p (p > 1). If the number of components iteratively computed directly is v (1 ≤ v ≤ p), then the number of components iteratively computed indirectly is w = p - v (0 ≤ w < p). It is possible that all components are iteratively computed directly (in this case v = p > 1 and w = 0). However, regardless of whether the result of the autocorrelation is needed and accessed in a specific round, the components iteratively computed directly must be computed.

[0044] For a given algorithm, if a component is directly iteratively computed, then the component must be computed (i.e., whenever an existing data element is removed from the computation window and whenever a data element is added to the computation window). However, if a component is indirectly iteratively computed, then the component can be computed as needed, i.e., only when autocorrelation needs to be computed and accessed, by using one or more other components outside of the component. Thus, when autocorrelation is not accessed in a particular computation round, only a small number of components need to be iteratively computed. An indirectly iteratively computed component may be used in the direct iterative computation of a component, in which case the computation of the indirectly iteratively computed component cannot be omitted.

[0045] The implementation of the present invention includes two or more (p (p>2)) components that iteratively compute autocorrelation based on two or more (p (p>2)) components computed for a previous computation window.

[0046] The computing system initializes two or more (p (p>2)) components for the autocorrelation of a given delay l (l≥1) of a computation window of a given size n (n>1). The initialization of the two or more components includes computing based on the data elements in the computation window according to their definitions or accessing or receiving the already computed components from one or more computer-readable media.

[0047] The computing system accesses a data element to be removed from the computation window and a data element to be added to the computation window.

[0048] The computing system adjusts the computation window by: removing the data element to be removed from the computation window and adding the data element to be added to the computation window.

[0049] The computing system iteratively computes a sum or an average or a sum and an average of the adjusted computation window.

[0050] The computing system directly iteratively computes one or more (let v (1≤v<p) ones) components other than the sum and the average for the autocorrelation of a specified delay of the adjusted computation window. Directly iteratively computing v components at a given delay l includes: accessing the removed data element, the l data elements adjacent to the removed data element in the pre-adjustment computation window, the added data element, the l data elements adjacent to the added data element in the post-adjustment computation window, and the v components for the given delay l computed for the pre-adjustment computation window; mathematically removing any contribution of the removed data element from each of the accessed components; and mathematically adding any contribution of the added data element to each of the accessed components.

[0051] The computing system iteratively calculates, as needed, w = p - v components of the autocorrelation at a given delay l for the adjusted computing window indirectly. Iteratively calculating the w components of the autocorrelation at a given delay l indirectly includes iteratively calculating each of the w components at the given delay l one by one and indirectly. Iteratively calculating one component at a given delay l indirectly includes accessing one or more components at the given delay l outside of that component and calculating that component based on the accessed components. These one or more components at the given delay l may have been initialized, directly iteratively calculated, indirectly iteratively calculated, or calculated in any other way.

[0052] The computing system generates the autocorrelation at delay l using one or more components at delay l that have been initialized or iteratively calculated, as needed.

[0053] The computing system may continuously access data elements to be removed from and data elements to be added to the computing window, adjust the computing window, iteratively calculate a sum or an average or a sum and an average of the adjusted computing window, directly iteratively calculate one or more, i.e., v, components at specified delays, iteratively calculate w = p - v components at specified delays indirectly as needed, generate the autocorrelation at a given delay using one or more iteratively calculated components as needed, and repeat this process as needed multiple times.

[0054] Embodiments of the present invention may include or utilize a computing device hardware that includes, for example, one or more processors and a storage device as described in more detail below, a special-purpose or general-purpose computing device. The scope of the embodiments of the present invention also includes physical and other computing device-readable media for carrying or storing computing device-executable instructions and / or data structures. These computing device-readable media may be any media accessible by a general-purpose or special-purpose computing device. A computing device-readable media that stores computing device-executable instructions is a storage medium (device). A computing device-readable media that carries computing device-executable instructions is a transmission medium. Thus, by way of example and not limitation, embodiments of the present invention may include at least two different types of computing device-readable media: storage media (devices) and transmission media.

[0055] Storage media (devices) include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), solid state drives (SSD), flash memory, phase change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other media that can be used to store program code in the form of computing device-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computing device.

[0056] A "network" is defined as one or more data links that enable computing devices and / or modules and / or other electronic devices to transfer electronic data. When information is transferred or provided to a computing device by a network or by another communication connection (wired, wireless, or a combination of wired and wireless), the computing device treats the connection as a transmission medium. A transmission medium can include a network and / or a data link that is used to carry program code in the form of computer-executable instructions or data structures that are required and that can be accessed by a general or special purpose computing device. Combinations of the above should also be included within the scope of computer-readable media.

[0057] In addition, when applying different computing device components, program code in the form of computer-executable instructions or data structures can be automatically transferred from a transmission medium to a storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received from a network or a data link can be temporarily stored in random access memory in a network interface module (e.g., NIC), and then ultimately transferred to the random access memory of the computing device and / or to a smaller volatile storage medium (device) of the computing device. Therefore, it should be understood that storage media (devices) can be included in computing device components that also (or even primarily) utilize a transmission medium.

[0058] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general or special purpose computing device to perform a particular function or a set of functions. Computer-executable instructions can be, for example, binary, intermediate format instructions such as assembly code, or even source code. Although the subject matter is described in specific language of structural features and / or method acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or acts described above. Rather, the described features or acts are disclosed only as examples for implementing the claims.

[0059] Embodiments of the present invention can be implemented in a network computing environment configured with various types of computing devices, including personal computers, desktops, laptops, information processors, handheld devices, multiprocessing systems, microprocessor-based or programmable consumer electronics, network computers, minicomputers, mainframe computers, supercomputers, mobile phones, palmtop computers, tablet computers, pagers, routers, switches, and the like. Embodiments of the present invention can also be applied to a distributed system environment consisting of local or remote computing devices that perform tasks interconnected by a network (i.e., via a wired data link, a wireless data link, or a combination of a wired data link and a wireless data link). In a distributed system environment, program modules can be stored on local or remote storage devices.

[0060] Embodiments of the present invention can also be implemented in a cloud computing environment. In this description and the following claims, "cloud computing" is defined as a model that enables on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be exploited by the market to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. The shared pool of configurable computing resources can be quickly provisioned through virtualization and provided with low administrative overhead or low service provider interaction, and then adjusted accordingly.

[0061] The cloud computing model can include various features such as, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, etc. The cloud computing model can also be embodied in various service models, such as, Software as a Service ("SaaS"), Platform as a Service ("PaaS"), and Infrastructure as a Service ("IaaS"). The cloud computing model can also be deployed through different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, etc.

[0062] Since the present invention effectively reduces the requirements for computing power, its embodiments can also be applied to edge computing.

[0063] Several examples will be given in the following sections.

[0064] Figure 1 A high-level overview of an example computing system 100 for autocorrelation of big data iterative computing is illustrated. Referring Figure 1 , the computing system 100 includes multiple devices connected by different networks, such as local area network 1021, wireless network 1022, and Internet 1023, etc. The multiple devices include, for example, a data analysis engine 1007, a storage system 1011, a real-time data stream 1006, and multiple distributed computing devices that can schedule data analysis tasks and / or query data analysis results, such as personal computers 1016, handheld devices 1017, and desktop computers 1018, etc.

[0065] The data analysis engine 1007 can include one or more processors, such as CPU 1009 and CPU 1010, one or more system memories, such as system memory 1008, and component computing module 131 and autocorrelation computing module 192. The details of module 131 will be illustrated in more detail in other diagrams (for example, Figure 1-1 and Figure 1-2 ). The storage system 1011 can include one or more storage media, such as storage media 1012 and storage media 1014, which can be used to store large data sets. For example, 1012 and / or 1014 can include data set 123. The data sets in the storage system 1011 can be accessed by the data analysis engine 1007.

[0066] Generally, the data stream 1006 may include streaming data from different data sources, such as stock prices, audio data, video data, geospatial data, Internet data, mobile communication data, online game data, bank transaction data, sensor data, and / or closed caption data, etc. Several are depicted here by way of example. The real-time data 1000 may include data collected in real time from sensors 1001, stocks 1002, communications 1003, banks 1004, and so on. The data analysis engine 1007 may receive data elements from the data stream 1006. Data from different data sources may be stored in the storage system 1011 and accessed for big data analysis. For example, the data set 123 may come from different data sources and be accessed for big data analysis.

[0067] Please understand Figure 1 is to introduce some concepts in a very simplified form. For example, the distribution devices 1016 and 1017 may be connected to the data analysis engine 1007 via a firewall. The data accessed or received by the data analysis engine 1007 from the data stream 1006 and / or the storage system 1011 may be filtered by a data filter, and so on.

[0068] Figure 1-1 Illustrated is an example computing system architecture 100A for iteratively computing the autocorrelation for a large data set, where all (v = p > 1) components are directly iteratively computed. Regarding the computing system architecture 100A, only the functions and interrelationships of the main components in this architecture will be introduced here first, and the process of how these components cooperate to jointly complete the iterative autocorrelation calculation will be introduced later in conjunction with Figure 2 the processes described in Figure 1 will be introduced. Figure 1-1 Illustrated is Figure 1 shown 1006 and 1007. Refer to Figure 1-1, the computing system architecture 100A includes a component computing module 131 and an autocorrelation computing module 192. The component computing module 131 can be tightly coupled to one or more storage media through a high-speed data bus or loosely coupled to one or more storage media managed by a storage system through a network, such as a local area network, a wide area network, or even the Internet. Accordingly, the component computing module 131 and any other connected computing devices and their components can send and receive message-related data over the network (e.g., Internet Protocol ("IP") datagrams and other higher-layer protocols that use IP datagrams, such as User Datagram Protocol ("UDP"), Real-Time Streaming Protocol ("RTSP"), Real-Time Transport Protocol ("RTP"), Microsoft Media Server ("MMS"), Transmission Control Protocol ("TCP"), Hypertext Transfer Protocol ("HTTP"), Simple Mail Transfer Protocol ("SMTP"), etc.). The output of the component computing module 131 is used as the input to the autocorrelation computing module 192, which can generate an autocorrelation 193.

[0069] Typically, the storage medium 121 can be a single local storage medium or a complex storage system composed of multiple physically distributed storage devices managed by a storage management system.

[0070] The storage medium 121 contains a data set 123. Typically, the data set 123 can contain data from different types, such as stock prices, audio data, video data, geospatial data, Internet data, mobile communication data, online game data, bank transaction data, sensor data, closed caption data, and real-time text, etc.

[0071] As shown in the figure, the data set 123 includes multiple data elements stored in multiple storage units of the storage medium 121. For example, the data elements 101, 102, 103, 104, 105, 106, 107, 108, 109, and 110 are stored in the storage units 121A, 121B, 121C, 121D, 121E, 121F, 121G, 121H, 121I, and 121J, respectively, and so on... There are also multiple data elements stored in other storage units.

[0072] Referring to the computing system architecture 100A, typically the component computing module 131 includes v component computing modules of v components that directly iterate to compute a set of n data elements for a computing window. v is the number of components that directly iterate in a given algorithm for computing the autocorrelation with a given latency, and it varies depending on the iterative algorithm used. As Figure 1-1 shown, the component computing module 131 includes a component Cd1 computing module 161 and a component Cd vComputing modules 162, and there are v - 2 other component computing modules between them, which can be component Cd2 computing module, component Cd3 computing module, ……, and component Cd v-1 Computing module. Each component computing module computes a specific component for a given latency. Each component computing module includes an initialization module that initializes a component for the first computing window and an algorithm that iteratively computes the component for the adjusted computing window. For example, component Cd1 computing module 161 includes initialization module 132 to initialize component Cd1 for a given latency and iterative algorithm 133 to iteratively compute component Cd1 for a given latency, and component Cd v Computing module 162 includes initialization module 138 to initialize component Cd v and iterative algorithm 139 to iteratively compute component Cd v .

[0073] Initialization module 132 can be used when initializing component Cd1 or when the autocorrelation calculation is reset. Similarly, initialization module 138 can be used when initializing component Cd v or when the autocorrelation calculation is reset.

[0074] Reference Figure 1-1 , computing system architecture 100A further includes an autocorrelation calculation module 192. The autocorrelation calculation module 192 can calculate the autocorrelation for a given latency based on one or more iteratively calculated components for a given latency as needed.

[0075] Figure 1-2 Illustrated is an example computing system architecture 100B that iteratively calculates the autocorrelation for a large data set, with some (v (1 ≤ v < p)) components directly iteratively calculated and some (w = p - v) components indirectly iteratively calculated. In some implementations, the difference between computing system architectures 100B and 100A is that architecture 100B includes component computing module 135. Other than that, the parts with the same reference numbers as in 100A work in the same way. To avoid repeating what was explained in the description of 100A before, only the different parts will be discussed here. The number v in 100B may be different from the number v in 100A because some components that were directly iteratively calculated in 100A are indirectly iteratively calculated in 100B. In 100A, v = p > 1, but in 100B, 1 ≤ v < p. Reference Figure 1-2, the computing system architecture 100B includes a component computing module 135. The output of the component computing module 131 can be used as the input of the component computing module 135, and the outputs of the computing modules 131 and 135 can be used as the input of the autocorrelation computing module 192, which can generate an autocorrelation 193. The component computing module 135 generally includes w = p - v component computing modules to indirectly iteratively compute w components. For example, the component computing module 135 includes a component computing module 163 for indirectly iteratively computing the component Ci1, and a component computing module 164 for indirectly iteratively computing the component Ci w , and the other w - 2 component computing modules therebetween. Indirectly iteratively computing w components includes indirectly iteratively computing each of the w components one by one. Indirectly iteratively computing a component includes accessing and using one or more components other than the component itself. The one or more components can be initialized, directly iteratively computed, indirectly iteratively computed, or computed in any other way.

[0076] Figure 2 The flowchart of an example method 200 for iteratively computing autocorrelation for big data is illustrated. The method 200 will be described in conjunction with the components and data of the computing system architectures 100A and 100B respectively.

[0077] Method 200 includes initializing v (1 ≤ v ≤ p, p > 1) components of an autocorrelation of a computing window with a specified scale of n (n > 1) and a specified delay of l (0 < l < n) for a data set (201). For example, for the computing devices 100A and 100B, method 200 can initialize the v components of the computing window according to the definition of the components by accessing and utilizing the data elements in a computing window of the data set 123 stored on the storage medium 121. The component computing module 131 can access the data elements 101, 102, 103, 104, 105, 106, 107, and 108 in the computing window 122. The initialization module 132 can initialize the component Cd1 141 with a given delay using the data elements 101 to 108. As shown, the component Cd1 141 includes contributions 151, 152, and other contributions 153. Contribution 151 is the contribution of the data element 101 to the component Cd1 141 with a given delay. Contribution 152 is the contribution of the data element 102 to the component Cd1 141 with a given delay. Other contributions 153 are the contributions of the data elements 103 to 108 to the component Cd1 141 with a given delay. Similarly, the initialization module 138 can initialize the component Cd v 145 with 101 to 108. As shown, the component Cd v 145 includes contributions 181, 182, and other contributions 183. Contribution 181 is the contribution of the data element 101 to the component Cd vThe contribution of 145. Contribution 182 is the contribution of data element 102 to component Cd with a given latency v The contribution of 145. Another contribution 183 is the contribution of data elements 103 to 108 to component Cd with a given latency v The contribution of 145.

[0078] Method 200 includes, when v < p, i.e., when not all components are directly iteratively calculated, indirectly iteratively calculating each of the w = p - v components one by one as needed based on one or more components outside the component to be calculated. These w components are only calculated when the autocorrelation is accessed. For example, referring to Figure 1-2 in which some of its components are directly iteratively calculated and some are indirectly iteratively calculated, calculation module 163 can indirectly iteratively calculate component Ci1 based on one or more components outside component Ci1, and calculation module 164 can be based on component Ci w outside one or more components to indirectly iteratively calculate component Ci w These one or more components can be initialized, directly iteratively calculated, indirectly iteratively calculated, or calculated in any other way.

[0079] Method 200 includes calculating the autocorrelation with latency l (210) of components with latency l using one or more initialized or iteratively calculated components as needed, otherwise only those v components are iteratively calculated.

[0080] Method 200 includes accessing a data element to be removed from the calculation window and a data element to be added to the calculation window (202). For example, referring to 100A and 100B, data element 101 and data element 109 can be accessed after data elements 101 - 108 are accessed. Data element 101 is accessed from location 121A of storage medium 121. Data element 109 is accessed from location 121I of storage medium 121.

[0081] Method 200 includes adjusting the calculation window, including: removing the data element to be removed from the calculation window and adding the data element to be added to the calculation window (203). For example, data element 101 is removed from calculation window 122, data element 109 is added to calculation window 122, and then calculation window 122 is transformed into adjusted calculation window 122A.

[0082] Method 200 includes directly iteratively calculating v components of autocorrelation with a delay of l for an adjusted calculation window (204), including: accessing l data elements adjacent to the removed data element and l data elements adjacent to the added data element in the calculation window (205); accessing v components of autocorrelation with a delay of l (206); mathematically removing any contribution of the removed data element from each of the v components (207); and mathematically adding any contribution of the added data element to each of the v components (208). Details are described below.

[0083] Directly iteratively calculating v components of autocorrelation with a specified delay l for an adjusted calculation window includes accessing l data elements adjacent to the removed data element and l data elements adjacent to the added data element in the calculation window (205). For example, if the specified delay l = 1, the iterative algorithm 133 can access the data element 102 which is adjacent to the removed data element 101 and the data element 108 which is adjacent to the added data element 109. If the specified delay l = 2, the iterative algorithm 133 can access the data elements 102 and 103 which are adjacent to the removed data element 101 and the data elements 107 and 108 which are adjacent to the added data element 109... Similarly, if the specified delay l = 1, the iterative algorithm 139 can access the data element 102 which is adjacent to the removed data element 101 and the data element 108 which is adjacent to the added data element 109. If the specified delay l = 2, the iterative algorithm 139 can access the data elements 102 and 103 which are adjacent to the removed data element 101 and the data elements 107 and 108 which are adjacent to the added data element 109...

[0084] Directly iteratively calculating v components of autocorrelation with a delay of l for an adjusted calculation window includes accessing v (1 ≤ v ≤ p) components of autocorrelation with a delay of l of the pre-adjustment calculation window (206). For example, if the specified delay l = 1, the iterative algorithm 133 can access the component Cd1 141 with a delay of 1. If the specified delay l = 2, the iterative algorithm 133 can access the component Cd1 141 with a delay of 2... Similarly, if the specified delay l = 1, the iterative algorithm 139 can access the component Cd v 145. If the specified delay l = 2, the iterative algorithm 139 can access the component Cd v 145...

[0085] The v components for directly iteratively calculating the autocorrelation of a specified delay l for the adjusted calculation window include mathematically removing any contribution of the removed data element from each of the v components (207). For example, if the specified delay l = 2, the component Cd1 143 for directly iteratively calculating the component with a delay of 2 may include the contribution removal module 133A mathematically removing the contribution 151 from the component Cd1 141 with a delay of 2. Similarly, the component Cd v 147 for directly iteratively calculating the component with a delay of 2 may include the contribution removal module 139A mathematically removing the contribution 181 from the component Cd v 145 with a delay of 2. The contributions 151 and 181 come from the data element 101.

[0086] The v components for directly iteratively calculating the autocorrelation of a delay of l for the adjusted calculation window include mathematically adding any contribution of the added data element to each of the v components (208). For example, if the specified delay l = 2, the component Cd1 143 for directly iteratively calculating the component with a delay of 2 may include the contribution addition module 133B mathematically adding the contribution 154 to the component Cd1 141 with a delay of 2. Similarly, the component Cd v 147 for directly iteratively calculating the component with a delay of 2 may include the contribution addition module 139B mathematically adding the contribution 184 to the component Cd v 145 with a delay of 2. The contributions 154 and 184 come from the data element 109.

[0087] As Figure 1-1 shown in FIGS. 1-2, the component Cd1 143 includes the contribution 152 (the contribution from the data element 102), the other contribution 153 (the contribution from the data elements 103-108), and the contribution 154 (the contribution from the data element 109). Similarly, the component Cd v 147 includes the contribution 182 (the contribution from the data element 102), the other contribution 183 (the contribution from the data elements 103-108), and the contribution 184 (the contribution from the data element 109).

[0088] When the autocorrelation is accessed and v < p (i.e., not all components are directly iteratively calculated), the method 200 includes indirectly iteratively calculating w = p - v components with a delay of l one by one as needed using one or more other components in addition to the component itself (209). These w components are only calculated when the autocorrelation is accessed. For example, referring to Figure 1-2 which has some components directly iteratively calculated and some components indirectly iteratively calculated, the calculation module 163 can indirectly iteratively calculate the component Ci1 based on one or more components other than the component Ci1, and the calculation module 164 can be based on the component Ci wone or more components other than to indirectly iteratively compute component Ci w These one or more components may have been initialized, directly iteratively computed, indirectly iteratively computed, or computed in any other way.

[0089] Method 200 includes computing autocorrelation on an as-needed basis. When the autocorrelation is accessed, the autocorrelation is computed based on one or more iteratively computed components; otherwise only v components are directly iteratively computed. When the autocorrelation is accessed, Method 200 includes indirectly iteratively computing w components (209) with a latency of l as needed. For example, in architecture 100A, autocorrelation module 192 may compute autocorrelation 193 for a given latency. In architecture 100B, computing module 163 may indirectly iteratively compute Ci1 based on one or more components other than component Ci1, and computing module 164 may indirectly iteratively compute Ci based on one or more components other than component Ci w one or more components other than to indirectly iteratively compute Ci w , ……, autocorrelation computing module 192 may compute autocorrelation 193 for a given latency (210). Once the autocorrelation for a given latency is computed, Method 200 includes accessing the next data element to be removed and the next data element to be added to begin the next iteration of the computation. Each time a new iteration of the computation is started, the adjusted computation window from the previous iteration becomes the pre-adjusted computation window for the new iteration.

[0090] As more data elements are accessed, 202-208 may be repeated, and 209-210 may be repeated as needed. For example, after data elements 101 and 109 are accessed or received and the components within the range of components Cd1 143 to component Cd v 147 are computed, data elements 102 and 110 may be accessed (202). Once a data element to be removed and a data element to be added are accessed or received, Method 200 includes removing the data element to be removed from the computation window and adding the data element to be added to the computation window to adjust the computation window (203). For example, computation window 122A may be transformed into computation window 122B after removing data element 102 and adding data element 110.

[0091] Method 200 includes directly iteratively calculating, for an adjusted calculation window, v components of an autocorrelation with a latency of l based on v components of a pre-adjustment calculation window (204), which includes accessing or receiving l data elements adjacent to a removed data element and l data elements adjacent to an added data element in the calculation window (205), accessing the v components (206), mathematically removing, from each of the v components, any contribution of the removed data element (207), and mathematically adding, to each of the v components, any contribution of the added data element (208). For example, referring to 100A and 100B, at a specified latency such as l = 1, iterative algorithm 133 can be used to directly iteratively calculate, for calculation window 122B, a component Cd1 144 with a latency of 1 based on a component Cd1 143 with a latency of 1 calculated for calculation window 122A (204). Iterative algorithm 133 can access data element 103 which is adjacent to removed data element 102 and data element 109 which is adjacent to added data element 110 (205). Iterative algorithm 133 can access component Cd1 143 with a latency of 1 (206). Directly iteratively calculating component Cd1 144 with a latency of 1 includes contribution removal module 133A mathematically removing contribution 152, i.e., the contribution of data element 102, from component Cd1 143 with a latency of 1 (207). Directly iteratively calculating component Cd1 144 with a latency of 1 includes contribution addition module 133B mathematically adding contribution 155, i.e., the contribution of data element 110, to component Cd1 143 with a latency of 1 (208). Similarly, at a specified latency such as l = 1, iterative algorithm 139 can be used to directly iteratively calculate, for calculation window 122B, a component Cd v 148 based on a component Cd with a latency of 1 calculated for calculation window 122A v 147. Iterative algorithm 139 can access data element 103 which is adjacent to removed data element 102 and data element 109 which is adjacent to added data element 110. Iterative algorithm 139 can access component Cd v 147. Directly iteratively calculating component Cd v 148 includes contribution removal module 139A mathematically removing contribution 182, i.e., the contribution of data element 102, from component Cd v 147. Directly iteratively calculating component Cd v 148 includes contribution addition module 139B mathematically adding contribution 185, i.e., the contribution of data element 110, to component Cd v 147.

[0092] As shown in the figure, the component Cd1 144 with a delay of l includes other contributions 153 (contributions from data elements 103 - 108), contribution 154 (contribution from data element 109), and contribution 155 (contribution from data element 110), and the component Cd with a delay of l v 148 includes other contributions 183 (contributions from data elements 103 - 108), contribution 184 (contribution from data element 109), and contribution 185 (contribution from data element 110).

[0093] Method 200 includes indirectly iteratively calculating the w components and autocorrelation with a given delay as needed.

[0094] Method 200 includes, when needed, i.e., only when the autocorrelation is accessed, indirectly iteratively calculating the w components and autocorrelation with a given delay. If the autocorrelation is not accessed, method 200 includes continuing to access or receive the next data element to be removed and the next data element to be added for the next calculation window (202). If the autocorrelation is accessed, method 200 includes indirectly iteratively calculating the w components with a given delay (209) and calculating the autocorrelation with a given delay based on one or more iteratively calculated components with a given delay (210).

[0095] When the next data element to be removed and the data element to be added are accessed, the component Cd1 144 can be used to directly iteratively calculate the next component Cd1, and the component Cd v 148 can be used to directly iteratively calculate the next component Cd v .

[0096] Figure 3-1 Illustrated are the data elements removed from the calculation window 300A and the data elements added to the calculation window 300A when iteratively calculating the autocorrelation on big data. The calculation window 300A moves to the right. Refer to Figure 3-1 , an existing data element is always removed from the left side of the calculation window 300A, and a data element is always added to the right side of the calculation window 300A.

[0097] Figure 3-2Illustrated is the data accessed in computation window 300A when iteratively computing autocorrelation on big data. For computation window 300A, the first n data elements are accessed to initialize two or more components for a given latency for the first computation window, and then, as needed, w = p - v components and autocorrelation are iteratively computed indirectly. As time progresses, an oldest data element such as the (m + 1)-th data element is removed from computation window 300A, and a data element such as the (m + n + 1)-th data element is added to computation window 300A. One or more components for the given latency of the adjusted computation window are then iteratively computed directly based on the two or more components computed for the first computation window. If the specified latency is 1, a total of 4 data elements are accessed, including the removed data element, a data element adjacent to the removed data element, the added data element, and a data element adjacent to the added data element. If the specified latency is 2, a total of 6 data elements are accessed, including the removed data element, 2 data elements adjacent to the removed data element, the added data element, and 2 data elements adjacent to the added data element. If the specified latency is l, a total of 2*(l + 1) data elements are accessed, including the removed data element, l data elements adjacent to the removed data element, the added data element, and l data elements adjacent to the added data element. Then, as needed, w = p - v components and autocorrelation for the given latency are iteratively computed indirectly. Then computation window 300A is adjusted again by removing an old data element and adding a new data element,.... For a given iterative algorithm, v is a constant, and the number of operands for iteratively computing w = p - v components indirectly is also a constant, so for a given latency, the amount of data access and computation is reduced and is constant. The larger the computation window size n, the more significant the reduction in the amount of data access and computation.

[0098] Figure 3-3 Illustrated are the data elements removed from computation window 300B and the data elements added to computation window 300B when iteratively computing autocorrelation on big data. Computation window 300B moves to the left. Refer to Figure 3-3 , a new data element is always removed from the right side of computation window 300B, and an old data element is always added to the left side of computation window 300B.

[0099] Figure 3-4Illustrated is the data accessed in computation window 300B when iteratively computing autocorrelation on big data. For computation window 300B, the first n data elements are accessed to initialize two or more components at a given latency for the first computation window, and then w = p - v components and autocorrelation are iteratively computed indirectly as needed. As time progresses, a data element such as the (m + n)-th data element is removed from computation window 300B, and a data element such as the m-th data element is added to computation window 300B. One or more components at the given latency of the adjusted computation window are then iteratively computed directly based on the two or more components computed for the first computation window. If the specified latency is 1, a total of 4 data elements are accessed, including the removed data element, a data element adjacent to the removed data element, the added data element, and a data element adjacent to the added data element. If the specified latency is 2, a total of 6 data elements are accessed, including the removed data element, 2 data elements adjacent to the removed data element, the added data element, and 2 data elements adjacent to the added data element. If the specified latency is l, a total of 2*(l + 1) data elements are accessed, including the removed data element, l data elements adjacent to the removed data element, the added data element, and l data elements adjacent to the added data element. Then w = p - v components at the given latency and autocorrelation are iteratively computed indirectly as needed. Then computation window 300B is adjusted again by removing a new data element and adding an old data element, ……. For a given iterative algorithm, v is a constant, and the number of operands for iteratively computing w = p - v components is also a constant, so for a given latency, the data access amount and computation amount are reduced and are constants. The larger the computation window size n, the more significant the reduction in data access amount and computation amount.

[0100] Figure 4-1 Illustrated is the definition of autocorrelation. Assume X = (x m+1 , x m+2 , ……, x m+n ) is a computation window of size n containing data involved in autocorrelation computation of a big data set. This computation window can move in two directions, right or left. For example, when computing the autocorrelation of the latest data, this computation window moves to the right. At this time, a data is removed from the left side of this computation window, and a data is added to the right side of this computation window. When reviewing the autocorrelation of old data, this computation window moves to the left. At this time, a data is removed from the right side of this computation window, and a data is added to the left side of this computation window. The equations for iteratively computing components in these two cases are different. To distinguish them, define the adjusted computation window in the former case as X I , and the adjusted computation window in the latter case as X IIEquation 401 is a traditional equation for calculating the sum S of all data elements in the calculation window X of size n for the k-th round. k Equation 402 is a traditional equation for calculating the average value of all data elements in the calculation window X for the k-th round. Equation 403 is a traditional equation for calculating the autocorrelation ρ with a given delay l in the calculation window X for the k-th round. (k,l) Equation 404 is a traditional equation for calculating the sum S of all data elements in the adjusted calculation window X of size n for the (k + 1)-th round. I I k+1 Equation 405 is a traditional equation for calculating the average value of all data elements in the adjusted calculation window X for the (k + 1)-th round. I Equation 406 is a traditional equation for calculating the autocorrelation ρ with a given delay l in the adjusted calculation window X for the (k + 1)-th round. I I (k+1,l) As mentioned before, when the calculation window moves to the left, the adjusted calculation window is defined as X II Equation 407 is a traditional equation for calculating the sum S of all data elements in the adjusted calculation window X of size n for the (k + 1)-th round. II II k+1 Equation 408 is a traditional equation for calculating the average value of all data elements in the adjusted calculation window X for the (k + 1)-th round. II Equation 409 is a traditional equation for calculating the autocorrelation ρ with a given delay l in the adjusted calculation window X for the (k + 1)-th round. II II (k+1,l)

[0101] To show how to use component iteration to calculate autocorrelation, three different iterative autocorrelation algorithms are provided as examples. A new round of calculation starts whenever there is a data change in the calculation window (e.g., 122 → 122A → 122B). A sum or average value is a basic component for calculating autocorrelation. The equations for iteratively calculating a sum or average value are iterative component equations used by all example iterative autocorrelation calculation algorithms.

[0102] Figure 4-2 Illustrate the first example iterative autocorrelation calculation algorithm (Iterative Algorithm 1). Equation 401 and 402 can be used to initialize the components S k and / or Equations 410, 411, and 412 can be used to initialize the components SS k , SX k , and covX (k,l) ​​​​​​​。Equation 413 can be used to calculate the autocorrelation ρ (k,l) 。When the calculation window moves to the right, iterative algorithm 1 includes component S I k+1 or SS I k+1 , SX I k+1 , and covX I (k+1,l) for iterative calculation. Once components SX I k+1 and covX I (k+1,l) are calculated, the autocorrelation ρ I (k+1,l) can be calculated based on them. Once component S k and / or is available, equations 414 and 415 can be used to iteratively calculate the components of the adjusted calculation window X I respectively I k+1 and Once component SS k is available, equation 416 can be used to directly iteratively calculate the component SS of the adjusted calculation window X I I k+1 。Once component S I k+1 or and SS I k+1 are available, equation 417 can be used to indirectly iteratively calculate the component SX of the adjusted calculation window X I I k+1 。Once component covX (k,l) , SS I k+1 , S k or and S I k+1 or is available, equation 418 can be used to directly iteratively calculate the component covX of the adjusted calculation window X I I (k+1,l) 。414, 415, 417, and 418 each contain multiple equations but only one of them is required respectively depending on whether and or average or both are available. Once components covX I (k+1,l) and SX I k+1 are calculated, equation 419 can be used to indirectly iteratively calculate the adjusted calculation window X I ​​​The autocorrelation ρ with a given delay of l I (k+1,l) . When the calculation window is shifted to the left, the iterative algorithm 1 includes components S II k+1 or SS II k+1 , SX II k+1 , and cov XII (k+1,l) . Once the components SX II k+1 and covX II (k+1,l) are calculated, the autocorrelation ρ II (k+1,l) can be calculated based on them. Equations 420 and 421 can be used to iteratively calculate the components S II of the adjusted calculation window X II k+1 and Once the component S k and / or are available. Equation 422 can be used to directly iteratively calculate the component SS II of the adjusted calculation window X II k+1 Once the component SS k is available. 423 can be used to indirectly iteratively calculate the component SX II of the adjusted calculation window X II k+1 Once the component S II k+1 or and SS II k+1 are available. Equation 424 can be used to directly iteratively calculate the component covX II of the adjusted calculation window X II (k+1,l) Once the component covX (k,l) , SS II k+1 , S k or and S II k+1 or are available. 420, 421, 423, and 424 each contain multiple equations but only one of them is required respectively depending on whether and or average or both are available. Equation 425 can be used to indirectly iteratively calculate the autocorrelation ρ II with a given delay of l of the adjusted calculation window X II (k+1,l) Once the component ccovX II (k+1,l)and SX II k+1 are calculated.

[0103] Figure 4-3 Illustrates the second example of an iterative autocorrelation calculation algorithm (Iterative Algorithm 2). Equations 401 and 402 can be used to initialize components S k and / or Equations 426 and 427 can be used to initialize components SX k and covX (k,l) . Equation 428 can be used to calculate the autocorrelation ρ (k,l) . As the calculation window moves to the right, Iterative Algorithm 2 includes the iterative calculation of components S I k+1 or SX I k+1 , and covX I (k+1,l) . Once components SX I k+1 and covX I (k+1,l) are calculated, the autocorrelation ρ I (k+1,l) can be calculated based on them. Once components S k and / or are available, Equations 429 and 430 can be used to iteratively calculate the components of the adjusted calculation window X I of S I k+1 and . Once components SX k , S I k+1 and / or are available, Equation 431 can be used to directly iteratively calculate the components of the adjusted calculation window X I of SX I k+1 . Equation 432 can be used to directly iteratively calculate the components of the adjusted calculation window X I of covX I (k+1,l) . Once components covX (k,l) , S k or and S I k+1 or are available. 429, 430, 431, and 432 each contain multiple equations but only one of them is required respectively depending on whether and or average or both are available. Once components covX I (k+1,l) and SX I k+1being calculated, Equation 433 can be used for indirect iterative calculation of the adjusted calculation window X I the autocorrelation ρ with a given delay of l of I (k+1,l) . When the calculation window is shifted to the left, the iterative algorithm 2 includes the component S II k+1 or SX II k+1 , and covX II (k+1,l) for iterative calculation. Once the components SX II k+1 and covX II (k+1,l) are calculated, the autocorrelation ρ II (k+1,l) can be calculated based on them. Equations 434 and 435 can be used to iteratively calculate the component S of the adjusted calculation window X II once the component S II k+1 and are available. Equation 436 can be used for direct iterative calculation of the component SX of the adjusted calculation window X k and / or once the component SX II is available. Equation 437 can be used for direct iterative calculation of the component covX of the adjusted calculation window X II k+1 once the component covX II k , S II k+1 and / or are available. Equation 437 can be used for direct iterative calculation of the component covX of the adjusted calculation window X II once the component covX II (k+1,l) is available. Equation 438 can be used for indirect iterative calculation of the autocorrelation ρ with a given delay of l of the adjusted calculation window X (k,l) , S k or and S II k+1 or are available. 434, 435, 436, and 437 each contain multiple equations but only one of them is needed respectively depending on whether and / or the average value or both are available. Equation 438 can be used for indirect iterative calculation of the autocorrelation ρ with a given delay of l of the adjusted calculation window X II once the components covX II (k+1,l) and SX II (k+1,l) are calculated. II k+1

[0104] Figure 4-4 ​Describe the third example of the iterative autocorrelation calculation algorithm (Iterative Algorithm 3). Equations 401 and 402 can be used to initialize components S k and / or Equations 439 and 440 can be used to initialize components SX k and covX (k,l) . Equation 441 can be used to calculate the autocorrelation ρ (k,l) . When the calculation window moves to the right, Iterative Algorithm 3 includes the iterative calculation of components S I k+1 or SX I k+1 , and covX I (k+1,l) . Once components SX I k+1 and covX I (k+1,l) are calculated, the autocorrelation ρ I (k+1,l) can be calculated based on them. Equations 442 and 443 can be used to iteratively calculate the adjusted calculation window X I for components S I k+1 and Once components S k and / or are available. Equation 444 can be used to directly iteratively calculate the adjusted calculation window X I for components SX I k+1 Once components SX k , S k and / or and S I k+1 and / or are available. Equation 445 can be used to directly iteratively calculate the adjusted calculation window X I for components covX I (k+1,l) Once components covX (k,l) , S k or and S I k+1 or are available. 442, 443, 444, and 445 each contain multiple equations but only one of them is required respectively depending on whether and / or average or both are available. Equation 446 can be used to indirectly iteratively calculate the autocorrelation ρ I for the adjusted calculation window X I (k+1,l) with a given delay of l once components covX I (k+1,l) and SXI k+1 are calculated. When the calculation window is shifted to the left, the iterative algorithm 3 includes component S II k+1 or SX II k+1 , and covX II (k+1,l) of the iterative calculation. Once components SX II k+1 and covX II (k+1,l) are calculated, the autocorrelation ρ II (k+1,l) can be calculated based on them. Equations 447 and 448 can be used to iteratively calculate component S of the adjusted calculation window X II respectively II k+1 and Once component S k and / or is available. Equation 449 can be used to directly iteratively calculate component SX of the adjusted calculation window X II respectively II k+1 Once component SX k , S k and / or and S II k+1 and / or is available. Equation 450 can be used to directly iteratively calculate component covX of the adjusted calculation window X II respectively II (k+1,l) Once component covX (k,l) , S k or and S II k+1 or is available. 447, 448, 449, and 450 each contain multiple equations but only one of them is required respectively depending on whether and or average or both are available. Once components covX II (k+1,l) and SX II k+1 are calculated, Equation 451 can be used to indirectly iteratively calculate the autocorrelation ρ of the adjusted calculation window X II with a given delay of l II (k+1,l) .

[0105] To demonstrate iterative autocorrelation algorithms and their comparison with traditional algorithms, three examples are given below. Three moving computing windows of data are used. For the traditional algorithm, the calculation processes for all three computing windows are exactly the same. For the iterative algorithm, the first computing window initializes two or more components, and the second and third computing windows perform iterative calculations.

[0106] Figure 5-1 , Figure 5-2 , Figure 5-3 respectively show the first computing window, the second computing window, and the third computing window for one computing instance. Computing window 503 includes the first 4 data elements of the large dataset 501: 8, 3, 6, 1. Computing window 504 includes 4 data elements of the large dataset 501: 3, 6, 1, 9. Computing window 505 includes 4 data elements of the large dataset 501: 6, 1, 9, 2. This computing instance assumes that the computing window moves from left to right. The large dataset 501 can be big data or streaming data. The computing window size 502(n) is 4.

[0107] First, calculate the autocorrelation with a delay of 1 for computing windows 503, 504, and 505 respectively using the traditional algorithm.

[0108] To calculate the autocorrelation with a delay of 1 for computing window 503:

[0109]

[0110] Without any optimization, calculating the autocorrelation with a delay of 1 for a computing window of size 4 involves 2 divisions, 7 multiplications, 8 additions, and 10 subtractions.

[0111] The same equations and processes can be used to calculate the autocorrelation with a delay of 1 for Figure 5-2 the shown computing window 504 and for Figure 5-3 the shown computing window 505. The autocorrelation with a delay of 1 for computing window 504 The autocorrelation with a delay of 1 for computing window 505 Each of these two calculations involves 2 divisions, 7 multiplications, 8 additions, and 10 subtractions without optimization. The traditional algorithm generally needs to complete 2 divisions, 2n - l multiplications, 3n - (l + 3) additions, and 3n - 2l subtractions when calculating the autocorrelation with a given delay of l for a computing window of size n without optimization.

[0112] Next, calculate the autocorrelation with a delay of 1 for computing windows 503, 504, and 505 respectively using iterative algorithm 1.

[0113] Calculate the autocorrelation with a delay of 1 for window 503:

[0114] 1. Initialize the components of the first round with equations 402, 410, 411, and 412 respectively SS1, SX1, and covX (1,1) :

[0115]

[0116] 2. Calculate the autocorrelation ρ of the first round with equation 413 (1,1) :

[0117]

[0118] There are 2 divisions, 9 multiplications, 8 additions, and 7 subtractions in total when calculating the autocorrelation with a delay of 1 for window 503.

[0119] Calculate the autocorrelation with a delay of 1 for window 504:

[0120] 1. Iteratively calculate the components of the second round with equations 415, 416, 417, and 418 respectively SS2, SS2, and conX (2,1) :

[0121]

[0122] SS2 = SS1 + x m+1+4 2 -x m+1 2 = 110 + 9 2 -8 2 = 110 + 81 - 64 = 127

[0123]

[0124] 2. Calculate the autocorrelation ρ of the second round with equation 419 (2,1) :

[0125]

[0126] There are 2 divisions, 10 multiplications, 8 additions, and 7 subtractions in total when iteratively calculating the autocorrelation with a delay of 1 for window 504.

[0127] Calculate the autocorrelation with a delay of 1 for window 505:

[0128] 1. Iteratively calculate the components of the third round with equations 415, 416, 417, and 418 respectively SS3, SX3, and covX (3,1) :

[0129]

[0130] SS3 = SS2 + x m+1+4 2 -x m+1 2 = 127 + 2 2 -3 2 = 127 + 4 - 9 = 122

[0131]

[0132] 2. Calculate the autocorrelation ρ of the third round using Equation 419 (3,1) :

[0133]

[0134] When calculating the autocorrelation with a delay of 1 for the calculation window 505, there are a total of 2 divisions, 10 multiplications, 8 additions, and 7 subtractions.

[0135] Next, use the iterative algorithm 2 to calculate the autocorrelation with a delay of 1 for the calculation windows 503, 504, and 505 respectively.

[0136] To calculate the autocorrelation with a delay of 1 for the calculation window 503:

[0137] 1. Initialize the components of the first round using Equations 402, 426, and 427 SX1, and covX (1,1) :

[0138]

[0139] 2. Calculate the autocorrelation ρ of the first round using Equation 428 (1,1) :

[0140]

[0141] When calculating the autocorrelation with a delay of 1 for the calculation window 503, there are a total of 2 divisions, 7 multiplications, 8 additions, and 10 subtractions.

[0142] To calculate the autocorrelation with a delay of 1 for the calculation window 504:

[0143] 1. Iteratively calculate the components of the second round using Equations 430, 431, and 432 respectively SX2, and covX (2,1) :

[0144]

[0145] 2. Calculate the autocorrelation ρ of the second round using Equation 433 (2,1) :

[0146]

[0147] There are 2 divisions, 7 multiplications, 10 additions, and 7 subtractions in total when calculating the autocorrelation with a delay of 1 for the calculation window 504.

[0148] To calculate the autocorrelation with a delay of 1 for the calculation window 505:

[0149] 1. Iteratively calculate the components of the third round using equations 430, 431, and 432 SX3, and covX (3,1) :

[0150]

[0151] 2. Calculate the autocorrelation ρ of the third round using equation 433 (3,1) :

[0152]

[0153] There are 2 divisions, 7 multiplications, 10 additions, and 7 subtractions in total when calculating the autocorrelation with a delay of 1 for the calculation window 505.

[0154] Next, use the iterative algorithm 3 to calculate the autocorrelation with a delay of 1 for the calculation windows 503, 504, and 505 respectively.

[0155] To calculate the autocorrelation with a delay of 1 for the calculation window 503:

[0156] 1. Initialize the components of the first round using equations 402, 439, and 440 SX1, and covX (1,1) :

[0157]

[0158] 2. Calculate the autocorrelation ρ of the first round using equation 441 (1,1) :

[0159]

[0160] There are 2 divisions, 7 multiplications, 8 additions, and 10 subtractions in total when calculating the autocorrelation with a delay of 1 for the calculation window 503.

[0161] To calculate the autocorrelation with a delay of 1 for the calculation window 504:

[0162] 1. Iteratively calculate the components of the second round using equations 443, 444, and 445 SX2, and covX (2,1) :

[0163]

[0164] 2. Calculate the autocorrelation ρ of the second round using Equation 446 (2,1) :

[0165]

[0166] When calculating the autocorrelation with a delay of 1 for the window 504 iteratively, there are a total of 2 divisions, 7 multiplications, 9 additions, and 8 subtractions.

[0167] For calculating the autocorrelation with a delay of 1 for the window 505:

[0168] 1. Iteratively calculate the components of the third round using Equations 443, 444, and 445 respectively SX3, and covX (3,1) :

[0169]

[0170]

[0171] 2. Calculate the autocorrelation ρ of the third round using Equation 446 (3,1) :

[0172]

[0173] When calculating the autocorrelation with a delay of 1 for the window 505 iteratively, there are a total of 2 divisions, 7 multiplications, 9 additions, and 8 subtractions.

[0174] In the above three examples, the average value is used for iterative autocorrelation calculation. And can also be used for iterative autocorrelation calculation, just with different operands. In addition, the calculation window in the above three examples moves from left to right. When the calculation window moves from right to left, the calculation process is similar but a different set of equations is applied.

[0175] Figure 6-1 Illustrates the comparison of the computational complexity between the traditional autocorrelation algorithm and the iterative autocorrelation algorithm when n = 4 and the delay is 1. As shown in the figure, the division operations, multiplication operations, addition operations, and subtraction operations of any iterative algorithm and the traditional algorithm are similar.

[0176] Figure 6-2It illustrates the comparison of the computational complexity between the traditional autocorrelation algorithm and the iterative autocorrelation algorithm when n = 1,000,000 and the delay is 1. As shown in the figure, any iterative algorithm requires far fewer multiplication operations, addition operations, and subtraction operations than the traditional algorithm. The iterative autocorrelation algorithm can complete the data processing that originally needed to be processed on thousands of computers on a single machine. This significantly improves the computational efficiency, reduces computing resources, and lowers the energy consumption of computing devices, making it possible to efficiently and with low energy consumption determine in real time the repeatability of big data with a given delay, and enabling some scenarios of real-time determination of the repeatability of big data with a given delay that were previously impossible to become possible.

[0177] The present invention can be implemented in other specific ways without departing from its spirit or essential characteristics. The implementation described in this application is exemplary rather than restrictive in all aspects. Therefore, the scope of the present invention is defined by the appended claims rather than the preceding description. All changes equivalent to the meaning and scope of the claims in the claims are included within their scope.

Claims

1. A method for real-time determination of the given delay repeatability of big data, characterized in that: A computing system composed of one or more computing devices initializes a delay l, where 0 < l < n, and two or more components of autocorrelation with a delay of l for an adjusted calculation window with a specified size of n for a data set, where n > 1; The computing system based on the computing device accesses a data element to be removed from the adjusted calculation window and a data element to be added to the adjusted calculation window; The computing system based on the computing device adjusts the adjusted calculation window by: Removing the data element to be removed from the adjusted calculation window; And Adding the data element to be added to the adjusted calculation window; The computing system based on the computing device iteratively calculates a sum, an average value, or a sum and an average value for the adjusted calculation window; The computing system based on the computing device iteratively calculates two or more components of autocorrelation with a delay of l for the adjusted calculation window based on at least two or more components of autocorrelation with a delay of l in the adjusted calculation window, and avoids accessing and using all data elements in the adjusted calculation window during the iterative calculation of the two or more components to reduce data access latency, improve computing efficiency, save computing resources, and reduce the energy consumption of the computing system; And The computing system based on the computing device generates an autocorrelation with a delay of l for the adjusted calculation window based on one or more components iteratively calculated for the adjusted calculation window.

2. The method according to claim 1, characterized in that: The accessing of a data element to be removed and a data element to be added includes accessing multiple data elements to be removed from the adjusted calculation window and multiple data elements to be added to the adjusted calculation window. The method further includes adjusting the calculation window for each data element among the multiple data elements to be removed and each data element among the multiple data elements to be added, iteratively calculating two or more components, and generating an autocorrelation with a delay of l for the adjusted calculation window.

3. The method according to claim 2, characterized in that: The generating of an autocorrelation with a delay of l for the adjusted calculation window is performed if and only if the autocorrelation is accessed.

4. The method according to claim 3, characterized in that: The generating of an autocorrelation with a delay of l for the adjusted calculation window further includes the computing system based on the computing device indirectly iteratively calculating one or more components of autocorrelation with a delay of l for the adjusted calculation window. The indirect iterative calculation of the one or more components includes calculating each of the one or more components separately based on one or more components other than the component to be calculated.

5. A computing system, characterized in that: One or more processors; One or more storage media, where at least one storage media stores a data set; and One or more computing modules, which, when executed by at least one of the one or more processors, execute a method for real-time determination of the given delay repeatability of big data, the method including: a. Initializing a delay l, where 0 < l < n, and two or more components of autocorrelation with a delay of l for an adjusted calculation window with a specified size of n for the data set, where n > 1; b. Access a data element to be removed from the pre-adjustment calculation window and a data element to be added to the pre-adjustment calculation window; c. Adjust the pre-adjustment calculation window, including: Removing the data element to be removed from the pre-adjustment calculation window; and Adding the data element to be added to the pre-adjustment calculation window; d. Iteratively calculate more than two components of the autocorrelation with a delay of l for the post-adjustment calculation window based on at least more than two components of the autocorrelation with a delay of l of the pre-adjustment calculation window, and avoid accessing and using all data elements in the post-adjustment calculation window during the iterative calculation of the more than two components to reduce data access latency, improve calculation efficiency, save calculation resources, and reduce the energy consumption of the calculation system; and e. Generate the autocorrelation with a delay of l for the post-adjustment calculation window based on one or more components iteratively calculated for the post-adjustment calculation window.

6. The computing system according to claim 5, wherein: The one or more calculation modules, when executed by at least one of the one or more processors, execute b, c, d, and e multiple times.

7. The computing system according to claim 6, characterized in that: Execute e if and only if the autocorrelation with a delay of l of the post-adjustment calculation window is accessed.

8. The computing system according to claim 7, wherein: The e further includes indirectly iteratively calculating one or more components of the autocorrelation with a delay of l for the post-adjustment calculation window by the calculation system, and indirectly iteratively calculating the one or more components includes calculating the one or more components separately one by one based on one or more components other than the components to be calculated.

9. A calculation system program product, running on a calculation system including one or more calculation devices, the calculation system including one or more processors and one or more storage media storing a data set, the calculation system program product including multiple calculation device-executable instructions, when these calculation device-executable instructions are run by the calculation system, execute a method for real-time judging the given-delay repeatability of big data itself, characterized in that: Initialize a delay l, 0 < l < n, and more than two components of the autocorrelation with a delay of l for a pre-adjustment calculation window with a specified size of n for a data set, n > 1; Access a data element to be removed from the pre-adjustment calculation window and a data element to be added to the pre-adjustment calculation window; Adjust the pre-adjustment calculation window by: Removing the data element to be removed from the pre-adjustment calculation window; And Adding the data element to be added to the pre-adjustment calculation window; Iteratively calculate more than two components of the autocorrelation with a delay of l for the post-adjustment calculation window based on at least more than two components of the autocorrelation with a delay of l of the pre-adjustment calculation window, and avoid accessing and using all data elements in the post-adjustment calculation window during the iterative calculation of the more than two components to reduce data access latency, improve calculation efficiency, save calculation resources, and reduce the energy consumption of the calculation system; And Generate the autocorrelation with a delay of l for the post-adjustment calculation window based on one or more components iteratively calculated for the post-adjustment calculation window.

10. A computer system program product comprising a plurality of computer-executable instructions that, when executed by at least one computer device in a computer system comprising one or more computer devices, cause the computer system to implement the method according to any one of claims 2-4.

Citation Information

Patent Citations

  • Code rate control method based on variable bit rate in low latency video coding

    CN103686172A

  • Iterative simple linear regression coefficient calculation for streamed data using components

    US9928215B1