Data recording and analysis system

The data recording and analysis system efficiently clusters and identifies target data segments in real-time, addressing the challenge of searching large data sets by using similarity protocols and maintaining reference segments, thereby reducing analysis time.

JP7857343B2Active Publication Date: 2026-05-12KEYSIGHT TECHNOLOGIES INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KEYSIGHT TECHNOLOGIES INC
Filing Date
2024-06-07
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The challenge of efficiently searching for target data patterns in large data sets exceeding terabytes, which requires several hours to read, necessitates a method to quickly identify and cluster similar data segments in real-time without prior knowledge of the signals.

Method used

A data recording and analysis system that includes an input port, buffer, and controller to extract and cluster data segments using similarity protocols, allowing real-time identification and classification of data segments without requiring detailed knowledge of the signals, and maintaining reference data segments in memory for rapid retrieval.

Benefits of technology

Enables rapid clustering and identification of target data segments, reducing the time required to analyze large data sets by identifying and grouping similar segments in real-time, facilitating efficient data search and retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007857343000003
    Figure 0007857343000003
  • Figure 0007857343000004
    Figure 0007857343000004
  • Figure 0007857343000001
    Figure 0007857343000001
Patent Text Reader

Abstract

To provide a system for recording and analyzing a data stream, a method for analyzing a data stream, and a computer readable memory that stores instructions that cause a computer to practice a method of analyzing a data stream.SOLUTION: The system comprises: an input port for receiving a data stream; an output port for outputting a compressed data stream to a disc; a FIFO buffer; and a controller. The controller identifies a segment, referred to as a new extracted data segment (EDS) of the data stream stored in the FIFO buffer, The new EDS satisfies an extraction protocol. The controller also compares the new EDS to each of reference data segments (RDSs) using a similarity protocol; if the new EDS is not similar to an existing EDS, a new RDS is created; and if the new EDS is similar to the existing EDS, the RDS stores information for identifying the new EDS in an RDS database.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Data recording systems can now record large amounts of data, and the amount is large Therefore, the time required to search for recorded data by sequentially reading stored data becomes extremely long. Data sets exceeding terabytes are routinely recorded. From a conventional disk drive The time required to read terabytes of data is several hours. Therefore, quickly searching for recorded data to find a target pattern presents a challenge.

Summary of the Invention

[0002] The present invention includes a system for recording and analyzing a data stream, a method for analyzing a data stream, and a computer-readable memory storing instructions for causing a computer to execute a method for analyzing a data stream. The system includes an input port, an output port, a buffer, and a controller. The input port is adapted to receive a data stream, and the data stream includes an ordered sequence of data values. The output port is adapted to communicate the data stream to a mass storage device. The buffer is connected to the input port to temporarily store a predetermined portion of the data stream when the data stream is received by the system. The controller identifies a new extracted data segment (EDS) of the data stream stored in the buffer, and the new EDS satisfies an extraction protocol. The controller uses a first similarity protocol to compare the new EDS with a plurality of reference data segments (RD ​​​​​​​​The controller compares each of the S (reference data segments) to the first similarity protocol. If Col indicates that the new EDS is similar to one of the RDS, then the new ED The controller stores information identifying S in the RDS database. The controller then checks if a new EDS is R If it is not similar to any of the existing RDSs, a new RDS is created. Each RDS is EDS was found to be similar to RDS, and the controller generated a new RDS. This includes a list of new EDSs.

[0003] In one embodiment, the buffer includes a FIFO buffer.

[0004] In another aspect, the extraction protocol uses data values ​​in the buffer in which a new EDS is initiated. Then, the new EDS identifies the data value in the buffer when it finishes.

[0005] In another aspect, the data value at which a new EDS ends is the data value at which the new EDS started. These are a certain number of sample values.

[0006] In another aspect, the first similarity protocol is a measure of the distance between two data segments. The similarity threshold is calculated, and if the distance has a predetermined relationship with the similarity threshold, the two data segments A ment is defined as something similar.

[0007] In another embodiment, the controller is a second type that is less restrictive than the first similarity protocol. When RDS instances are similar to each other as determined by the similarity protocol, user input is used. Accordingly, two of the RDS instances will be joined together.

[0008] In another embodiment, the controller is a second similarity protocol that is more restrictive than the first similarity protocol. Using a sex protocol, compare existing RDSs with their associated EDSs. This will generate multiple new RDS instances.

[0009] In another embodiment, the controller finds that each EDS is similar to its own EDS. A compressed data stream is generated by replacing it with a symbol representing RDS.

[0010] In another aspect, the controller processes each sequence of data values ​​that are not part of the EDS. Replace this with a count indicating the number of symbols in the sequence.

[0011] The present invention also provides a data stream containing an ordered sequence of data values ​​as a signal. This also includes a method for operating a data processing system to perform analysis on a cluster. , receiving the data stream sequentially, and inputting each data value when a data value is received. This includes allocating a DEX. A portion of the incoming data stream is buffered. It is stored in the buffer, and a new EDS that satisfies the extraction protocol is extracted from that buffer. The data processing system uses the first similarity protocol to generate new EDSs from multiple RDs. By comparing each of S, the data processing system determines that the first similarity protocol is a new E If the DS indicates that it is similar to one of the RDSs, then information to identify the new EDS. Store this in the RDS database, and if the new EDS is not similar to any of the RDSs... If not, a new RDS instance will be created.

[0012] In another aspect, the extraction protocol uses data values ​​in the buffer in which a new EDS is initiated. Then, the new EDS identifies the data value in the buffer when it finishes.

[0013] In another aspect, the data value at which the new EDS ends is a fixed number of sample values from the data value at which the new EDS starts.

[0014] In another aspect, the data processing system calculates a measure of the distance between two data segments and a similarity threshold, and if the distance has a predetermined relationship with the similarity threshold, the two data segments are defined as being similar.

[0015] In another aspect, if the RDSs are similar to each other as determined by a second similarity protocol that is less restrictive than the first similarity protocol, the data processing system combines two of the RDSs in response to a user [[ID=​​​​​​​​​​​​​​​​​​​​​​​​​​​​​Instructions to perform a method of analyzing a data stream containing signals against a cluster of signals. This includes receiving the data stream sequentially and when data values ​​are received. Assign an index to each data value, and a portion of the received data stream This includes storing the data in a memory buffer. From that buffer, the extraction protocol is satisfied. A new EDS is extracted. The new EDS is then analyzed using the first similarity protocol. The data processing system is compared with each of the RDSs, and the first similarity protocol is, Identify the new EDS if it indicates that it is similar to one of the RDSs. The information is stored in the RDS database, and the new EDS is similar to any of the RDSs. If not already done, a new RDS instance will be created.

[0020] In another embodiment, the data processing system may make each EDS similar to that EDS. By replacing it with the symbol representing the known RDS, a compressed data stream is generated. ru.

[0021] In another embodiment, the data processing system processes each sequence of data values ​​that are not part of the EDS. Replace this with a count indicating the number of symbols in the sequence. [Brief explanation of the drawing]

[0022] [Figure 1] This figure shows a data recording device according to one embodiment of the present invention. [Figure 2] This figure shows an illustrative plot of the distance distribution as a function of distance from RDS. [Modes for carrying out the invention]

[0023] The present invention provides the advantage of a method in which signals in an incoming data channel are digitized. And, in relation to data logging systems that store data in memory devices such as disk drives, This makes it easier to understand. The data stream is extracted using an "extraction algorithm". The target signal is defined as such, and in the following discussion, we will refer to it as the idle signal, which is the signal between the target signals. It can be considered to include.

[0024] Generally, users of recorded data need to understand the various signals in the data and search for the target signal. It is necessary to be able to do so. For the purpose of this study, the user should record the data streams to be recorded. Assume you do not have detailed knowledge of all the signals in the data stream. It is assumed that there are too many options for a user to consider one at a time. Therefore, the user This allows for a complete understanding of key characteristics of a signal without having to view the entire data stream. It is necessary. For this purpose, defining clusters of similar signals is effective. By examining representative members of the cluster, the user can determine which signals are being recorded. It is possible to acquire sufficient knowledge and specify the parameters necessary to search for the target signal. .

[0025] This invention provides a similarity algorithm that allows a user to calculate a similarity score related to the similarity between two signals. Based on golism, it allows defining clusters in a collection of recorded signals. We provide users with a tool to cluster objects based on similarity. Algorithms are known in this field. Unfortunately, these algorithms The inherent computational workload when applying many of these is N² or higher. Several terabytes of recorded data If a data stream can have over a million signals, then the user will have to signal Clustering recorded signals in the few minutes it takes to explore them is often impractical.

[0026] As will be explained in more detail later, the present invention relates to the recording process of small clusters of the target signal Detect. Then, these small clusters are combined in the input data stream. It provides a larger cluster that matches the cluster of signals. It is constructed without requiring a predetermined description of the signal to be performed. Ideally, these classes Each of these represents a small portion of a single cluster of the underlying signal present in the input stream. Includes. Each cluster starts with the observed signal in the input stream, as described below. The cluster size determines whether the second signal should be included in the same cluster as the first signal. It is determined by a similarity algorithm that includes a threshold for judgment. Clusters are joined, The method by which clusters are broken down into smaller clusters will be described in more detail later.

[0027] This invention inspects a digitized data stream to obtain detailed information about the data segments in advance. Detects data segments within a target data stream without requiring any knowledge. The data segment is the data that the data stream travels to on its way to the mass storage device. It is identified in real time as it passes through the logger. The data stream is mainly the target data It is assumed that it consists of individual signals separated by regions that do not contain the segment. Extraction A A data stream segment that satisfies the striction is the extracted data segment (EDS). ) is called a data stream segment containing signals that do not satisfy the extraction algorithm, and These are called idle data segments (IDS).

[0028] Ideally, each EDS should provide a data sample corresponding to one target signal without any background samples. Includes a sample. However, the need to identify EDS in a short period of time is necessary for extraction. The algorithm is constrained. In order to find the exact target signal segment, a defined threshold must be used. Easily detectable events such as rising or falling edges that cross a value level The system detects the start of a signal and defines the end of the signal as a certain number of samples from the start of the signal. It takes significantly longer than correcting the facts. If the two signals were actually the same. Therefore, in one aspect of the present invention, The extraction algorithm specifies trigger conditions that define the start of EDS, and the end of EDS is... It is defined as a certain number of input samples relative to the start of EDS. This approximation is the final When interfering with clustering, as described later, the EDS is searched from long-term storage. This allows for clustering based on more precise signal termination.

[0029] When EDS is encountered, the EDS is copied to a buffer for further examination, and the data... An index value is assigned to uniquely identify the EDS in relation to its position within the trim. It can be done. The similarity algorithm is used to determine the "similarity measure" for EDS. Similarity is also defined as the degree of similarity between any two extracted data segments. This reflects the similarity. Based on the similarity, the system of this disclosure classifies extracted data segments as similar to each other. They can be grouped into clusters of EDS. In one aspect of the present invention, similarity A A similarity algorithm includes a threshold. If the similarity has a predetermined relationship with the threshold, then the two EDSs are mutual. It is defined as being similar to the same thing. For example, if the similarity is below a threshold, then the two EDs S can be defined as being similar to each other.

[0030] When a new EDS is found, the system identifies one of the clusters where the EDS has already been found. Determine whether it is a part or not. If EDS is part of an existing cluster, then the existing cluster This will be updated to reflect the addition of new EDSs. EDSs are part of the existing cluster If none of them are sufficiently similar, a new cluster is defined, and the EDS is assigned to that cluster. It will be added.

[0031] Each cluster is represented by a reference data segment (RDS). Extraction and clustering The ringing is performed in real time during recording, so the user can see the newly recorded data. Without needing to recover EDS from the data stream, the EDS present in the data stream Clusters can be viewed. Data stream during data recording and initial clustering. Only newly identified EDSs are retained in memory. This facilitates clustering operations. To achieve this, RDS is maintained in system memory. Once the recording of the data stream is complete... Afterward, the clustered EDS can be recovered and used for further classification.

[0032] Here, we refer to Figure 1, which shows a data recording device according to one embodiment of the present invention. The incoming data stream is digitized by the ADC (Analog-to-Digital Converter) 11, The output of DC is stored in the local FIFO buffer 12. The FIFO buffer 12 is It should be noted that it can be implemented in the cal memory 16. Clock 13 or For each clock cycle, one sample is digitized. Controller 1 5 maintains an internal register that is incremented in each clock cycle, FIF Identify the data segment that begins with the data sample just transferred to O buffer 12. It provides a unique index for this purpose. New data entries are placed in the FIFO buffer. It is transferred to 12, and in each cycle of clock 13, the most in FIFO buffer 12 Older entries are also read. In each clock cycle, the controller 15, Determine whether the elephant data segment has started or is complete at that point. R15 may include hardware to detect the start of the target data segment, or The controller 15 then inspects the contents of the FIFO buffer 12 and determines the target data segment It is possible to determine whether it has started or is completed at that point. Hardware triggers are used in the technical field, and this is known to those skilled in the art. At that point, if the target data sequence is in the FIFO buffer 12, the controller 15 This copies the data sequence from the FIFO buffer to a new EDS buffer 17. The location of the new EDS in the data stream is identified, and that information is stored in the EDS database. Enter it into S19.

[0033] To facilitate searching for EDS from disk 14, disk database 22 is provided. The correspondence between records on disk 14 and the index assigned to the start of each EDS It records the data. Generally, disk 14 contains multiple disks that can be accessed randomly. It is compiled as a disk record. Controller 15 stores it on disk 14. If you need to recover EDS, use disk database 22 to retrieve EDS related information. The disk record number where the index begins is required.

[0034] If the target data sequence has just started with a preceding sample, the controller 15 records the sample index in which the data sequence started in the EDS database. do.

[0035] As mentioned above, a predetermined extraction algorithm defines the data segments to be extracted. It must be. Generally, the extraction algorithm is the data that should become the extracted data segment. Defines the start and end of the sequence. The controller that executes the extraction algorithm is The data sequence must be able to be identified before it leaves the FIFO buffer 12. No. The extraction algorithm must operate in real time. To the oscilloscope. This technology is a real-time trigger algorithm that identifies the start of a target sequence in the input. It is known in the field of technology. The trigger algorithm uses simple features such as rising edges. Or it identifies complex features, such as specific signals. The system of this disclosure uses the target data sheet Since the exact properties of Kens are not known beforehand, the extraction algorithm preferably uses a wide range of methods. A real-time trigger algorithm that selects signals and therefore identifies signals of a broader classification. This is preferable. The start of the data sequence that will become the extracted data segment is in real time. Please note that it does not need to occur in the sample that triggered the event. For example, The extracted data segment is identified by a real-time trigger a predetermined number of times prior to the sample. You can start with this sample.

[0036] The extraction algorithm must also specify the end of the target data sequence. In one exemplary embodiment, the extraction algorithm is triggered and placed in the FIFO buffer 12. Specify the window. In this example, the extracted data segment ends at the end of the window, and the target signal is Even if it is possible that the process will end before the last data value in the window, the sample within the specified window All of the elements are part of EDS.

[0037] In another exemplary embodiment, the extraction algorithm determines the end of the data sequence to be extracted. Specify a trigger to notify. For example, the extraction algorithm terminates when the value falls below a certain threshold. The system then defines a falling edge that remains below a certain value for a specified number of samples. The resulting data value may be required to signal the end of the target data segment. Therefore, the EDS database also contains the length of the EDS, or the last EDS. This also includes equivalent information such as the index of the data sample.

[0038] In one aspect of the present invention, information specifying the termination of the EDS is also included in the EDS database 1. It is included in 9.

[0039] When a new EDS is extracted, that EDS is placed in each of the dynamically generated reference libraries. It is compared to RDS. The RDS library stores information about each RDS within the library. This includes RDS database 18. The new EDS is sufficiently similar to one of the RDS databases. If this occurs, a new EDS entry in the EDS database will be updated to indicate the relationship. And the RDS database is newly recognized as part of the cluster associated with that RDS. The EDS will be updated to indicate the identification of the EDS. The new EDS is sufficient for one of the RDSs. If they are not similar, after comparing the new EDS with all RDS in the RDS database, sufficient processing is performed. If processing time remains, use a new EDS as RDS and perform the RDS database processing. Enter the linked data and a new RDS will start. If sufficient processing time is not available... New EDS entries in the EDS database will be treated as if they were not assigned. It is then marked. For example, matching EDS to RDS before all RDS is considered. There is a possibility that a new EDS may be discovered inside, and therefore, controller 15 will... A new EDS buffer must be used for the EDS.

[0040] At the start of processing the data stream, the controller 15 checks between the two data segments. The system receives a similarity measurement algorithm that measures the similarity of the two elements. In one aspect of the present invention, the system measures the similarity of the two elements. The similarity algorithm uses a threshold to determine whether two data segments are similar or not. The algorithm generates similarity scores to be compared. This algorithm is controlled by controller 15 using EDS. It is used to measure the similarity between RDS and other RDS within the RDS library. Similarity Algorithm The SM can be more easily understood by considering four types of algorithms. Yes, it is. The first three types of algorithms operate on the data values ​​themselves. The fourth is The algorithm of this type moves the "signature" derived from each data sequence. To make.

[0041] The first type of similarity algorithm directly compares data segments to determine their similarity. In the simplest case, the two data segments have the same length, and the similarity function measures the distance between two vectors whose components are data values. For example, if EDS has sample values ​​p(i) for i=1 to N and RDS has sample values ​​q(i) for i=1 to N, then the Euclidean distance...

number

[0042] The second type of similarity function measures the distance between data segments before measuring the distance between data segments. Standardize the data segment. In some applications, the shape of the data segment is the data segment. More important than a perfect match is the same shape as the data segments, even if they have different amplitudes. This can represent two signals having the same properties. That is, p(i) = Kq(i). If the objective is to find signals that have the same shape regardless of the signal amplitude, then each data The segment is first determined by the average amplitude before calculating the distance between segments. It is divided by a constant. In one example, the constant is the maximum value of the data segment. In this example, the constant is the average of the absolute values ​​of the data values ​​in the data segment.

[0043] The third type of similarity function is used for relatively small data segments and relatively large data segments. Search for a match with the segment. This is because the user has some relatively small sequence. This is useful when you want to find the data segment that contains This occurs when the lengths are different. Basically, the user is using a relatively small data sequence. We want to find a relatively large data sequence that contains a sequence similar to S. In this example, the correspondence between a relatively small data segment and a relatively large data segment is shown. The distance between the parts is measured. Relatively small data segments are compared to i=1~m. Then p(i), and the relatively large data segment is q(i) for i=1 to N. In some cases, for k = 0 to (Nm-1), the distance function

number

[0044] The similarity function described above acts directly on the data segments being compared. The similarity function is intuitively understandable even to those who are not experts in clustering analysis. However, the workload of calculating similarity when classifying EDS is high when the EDS is large. There is great potential. Furthermore, the similarity ties that users want to use to classify EDS Depending on the context, a fourth class of similarity function may be preferable.

[0045] In the fourth class similarity analysis, a signature vector is derived from each data segment. And, using the distance between signature vectors, in a similar manner to that described above. Similarity can be measured using this method. This type of similarity measurement is applied to all of the EDS. Signature vectors have the same components even if the data segments have different lengths. Generally speaking, the number of components in a signature vector is greater than the number of data values ​​in an EDS. It is much smaller, and therefore the computational workload for performing distance measurements is significantly reduced, however The savings come from the computational work of deriving the components of the signature vector from the corresponding data segments. It is offset by the load. Generally, the components of the signature vector are its data segment. Let be any function of the data segment that is likely to be identified from other data segments. It is possible. When the extraction algorithm generates data segments of different lengths, One component of the necha vector can be the length of the data segment. Other components This can be derived from a finite impulse response filter applied to the data segment. For example, a component representing the amplitude of the frequency components of a data segment can be used.

[0046] Update the RDS library to identify EDSs and account for each new EDS found. The process is preferably performed in real time. For the purposes of this disclosure, the process is This invention completes the process without reducing the speed at which the data stream enters the data logger. If possible, the process will run in real time. (Data extraction portion of the process) In this case, the input data stream passes through the FIFO and then exits to the disk storage device. Therefore, through the extraction process, i.e., the identification of a new EDS, the controller, Identify the data segments that satisfy the extraction algorithm and note those data segments. Move to the buffer within the FIFO buffer before a portion of that data segment leaves the FIFO buffer. It must be possible to make it happen.

[0047] The time it takes to complete preliminary classification and update the RDS library depends on the amount of memory and available parallelism. It depends on the extent of processing. In one mode, the new EDS is the EDS batch in memory. Move to F17 and compare with RDS in the library. A new EDS is added to the RDS. The time required for the test can be improved by maintaining RDS in memory during the comparison. It is possible.

[0048] Furthermore, the time it takes to find a match reflects the likelihood of finding a match in an existing RDS. This can be improved by performing a comparison. The RDS database is its R Includes the count of EDS matches already found for S. Those counts correspond to RDS is a measure of likelihood of matching the next EDS. Therefore, the count associated with each RDS Performing the matching in the order of T improves the speed at which matches (if any) are found.

[0049] When the likelihood changes over time, we can use a separate likelihood variable that decays over time. Each time an EDS is assigned to an RDS, the likelihood count for that RDS is set to 1. It is reduced. Periodically, the likelihood count is multiplied by a decay factor that is less than 1. It decreases by doing so. The search for matches is performed in an order defined by the likelihood count. It will be done.

[0050] Finally, it should be noted that parallel processing can shorten the matching process time. The new EDS matching to one of the RDSs is the same as matching to another of the RDSs. This allows the EDS matching process to proceed in parallel. Therefore, the matching time can be reduced by approximately 1 / M. This can be done, where M is the number of available parallel processors. Distance calculation is also high performance This can also be done using the graphics processor core of a graphics display card. Therefore, the speed improvement achieved through parallel processing can exceed 1 / 1000th.

[0051] In the matching process, the controller needs the average to find and extract the EDS. It should also be noted that, on average, one EDS needs to be processed per hour. If there is enough buffer to store the new EDS waiting, the system will match The EDS is processed based on the average time to find a match, not the longest time to find one. Just do it.

[0052] The EDS matching against RDS within the RDS library is still waiting for a match. It may fail to complete before the buffer capacity to hold the new EDS is exceeded. For EDS entries that were not matched, no match was found. It is marked as such, and the process proceeds to the next EDS waiting to be matched. Therefore, the buffer that held the unmatched EDS was used by the new EDS. To be used. Unmatched EDSs will be at the end of the recording period, or during the recording period. Between the subsequent parts, the discovery of new EDSs is slow, so buffer space is available. Then it can be processed.

[0053] In one aspect of the system, the reference database is empty at the start of the data recording operation. When a new extracted data segment is encountered, how many of the new extracted data segments are selected? This becomes a reference data segment. For example, the first extracted data segment is the reference data This becomes a segment. The second new extracted data segment is a new reference data segment. This can be, or simply represented by a previously generated reference data segment. It can be labeled as part of a cluster.

[0054] In another aspect of the present invention, the user has one or more reference data segments to use for comparison. You can enter data. Reference segments are generated by the user, or data Analyzed by a device similar to the present invention, provided by a manufacturer of logging equipment. It may have been found in a different data stream.

[0055] In one aspect of the present invention, the RDS database held in memory during recording and initial processing The RDS and related RDS help users understand the data streams being logged. In this way, the user can see during recording. In one embodiment, the user can access RDS A list was presented, and it was found that the list of RDSs was similar to that RDS. They are ordered by the count of the number of S. And the user is one to display You can choose from the above RDS options.

[0056] As mentioned above, RDS databases are similar to each RDS, in particular. Includes entries listing the identification of each EDS that has been found to be The EDS is the index in the data stream where it was found. The list is compiled when the measure of similarity between the EDS and RDS meets some predetermined threshold condition. This is because the condition is met. For example, the distance between EDS and RDS is less than a certain threshold. If the threshold conditions are too lenient, a large number of EDSs will be associated with RDSs. This involves a single RDS having two or more different classes of signals in the input data stream. It may contain signals from the terminal. As will be explained in more detail later, these RDSs are It should be avoided.

[0057] If the threshold conditions are too strict, there will be even more RDSs, and each RDS will have its own threshold. The size of the cluster being processed will be smaller. In principle, during the post-recording process, relatively Smaller RDS clusters can be combined to provide a larger cluster. However, However, having a large number of small RDS clusters during data collection makes it possible to create new EDS clusters. The computational workload associated with matching against the DS cluster increases substantially. There is a trade-off between the threshold condition and the specificity of the RDS cluster.

[0058] In one aspect of the present invention, the RDS database entry for each RDS is also its This also includes a measure of actual similarity to each EDS related to RDS. The RAM can optionally be controlled by the controller during the recording process and post-recording processing, for the user It is provided to the user upon request. By looking at the histogram, especially a large amount of ED When S is associated with RDS, the user can gain insights into the cluster structure. Yes, it is possible. For example, if the histogram reveals multiple peaks, the cluster is determined by the input data. The data stream may include EDS from multiple clusters of signals, and therefore Therefore, as will be discussed later, it may be necessary for RDS to be expanded into multiple new RDS instances. There is.

[0059] At the end of the recording stage, the present invention will have generated two databases. The first The database identifies all data segments that satisfy the extraction algorithm. The database contains each EDS in the recorded data stream, and similar EDSs. Includes all RDS locations. Using this database, the controller can access any RDS It is possible to access any EDS related to S. The second database is a record Identify all RDS generated during the process. Information in the RDS database is All EDS related to the given RDS and the RDS started in the recorded data stream. Identify the location of the EDS and other information about the RDS as described above.

[0060] In some cases, it may be useful to test for one or more IDSs. For example, if the extraction algorithm defines a fixed window for the trigger position, that window is the trigger It may be too small to capture all the signals related to the EDS. The dollar data segment can provide the truncated missing portion of the EDS. The disk index of the IDS after the DS starts from the last index of that EDS. It can be calculated.

[0061] As mentioned above, there wasn't enough time available to perform all the comparisons. There may be EDS databases that failed to classify RDS libraries. The system tags any of these EDS. In post-processing, these failed EDS are tagged. DS can be re-examined. The location of each of these EDS is recorded in the EDS database. It is recorded. The EDS is taken from the recorded data stream, and its position in the data stream. Since the location is known, it can be searched. Furthermore, the recorded data stream is on disk. If it is on a live or similar random access storage device, in order to reach that EDS Therefore, it is not necessary to play back the entire recorded data stream. For this reason, search the EDS and select the current R It can be compared against the DS library. At this point, a similarity algorithm is used. Therefore, EDS can be associated with one or more of the RDS, or sufficiently similar R If a DS is not found, a new RDS can be defined for that EDS.

[0062] One of the objectives of the data logging processes described herein is to obtain similar information from one another. Registering a number allows users to understand the various signal types in the recorded data. The goal is to enable this. Each RDS represents a cluster of EDS, but a collection of RDSs does not necessarily mean This also allows users to fully understand the clustering of the underlying EDS set. Not necessarily. For example, a much larger number of clusters than those in the underlying EDS set. RDS may exist. This invention provides insights into the clustering of underlying signals. We provide two tools.

[0063] The first tool finds groups of RDS that are part of the same underlying signal cluster. It acts on RDS in such a way. Each RDS entry in the RDS database is E This includes EDS, which represents a smaller group of DS. Therefore, selected RDS also By clustering them, users can access larger clusters of similar EDSs. It can be built. Because the number of RDS is effectively less than the number of EDS, the RDS cluster Stirring can be performed with substantially less computational workload. See the following discussion. To simplify terminology, we will refer to RDS clusters as groups.

[0064] The purpose of clustering RDS can be more easily understood by referring to a simple example. Yes, it is possible. It is located at or near the center of the cluster of signals in the input data stream. Consider RDS. The similarity algorithm measures the distance between two data segments. It is assumed that this will happen. In particular, from EDS, which forms the basis of RDS, to the EDS library Consider the distance to each of the other EDSs. Figure 2 shows the distance from the RDS as a function of the distance. An illustrative plot of this distance distribution is shown. In the example shown in Figure 2, T1 is this RD The cutoff distance used to define the EDS corresponding to S is shown, and RDS is the input It includes only the EDS corresponding to the first cluster 31 of the signals. Ideally, the result is This RDS is designed so that the group has an effective cutoff distance as shown in T3. It should be integrated with RDS.

[0065] As mentioned above, if the original cutoff distance is too large as shown in T2, RDS will, Includes an EDS corresponding to the second cluster 32 in the input signal. Such an RDS is another R When coupled with DS, the resulting RDS also has two clusters in the input signal. It includes EDS belonging to the sta, and therefore the resulting group is in the input signal It is not limited to a single cluster. As mentioned above, the frequency distribution is as shown in Figure 2. This could be useful in identifying RDS instances that are too large.

[0066] RDS is glued in a similar way to the method used to generate RDS from EDS. The group is formed. When forming the group, the similarity relationship and threshold are similar to the method described above. These definitions are defined by law. These definitions are by the user, user interface 21 or system It can be provided through the service itself. In the simplest case, it is used to generate RDS. Using the same similarity relationships used, the similarity threshold is set to group the candidate RDS. By changing the process to be less selective, groups are generated. However, different similarity relationships can be utilized.

[0067] Initially, there were no groups, and therefore, the first group was from the first RDS to be tested. In one aspect of the present invention, to initiate this first group, the most EDS An RDS having is selected. In this embodiment, the underlying cluster in EDS is an RDS It is based on a model that places the center in one or near one of them. Therefore, it has the highest count. RDS is likely to be located at or near the center of these clusters. If the similarity score for two RDSs indicates that those two RDSs satisfy the similarity condition, This involves checking the remaining RDS that have not yet been assigned to a group against that group. This fills the gap. And the group has the maximum count that has not yet been assigned. The process is repeated by the RDS. If there are no more unallocated RDS, the process The task is complete.

[0068] The user, in accordance with the appropriate command given to the data processing system, will perform the following for each group. You can see the data segments corresponding to each RDS. This display shows the data segments within each RDS. It is possible to limit it to EDS defined as mind, or to all EDS related to a group. These displays allow the user to see that signals grouped using similarity relationships are... In fact, it is possible to determine whether or not the user appears similar to the actual user. Finally, the user To help determine whether the grouping process was performed to an extreme, see Figure 2. You can see a frequency distribution like the one shown.

[0069] If the number of groups is still too large, use the same similarity algorithm, but similar Using different thresholds that are not so strict in order to find sex, the process The process can be repeated. Furthermore, the process can be repeated using different similarity algorithms. It can be reversed. The restriction on the similarity algorithm is that for any two EDSs It simply has to be able to function. For example, EDs of different lengths The similarity algorithm that operates on S is in the case where the lengths of the two EDS are not substantially the same. It can be constructed by setting the similarity metric so that the two EDSs are dissimilar. If the lengths of the two EDSs are substantially the same, the distance function is calculated and compared to a threshold. Then, it is determined whether the two EDSs are similar or not. In another example, the similarity algorithm First, Zum derives a signature for each of the EDS to be compared, and then the signature By measuring the distance between the two EDSs, it is possible to determine whether the two EDSs are similar or not.

[0070] The above description refers to a specific type of clustering algorithm for reclustering RDS. We assume that... However, we will use other clustering techniques with the first tool. It is possible.

[0071] As mentioned above, the similarity criteria used to generate one of the RDSs were too lenient. In that case, the RDS may contain a very large number of EDS. Furthermore, the RDS The input signal cluster may span two or more clusters. Therefore, Replace the RDS with multiple RDS instances, each with a smaller number of associated EDS instances. This is useful. In one aspect of the present invention, the RDS is the entire EDS associated with the RDS. Search for them and reclassify their EDS using a more restrictive similarity cutoff threshold. By clustering, it can be divided into smaller RDSs. The process proceeds in a manner similar to the method described above for the original clustering of EDS. A new RDS of 1 is defined to contain the first EDS of the group of extracted EDSs. Then, each consecutive EDS is compared to a new RDS. When determining whether an EDS is similar to an RDS, that EDS is included in the new RDS. If the EDS is not sufficiently similar to one of the new RDSs, another new An RDS is defined, and its EDS is used to start that RDS. When a set of RDS objects is included in the RDS library, it is possible to repeat the grouping of RDS objects. ru.

[0072] RDS clustering after recording is when selecting EDS grouped into RDS. Based on the extraction algorithm used. First, as mentioned above, each RDS has multiple queries. It can include rasters. Such RDS can be regenerated as described above. By doing so, or if computing resources allow, load all EDS associated with RDS, and By directly applying the clustering algorithm to the EDS, smaller clusters can be obtained. It can be divided into rasters. Alternatively, for example, the extraction algorithm can be triggered by a start condition. In contrast, if it works by selecting all samples within a window of a fixed size and position, The resulting EDS simply approximates the data segment containing only the target signal. If the background sample size is too large, EDS may distort the distance calculation by using a significant number of background samples. This will include... Similarly, if the window is too small, a portion of the target signal will be cut off. As mentioned above, you can access the IDS following the EDS, and in a fixed window... Therefore, the lost portion of the signal that was cut off can be recovered. This is merely an approximation of the data signal.

[0073] The second tool mentioned above allows users to correct these approximations, and therefore, clusters The ring can be improved. For example, if the EDS extraction algorithm is based on a fixed window... In addition, the data processing system determines the stop position related to the EDS from the end of the fixed window to the target signal. A trimming algorithm can be executed that modifies the image to the position that matches the logical end. For example, the last part of the EDS is a sample stream representing the background level in the data channel. If it is a ng, the end of the EDS is defined as the position of the last data value above the background. It is possible. Similarly, the EDS is cut off by a window, and adjacent to the EDS is If there is an idle data segment, the end of the data in the idle data segment As shown, the end of the EDS can be changed.

[0074] After the EDS is updated to correct these approximations, the same similarity algorithm or a different one is used. Using a similarity algorithm, the new collection of EDSs is clustered into RDS. It is possible. And, as mentioned above regarding the first tool, a new set of RDS These can be clustered into groups. If sufficient computing resources are available, Reclass to provide a new set of RDS instances to replace the original large RDS instance. Assemble all EDS corresponding to the tagged RDS as a group and reclass them. It can be tagged.

[0075] During post-processing using the two tools mentioned above, the user will be responsible for one or more clusters. You can see the actual EDS. The EDS within the cluster is sufficiently similar to the user. If it does not appear that there are similarities, the user will not know the similarity algorithm used to determine the similarity. Both or either the m and the threshold can be changed. In one aspect of the present invention, the user You can select a similarity algorithm from a predetermined list of similarity algorithms.

[0076] In the embodiment described above, the RDS database starts empty and is filled as records are added. However, embodiments in which one or more RDSs are defined before the start of recording are also possible. It can be built. With these initial RDS, users can receive data streams. The process of searching for a specific signal while still being aware of the contents of the data stream during the capture process. This is possible. Another analysis performed by an apparatus similar to the present invention, or generated by the user. In some cases, a reference data segment may be found within the data stream.

[0077] EDS and RDS are standardized before measuring the similarity between the two data segments. Similarity algorithms can also be used. For example, each data segment can be compared To measure similarity in shape, the maximum value of the samples in the data segment is used. Therefore, it can be divided. In another example, EDS is multiplied by a constant and the similarity is calculated. The process can be completed for different predetermined constants, and the best similarity is used. It is possible.

[0078] In one aspect of the present invention, while recording the data stream in real time, as far as possible The process is organized to provide as much preliminary data as possible. And in the background, Alternatively, higher-level processing is performed after data recording is complete. Recording large amounts of data... Processing at that point extracts data segments that satisfy the extraction criteria, extracted data segments Preliminary classification of similar data segments based on performance, and real-time operational performance. This requires the detection and registration of reference data segments while recording input data. This becomes possible. First, using the results of the preliminary classification, the preliminaryly classified reference data segment Classification can be performed by cluster analysis on the ment. Therefore, user However, without having to wait for data to be recovered from long-term recording devices, analysis processing time and user - It can provide a response to the query.

[0079] In other words, the processing prioritizes high speed and therefore has the best similarity. It does not identify DS, but instead identifies RDS that are similar to EDS with a predetermined threshold of accuracy. As soon as an RDS is found, it is adopted as an EDS tag, and processing is terminated. Until the detailed classification of data segments is performed by clustering, the complete The process of obtaining a classification decision will be postponed.

[0080] The number of data segments and RDSs to be stored sets classification thresholds for similarity evaluation. This can be adjusted by sacrificing classification errors during preliminary processing. This can reduce the time required for preliminary classification.

[0081] As mentioned above, processing time can be reduced by using parallel processing as described above. Furthermore, by inspecting the RDS in an order that reflects the number of EDS associated with each RDS, Furthermore, processing time can be further reduced. In addition, similarity assessment used in preliminary classification The valence function can be chosen to have a low computational load, thereby allowing for subsequent processing. Using more complex similarity measures can improve the accuracy of clustering.

[0082] The above-described embodiment uses a data logger as an example, but the present invention is similar in that two signals We define an extraction algorithm in conjunction with a similarity algorithm that determines whether or not they are similar. It can be applied to a wide range of data signals.

[0083] In the embodiments described above, the input data stream was essentially a scalar. That is, The input data stream consists of a single value in each clock cycle. However, The teachings of the present invention can be applied to vector input data streams. In the data stream, each channel provides an input vector in each clock cycle. There are multiple input data channels that are processed by the ADC. Start of a new EDS. The trigger circuit that defines it operates for one or more of the channels. It is possible to apply the above teachings to such vector data streams. can.

[0084] In the embodiment described above, the original data stream digitizes the original analog signal. Except for any quantization errors introduced by the process, the disk or other long-term storage can be stored without loss. It can be recovered from the device. As mentioned above, for this original data stream The memory requirements can be tens of terabytes. In some applications, lossy compression is used. The ability to provide compressed data streams using algorithms is an advantage. There are two types of approximations that can be used to provide a compressed data stream. The first approximation replaces IDS with a count of the number of data samples in each IDS. This reduces each IDS to a code and count that indicates it is an IDS. do.

[0085] The second approximation replaces each EDS with the EDS in the RDS that contains that EDS. RDS includes representative EDS, and the remaining EDS related to that RDS are similar to that EDS. Therefore, each EDS in the data stream is located in the RDS where that EDS is located. It is replaced with identification information. However, in this approximation, the data stream is RDS The library must be included at some point. However, the average number of RDS is EDS Assuming the number is much smaller, the level of compression is remarkable. Each representative EDS is It can be replaced with a compressed version of EDS. When compressing a typical EDS, Lossless compression algorithms such as Tropi coding can be used. Alternatively, representative E DS is one of the lossy data compression algorithms known in the field of data compression technology. It can be compressed using one of these methods. These conventional data compression techniques are lossless compression. It should be noted that this can include both gorymastic and lossy compression algorithms. be.

[0086] The present invention also provides an analytical tool for understanding signals in pre-recorded datasets. It can also be used. In this case, the pre-recorded dataset is similar to the device shown in Figure 1. The data is input to the device. If the recorded data is already in digital format, ADC11 is omitted. This is possible. In such applications, the controller 15 can optionally process the data. The speed at which the input to the system can be controlled. Therefore, the next new EDS can be processed. If there isn't enough time to compare the new EDS with each of the RDS before it becomes necessary The controller 15 simply stops inputting data so that the system can catch up. It is possible.

[0087] If a compressed version of the pre-recorded data is desired, use the compressed data stream. After determining the RDS, the dataset can be read in the second time. Then, A compressed data stream can be output to disk 14.

[0088] As described above, the controller of the present invention is a conventional computer or multiprocessor This can be done. Matching EDS to RDS can be done by using a multiprocessor. This process can increase the speed, and the reason is the new EDS and RDS The result of matching with one of them is parallel to the matching of that EDS with another RDS. This is because it can be executed in a shorter time. Multiprocessors are different from conventional multicore computers. Alternatively, it can be a graphics processing board with thousands of cores.

[0089] The present invention also includes a computer that stores instructions for a data processing system to perform the method of the present invention. This includes computer-readable media. Computer-readable media are patentable under Section 101 of the United States Patent Act. It is defined as any medium that constitutes a subject matter that may be patented, and under Section 101 of the United States Patent Act Any medium that does not constitute a subject that could be permitted is excluded. Examples of such mediums include A computer or data processing system stores information in a format that is readable by the computer. Examples include non-temporary media such as data memory devices.

[0090] The embodiments described above are provided to illustrate various aspects of the present invention. However, it will be understood that other embodiments of the present invention can be provided by combining different aspects of the present invention shown in different specific embodiments. Furthermore, various modifications to the present invention will become apparent from the above description and accompanying drawings. Accordingly, the present invention is limited only to the scope of the appended claims. The original claims of the parent application are as follows: Claim 1: A system for recording and analyzing data streams, An input port adapted to receive the aforementioned data stream, wherein the data stream includes an ordered sequence of data values, An output port adapted for communicating the aforementioned data stream to a mass storage device, A buffer connected to the input port is used to temporarily store a predetermined portion of the data stream when the data stream is received by the system. A controller that identifies a new segment called an extracted data segment (EDS) from the data stream stored in the buffer that satisfies an extraction protocol, and compares the new EDS to each of a plurality of reference data segments (RDS) using a first similarity protocol, wherein the controller stores information identifying the new EDS in an RDS database if the first similarity protocol indicates that the new EDS is similar to one of the RDS, and if the controller is not similar to any of the RDS, it generates a new RDS, each RDS including a list of the EDS that was found to be similar to that RDS and the new EDS that caused the controller to generate the new RDS. A system that includes these features. Claim 2: The system according to claim 1, wherein the first similarity protocol calculates a measure of distance and a similarity threshold between two data segments, and the two data segments are defined as similar if the distance has a predetermined relationship with the similarity threshold. Claim 3: The system according to claim 2, wherein the controller generates a plurality of new RDSs by comparing the EDSs associated with an existing RDS with each other using a second similarity protocol that is more restrictive than the first similarity protocol. Claim 4: A method for operating a data processing system to analyze a data stream containing an ordered sequence of data values ​​against a cluster of signals, The aforementioned data stream is received sequentially, and an index is assigned to each data value as it is received. The process involves storing a portion of the received data stream in a buffer, Extracting a new EDS that satisfies the extraction protocol from the aforementioned buffer, The data processing system compares the new EDS to each of a plurality of RDSs using a first similarity protocol, wherein if the first similarity protocol indicates that the new EDS is similar to one of the RDSs, the data processing system stores information identifying the new EDS in the RDS database, and if the new EDS is not similar to any of the RDSs, the data processing system generates a new RDS. Methods that include... Claim 5: The extraction protocol identifies a data value in the buffer at which the new EDS starts and a data value in the buffer at which the new EDS ends, wherein the data value at which the new EDS ends is a certain number of sample values ​​from the data value at which the new EDS started, according to the method of claim 4 or the system according to claim 1. Claim 6: The method according to claim 4 or the system according to claim 1, wherein the data processing system calculates a measure of distance and a similarity threshold between two data segments, defines the two data segments as similar if the distance has a predetermined relationship with the similarity threshold, and the data processing system combines two of the RDSs in response to user input if the RDSs are similar to each other as determined by a second similarity protocol which is less restrictive than the first similarity protocol. Claim 7: The method according to claim 4 or the system according to claim 1, wherein the data processing system calculates a measure of distance and a similarity threshold between two data segments, defines the two data segments as similar if the distance has a predetermined relationship with the similarity threshold, and generates a plurality of new RDSs from an existing RDS by comparing the EDS associated with that RDS with each other using a second similarity protocol which is more restrictive than the first similarity protocol. Claim 8: A computer-readable memory including instructions for a data processing system to perform a method of analyzing a data stream containing an ordered sequence of data values ​​against a cluster of signals, wherein the method is: The aforementioned data stream is received sequentially, and an index is assigned to each data value as it is received. The process involves storing a portion of the received data stream in a memory buffer, Extracting a new EDS that satisfies the extraction protocol from the aforementioned buffer, The data processing system compares the new EDS to each of a plurality of RDSs using a first similarity protocol, wherein if the first similarity protocol indicates that the new EDS is similar to one of the RDSs, the data processing system stores information identifying the new EDS in the RDS database, and if the new EDS is not similar to any of the RDSs, the data processing system generates a new RDS. Computer-readable memory, including [specific data / data]. Claim 9: The computer-readable memory according to claim 8, the system according to claim 1, or the method according to claim 4, wherein the data processing system generates a compressed data stream by replacing each EDS with a symbol representing the RDS which is found to be similar to the EDS. Claim 10: The computer-readable memory according to claim 8, the system according to claim 1, or the method according to claim 4, wherein the data processing system generates a compressed data stream by replacing each EDS with a symbol representing the RDS which is found to be similar to the EDS, and the data processing system replaces each sequence of data values ​​that are not part of the EDS with a count indicating the number of symbols in the sequence. [Explanation of symbols]

[0091] 11. Analog-to-Digital Converter (ADC) 12 FIFO buffers 13 clock 14 discs 15 Controllers 16 Local memory 17 EDS buffer 18 RDS Databases 19 EDS Database 21 User Interface 22 Disk Databases

Claims

1. A system for recording and analyzing data streams, An input port adapted to receive the aforementioned data stream, wherein the data stream includes an ordered sequence of data values, An output port adapted for communicating the aforementioned data stream to a mass storage device, A buffer connected to the input port is used to temporarily store a predetermined portion of the data stream when the data stream is received by the system. A controller that identifies a new segment called an extracted data segment (EDS) from the data stream stored in the buffer that satisfies an extraction protocol, and compares the new EDS to each of a plurality of reference data segments (RDS) using a first similarity protocol, wherein the controller stores information identifying the new EDS in an RDS database if the first similarity protocol indicates that the new EDS is similar to one of the RDS, and the controller generates a new RDS if the new EDS is not similar to any of the RDS, each RDS including a list of the EDS that was found to be similar to that RDS and the new EDS that caused the controller to generate the new RDS. A system equipped with these features.

2. The system according to claim 1, wherein the first similarity protocol calculates a distance measurement between two data segments and a similarity threshold, and the two data segments are defined as similar if the distance has a predetermined relationship with the similarity threshold.

3. The system according to claim 2, wherein the controller generates a plurality of new RDSs by comparing the EDSs associated with an existing RDS with each other using a second similarity protocol that is more restrictive than the first similarity protocol.

4. The extraction protocol identifies the data value in the buffer at which the new EDS starts and the data value in the buffer at which the new EDS ends. The system according to claim 1, wherein the data value at the end of the new EDS is a certain number of sample values ​​from the data value at the start of the new EDS.

5. The first similarity protocol calculates a distance measurement between two data segments and a similarity threshold, and defines the two data segments as similar if the distance measurement has a predetermined relationship with the similarity threshold. The system according to claim 1, wherein the controller combines two of the RDSs in response to user input if the RDSs are similar to each other when determined by a second similarity protocol which is less restrictive than the first similarity protocol.

6. The first similarity protocol calculates a distance measurement between two data segments and a similarity threshold, and defines the two data segments as similar if the distance measurement has a predetermined relationship with the similarity threshold. The system according to claim 1, wherein the controller generates a plurality of new RDSs from an existing RDS by comparing the EDSs associated with that RDS with one another, using a second similarity protocol that is more restrictive than the first similarity protocol.

7. The system according to claim 1, wherein the controller generates a compressed data stream by replacing each EDS with a symbol representing the RDS which has been found to be similar to that EDS.

8. The controller generates a compressed data stream by replacing each EDS with a symbol representing the RDS that is found to be similar to that EDS. The system according to claim 1, wherein the controller replaces each sequence of data values ​​that are not part of the EDS with a count indicating the number of symbols in the sequence.

9. A method for operating a data processing system to analyze a data stream containing an ordered sequence of data values ​​against a cluster of signals, The aforementioned data stream is received sequentially, and an index is assigned to each data value as it is received. The process involves storing a portion of the received data stream in a buffer, From the aforementioned buffer, extract a new extracted data segment (EDS) that satisfies the extraction protocol, The data processing system compares the new EDS to each of a plurality of reference data segments (RDS) using a first similarity protocol, wherein if the first similarity protocol indicates that the new EDS is similar to one of the RDS, the data processing system stores information identifying the new EDS in the RDS database, and if the new EDS is not similar to any of the RDS, the data processing system generates a new RDS. A method that includes this.

10. The extraction protocol identifies the data value in the buffer at which the new EDS starts and the data value in the buffer at which the new EDS ends. The method according to claim 9, wherein the data value at which the new EDS ends is a certain number of sample values ​​from the data value at which the new EDS started.

11. The data processing system calculates a measured distance between two data segments and a similarity threshold. If the measured distance has a predetermined relationship with the similarity threshold, the two data segments are defined as similar. The method according to claim 9, wherein the data processing system combines two of the RDSs in response to user input if the RDSs are similar to each other when determined by a second similarity protocol which is less restrictive than the first similarity protocol.

12. The data processing system calculates a measured distance between two data segments and a similarity threshold. If the measured distance has a predetermined relationship with the similarity threshold, the two data segments are defined as similar. The method according to claim 9, wherein the data processing system generates a plurality of new RDSs from an existing RDS by comparing EDSs associated with that RDS with one another, using a second similarity protocol that is more restrictive than the first similarity protocol.

13. The method according to claim 9, wherein the data processing system generates a compressed data stream by replacing each EDS with a symbol representing the RDS which is found to be similar to that EDS.

14. The data processing system generates a compressed data stream by replacing each EDS with a symbol representing the RDS that is found to be similar to that EDS. The method according to claim 9, wherein the data processing system replaces each sequence of data values ​​that are not part of the EDS with a count indicating the number of symbols in the sequence.

15. A computer-readable memory comprising instructions for causing a data processing system to perform a method of analyzing a data stream containing an ordered sequence of data values ​​against a cluster of signals, wherein the method is: The aforementioned data stream is received sequentially, and an index is assigned to each data value as it is received. The process involves storing a portion of the received data stream in a memory buffer, From the aforementioned buffer, extract a new extracted data segment (EDS) that satisfies the extraction protocol, The data processing system compares the new EDS to each of a plurality of reference data segments (RDS) using a first similarity protocol, wherein if the first similarity protocol indicates that the new EDS is similar to one of the RDS, the data processing system stores information identifying the new EDS in the RDS database, and if the new EDS is not similar to any of the RDS, the data processing system generates a new RDS. Computer-readable memory, including [specific data / data].

16. The computer-readable memory according to claim 15, wherein the data processing system generates a compressed data stream by replacing each EDS with a symbol representing the RDS which is found to be similar to that EDS.

17. The data processing system generates a compressed data stream by replacing each EDS with a symbol representing the RDS that is found to be similar to that EDS. The computer-readable memory according to claim 15, wherein the data processing system replaces each sequence of data values ​​that are not part of the EDS with a count indicating the number of symbols in the sequence.