Methods, systems, and computer program products for efficient content-based time series inspection
By calculating the paired distance matrix between the time series and the learned template and processing these matrices using a residual network, the eigenvectors of the time series are generated, and the problem of low calculation efficiency of time series similarity scores in the prior art is solved, and an efficient time series retrieval system is realized.
Patent Information
- Application Number
- CN202480004422.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-01
- Filing Date
- 2024-05-31
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-05-31
AI Technical Summary
When existing content-based time series retrieval systems process multi-domain time series data, it is difficult to efficiently calculate the similarity score between time series, especially in real-time interaction scenarios.
By obtaining a plurality of known time series from the database using at least one processor, calculating pairwise distance matrices between the time series and the learned templates, stacking the matrices to generate tensors, and processing the tensors using a residual network to generate eigenvectors of the time series.
It realizes efficient calculation of similarity scores between time series, improves the system's performance in real-time interactive scenarios, and can effectively process multi-domain time series data.
Smart Images

Figure CN120112902A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 505,570, filed on June 1, 2023, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates generally to time series data and, in some non-limiting embodiments or aspects, to methods, systems, and computer program products for efficient content-based time series retrieval. Background Art
[0004] The Content-based Time Series Retrieval (CTSR) system is an information retrieval system for users to interact with time series emerging from multiple domains, such as finance, healthcare, manufacturing, etc. For example, a user who wishes to learn more about the source of a time series can submit the time series as a query to the CTSR system and retrieve a list of related time series with associated metadata. By analyzing the retrieved metadata, the user can gather more information about the source of the time series. Since the CTSR system can process time series data from different domains, the CTSR system can use high-capacity models to effectively measure the similarity between different time series. In addition, when the user interacts with the system in real time, the user may need the model within the CTSR system to calculate the similarity score in an efficient manner. Summary of the invention
[0005] Accordingly, improved methods, systems, and computer program products for content-based time series retrieval are provided.
[0006] According to some non-limiting embodiments or aspects, a method is provided, comprising: obtaining a plurality of known time series from at least one database using at least one processor; for each known time series in the plurality of known time series: calculating a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates using the at least one processor to generate a plurality of pairwise distance matrices; stacking the plurality of pairwise distance matrices together using the at least one processor to generate a tensor; and processing the tensor using a residual network using the at least one processor, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; and providing the feature vector of each known time series in the plurality of known time series using the at least one processor.
[0007] In some non-limiting embodiments or aspects, the method further includes: using the at least one processor to store the feature vector of each known time series in the plurality of known time series in the at least one database.
[0008] In some non-limiting embodiments or aspects, the method further includes: obtaining an unknown time series using the at least one processor; calculating a pairwise distance matrix between the unknown time series and each of the multiple learned templates using the at least one processor to generate another plurality of pairwise distance matrices; stacking the other plurality of pairwise distance matrices together using the at least one processor to generate another tensor; processing the other tensor using the residual network using the at least one processor, wherein the residual network receives the other tensor as input and provides a feature vector of the unknown time series as output; for each known time series of the multiple known time series stored in the database, determining the distance between the known time series and the unknown time series using the at least one processor based on the stored feature vector of the known time series and the feature vector of the unknown time series; and identifying at least one known time series determined to be similar to the unknown time series based on the distance between each known time series and the unknown time series using the at least one processor.
[0009] In some non-limiting embodiments or aspects, the residual network is trained using a loss function defined according to the following equation:
[0010]
[0011] in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
[0012] In some non-limiting embodiments or aspects, the plurality of known time series includes a plurality of known transaction time series associated with a plurality of merchants, and wherein each known time series is associated with metadata associated with the merchant associated with the known time series.
[0013] In some non-limiting embodiments or aspects, the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four.
[0014] In some non-limiting embodiments or aspects, the residual network comprises a two-dimensional residual network.
[0015] According to some non-limiting embodiments or aspects, a system is provided, comprising: at least one processor coupled to a memory and configured to: obtain a plurality of known time series from at least one database; for each known time series in the plurality of known time series: calculate a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates to generate a plurality of pairwise distance matrices; stack the plurality of pairwise distance matrices together to generate a tensor; and process the tensor with a residual network, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; and provide the feature vector of each known time series in the plurality of known time series.
[0016] In some non-limiting embodiments or aspects, the at least one processor is further configured to: store the feature vector of each known time series in the plurality of known time series in the at least one database.
[0017] In some non-limiting embodiments or aspects, the at least one processor is further configured to: obtain an unknown time series; calculate a pairwise distance matrix between the unknown time series and each of the multiple learned templates to generate another plurality of pairwise distance matrices; stack the other plurality of pairwise distance matrices together to generate another tensor; process the other tensor with the residual network, wherein the residual network receives the other tensor as input and provides a feature vector of the unknown time series as output; for each known time series of the multiple known time series stored in the database, determine the distance between the known time series and the unknown time series based on the stored feature vector of the known time series and the feature vector of the unknown time series; and identify at least one known time series determined to be similar to the unknown time series based on the distance between each known time series and the unknown time series.
[0018] In some non-limiting embodiments or aspects, the residual network is trained using a loss function defined according to the following equation:
[0019]
[0020] in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
[0021] In some non-limiting embodiments or aspects, the plurality of known time series include a plurality of known transaction time series associated with a plurality of merchants, and wherein each known time series is associated with metadata associated with the merchant associated with the known time series.
[0022] In some non-limiting embodiments or aspects, the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four.
[0023] In some non-limiting embodiments or aspects, the residual network comprises a two-dimensional residual network.
[0024] According to some non-limiting embodiments or aspects, a computer program product is provided, comprising a non-transitory computer-readable medium, the non-transitory computer-readable medium comprising program instructions, which, when executed by at least one processor, causes the at least one processor to: obtain a plurality of known time series from at least one database; for each known time series in the plurality of known time series: calculate a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates to generate a plurality of pairwise distance matrices; stack the plurality of pairwise distance matrices together to generate a tensor; and process the tensor with a residual network, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; and provide the feature vector of each known time series in the plurality of known time series.
[0025] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, further cause the at least one processor to: store the feature vector of each known time series in the plurality of known time series in the at least one database.
[0026] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, further cause the at least one processor to: obtain an unknown time series; calculate a pairwise distance matrix between the unknown time series and each of the multiple learned templates to generate another plurality of pairwise distance matrices; stack the other plurality of pairwise distance matrices together to generate another tensor; process the other tensor with the residual network, wherein the residual network receives the other tensor as input and provides a feature vector of the unknown time series as output; for each known time series of the multiple known time series stored in the database, determine the distance between the known time series and the unknown time series based on the stored feature vector of the known time series and the feature vector of the unknown time series; and identify at least one known time series determined to be similar to the unknown time series based on the distance between each known time series and the unknown time series.
[0027] In some non-limiting embodiments or aspects, the residual network is trained using a loss function defined according to the following equation:
[0028]
[0029] in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
[0030] In some non-limiting embodiments or aspects, the plurality of known time series include a plurality of known transaction time series associated with a plurality of merchants, and wherein each known time series is associated with metadata associated with the merchant associated with the known time series.
[0031] In some non-limiting embodiments or aspects, the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four, and wherein the residual network comprises a two-dimensional residual network.
[0032] Other non-limiting embodiments or aspects are set forth in the following numbered clauses:
[0033] Item 1: A method comprising: obtaining a plurality of known time series from at least one database using at least one processor; for each known time series in the plurality of known time series: calculating a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates using the at least one processor to generate a plurality of pairwise distance matrices; stacking the plurality of pairwise distance matrices together using the at least one processor to generate a tensor; and processing the tensor using a residual network using the at least one processor, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; and providing the feature vector of each known time series in the plurality of known time series using the at least one processor.
[0034] Item 2: The method according to Item 1 further includes: using the at least one processor to store the feature vector of each known time series in the multiple known time series in the at least one database.
[0035] Item 3: The method according to item 1 or 2 further includes: obtaining an unknown time series using the at least one processor; calculating the pairwise distance matrix between the unknown time series and each of the multiple learned templates using the at least one processor to generate another plurality of pairwise distance matrices; stacking the other plurality of pairwise distance matrices together using the at least one processor to generate another tensor; processing the other tensor with the residual network using the at least one processor, wherein the residual network receives the other tensor as input and provides a feature vector of the unknown time series as output; for each known time series among the multiple known time series stored in the database, determining the distance between the known time series and the unknown time series using the at least one processor based on the stored feature vector of the known time series and the feature vector of the unknown time series; and identifying at least one known time series determined to be similar to the unknown time series based on the distance between each known time series and the unknown time series using the at least one processor.
[0036] Clause 4: A method according to any one of clauses 1 to 3, wherein the residual network is trained using a loss function defined according to the following equation:
[0037]
[0038] in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series ti+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
[0039] Item 5: A method according to any one of items 1 to 4, wherein the multiple known time series include multiple known transaction time series associated with multiple merchants, and wherein each known time series is associated with metadata, and the metadata is associated with the merchant associated with the known time series.
[0040] Item 6: A method according to any one of items 1 to 5, wherein the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four.
[0041] Clause 7: A method according to any one of clauses 1 to 6, wherein the residual network comprises a two-dimensional residual network.
[0042] Item 8: A system comprising: at least one processor coupled to a memory and configured to: obtain a plurality of known time series from at least one database; for each known time series in the plurality of known time series: calculate a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates to generate a plurality of pairwise distance matrices; stack the plurality of pairwise distance matrices together to generate a tensor; and process the tensor with a residual network, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; and provide the feature vector for each known time series in the plurality of known time series.
[0043] Clause 9: The system of clause 8, wherein the at least one processor is further configured to: store the feature vector for each known time series in the plurality of known time series in the at least one database.
[0044] Item 10: A system according to item 8 or 9, wherein the at least one processor is further configured to: obtain an unknown time series; calculate a pairwise distance matrix between the unknown time series and each of the multiple learned templates to generate another plurality of pairwise distance matrices; stack the other plurality of pairwise distance matrices together to generate another tensor; process the other tensor with the residual network, wherein the residual network receives the other tensor as input and provides a feature vector of the unknown time series as output; for each known time series of the multiple known time series stored in the database, determine the distance between the known time series and the unknown time series based on the stored feature vector of the known time series and the feature vector of the unknown time series; and identify at least one known time series determined to be similar to the unknown time series based on the distance between each known time series and the unknown time series.
[0045] Clause 11: A system according to any one of clauses 8 to 10, wherein the residual network is trained using a loss function defined according to the following equation:
[0046]
[0047] in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
[0048] Clause 12: A system according to any one of clauses 8 to 11, wherein the plurality of known time series comprises a plurality of known transaction time series associated with a plurality of merchants, and wherein each known time series is associated with metadata, the metadata being associated with the merchant associated with the known time series.
[0049] Clause 13: A system according to any one of clauses 8 to 12, wherein the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four.
[0050] Clause 14: A system as described in any of clauses 8 to 13, wherein the residual network comprises a two-dimensional residual network.
[0051] Item 15: A computer program product comprising a non-transitory computer-readable medium, the non-transitory computer-readable medium comprising program instructions, which, when executed by at least one processor, cause the at least one processor to: obtain a plurality of known time series from at least one database; for each known time series in the plurality of known time series: calculate a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates to generate a plurality of pairwise distance matrices; stack the plurality of pairwise distance matrices together to generate a tensor; and process the tensor with a residual network, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; and provide the feature vector for each known time series in the plurality of known time series.
[0052] Clause 16: A computer program product according to clause 15, wherein the program instructions, when executed by the at least one processor, further cause the at least one processor to: store the feature vector of each known time series in the plurality of known time series in the at least one database.
[0053] Item 17: A computer program product according to item 15 or 16, wherein the program instructions, when executed by the at least one processor, further cause the at least one processor to: obtain an unknown time series; calculate a pairwise distance matrix between the unknown time series and each of the multiple learned templates to generate another plurality of pairwise distance matrices; stack the other plurality of pairwise distance matrices together to generate another tensor; process the another tensor with the residual network, wherein the residual network receives the another tensor as input and provides a feature vector of the unknown time series as output; for each known time series of the multiple known time series stored in the database, determine the distance between the known time series and the unknown time series based on the stored feature vector of the known time series and the feature vector of the unknown time series; and identify at least one known time series determined to be similar to the unknown time series based on the distance between each known time series and the unknown time series.
[0054] Clause 18: A computer program product according to any one of Clauses 15 to 17, wherein the residual network is trained using a loss function defined according to the following equation:
[0055]
[0056] in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
[0057] Clause 19: A computer program product according to any one of clauses 15 to 18, wherein the plurality of known time series comprises a plurality of known transaction time series associated with a plurality of merchants, and wherein each known time series is associated with metadata, the metadata being associated with the merchant associated with the known time series.
[0058] Item 20: A computer program product according to any one of items 15 to 19, wherein the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four, and wherein the residual network comprises a two-dimensional residual network.
[0059] These and other features and characteristics of the present disclosure, as well as methods of operation and functions of related structural elements and combinations of parts and economies of manufacture, will become more apparent when considering the following description and appended claims with reference to the accompanying drawings, all of which form a part of this specification, wherein like reference numerals designate corresponding parts in the various figures. It is to be expressly understood, however, that the drawings are for purposes of illustration and description only and are not intended as a definition of the limits of the disclosed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Additional advantages and details are explained in more detail below with reference to a non-limiting exemplary embodiment shown in the schematic drawings, in which:
[0061] Figure 1 is a schematic diagram of an electronic payment processing network according to some non-limiting embodiments or aspects;
[0062] Figure 2 According to some non-limiting embodiments or aspects Figure 1 schematic diagrams of example components of one or more devices;
[0063] Figure 3A and Figure 3B is a flow chart of a method for efficient content-based time series retrieval according to some non-limiting embodiments or aspects;
[0064] Figure 4An electronic payment network use case showing content-based time series retrieval (CTSR) and a database including transaction time series;
[0065] Figure 5 A multi-domain use case showing CTSR and a database including time series from multiple domains;
[0066] Figure 6 Showing examples of feature extractors and distance functions;
[0067] Figure 7 An algorithm for calculating the dynamic time warping (DTW) distance is shown;
[0068] Figure 8 are building blocks and network diagrams of a residual network 2D (RN2D) model according to some non-limiting embodiments or aspects;
[0069] Fig. 9 is a network diagram of a residual network 2DRN2Dw / T with a template learning model according to some non-limiting embodiments or aspects;
[0070] Fig.10 It is the performance measurement table of the experiment;
[0071] Fig.11 is the critical difference (CD) plot for comparing experimental performance;
[0072] Fig.12 is a graph of the performance measurement results of the experiment;
[0073] Fig.13 The time series of the first eight captures of the experiment are shown;
[0074] Fig.14 is a table of further performance measures of the experiment;
[0075] Fig.15 is the CD plot of further performance of the comparative experiment;
[0076] Fig.16 is a graph of the results of another performance measure of the experiment; and
[0077] Fig.17 is the query schedule for the experiment. DETAILED DESCRIPTION
[0078] For the purposes of the following description, the terms "end," "upper," "lower," "right," "left," "vertical," "horizontal," "top," "bottom," "lateral," "longitudinal," and their derivatives shall relate to the orientation of the embodiments in the accompanying drawings. However, it shall be understood that the present disclosure may employ various alternative variations and step sequences, unless expressly specified to the contrary. It shall also be understood that the specific devices and processes shown in the drawings and described in the following specification are merely exemplary and non-limiting embodiments or aspects of the disclosed subject matter. Accordingly, specific dimensions and other physical characteristics relating to the embodiments or aspects disclosed herein shall not be considered limiting.
[0079] Some non-limiting embodiments or aspects may be described herein in conjunction with a threshold value. As used herein, satisfying a threshold value may refer to a value greater than a threshold value, more than a threshold value, above a threshold value, greater than or equal to a threshold value, less than a threshold value, less than a threshold value, below a threshold value, less than or equal to a threshold value, equal to a threshold value, etc.
[0080] Aspects, parts, elements, structures, actions, steps, functions, instructions and / or the like used herein should not be understood as critical or necessary unless explicitly described as such. Moreover, as used herein, the article "one" is intended to include one or more items, and can be used interchangeably with "one or more" and "at least one". In addition, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related items and unrelated items, etc.), and can be used interchangeably with "one or more" or "at least one". In the case of wishing to have only one item, the term "one" or similar language is used. Moreover, as used herein, the term "having" and / or the like is intended to be an open term. In addition, unless otherwise explicitly stated, the phrase "based on" is intended to mean "based at least in part". In addition, the reference to the action of "based on" a condition may refer to the action being "in response to" the condition. For example, in some non-limiting embodiments or aspects, the phrases "based on" and "in response to" may refer to the condition of automatically triggering an action (e.g., a specific operation of an electronic device such as a computing device, a processor).
[0081] As used herein, the term "communication" may refer to the reception, acceptance, transmission, transmission, provision, etc. of data (e.g., information, signals, messages, instructions, commands, etc.). A unit (e.g., a device, a system, a component of a device or system, a combination thereof, etc.) communicating with another unit means that the unit is able to directly or indirectly receive information from the other unit and / or transmit information to the other unit. This may refer to a direct or indirect connection (e.g., a direct communication connection, an indirect communication connection, and / or its analogue) that is wired and / or wireless in nature. In addition, although the information sent may be modified, processed, relayed, and / or routed between the first unit and the second unit, the two units may also communicate with each other. For example, even if the first unit passively receives information and does not actively send information to the second unit, the first unit may communicate with the second unit. As another example, if at least one intermediate unit processes the information received from the first unit and transmits the processed information to the second unit, the first unit may communicate with the second unit. In some non-limiting embodiments or aspects, a message may refer to a network packet (e.g., a data packet and / or its analogue) comprising data. It will be appreciated that many other arrangements are possible.
[0082] As used herein, the term "computing device" may refer to one or more electronic devices configured to process data. In some examples, a computing device may include the necessary components to receive, process, and output data, such as a processor, a display, a memory, an input device, a network interface, and / or the like. The computing device may be a mobile device. As an example, a mobile device may include a cellular phone (e.g., a smartphone or a standard cellular phone), a portable computer, a wearable device (e.g., a watch, glasses, lenses, clothing, and / or the like), a personal digital assistant (PDA), and / or other similar devices. The computing device may also be a desktop computer or other form of non-mobile computer.
[0083] As used herein, the term "server" may refer to or include one or more computing devices operated by or facilitating communications and processing by multiple parties in a network environment such as the Internet, but it should be understood that communications may be facilitated through one or more public or private network environments, and that various other arrangements are possible. In addition, multiple computing devices (e.g., servers, point of sale (POS) devices, mobile devices, etc.) that communicate directly or indirectly in a network environment may constitute a "system."
[0084] As used herein, the term "system" may refer to one or more computing devices or a combination of computing devices (e.g., processors, servers, client devices, software applications, components of such devices, and / or the like). As used herein, references to a "device," "server," "processor," and / or the like may refer to a previously described device, server, or processor that is described as performing a previous step or function, a different device, server, or processor, and / or a combination of devices, servers, and / or processors. For example, as used in the specification and claims, a first device, first server, or first processor that is described as performing a first step or a first function may refer to the same or a different device, server, or processor that is described as performing a second step or a second function.
[0085] As used herein, the term "real-time" refers to performing one or more tasks during or before another process is completed. For example, real-time reasoning can be reasoning obtained from a model before a payment transaction is authorized, completed, etc.
[0086] Time series are a common data type analyzed for a variety of applications. For example, engineers may examine time series from different sensors on manufacturing machines to identify ways to improve factory efficiency, doctors may study various biometric time series for medical research, and multiple time series streams from operating payment networks may be monitored to discover abnormal activities. With the large amount of time series data available from various sources, an effective content-based time series retrieval (CTSR) system is needed to help users navigate the time series database.
[0087] Figure 4 An electronic payment network use case showing CTSR and a database including transaction time series. To understand what the CTSR system is and how it can help users, consider Figure 4 Here, one of the merchants using the electronic payment network fails to provide accurate business type information. When a merchant uses the electronic payment network, the payment processing company can obtain a time series identification signature about the merchant. Subsequently, an investigator from the company can use the CTSR system with time series identification signatures from various merchants to identify the correct business type of the relevant merchant. The CTSR system can help the investigator correct the information in a timely manner.
[0088] exist Figure 4 In the above examples shown, the CTSR system only includes transaction time series. However, it is also possible to build a CTSR system with time series from various domains, such as Figure 5, which shows a multi-domain use case of CTSR and a database that includes time series from multiple domains. Assume that a user encounters a time series that does not have any associated metadata. The time series may be a power consumption time series or a data record from another sensor. The user may want to identify the possible source of the time series and recover the lost information. To solve this problem, the user may query the CTSR system with time series (which may or may not exist in the CTSR system's database), and the system may return a sorted list of similar time series with associated metadata. In this example, five of the first six returned time series are power consumption signatures for microwave ovens. Therefore, the user may be able to infer that the unknown time series is most likely a power consumption signature for microwave ovens. Therefore, the CTSR system can help the user recover the lost information about the time series.
[0089] Design goals when building a CTSR system may include: 1) effectively capturing various concepts in time series from different domains, and 2) remaining efficient during inference when the user interacts with the system in real time. The reason for the difference in inference time between CTSR systems may be the difference in the role of the neural network model. Figure 6 An example of a feature extractor and a distance function is shown. In the faster method, the neural network acts only as a feature extractor, and the distance is calculated using the Euclidean distance function, such as Figure 6 As shown in the example (a) of . Therefore, before inference, each time series in the database can be projected into Euclidean space only once using a neural network. During query time, the neural network model may only need to project the query time series into the same Euclidean space, and can efficiently perform distance calculations in this space. On the other hand, the existing residual network 2D (RN2D) model acts as both a feature extractor and a distance function, such as Figure 6 As shown in the example (b) of . Therefore, the RN2D model is called every time the distance is calculated. In the case where there is a time series in the database, the faster method only needs to call the neural network model once for query. In contrast, the RN2D model is called multiple times, which greatly increases the running time.
[0090] Non-limiting embodiments or aspects of the present disclosure provide a method, system, and computer program product for content-based time series retrieval, which perform the following operations: obtain multiple known time series from at least one database; for each known time series in the multiple known time series: calculate a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates to generate multiple pairwise distance matrices; stack the multiple pairwise distance matrices together to generate a tensor; process the tensor with a residual network, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; and provide a feature vector for each known time series in the multiple known time series. Therefore, non-limiting embodiments or aspects of the present disclosure provide methods, systems and computer program products for content-based time series retrieval, which perform the following operations: being able to obtain an unknown time series; calculating a pairwise distance matrix between the unknown time series and each of a plurality of learned templates to generate another plurality of pairwise distance matrices; stacking the other plurality of pairwise distance matrices together to generate another tensor; processing the other tensor with a residual network, wherein the residual network receives the other tensor as input and provides a feature vector of the unknown time series as output; for each known time series in a plurality of known time series stored in a database, determining the Euclidean distance between the known time series and the unknown time series based on the stored feature vector of the known time series and the feature vector of the unknown time series; and identifying at least one known time series determined to correspond to the unknown time series based on the Euclidean distance between each known time series and the unknown time series.
[0091] In this way, non-limiting embodiments or aspects of the present disclosure may provide an improved model architecture based on an RND2D model with improved efficiency, which may be referred to herein as a residual network 2D with template learning (RN2Dw / T). Figure 6 As shown in example (c) of FIG. 1 , the RN2Dw / T model according to a non-limiting embodiment or aspect may incorporate a template (landmark) learning mechanism into the input ( Figure 6 ) and modify the model to output a feature vector instead of a distance value. The RN2Dw / T model according to a non-limiting embodiment or aspect may use the learned landmarks as a reference to generate a feature vector for an input time series. Unlike the residual network 2D (RN2D) method introduced herein below, the RN2Dw / T model according to a non-limiting embodiment or aspect may only act as a feature extractor. Non-limiting embodiments or aspects of the RN2Dw / T model may achieve comparable effectiveness to the RN2D method while achieving an average query time of less than 0.04 seconds (see, e.g., Fig.10). Thus, non-limiting embodiments or aspects of the present disclosure enable an effective and efficient CTSR system that can be a valuable tool for businesses in various industries.
[0092] Reference Figure 1 , Figure 1 An electronic payment processing network 100 is shown according to a non-limiting embodiment or aspect. The payment processing network can be used in conjunction with the systems and methods described herein. It should be understood that the specific arrangement of the electronic payment processing network 100 shown is for exemplary purposes only, and various arrangements are possible. The transaction processing system 101 (e.g., a transaction processor) is shown as communicating with one or more issuer systems (e.g., such as an issuer system 106) and one or more acquirer systems (e.g., such as an acquirer system 108). Although only a single issuer system 106 and a single acquirer system 108 are shown, it should be understood that the transaction processing system 101 can communicate with multiple issuer systems and / or acquirer systems. In some embodiments, the transaction processing system 101 can also work as an issuer system, so that the transaction processing system 101 and the issuer system 106 are both single systems and / or controlled by a single entity.
[0093] In some non-limiting embodiments or aspects, the transaction processing system 101 can communicate directly with the merchant system 104 through a public or private network connection. Additionally or alternatively, the transaction processing system 101 can communicate with the merchant system 104 through a payment gateway 102 and / or an acquirer system 108. In some non-limiting embodiments or aspects, the acquirer system 108 associated with the merchant system 104 can operate as a payment gateway 102 to facilitate the transmission of transaction requests from the merchant system 104 to the transaction processing system 101. The merchant system 104 can communicate with the payment gateway 102 through a public or private network connection. For example, a merchant system 104 including a physical POS device can communicate with the payment gateway 102 through a public or private network to conduct card-present transactions. As another example, a merchant system 104 including a server (e.g., a web server) can communicate with the payment gateway 102 through a public or private network such as a public Internet connection to conduct card-not-present transactions.
[0094] In some non-limiting embodiments or aspects, upon receiving a transaction request from a merchant system 104 that identifies an account identifier of a payee (e.g., such as an account holder) associated with an issued consumer device 110, the transaction processing system 101 may generate an authorization request message to be transmitted to the issuer system 106 that issued the consumer device 110 and / or the account identifier. The issuer system 106 may then approve or deny the authorization request and, based on the approval or denial, generate an authorization response message that is transmitted to the transaction processing system 101. The transaction processing system 101 may transmit the approval or denial to the merchant system 104. When the issuer system 106 approves the authorization request message, it may then clear and settle the payment transaction between the issuer system 106 and the acquirer system 108.
[0095] Figure 1 The number and arrangement of systems and / or devices shown in A are provided as examples. Figure 1 There may be additional systems and / or devices, fewer systems and / or devices, different systems and / or devices, and / or systems and / or devices arranged in a different manner than the systems and / or devices shown in the drawings. Furthermore, a single system and / or device may be implemented Figure 1 Two or more systems or devices shown in Figure 1 The single system or device shown in the embodiment may be implemented as multiple distributed systems or devices. Additionally or alternatively, a group of systems (e.g., one or more systems) and / or a group of devices (e.g., one or more devices) of system 100 may perform one or more functions described as being performed by another group of systems or another group of devices of system 100.
[0096] Reference Figure 2 , which shows a diagram of example components of a device 200 according to a non-limiting embodiment. As an example, the device 200 may correspond to a transaction processing system 101, a payment gateway 102, a merchant system 104, an issuer system 106, an acquirer system 108, and / or a consumer device 110. In some non-limiting embodiments, such a system or device may include at least one device 200 and / or at least one component of the device 200. The number and arrangement of the components shown are provided as examples. In some non-limiting embodiments, the device 200 may include additional components, fewer components, different components, or differently arranged components compared to those shown. Additionally or alternatively, a group of components (e.g., one or more components) of the device 200 may perform one or more functions described as being performed by another group of components of the device 200.
[0097] like Figure 2As shown, the device 200 may include a bus 202, a processor 204, a memory 206, a storage component 208, an input component 210, an output component 212, and a communication interface 214. The bus 202 may include components that allow communication between components of the device 200. In some non-limiting embodiments, the processor 204 may be implemented in hardware, firmware, or a combination of hardware and software. For example, the processor 204 may include a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), etc.), a microprocessor, a digital signal processor (DSP), and / or any processing component that can be programmed to perform a function (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.). The memory 206 may include a random access memory (RAM), a read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, optical memory, etc.) that stores information and / or instructions for use by the processor 204.
[0098] Continue to refer Figure 2 , the storage component 208 may store information and / or software related to the operation and use of the device 200. For example, the storage component 208 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, a solid-state disk, etc.) and / or another type of computer-readable medium. The input component 210 may include a component that allows the device 200 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, etc.). In addition or alternatively, the input component 210 may include a sensor for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, an actuator, etc.). The output component 212 may include a component that provides output information from the device 200 (e.g., a display, a speaker, one or more light-emitting diodes (LEDs), etc.). The communication interface 214 may include a transceiver-type component (e.g., a transceiver, a separate receiver and transmitter, etc.) that enables the device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. The communication interface 214 may allow the device 200 to receive information from another device and / or provide information to another device. For example, the communication interface 214 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, Interface, cellular network interface, etc.
[0099] Device 200 can perform one or more processes described herein. Device 200 can perform these processes based on processor 204 executing software instructions stored by computer-readable media such as memory 206 and / or storage component 208. Computer-readable media may include any non-transient memory device. Memory device includes a memory space located in a single physical storage device or a memory space extended across multiple physical storage devices. Software instructions can be read from another computer-readable medium or from another device to memory 206 and / or storage component 208 via communication interface 214. When executed, software instructions stored in memory 206 and / or storage component 208 can enable processor 204 to perform one or more processes described herein. In addition or alternatively, hard-wired circuit system can be used in place of or in combination with software instructions to perform one or more processes described herein. Therefore, the embodiments described herein are not limited to any specific combination of hardware circuit system and software. As used herein, the term "configured to" can refer to the arrangement of software, equipment and / or hardware for performing and / or realizing one or more functions (such as actions, processes, steps of processes and / or the like). For example, "a processor configured to" may refer to a processor executing software instructions (eg, program code) that cause the processor to perform one or more functions.
[0100] The following conventions may be used in this document for notation: lowercase letters (e.g., x) may denote scalars, and bold lowercase letters (e.g., ) denotes a vector, an uppercase letter (e.g., X) denotes a matrix, and bold uppercase letters (e.g., ) can represent tensors, and calligraphic letters (e.g., ) can represent a set.
[0101] The content-based time series retrieval (CTSR) problem can be formulated as follows: given a set of time series and any query time series In the case of , we obtain the correlation score function (·,·), which satisfies if Compare and More relevant, The scoring function can be a predefined similarity / distance function, or use the same Metadata associated with each time series in an optimized trainable function.
[0102] The time series extraction problem can be formulated in two ways. The first is also called the time series similarity search problem, where the goal is to find the previous time series that is most similar to a given query based on a fixed distance function. Because the distance function is fixed, this type of research focuses on efficiency, with acceleration achieved through techniques such as lower bounds, early abandonment, and / or indexing. If this problem is compared to the above problem statement, it can be seen that the technical goals for solving the time series similarity search problem are different from the technical goals for solving the above problem statement.
[0103] The second type of problem formulation is more consistent with the problem formulation used to solve the above problem statement, where the goal is to develop a model or scoring function to help users retrieve relevant time series from a database based on a submitted query time series. However, existing models for solving this second type of problem formulation are designed to solve multivariate time series, which, if applied to the above problem statement, will simply reduce to a standard long short-term memory network.
[0104] Euclidean distance and dynamic time warping distance are popular and simple tools for analyzing time series data. They are widely used in various tasks such as similarity search, classification, and anomaly detection, and both distance functions can be easily applied to the above problems. Another class of methods that can be applied to the above problems is neural networks, especially neural networks that can model sequence data. For example, long short-term memory networks, gated recurrent unit networks, transformers, and convolutional neural networks have shown effectiveness in tasks such as time series classification, prediction, and anomaly detection.
[0105] Six existing baseline methods are now presented. Thereafter, the previously mentioned RN2D method is introduced, and the benefits of the RN2D method are compared with the benefits of the other baseline methods. After the RN2D method is introduced, further details are provided about the residual network 2D with template learning (RN2Dw / T) method according to non-limiting embodiments or aspects, which solves the efficiency issues associated with the design of RN2D.
[0106] The six existing baseline methods considered include Euclidean distance (ED), dynamic time warping (DTW), long short-term memory (LSTM), gated recurrent unit (GRU), transformer (TF), and residual network 1D (RN1D).
[0107] The Euclidean distance between the query time series and the time series in the ensemble can be calculated. The ensemble can then be classified based on the distance. This is probably the simplest way to solve the CTSR problem.
[0108] DTW is similar to the ED baseline, but uses the DTW distance instead. The DTW distance is considered a simple but effective baseline for time series classification problems.
[0109] LSTM is one of the most popular recurrent neural networks (RNNs) for modeling sequence data. LSTM models can be optimized using the Siamese network architecture (see, for example, Figure 4 Example (a) of the Siamese network. The Siamese network takes two input time series and each input is first processed with a 1D convolutional layer to extract local features. The output is then fed into a bidirectional LSTM model to obtain a hidden representation. Next, the hidden representation of the last time step is passed through a linear layer to obtain the final representation of the input time series. The correlation score between the two inputs is calculated using the Euclidean distance between the final representations.
[0110] GRU is another popular RNN architecture widely used to model sequential data. To optimize the GRU model, a similar approach to the LSTM model can be applied, where the LSTM units in the RNN architecture are replaced by GRU units.
[0111] TF is an alternative to RNN for sequence modeling. To learn hidden representations of input time series, we can use the TF-CNN framework by Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan NGomez, The transformer encoder was proposed by Kaiser and Illia Polosukhin in the paper titled “Attention is all you need. Advances in neural information processing systems” in 2017. The RNN used in the previous two methods (i.e., LSTM and GRU) can be replaced with the transformer encoder, resulting in a transformer-based twin network architecture instead of an RNN-based architecture.
[0112] RN1D is a time series classification model inspired by the success of residual networks in computer vision. RN1D uses 1D convolutional layers instead of 2D convolutional layers. Extensive evaluation has proven that the RN1D design is one of the strongest models for time series classification. The RN1D model can also be optimized in a Siamese network (see, e.g., Figure 4 Example (a)).
[0113] Each of the ED and DTW methods does not require a training phase because there are no parameters to optimize in either method. The DTW method is the more efficient of the two methods for time series data because the DTW method considers all alignments between the input time series. The calculation of the DTW distance can be abstracted as a two-stage process, such as Figure 7 In the first stage (lines 2 to 5), the input time series (W is length) and (where h is length) to calculate the pairwise distance matrix Because D[i,j]=|a i -b j |. In the second stage (lines 6 to 8), for each element in D, a fixed loop function is applied to D (i.e., D[i,j]←D[i,j]+min(D[i-1,j],D[i,j-1],D[i-1,j-1])). Therefore, the DTW method can be viewed as running a predefined function on the pairwise distance matrix between the input time series.
[0114] The remaining four baseline methods use the Siamese network distance learning framework (see e.g., Figure 4 Example (a) in ), and a high-capacity (e.g., high-expressiveness, etc.) neural network model (i.e., LSTM, GRU, TF, and RN1D) is used to learn hidden representations of the input time series. These representations are used to calculate the distance between two time series, and the model is learned using the optimization procedure described in this article. Once the model is optimized, the hidden representation of each time series in the database is extracted before deployment. When a user submits a query, the model can be applied only to the query time series to extract its hidden representation, because the hidden representation of each time series in the database may have been extracted before the query time. Then, the Euclidean distance can be used to calculate the distance between the query and each item in the database.
[0115] Reference Figure 8 , Figure 8 are building blocks and network diagrams of RN2D models according to some non-limiting embodiments or aspects. The RN2D model utilizes rich alignment information from a pairwise distance matrix, similar to the DTW method. However, instead of using a fixed function, the RN2D model can use a high-capacity neural network as a function, thereby using a similar representation model to the four neural network baseline.
[0116] The design of RN2D is motivated by deep residual networks used in computer vision. Figure 8As shown in , non-limiting embodiments or aspects of the RN2D model may employ a bottleneck building block design as described in a paper titled “Deep residual learning for image recognition” by Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, 2016, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770-778, the entire disclosure of which is incorporated herein by reference in its entirety. Given an input dimension n in , bottleneck dimension n neck and output dimension n out In the case of We can first use a 1×1 convolutional layer to project The tensor is then passed through a ReLU layer and then further transformed to a 3×3 convolutional layer with a stride of two. After another ReLU layer, the intermediate representation can be projected to space, where the output of the 1×1 convolutional layer is called X out . Due to X in and X out The sizes of X do not match, so in With X out may not be added directly for skip connections, and X in It can be added to X out is processed with a 1×1 convolutional layer before. After adding, the combined representation can be processed with ReLU and exit the building block. If the input is space, the output will be In space.
[0117] Still refer to Figure 8 , the overall network design of RN2D can also be similar to the network design described in the paper entitled “Deep residual learning for image recognition” by Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778, the entire disclosure of which is incorporated herein by reference in its entirety. Given two input time series a=[a1 ,···,a w ] and b=[b 1 ,···,b h ], the pairwise distance matrix can be calculated similarly to the DTW method The i-th and j-th positions of D are available [i,j] = |a i -b j | Calculation. Before applying the convolutional layer, the shape of D can be transformed to w×h×1 by adding an extra dimension. Next, a 7×7 convolutional layer with a stride of 2 can be used to project D onto space. After the ReLU layer, the intermediate representation may be passed through eight building blocks with a 64→16→64 setting. A global average pooling layer may then be applied to reduce the spatial dimension, and the output of the global average pooling layer may include a vector of size sixty-four. Finally, the linear layer may project the vector to a scalar number, which may include a correlation score between the two input time series. In some non-limiting embodiments or aspects, the plurality of learned templates includes thirty-two learned templates, wherein the plurality of pairwise distance matrices includes thirty-two pairwise distance matrices, wherein the tensor includes an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series includes a vector of size sixty-four.
[0118] like Figure 6 As shown in example (b), unlike the method using the twin network framework, when using RN2D, the hidden representation of each time series in the database may not be extracted before deployment. If there are n time series in the database, RN2D may be run n times during the query time to calculate the distance between the query time series and each time series in the database. In contrast, the method using the twin network framework only needs to run the model once during the query time, making the RN2D method an order of magnitude slower. In order to solve the efficiency problem of RN2D, non-limiting embodiments or aspects of the present disclosure provide residual network 2D and RN2Dw / T.
[0119] Reference Fig. 9 , Fig. 9 is a network diagram of an RN2Dw / T model according to some non-limiting embodiments or aspects. The RN2Dw / T model according to some non-limiting embodiments or aspects can solve the efficiency problem of RN2D. The RN2Dw / T method according to some non-limiting embodiments or aspects can be designed to be as effective as the RN2D method while being an order of magnitude faster.
[0120] like Fig. 9As shown in , the RN2Dw / T model according to some non-limiting embodiments or aspects may differ from the RN2D model in the following four ways: (1) the last linear layer of the RN2Dw / T model may output a vector instead of a scalar as in the RN2D model; (2) the RN2Dw / T model may take a single time series as input, while the RN2D model takes a pair of time series as input; (3) for the RN2Dw / T model, multiple pairwise distance matrices (e.g., 32 pairwise distance matrices, etc.) between the input time series and multiple templates (e.g., 32 templates, etc.) may be calculated, while for the RN2D model, only one pairwise distance matrix between two input time series is calculated; and (4) the input dimension of the first 2D convolutional layer in the RN2Dw / T model may include multiple dimensions (e.g., 32 dimensions, etc.), while in the RN2D model, the input dimension is only one.
[0121] Here, the first two differences between the models may exist because the RN2Dw / T model aims to extract the feature vector of the input time series, while the RN2D model calculates the correlation score between the two input time series.
[0122] The third difference is in the pairwise distance matrix calculation step, which is also the reason why the RN2Dw / T model according to some non-limiting embodiments or aspects is much faster than the RN2D model. The pairwise distance matrix can be calculated as follows: Given an input time series a = [a 1 ,···,a w ] and the kth template tk = [tk,1,···,tk,w], the kth pairwise distance matrix We can use Dk[i,j]=|a i –t k,j | to calculate. A pairwise distance matrix for each of a plurality of templates (e.g., 32 templates, etc.) may be calculated, thereby generating a plurality of w×h matrices (e.g., 32 w×h matrices, etc.). A plurality of templates (e.g., 32 templates, etc.) may be learned during the training phase and may include a reference time series that helps the model project the input time series into Euclidean space using a 2D convolutional design. Then, a plurality of w×h matrices (e.g., 32 w×h matrices, etc.) may be stacked together to form a w×h×32 tensor of the first 2D convolutional layer. The w×h×32 tensor may be the output of the pairwise distance matrix calculation step of the RN2Dw / T model.
[0123] The fourth difference between the two models may correspond to the fact that the input tensor of the first convolutional layer of the RN2Dw / T model may be w×h×32, while the input tensor of the first convolutional layer in the RN2D model is w×h×1.
[0124] like Figure 6As shown in example (c) of , the RN2Dw / T model can be used to extract feature vectors for each time series in the database before query time. In this way, when a user submits a query time series, the non-limiting embodiments or aspects of the present disclosure may only need to run the model once for the time series. Although each of the RN2Dw / T and RN2D models has similar capacity, the RN2Dw / T model enables a more efficient query mechanism, which is advantageous when designing real-world CTSR systems.
[0125] In some non-limiting embodiments or aspects, a Bayesian personalized ranking loss as described in a paper titled “BPR: Bayesian personalized ranking from implicit feedback” by Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme in 2009, Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, pages 452-461, the entire disclosure of which is incorporated herein by reference in its entirety, may be used to train or optimize the RN2Dw / T model. The Bayesian personalized ranking loss is applicable to the CTSR problem because the CTSR problem is a “Learning to Rank” problem. Given a batch of training data, In the case of , the loss function can be defined according to the following equation (1):
[0126]
[0127] in For training data batches m is the batch size, and each sample in the batch includes the query (or anchor) time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ(·,·) is a residual network or model. In some non-limiting embodiments or aspects, the AdamW optimizer as described in a paper titled “Decoupled Weight Decay Regularization” by Ilya Loshchilov and Frank Hutter at the International Conference on Learning Representations in 2018, the entire disclosure of which is incorporated herein by reference in its entirety, may be used to train RN2Dw / T using a Bayesian personalized ranking loss.
[0128] Reference Figure 3A and Figure 3B , which shows a flowchart of a method 300 for efficient content-based time series retrieval according to some non-limiting embodiments or aspects. Figure 3A and Figure 3B The steps shown in are for illustrative purposes only. It will be appreciated that in some non-limiting embodiments or aspects, additional, fewer, different and / or different order of steps may be used. In some non-limiting embodiments or aspects, steps may be automatically performed in response to the execution and / or completion of previous steps.
[0129] like Figure 3A As shown in FIG. 3 , at step 302 , method 300 includes obtaining a plurality of known time series from at least one database. For example, transaction processing system 101 may obtain a plurality of known time series from at least one database. The plurality of known time series may be associated with or appear from a plurality of different data domains.
[0130] In some non-limiting embodiments or aspects, the plurality of known time series include a plurality of known transaction time series associated with a plurality of merchants, and each known time series is associated with metadata associated with the merchant associated with the known time series. For example, the known time series may include a time series identification mark indicating that the merchant system 104 uses the electronic payment processing network 100. As an example, the known (or unknown) time series may include transaction data associated with a plurality of transactions and / or a plurality of time points. As an example, a payment transaction may include transaction parameters and / or features associated with the payment transaction. The transaction parameters and / or features (e.g., category features, numerical features, local features, graphical features, or embeddings, etc.) associated with the payment transaction may include transaction parameters of the transaction, features determined based on the transaction parameters (e.g., using feature engineering, etc.), such as an account identifier (e.g., PAN, etc.), a transaction amount, a transaction date and / or time, a type of product and / or service associated with the transaction, a currency conversion rate, a currency type, a merchant type, a merchant name, a merchant location, etc. However, non-limiting embodiments or aspects are not so limited, and transaction parameters and / or characteristics of a transaction may include any data, including any type of parameters associated with any type of transaction.
[0131] like Figure 3A As shown in FIG. 3 , in step 304, the method 300 includes: for each known time series in a plurality of known time series, calculating a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates to generate a plurality of pairwise distance matrices. For example, for each known time series in a plurality of known time series, the transaction processing system 101 may calculate a pairwise distance matrix between the known time series and each learned template in a plurality of learned templates to generate a plurality of pairwise distance matrices. As an example, the transaction processing system 101 may calculate the pairwise distance matrix as follows: Given an input time series a=[a 1 ,···,a w ] and the kth template tk = [tk,1,···,tk,w], the kth pairwise distance matrix We can use Dk[i,j]=|a i -t k,j | to calculate. A pairwise distance matrix for each of a plurality of templates (e.g., 32 templates, etc.) may be calculated, thereby generating a plurality of w×h matrices (e.g., 32 w×h matrices, etc.). A plurality of templates (e.g., 32 templates, etc.) may be learned during a training phase and may include a reference time series that helps the model project the input time series into a Euclidean space using a 2D convolutional design.
[0132] like Figure 3AAs shown in , at step 306, the method 300 includes: for each known time series in the plurality of known time series, stacking a plurality of pairwise distance matrices together to generate a tensor. For example, for each known time series in the plurality of known time series, the transaction processing system 101 may stack a plurality of pairwise distance matrices together to generate a tensor. As an example, the transaction processing system 101 may stack a plurality of w×h matrices (e.g., 32 w×h matrices, etc.) together to form a w×h×32 tensor of the first 2D convolutional layer. The w×h×32 tensor may be an output of the pairwise distance matrix calculation step of the RN2Dw / T model.
[0133] like Figure 3A As shown in , at step 308, the method 300 includes: for each known time series in a plurality of known time series, processing the tensor with a residual network. For example, for each known time series in a plurality of known time series, the transaction processing system 101 may process the tensor with a residual network, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output. As an example, the residual network may receive the tensor as input and may provide a feature vector of the known time series as output. In such an example, and again referring to Fig. 9 , the residual network may include a method for projecting D onto space (for example, project D onto The output of the global average pooling layer can be multi-dimensional (e.g., a vector of size sixty-four, etc.), and the last linear layer of the residual network can output a vector instead of a scalar as in the RN2D model.
[0134] like Figure 3A As shown in , at step 310, the method 300 includes providing and / or storing a feature vector for each of the plurality of known time series. For example, the transaction processing system 101 may provide and / or store a feature vector for each of the plurality of known time series. As an example, the transaction processing system 101 may provide a feature vector for each of the plurality of known time series. In such examples, the transaction processing system 101 may store the feature vector for each of the plurality of known time series in at least one database.
[0135] like Figure 3A As shown in , at step 312, method 300 includes obtaining an unknown time series. For example, transaction processing system 101 may obtain the unknown time series.
[0136] In some non-limiting embodiments or aspects, the unknown time series includes an unknown transaction time series associated with a merchant and / or includes metadata associated with the merchant. For example, the unknown time series may include a time series identification mark indicating that the merchant system 104 uses the electronic payment processing network 100. As an example, the unknown (or known) time series may include transaction data associated with multiple transactions and / or multiple time points. As an example, a payment transaction may include transaction parameters and / or features associated with the payment transaction. The transaction parameters and / or features associated with the payment transaction (e.g., category features, numerical features, local features, graphic features, or embedding, etc.) may include transaction parameters of the transaction, features determined based on it (e.g., using feature engineering, etc.), and / or the like, such as an account identifier (e.g., PAN, etc.), a transaction amount, a transaction date and / or time, a type of product and / or service associated with the transaction, a currency conversion rate, a currency type, a merchant type, a merchant name, a merchant location, and / or the like. However, the non-limiting embodiments or aspects are not limited thereto, and the transaction parameters and / or features of the transaction may include any data, including any type of parameters associated with any type of transaction.
[0137] like Figure 3B As shown in FIG. 3 , at step 314, the method 300 includes calculating a pairwise distance matrix between the unknown time series and each of the learned templates in the plurality of learned templates to generate another plurality of pairwise distance matrices. For example, the transaction processing system 101 may calculate a pairwise distance matrix between the unknown time series and each of the learned templates in the plurality of learned templates to generate another plurality of pairwise distance matrices. As an example, the transaction processing system 101 may calculate another plurality of pairwise distance matrices between the unknown time series and each of the learned templates as follows: Given as an input time series a=[a 1 ,···,a w ] and the kth template tk = [tk,1,···,tk,w], the kth pairwise distance matrix We can use Dk[i,j]=|a i -t k,j | to calculate. Another pairwise distance matrix may be calculated for each of a plurality of templates (eg, 32 templates, etc.), thereby generating another plurality of w×h matrices (eg, 32 w×h matrices, etc.).
[0138] like Figure 3BAs shown in , at step 316, the method 300 includes stacking another plurality of pairwise distance matrices together to generate another tensor. For example, the transaction processing system 101 may stack another plurality of pairwise distance matrices together to generate another tensor. As an example, the transaction processing system 101 may stack another plurality of w×h matrices (e.g., 32 w×h matrices, etc.) together to form another w×h×32 tensor of the first 2D convolutional layer. The other w×h×32 tensor may be the output of the pairwise distance matrix calculation step of the RN2Dw / T model.
[0139] like Figure 3B As shown in FIG. 3 , at step 318, method 300 includes processing another tensor with a residual network, wherein the residual network receives the another tensor as input and provides a feature vector of the unknown time series as output. For example, transaction processing system 101 may process another tensor with a residual network. As an example, the residual network may receive another tensor as input and may provide a feature vector of the unknown time series as output. In such an example, and again referring to Fig. 9 , the residual network may include a method for projecting D onto space (for example, project D onto The output of the global average pooling layer can be multi-dimensional (e.g., a vector of size sixty-four, etc.), and the last linear layer of the residual network can output a vector instead of a scalar as in the RN2D model.
[0140] like Figure 3B As shown in , in step 320, the method 300 includes: for each known time series in a plurality of known time series stored in a database, determining the distance between the known time series and the unknown time series based on the stored feature vector of the known time series and the feature vector of the unknown time series. For example, for each known time series in a plurality of known time series stored in a database, the transaction processing system 101 may determine the distance between the known time series and the unknown time series based on the stored feature vector of the known time series and the feature vector of the unknown time series. As an example, for each known time series in a plurality of known time series stored in a database, the transaction processing system 101 may determine the Euclidean distance between the known time series and the unknown time series based on the stored feature vector of the known time series and the feature vector of the unknown time series.
[0141] like Figure 3BAs shown in , at step 322, the method 300 includes identifying at least one known time series determined to correspond to the unknown time series based on the distance between each known time series and the unknown time series. For example, the transaction processing system 101 may identify at least one known time series determined to correspond to the unknown time series based on the distance between each known time series and the unknown time series. As an example, the transaction processing system 101 may identify a predetermined or expected number of known time series that are closest to the unknown time series based on the distance (e.g., Euclidean distance, etc.) between each known time series and the unknown time series. As an example, the transaction processing system 101 may identify one or more known time series that are within a threshold distance from the unknown time series based on the distance (e.g., Euclidean distance, etc.) between each known time series and the unknown time series. In such examples, the transaction processing system 101 may identify the type of merchant (e.g., restaurant, grocery store, etc.) associated with the unknown time series based on the at least one known time series determined to correspond to the unknown time series, for example, if the merchant fails to provide accurate business type information.
[0142] experiment
[0143] In this section, we present experimental results on the CTSR benchmark dataset created from the UCR archive and a transaction dataset based on a real business problem (see e.g. Figure 4 , etc.). Neural network-based methods were implemented using PyTorch, and the model with the best average NDCG@10 score on the validation data was selected for testing. SciPy was used to calculate ED, and Tslearn was used to calculate DTW.
[0144] The CTSR benchmark dataset is created from the UCR archive, which is a collection of 128 time series classification datasets from various domains such as sports, power demand, and traffic. The UCR archive is widely used to benchmark time series classification algorithms. In order to convert the UCR archive to the CTSR benchmark dataset, use the following steps.
[0145] (1) Extract all time series data from each dataset and merge any identical time series that appear in multiple datasets. This step is performed because the same time series may exist in multiple datasets. For example, the same time series exists in the DodgerLoopDay, DodgerLoopGame, and DodgerLoopWeekend datasets because these datasets are created from the same set of time series.
[0146] (2) To ensure consistency of length, the length of each time series is normalized to 512. For time series longer than 512 time steps, the resampling function from the SciPy library is used to shorten the longer time series. Conversely, for shorter time series, the shorter time series are zero-filled to 512 time steps.
[0147] (3) Apply z-normalization to all time series. Ignore padded zeros during normalization. The z-normalization step is a standard procedure for preparing time series data.
[0148] (4) The ground truth labels for each pair of time series are generated by determining whether the pair of time series are related or unrelated. If two time series belong to the same dataset and share the same class label in the original UCR archive, the two time series are considered to be related. Otherwise, the two time series are considered to be unrelated.
[0149] (5) The data is split into three groups: training, testing, and validation. Specifically, 10% of the time series are randomly selected as test queries, another 10% of the time series are used as validation queries, and the remaining time series are used as training data. In order to ensure that each test / validation query has a sufficient number of related time series in the training set, any query time series with less than two related time series is transferred to the training set. After this data splitting procedure, 136,377 training time series, 17,005 test queries, and 17,005 validation queries are obtained.
[0150] (6) To facilitate efficient evaluation, 1,000 time series are sampled from the training set for each test / validation query. For a given query, if the number of relevant time series for the query is less than 100, all relevant time series are selected. If the number of relevant time series exceeds 100, 100 relevant time series are randomly selected. Ensure that the number of relevant time series in each sampled set is less than or equal to 100 (i.e., 10% of 1,000). For the remaining time series, randomly sample the remaining time series from unrelated time series in the training set.
[0151] (7) To measure the performance of different retrieval methods, common information retrieval metrics are calculated for each query, including precision at k (Prec@k), average precision at k (AP@k), and normalized discounted cumulative gain at k (NDCG@k).
[0152] Fig.10 The performance measurements for each of the 17,005 test queries at k=10 are averaged and presented in Fig.10The results are shown in the table. When comparing performance, a two-sample t-test (with α=0.05) using non-aggregate performance measures was performed to test for statistical significance. The query time reported is the average time taken to calculate the correlation score between the query and the 136,377 time series in the training dataset. The average query time was calculated by using 1,000 different time series from the test data as queries. Fig.10 The table allows easy comparison of different methods based on different performance measures.
[0153] First, the following three performance measures are discussed: PREC@10, AP@10, and NDCG@10. When comparing the performance of two non-neural network baselines (ED and DTW), it is observed that DTW significantly outperforms ED in all three performance measures. This suggests that the use of alignment information helps solve the CTSR problem, and similar conclusions have been drawn for time series classification problems.
[0154] When considering the top four neural network baselines (i.e., LSTM, GRU, TF, and RN1D), each of them significantly outperforms the DTW method, indicating that using high-capacity models helps solve the CTSR problem. One possible reason for this is that the CTSR dataset consists of time series from many different domains, and a higher-capacity model is required to learn different patterns within the data. Among the four methods, LSTM significantly outperforms the second best in all three performance measures.
[0155] According to the t-test results, the RN2D method, a high-capacity model utilizing alignment information, significantly outperforms all other methods. When the RN2Dw / T method according to a non-limiting embodiment or aspect is compared with the RN2D method, the former achieves higher performance in all three performance measures, but the differences are not significant. Therefore, each of the RN2Dw / T method and the RN2D method according to a non-limiting embodiment or aspect can be regarded as a method with better performance for the CTSR dataset in terms of the three performance measures.
[0156] When considering query time, the eight tested methods can be grouped into two categories: slower methods (i.e., DTW and RN2D) with query time exceeding 30 seconds, and faster methods (i.e., ED, LSTM, GRU, TF, RN1D, and RN2Dw / T) with each query taking less than 100 milliseconds. The main difference between the faster group and the slower group is that all the fast methods compute the correlation scores in Euclidean space, while the slower methods compute scores in other spaces. In general, the RN2Dw / T method according to a non-limiting embodiment or aspect is the best method because it is efficient in retrieving relevant time series and is also efficient in terms of query time.
[0157] Fig.11is a critical difference (CD) plot comparing experimental performance. The CD plot is constructed to compare the performance of different methods and follows many previous works in time series classification. The CD plot shows the average rank of each method based on the performance measure and indicates whether two methods show significant performance differences based on the Wilcoxon signed rank test (α=0.05). The results show that, except for the RN2Dw / T method and the RN2D method according to non-limiting embodiments or aspects (their performance is not significantly different), almost all methods show significant differences in performance from each other. This conclusion is consistent with the conclusion in Fig.10 The findings are consistent with those presented in the table.
[0158] Fig.12 is a graph of the performance measurements of the experiment. The graph presents the performance measurements using various values of k ranging from 5 to 15. This is to ensure that Fig.10 Table and Fig.11 The conclusions drawn from the CD plots are not limited to a specific choice of k. To improve readability, ED and DTW are omitted from the plots as their performance is much worse than the other methods.
[0159] like Fig.12 As shown in FIG. 1 , the RN2D w / T method according to a non-limiting embodiment or aspect achieves the best performance at different values of all three performance measures. The remaining methods are ranked from best to worst: RN2D, LSTM, GRU, RN1D, and TF. These results are consistent with Fig.10 Table and Fig.11 This is consistent with the findings presented in the CD figure.
[0160] Fig.13 The time series of the first eight fetches of the experiment are shown. Two queries with different levels of complexity are selected from the test dataset. The simpler query consists of a single cycle of the pattern, while the more complex query contains a periodic signal. Periodic signals in complex queries usually require a shift-invariant distance measure to correctly fetch related items. Fig.13 This shows that the CTSR problem is challenging because even irrelevant time series are visually similar to the query. If the retrieved time series is relevant, it is plotted in light grey, and if it is irrelevant, it is plotted in black.
[0161] Can be checked by Fig.13The following observations are made based on the time series retrieved by the different methods shown in . The ED method has difficulty handling more complex queries because the ED method cannot align the query with the relevant time series. The DTW method is better than the ED method on complex queries, but the alignment freedom of the DTW method will impair the performance of the DTW method on simple queries. When considering two queries, the performance of the four neural network baselines (i.e., LSTM, GRU, TF, and RN1D) is better than both the ED and DTW methods. However, none of these baselines is better than the RN2Dw / T method and the RN2D method according to non-limiting embodiments or aspects for reliably retrieving related items.
[0162] In order to evaluate the effectiveness of different CTSR system designs in solving Figure 4 In order to test the effectiveness and efficiency of these CTSR solutions in terms of the business problems presented in the paper, a transaction time series dataset is constructed to test these CTSR solutions. The dataset includes 160,014 training time series, 19,993 test queries, and 19,992 validation queries, each of which represents a time series identification flag of a merchant with a length of 168. For the retrieved time series of a given query time series, if the retrieved time series belongs to a merchant with the same business type as the query time series, the retrieved time series is regarded as a relevant item from the database. If the retrieved time series is of another business type different from the query time series, the retrieved time series is regarded as an irrelevant item. Fig.14 is a table of other performance measurements from the experiment. Fig.14 Table 1, computing the performance measure at k=10 for each of the 19,993 test queries and presenting the average. Only the faster deep learning methods (i.e., LSTM, GRU, TF, RN1D, and RN2Dw / T according to a non-limiting embodiment or aspect) were tested for transaction time series, as these methods were the clear winners from experiments conducted on the UCR archived CTSR dataset in terms of effectiveness and efficiency.
[0163] like Fig.14 As shown in , the RN2Dw / T method according to a non-limiting embodiment or aspect is the best performing method, wherein the performance difference between the RN2Dw / T method according to a non-limiting embodiment or aspect and the second best RN1D method appears small. However, based on a two-sample t-test (where α=0.05), the difference is statistically significant. Fig.15 is a CD plot of further performance of the comparative experiment. The CD plots (which are similar to those of the UCR archived experiment) confirm these findings and are similar to those of Fig.14 The performance results are consistent with those presented in the table.
[0164] The performance differences between the test methods are examined under different settings of k, and the results are presented in Fig.16, which is a graph of the results of additional performance measurements of the experiment. Fig.16 As shown in , the RN2Dw / T method according to non-limiting embodiments or aspects consistently outperforms other methods at different values of k.
[0165] Fig.17 is the query schedule for the experiment. Fig.17 The table shows the average query time measured for each method. The query time for each test query is measured in milliseconds. Each of the exact and approximate nearest neighbor searches are used in the experiments.
[0166] To perform approximate nearest neighbor search, the nearest neighbor descent method is used to construct the k-nearest neighbor graph. The PyNNDescent library is used to implement the method. By replacing the exact nearest neighbor search with the approximate nearest neighbor search method, the query time is significantly reduced. In addition, the performance (i.e., PREC@10, AP@10, and NDCG@10) remains the same as Fig.14 The numbers presented in the table are exactly the same. A similar construction can be constructed for the results of the approximate nearest neighbor search Fig.15 The CD graph in is similar to Fig.16 The performance in is similar to that in the k-graph, and the conclusions remain the same.
[0167] Therefore, non-limiting embodiments or aspects of the present disclosure can provide effective and efficient CTSR models that are superior to alternative models while still providing reasonable inference runtimes. For example, non-limiting embodiments or aspects of the present disclosure can be superior to existing methods for time series retrieval in terms of effectiveness and efficiency. Non-limiting embodiments or aspects of the present disclosure can be used to identify business types in electronic payment networks, and / or to improve the efficiency of non-limiting embodiments or aspects of the present disclosure by incorporating low-bit representation techniques.
[0168] The described aspects include artificial intelligence or other operations, whereby the system uses apparent intelligence to process inputs and generate outputs. Artificial intelligence can be implemented in whole or in part by a model. The model can be implemented as a machine learning model. Learning can be supervised learning, unsupervised learning, reinforcement learning, or hybrid learning, whereby a variety of learning techniques are adopted to generate the model. Learning can be performed as part of training. Training the model can include obtaining a set of training data and adjusting the characteristics of the model to obtain the desired model output. For example, three characteristics can be associated with a desired project location. In this case, training can include receiving three characteristics as inputs to the model and adjusting the characteristics of the model so that for each set of three characteristics, the output device state matches the desired device state associated with the historical data.
[0169] In some embodiments, training can be dynamic. For example, the system can use a set of events to update the model. Detectable properties from the events can be used to adjust the model.
[0170] The model can be an equation, an artificial neural network, a recursive neural network, a convolutional neural network, a decision tree, or other machine-readable artificial intelligence structure. The characteristics of the structure that can be used to adjust during training can vary based on the selected model. For example, if a neural network is the selected model, the characteristics can include input elements, network layers, node density, node activation thresholds, weights between nodes, input or output value weights, etc. If the model is implemented as an equation (e.g., regression), the characteristics can include weights for input parameters, thresholds or limits for evaluating output values, or criteria for selecting from a set of equations.
[0171] Once the model is trained, retraining can be included to refine or update the model to reflect additional data or specific operating conditions. Retraining can be based on one or more signals detected by the apparatus described herein or as part of the methods described herein. Upon detection of the indicated signal, the system can activate the training process to adjust the model as described.
[0172] Further examples of machine learning and modeling features that may be included in the embodiments discussed above are described in “Asurvey of machine learning for big data processing” by Qiu et al., EURASIP Journal on Advances in Signal Processing (2016), which is hereby incorporated by reference in its entirety.
[0173] Although embodiments have been described in detail for purposes of illustration, it will be understood that such details are intended for that purpose only, and that the present disclosure is not limited to the disclosed embodiments or aspects, but, on the contrary, is intended to cover modifications and equivalent arrangements within the spirit and scope of the appended claims. For example, it will be understood that the present disclosure contemplates that, to the extent possible, one or more features of any embodiment or aspect may be combined with one or more features of any other embodiment or aspect. Indeed, any of these features may be combined in a manner not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may be directly dependent on only one claim, the disclosure of possible embodiments includes each dependent claim in combination with each other claim in the claim set.
Claims
1. A method comprising: obtaining a plurality of known time series from at least one database using at least one processor; For each known time series in the plurality of known time series: calculating, using the at least one processor, a pairwise distance matrix between the known time series and each of the learned templates to generate a plurality of pairwise distance matrices; stacking together the plurality of pairwise distance matrices to generate a tensor using the at least one processor; as well as processing the tensor with a residual network using the at least one processor, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; as well as The feature vector for each known time series of the plurality of known time series is provided using the at least one processor.
2. The method according to claim 1, further comprising: The feature vector of each known time series of the plurality of known time series is stored in the at least one database using the at least one processor.
3. The method according to claim 2, further comprising: obtaining, using the at least one processor, an unknown time series; calculating, using the at least one processor, a pairwise distance matrix between the unknown time series and each of the plurality of learned templates to generate another plurality of pairwise distance matrices; stacking together, with the at least one processor, the additional plurality of pairwise distance matrices to generate another tensor; processing the other tensor with the residual network using the at least one processor, wherein the residual network receives the other tensor as input and provides a feature vector for the unknown time series as output; for each known time series of the plurality of known time series stored in the database, determining, using the at least one processor, a distance between the known time series and the unknown time series based on a stored feature vector of the known time series and the feature vector of the unknown time series; as well as At least one known time series determined to correspond to the unknown time series is identified, using the at least one processor, based on the distance between each known time series and the unknown time series.
4. The method of claim 1, wherein the residual network is trained using a loss function defined according to the following equation: in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
5. The method of claim 1, wherein the plurality of known time series comprises a plurality of known transaction time series associated with a plurality of merchants, and wherein each known time series is associated with metadata associated with the merchant associated with the known time series.
6. A method according to claim 1, wherein the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four.
7. The method of claim 1, wherein the residual network comprises a two-dimensional residual network.
8. A system comprising: at least one processor coupled to the memory and configured to: Obtain a plurality of known time series from at least one database; For each known time series in the plurality of known time series: calculating a pairwise distance matrix between the known time series and each of the learned templates to generate a plurality of pairwise distance matrices; stacking the plurality of pairwise distance matrices together to generate a tensor; as well as processing the tensor with a residual network, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; as well as The feature vector of each known time series in the plurality of known time series is provided.
9. The system of claim 8, wherein the at least one processor is further configured to: The feature vector of each known time series of the plurality of known time series is stored in the at least one database.
10. The system of claim 9, wherein the at least one processor is further configured to: Get unknown time series; calculating a pairwise distance matrix between the unknown time series and each of the plurality of learned templates to generate another plurality of pairwise distance matrices; stacking the additional plurality of pairwise distance matrices together to generate another tensor; processing the further tensor with the residual network, wherein the residual network receives the further tensor as input and provides a feature vector of the unknown time series as output; For each known time series of the plurality of known time series stored in the database, determining a distance between the known time series and the unknown time series based on a stored feature vector of the known time series and the feature vector of the unknown time series; as well as At least one known time series determined to correspond to the unknown time series is identified based on the distance between each known time series and the unknown time series.
11. The system of claim 8, wherein the residual network is trained using a loss function defined according to the following equation: in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
12. The system of claim 8, wherein the plurality of known time series comprises a plurality of known transaction time series associated with a plurality of merchants, and wherein each known time series is associated with metadata associated with the merchant associated with the known time series.
13. The system of claim 8, wherein the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four.
14. The system of claim 8, wherein the residual network comprises a two-dimensional residual network.
15. A computer program product comprising a non-transitory computer readable medium comprising program instructions that, when executed by at least one processor, cause the at least one processor to: Obtain a plurality of known time series from at least one database; For each known time series in the plurality of known time series: calculating a pairwise distance matrix between the known time series and each of the learned templates to generate a plurality of pairwise distance matrices; stacking the plurality of pairwise distance matrices together to generate a tensor; as well as processing the tensor with a residual network, wherein the residual network receives the tensor as input and provides a feature vector of the known time series as output; as well as The feature vector of each known time series in the plurality of known time series is provided.
16. The computer program product of claim 15, wherein the program instructions, when executed by the at least one processor, further cause the at least one processor to: The feature vector of each known time series of the plurality of known time series is stored in the at least one database.
17. The computer program product of claim 16, wherein the program instructions, when executed by the at least one processor, further cause the at least one processor to: Get unknown time series; calculating a pairwise distance matrix between the unknown time series and each of the plurality of learned templates to generate another plurality of pairwise distance matrices; stacking the additional plurality of pairwise distance matrices together to generate another tensor; processing the further tensor with the residual network, wherein the residual network receives the further tensor as input and provides a feature vector of the unknown time series as output; For each known time series of the plurality of known time series stored in the database, determining a distance between the known time series and the unknown time series based on a stored feature vector of the known time series and the feature vector of the unknown time series; as well as At least one known time series determined to correspond to the unknown time series is identified based on the distance between each known time series and the unknown time series.
18. The computer program product of claim 15, wherein the residual network is trained using a loss function defined according to the following equation: in For training data batches m is the batch size, and each sample in the batch includes the query time series t i , positive time series t i+ and negative time series t i- Tuple σ(·) is a S-shaped function, and f θ (·,·) is the residual network.
19. The computer program product of claim 15, wherein the plurality of known time series comprises a plurality of known transaction time series associated with a plurality of merchants, and wherein each known time series is associated with metadata associated with the merchant associated with the known time series.
20. The computer program product of claim 15, wherein the plurality of learned templates comprises thirty-two learned templates, wherein the plurality of pairwise distance matrices comprises thirty-two pairwise distance matrices, wherein the tensor comprises an input dimension of thirty-two, and wherein the feature vector of each known time series in the plurality of known time series comprises a vector of size sixty-four, and wherein the residual network comprises a two-dimensional residual network.
Citation Information
Patent Citations
Multi-modal multivariable time sequence automatic classification method and device
CN114722950A
Generating input data for machine learning model
CN114970804A
Integrated circuit path delay prediction method based on feature selection and deep learning
CN115146580A
Single-view-angle three-dimensional human skeleton key point detection method and device, equipment and medium
CN115482481A
Data processing method and system based on Internet of Things technology
CN115834433A