Analysis method of time sequence and related device

By clustering multiple groups of time series and building regression models, the problem of indicator data analysis relying on manual labor is solved, and efficient and automated indicator data analysis is achieved.

CN120654005APending Publication Date: 2025-09-16HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410305827.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the existing technology, indicator data analysis relies heavily on manual work, which is inefficient and cannot effectively utilize the correlation between indicator data.

Method used

By clustering multiple groups of time series, building regression models, and analyzing the relationship between preset indicators and target indicators, automated analysis can be achieved.

Benefits of technology

It achieves efficient analysis of indicator data without manual intervention, improving analysis efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654005A_ABST
    Figure CN120654005A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence analysis method, which is applied to realizing efficient analysis of index data forming a time sequence. In the method, data of a plurality of preset indexes are sorted into a plurality of sequences according to a time sequence, and grouping is carried out according to an acquisition dimension of the data to obtain a plurality of groups of time sequences; and moreover, multiple groups of time sequences are clustered, so that the sequences with similar time sequence dependency relationships can be classified into the same category. Therefore, the regression model can be constructed based on the values of the preset indexes included in the sequences under the same category and the values of the target indexes corresponding to the sequences, and the regression model can indicate the relationship between the preset indexes and the target indexes under the specific time sequence dependency relationship. And finally, on the basis of the constructed regression model, how the value of the preset index affects the value of the target index under various conditions can be determined, so that efficient analysis of the index data is realized, and manual analysis is not needed in the whole process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a time series analysis method and related devices. Background Art

[0002] In every field of human research, there is often a large amount of indicator data, and these indicators often have certain correlations. For example, in the field of finance, financial indicators such as accounts payable, accounts receivable, inventory-to-revenue ratio, contract assets, and contract liabilities often have certain correlations. For another example, in the field of climate, climate indicators such as temperature, humidity, rainfall, UV intensity, and wind speed often have certain correlations.

[0003] Business experts in various fields can draw specific conclusions by studying the changes and correlations between indicator data within their respective fields. For example, they can deduce the value of a specific indicator based on the values ​​of a subset of indicator data. For example, in the financial field, financial experts often determine a company's operating strategy based on the values ​​of multiple financial indicators and the fluctuations between these indicators. They then study the impact of these operating strategies on specific financial indicators.

[0004] Currently, business experts in the field usually analyze indicator data based on experience and derive corresponding conclusions. As a result, the analysis of indicator data is heavily dependent on manual implementation, and the analysis efficiency of indicator data is low. Summary of the Invention

[0005] This application provides a time series analysis method that can achieve efficient analysis of indicator data.

[0006] The first aspect of the present application provides a time series analysis method, which is applied to analyze multiple groups of time series under different dimensions. The method comprises: first, obtaining multiple groups of time series under different dimensions, that is, multiple groups of time series from different sources. Moreover, each group of time series in the multiple groups of time series includes multiple sequences arranged in chronological order, for example, each sequence corresponds to a year, and the multiple sequences are sorted according to the order of the years. Each sequence includes the values ​​of multiple preset indicators, and each sequence also corresponds to the value of a target indicator, and the value of the target indicator is related to the values ​​of multiple preset indicators. That is, the target indicator and the preset indicator are not isolated, but have a certain correlation. The size of the multiple preset indicators will affect the specific value of the target indicator.

[0007] Then, multiple groups of time series are clustered to obtain multiple categories. Each category includes at least one time window, and each time window includes adjacent partial sequences in a group of time series. Clustering multiple groups of time series actually extracts time windows with similar temporal dependencies from multiple groups of time series and groups them into the same category, thereby obtaining time windows corresponding to each category under multiple categories. Time windows are extracted from sequences in a group of time series. Each time window actually includes multiple adjacent sequences, but the multiple sequences included in the time window are only partial sequences in a group of time series.

[0008] Secondly, based on the sequence under each category and the value of the target indicator corresponding to the sequence, a regression model is constructed to obtain multiple regression models corresponding to the multiple categories. The multiple regression models are used to indicate the relationship between the multiple preset indicators and the target indicator. Specifically, by using the preset indicator as the independent variable and the target indicator as the dependent variable, based on the values ​​of the multiple preset indicators included in the sequence under each category and the value of the target indicator corresponding to the sequence, a regression model under each category can be established.

[0009] Finally, the preset values ​​of the multiple preset indicators are input into the target model in the multiple regression models to obtain the target indicator value output by the target model. In this way, based on multiple regression models, it is possible to evaluate the impact of various values ​​of the preset indicators on the target indicator value, thereby enabling analysis between indicators and facilitating subsequent decision-making and planning by field personnel.

[0010] In this solution, multiple sets of time series are obtained by sorting the data of multiple preset indicators into multiple sequences in chronological order and grouping them according to the dimensions of data acquisition. In addition, by clustering multiple sets of time series, time windows with similar temporal dependencies can be clustered, that is, sequences with similar temporal dependencies are classified into the same category. In this way, based on the values ​​of the preset indicators included in the sequences under the same category and the values ​​of the target indicators corresponding to these sequences, a regression model can be constructed, which can indicate the relationship between the preset indicators and the target indicators under specific temporal dependencies. Finally, based on the constructed regression model, it is possible to determine how the values ​​of the preset indicators will affect the values ​​of the target indicators under various circumstances, thereby achieving efficient analysis of the indicator data, and the entire process does not require manual analysis.

[0011] In one possible implementation, the multiple preset indicators and target indicators are all indicators in the same field, for example, any of the following indicators: financial indicators, climate indicators, health indicators, or financial indicators. In addition to the various types of indicators listed above, the preset indicators and target indicators can also be indicators in other fields, as long as there is a correlation between the preset indicators and the target indicators.

[0012] In this solution, by applying the time series analysis method in various fields, it is possible to realize the automated analysis of time series in various fields, expand the application scope of the solution, and avoid spending a lot of manpower and material resources in various fields to analyze the indicator data in the time series.

[0013] In a possible implementation, when the plurality of preset indicators and the target indicator are all financial indicators, different groups of time series in the plurality of groups of time series correspond to different companies.

[0014] In other words, the multiple time series in different dimensions actually refer to multiple time series corresponding to different companies, essentially representing time series of financial indicators for different companies. Therefore, these multiple time series actually indicate how the financial indicators of different companies change over time. By obtaining the values ​​of multiple preset indicators for different companies in chronological order, the subsequent clustering process can cluster time windows with similar operating strategies across different companies, thereby increasing the richness of the indicator data.

[0015] In one possible implementation, a target model selected from multiple regression models corresponds to a target category within multiple categories, where the time windows included within the target category belong to one or more target companies. Specifically, based on the multiple categories obtained by classification, a time window that includes a specific target company (e.g., a company to be analyzed or referenced) is first found within the multiple categories, and then the target category is determined within the multiple categories. Thus, based on the target category, a target model can be further determined within the multiple regression models. This target model can then be considered to reflect the relationship between indicators for one or more specific target companies when adopting a specific business strategy.

[0016] That is to say, after determining the regression model under a certain business strategy, based on the regression model, the specific values ​​of the target indicators under various preset values ​​of multiple preset indicators can be obtained, which makes it easier for business personnel in the financial field to implement the company's business strategy or preset indicator planning.

[0017] In one possible implementation, when multiple preset indicators and target indicators are all financial indicators, the multiple preset indicators are accounts payable, accounts receivable, inventory-to-revenue ratio, contract assets and contract liabilities, and the target indicator is accounts payable turnover rate.

[0018] In one possible implementation, clustering multiple groups of time series to obtain multiple categories may specifically include: first, dividing each group of time series into time windows based on a preset time window length to obtain multiple time windows, where the time window length indicates the number of sequences included in the time window. Then, dividing each group of time series into time windows separately, ensuring that each divided time window only includes sequences within one group of time series and does not include sequences from different time series.

[0019] Then, clustering is performed on the multiple time windows to obtain multiple categories, where the time windows included in each category are part of the multiple time windows. In other words, after dividing the multiple time series into multiple time windows, clustering is performed on the multiple time windows, so that time windows with similar temporal dependencies are grouped into the same category.

[0020] In this solution, each group of time series is divided into multiple time windows independently in advance, and then all the divided time windows are clustered. This ensures that time windows including sequences in different time series will not be clustered during the clustering process, thereby ensuring the accuracy of clustering.

[0021] In one possible implementation, when multiple time series are concatenated, time windows are divided sequentially according to the concatenation order of the time series, thereby obtaining multiple ordered time windows. Furthermore, the sequences included in each time window must be sequences within the same set of time series.

[0022] Adjacent time windows divided from the same set of time series can contain at least one repeating sequence. That is, when dividing each set of time series into time windows, the sequences between the preceding and succeeding time windows can partially overlap, as well as partially not overlap. For example, for two adjacent time windows, except for the first sequence in the preceding window and the last sequence in the succeeding window, all other sequences are repeating.

[0023] It should be noted that only when two adjacent time windows belong to the same group of time series will the two adjacent time windows contain repeated sequences. In this way, when dividing the time windows, if the previous time window belongs to one group of time series and the next time window belongs to another group of time series, the two adjacent time windows will no longer have repeated sequences.

[0024] In this scheme, by setting the adjacent time windows obtained by division to include repeated sequences, the same sequence can appear in different time windows, thereby ensuring that a sufficient number of sequence combinations can be obtained, which is conducive to discovering the temporal dependency between sequences in the subsequent clustering process.

[0025] In another possible implementation, adjacent time windows divided from the same set of time series may also include non-repeated sequences. For example, for two adjacent time windows, the first time window includes the first to fifth sequences in the set of time series, while the second time window includes the sixth to tenth sequences in the same set of time series. For another example, for two adjacent time windows, the first time window includes the first to third sequences in the set of time series, while the second time window includes the fifth to seventh sequences in the same set of time series.

[0026] In one possible implementation, clustering is performed on multiple time windows, specifically including: using an improved Toeplitz Inverse Covariance-Based Clustering (TICC) method to cluster the multiple time windows to obtain multiple categories. The improved TICC cancels penalty terms between adjacent time windows in different time series groups while retaining penalty terms between adjacent time windows in the same time series group when solving an objective function to optimize time window clustering. The retained penalty terms are used to encourage adjacent retained time windows in the same time series group to be classified into the same category.

[0027] That is to say, when setting the objective function, if two adjacent time windows belong to the same group of time series, the penalty term between the two adjacent time windows can be retained; if the two adjacent time windows do not belong to the same group of time series, the penalty term between the two adjacent time windows is canceled.

[0028] In this scheme, based on the characteristics of multiple groups of time series as clustering objects, TICC is improved to adjust the penalty term of the objective function in TICC. This can avoid clustering time windows of different groups of time series into the same category, thereby ensuring the accuracy of clustering.

[0029] The second aspect of the present application provides a time series analysis device, including: an acquisition module, used to acquire multiple groups of time series under different dimensions, each group of time series includes multiple sequences arranged in chronological order, each sequence includes multiple values ​​of preset indicators, and each sequence also corresponds to the value of a target indicator, and the value of the target indicator is related to the values ​​of multiple preset indicators; a processing module, used to cluster the multiple groups of time series to obtain multiple categories, each category includes at least one time window, and each time window includes adjacent partial sequences in a group of time series; the processing module is also used to construct a regression model based on the sequence under each category and the value of the target indicator corresponding to the sequence, to obtain multiple regression models corresponding to multiple categories, and the multiple regression models are all used to indicate the relationship between multiple preset indicators and the target indicator; the processing module is also used to input the preset values ​​of the multiple preset indicators into the target model in the multiple regression models to obtain the value of the target indicator output by the target model.

[0030] In a possible implementation, the plurality of preset indicators and the target indicator are any one of the following indicators: a financial indicator, a climate indicator, a health indicator, or a financial indicator.

[0031] In a possible implementation, when the plurality of preset indicators and the target indicator are all financial indicators, different groups of time series in the plurality of groups of time series correspond to different companies.

[0032] In a possible implementation, the target model corresponds to a target category among the multiple categories, and the time windows included in the target category belong to one or more target companies.

[0033] In one possible implementation, when multiple preset indicators and target indicators are all financial indicators, the multiple preset indicators are accounts payable, accounts receivable, inventory-to-revenue ratio, contract assets and contract liabilities, and the target indicator is accounts payable turnover rate.

[0034] In one possible implementation, the processing module is further used to: perform time window division on each group of time series based on a preset time window length to obtain multiple time windows, wherein the time window length is used to indicate the number of sequences included in the time window; perform clustering on the multiple time windows to obtain multiple categories, and the time windows included in each category are part of the multiple time windows.

[0035] In a possible implementation, adjacent time windows obtained by dividing the same group of time series include at least one repeated sequence.

[0036] In one possible implementation, the processing module is further used to: use an improved Toeplitz inverse covariance matrix-based clustering method TICC to perform clustering on multiple time windows to obtain multiple categories; wherein, in the process of solving the objective function to optimize the time window clustering, the improved TICC cancels the penalty terms between adjacent time windows in different groups of time series, and retains the penalty terms between adjacent time windows in the same group of time series, and the retained penalty terms are used to encourage adjacent retained time windows in the same group of time series to be divided into the same category.

[0037] A third aspect of the present application provides a computing device cluster, comprising at least one computing device, each computing device including a processor and memory. The processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, causing the computing device cluster to perform the method described in the first aspect or any of the implementations of the first aspect. For details regarding the steps in each possible implementation of the first aspect performed by the computing device cluster, please refer to the first aspect and will not be repeated here.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium having instructions stored therein. When the instructions are executed on a computer, the computer can execute any of the above methods.

[0039] A fifth aspect of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the methods described above.

[0040] In a sixth aspect, the present application provides a chip comprising a processor and a communication interface, wherein the communication interface is used to communicate with modules outside the chip, and the processor is used to run computer programs or instructions so that a device in which the chip is installed can execute any of the methods described above.

[0041] Among them, the technical effects brought about by any design method in the second to sixth aspects can refer to the technical effects brought about by different implementation methods in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A schematic diagram of a system architecture 100 provided in an embodiment of the present application;

[0043] Figure 2 A schematic diagram of a time series analysis method provided in an embodiment of the present application;

[0044] Figure 3 A schematic diagram of multiple time series provided in an embodiment of the present application;

[0045] Figure 4A schematic diagram of clustering multiple time series provided in an embodiment of the present application;

[0046] Figure 5 A schematic diagram of processing multiple time series based on TICC provided in an embodiment of the present application;

[0047] Figure 6 A schematic diagram of a process for processing multiple time series of different companies provided in an embodiment of the present application;

[0048] Figure 7 A schematic diagram of integrating company indicator data provided in an embodiment of the present application;

[0049] Figure 8 A schematic diagram of the concatenation of indicator data of multiple companies provided in an embodiment of the present application;

[0050] Figure 9 A schematic diagram of a time window divided into categories provided in an embodiment of the present application;

[0051] Figure 10 A schematic diagram of the structure of a time series analysis device provided in an embodiment of the present application;

[0052] Figure 11 A schematic diagram of the structure of a computing device 1100 provided in an embodiment of the present application;

[0053] Figure 12 A schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0054] Figure 13 A schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application;

[0055] Figure 14 A schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0057] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.

[0058] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.

[0059] (1) Sequence

[0060] A sequence refers to a plurality of objects arranged in a row. In this embodiment, the order of the plurality of objects in the sequence is not important, but the values ​​of the plurality of objects are important.

[0061] (2) Time series

[0062] A time series (or dynamic series) is a sequence of values ​​of the same statistical indicator arranged in chronological order. A time series is a set of random variables ordered by time. It is typically the result of observing an underlying process at a given sampling rate over equally spaced time periods. Essentially, time series data reflects the changing trends of one or more random variables over time.

[0063] (3) Regression model

[0064] A regression model is a mathematical model that quantitatively describes statistical relationships. Specifically, a regression model is a predictive modeling technique that studies the relationship between dependent and independent variables.

[0065] For example, the mathematical model of multiple linear regression can be expressed as y = β0 + β1*x1 + β2*x2… + β p *x p ;In the formula, β0, β1,…,β p There are p+1 parameters to be estimated, y is the dependent variable; x1-x p is the independent variable, and βi is called the regression coefficient, which represents the degree of influence of the independent variable on the dependent variable.

[0066] (4) Clustering

[0067] Cluster analysis, also known as group analysis, is a statistical analysis method used to study the classification of samples or indicators. Clustering is essentially the process of dividing a collection of physical or abstract objects into clusters of similar objects. The resulting clusters are collections of objects that are similar to objects in the same cluster but different from objects in other clusters.

[0068] Currently, to analyze indicator data in various fields (such as finance, climate change, and finance), business experts in each field typically study the changes and correlations between indicator data based on their experience, drawing specific conclusions. This heavily relies on manual analysis, resulting in low efficiency.

[0069] Based on this, the present application provides a method for analyzing time series, by sorting the data of multiple preset indicators into multiple sequences in chronological order, and grouping them according to the dimensions of data acquisition, to obtain multiple groups of time series. In addition, by clustering multiple groups of time series, time windows with similar temporal dependencies can be clustered, that is, sequences with similar temporal dependencies are classified into the same category. In this way, based on the values ​​of the preset indicators included in the sequences under the same category, and the values ​​of the target indicators corresponding to these sequences, a regression model can be constructed, and the regression model can indicate the relationship between the preset indicators and the target indicators under this temporal dependency. Finally, based on the constructed regression model, it is possible to determine how the value of the preset indicator will affect the value of the target indicator under various circumstances, thereby achieving efficient analysis of the indicator data.

[0070] See also Figure 1 , Figure 1 A schematic diagram of a system architecture 100 provided in an embodiment of the present application. Figure 1As shown, in the system architecture 100, the execution device 110 can be implemented by at least one computing instance of a physical host (computing device), a virtual machine, or a container. When the execution device 110 is implemented by a virtual machine or a container, the execution device 110 actually exists in the form of a cloud computing product and can provide cloud services.

[0071] Optionally, the execution device 110 cooperates with other computing devices, such as data storage, routers, load balancers and other devices; the execution device 110 can be deployed on one physical site or distributed on multiple physical sites.

[0072] Optionally, in order to store data persistently, the system architecture 100 is further provided with a data storage system 120, which may be located outside the execution device 110 (e.g., Figure 2 As shown, the execution device 110 exchanges data with the execution device 110 via a network. Optionally, if the execution device 110 is a physical host, the data storage system 120 may be located within the execution device 110, such as if the data storage system 120 exchanges data with the processor via a bus. In this case, the data storage system 120 is represented by a hard disk. With the data storage system 120, the execution device 110 can use the data in the data storage system 120 or call program code in the data storage system 120 to implement the time series analysis method provided in the embodiments of the present application.

[0073] Optionally, users can operate their respective user devices (such as local device 101 and local device 102) to interact with execution device 110. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a laptop computer, and a smart car.

[0074] Each user's local device can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0075] In one implementation, execution device 110 is configured to implement the time series analysis method provided in the embodiments of the present application, cluster multiple groups of time series, and construct regression models corresponding to each category, thereby determining the influence relationship between indicators based on the regression models. Optionally, before or during execution device 110 implementing the time series analysis method provided in the embodiments of the present application, local device 101 or local device 102 provides various time series to execution device 110, thereby ensuring that execution device 110 can perform analysis based on the provided time series.

[0076] In another implementation, one or more aspects of the execution device 110 can be implemented by each local device. For example, the local device 101 can provide local data or feedback calculation results to the execution device 110, or execute the time series analysis method provided in the embodiment of the present application.

[0077] In general, the time series analysis method provided in the embodiments of the present application can be applied to electronic devices, such as the aforementioned execution device 110 , local device 101 , or local device 102 .

[0078] See also Figure 2 , Figure 2 This is a flow chart of a time series analysis method provided in an embodiment of the present application. Figure 2 As shown, the time series analysis method provided in the embodiment of the present application includes the following steps 201-204.

[0079] Step 201: Acquire multiple groups of time series in different dimensions, each group of time series includes multiple sequences arranged in chronological order, each sequence includes multiple values ​​of preset indicators, and each sequence also corresponds to a value of a target indicator, and the value of the target indicator is related to the values ​​of multiple preset indicators.

[0080] In this embodiment, multiple sets of time series under different dimensions may refer to sets of time series obtained from different dimensions, that is, sets of time series from different sources. Furthermore, each set of time series includes multiple sequences arranged in chronological order. For example, each sequence corresponds to a year, and the multiple sequences are sorted by year.

[0081] In addition, although the acquisition dimensions of multiple groups of time series are different, each group of time series is used to represent the same multiple preset indicators, that is, multiple groups of time series actually represent the specific values ​​of the same multiple preset indicators under various different sources.

[0082] It's important to note that each sequence also corresponds to a target indicator value. However, the sequence itself doesn't include the target indicator value; it only includes the value of the preset indicator. Since each sequence corresponds to a dimension and a time point, the target indicator value corresponding to the sequence represents the specific value of the target indicator at the dimension and time point corresponding to the sequence.

[0083] For example, see Figure 3 , Figure 3 Schematic diagram of multiple time series provided in the embodiment of this application. Figure 3 As shown, Figure 3The figure shows two sets of time series with different dimensions. The first set of time series is the time series under dimension 1, and the second set of time series is the time series under dimension 2. Furthermore, both the first and second sets of time series include five series arranged in chronological order, namely, five series from 2006 to 2010. Furthermore, each series is used to represent the values ​​of the same five preset indicators. For example, the first series in the first set of time series represents the specific values ​​of indicators A through E in 2006, and this series also corresponds to the value of a target indicator, namely, indicator M. In general, different series are actually used to represent the values ​​of the same five indicators (i.e., indicators A through E) under a specific dimension (e.g., dimension 1 or dimension 2) and a specific time (e.g., any year from 2006 to 2010).

[0084] For any sequence, the target indicator value is related to the values ​​of multiple pre-set indicators. That is, the target indicator and pre-set indicators are not isolated but rather correlated. The magnitude of the pre-set indicators affects the specific value of the target indicator.

[0085] Illustratively, in this embodiment, the plurality of preset indicators and target indicators are all indicators in the same field, for example, they are all any one of the following indicators: financial indicators, climate indicators, health indicators or financial indicators.

[0086] For example, if multiple preset indicators and target indicators are all financial indicators, the preset indicators are accounts payable, accounts receivable, inventory-to-revenue ratio, contract assets, and contract liabilities, and the target indicator is accounts payable turnover. Obviously, accounts payable turnover will be affected by these preset indicators. It should be noted that the above examples only illustrate some financial indicators. In actual scenarios, the multiple preset indicators and target indicators can also be other financial indicators, and this is not a specific limitation here.

[0087] Furthermore, when multiple preset indicators and target indicators are all financial indicators, different time series within the multiple time series groups correspond to different companies. In other words, the multiple time series groups under different dimensions actually refer to multiple time series groups corresponding to different companies, and are essentially time series of financial indicators for different companies. Therefore, these multiple time series groups actually indicate how the financial indicators of different companies change over time. By obtaining the values ​​of multiple preset indicators for different companies, arranged in chronological order, the subsequent clustering process can cluster together time windows with similar operating strategies across different companies, thereby increasing the richness of the indicator data.

[0088] For example, if multiple preset indicators and target indicators are climate indicators, the preset indicators can be indicators such as humidity, rainfall, wind speed, air pressure, and ultraviolet intensity, while the target indicator can be air temperature. Multiple time series in different dimensions can refer to multiple time series including climate indicators obtained at different locations, that is, how the climate indicators at multiple locations change over time.

[0089] For another example, if multiple preset indicators and target indicators are all health indicators, the multiple preset indicators can be heart rate, blood pressure, blood sugar, blood lipids, etc., and the target indicator can be health status. Multiple sets of time series under different dimensions can refer to multiple sets of time series including health indicators obtained for different individuals, that is, the changes in health indicators of multiple people over time.

[0090] For another example, if multiple preset indicators and target indicators are all financial indicators, the multiple preset indicators can be indicators such as stock price, stock price increase, number of stocks held, number of stock holders, etc., while the target indicator can be stock trading volume. Multiple sets of time series under different dimensions can refer to multiple sets of time series including financial indicators obtained for different stocks, that is, the changes in the financial indicators of multiple stocks over time.

[0091] In general, in addition to the financial indicators, climate indicators, health indicators, and financial indicators exemplified above, the preset indicators and target indicators can also be indicators in other fields, as long as there is a correlation between the preset indicators and the target indicators. This embodiment does not make any specific restrictions on this.

[0092] Step 202 : clustering the multiple groups of time series to obtain multiple categories, each category including at least one time window, and each time window including adjacent partial sequences in a group of time series.

[0093] In this embodiment, clustering multiple time series groups actually involves extracting time windows with similar temporal dependencies from the multiple time series groups and grouping them into the same category, thereby obtaining time windows corresponding to each category under multiple categories. A time window is extracted from a set of time series. Each time window actually includes multiple adjacent sequences, but the multiple sequences included in the time window are only a subset of the sequences in the set. For example, if a set of time series includes 10 sequences, a time window may only include 5 adjacent sequences from the 10 sequences. In other words, any time window is a subset of a set of time series.

[0094] Since each time window includes multiple adjacent series within the same set of time series (i.e., multiple temporally adjacent series), clustering multiple time series can be achieved by analyzing the temporal dependencies between indicators within each time window, thereby grouping time windows with similar temporal dependencies. Temporal dependencies within a time window can refer to the dependencies between different indicators within the same time window. For example, assuming a time window includes three series, each of which includes five indicators, then the temporal dependencies within the time window include: the relationship between the value of the first indicator in the first sequence and the value of the second indicator in the first sequence; the relationship between the value of the first indicator in the first sequence and the value of the second indicator in the second sequence; the relationship between the value of the first indicator in the first sequence and the value of the second indicator in the third sequence; and finally, the relationship between the value of the mth indicator in the nth sequence and the value of the pth indicator in the sth sequence.

[0095] For example, see Figure 4 , Figure 4 This is a schematic diagram of clustering multiple time series provided in an embodiment of the present application. Figure 4 As shown in the figure, after clustering the first and second time series, three categories are obtained, and each of the three categories includes two time windows. The two time windows belonging to category one are the time window including the series from 2006 to 2008 under dimension 1, and the time window including the series from 2007 to 2009 under dimension 1; the two time windows belonging to category two are the time window including the series from 2007 to 2009 under dimension 2, and the time window including the series from 2008 to 2010 under dimension 2; the two time windows belonging to category three are the time window including the series from 2008 to 2010 under dimension 1, and the time window including the series from 2006 to 2008 under dimension 2.

[0096] It should be noted that the above example is based on the example that each category includes 2 time windows. In actual scenarios, the number of time windows included in each category can be different or the same, and no specific limitation is made here.

[0097] Step 203 : Based on the sequence under each category and the value of the target indicator corresponding to the sequence, a regression model is constructed to obtain multiple regression models corresponding to multiple categories. The multiple regression models are used to indicate the relationship between multiple preset indicators and the target indicator.

[0098] In this embodiment, multiple regression models have a one-to-one correspondence with multiple categories, and each regression model is constructed based on the time series under the corresponding category. Since each category includes at least one time window, and each time window includes multiple sequences, each category actually includes multiple sequences. In this way, by using the preset indicator as the independent variable and the target indicator as the dependent variable, based on the values ​​of the multiple preset indicators included in the sequence under each category and the value of the target indicator corresponding to the sequence, a regression model under each category can be established. The regression model under each category is used to indicate the relationship between the preset indicator and the target indicator under that category.

[0099] For example, taking the simplest linear regression model as an example, assuming that a sequence includes 5 preset indicators, and the values ​​of the 5 preset indicators are represented by x1, x2, x3, x4, and x5 respectively, and the target indicator is represented by y, then the regression model can be expressed as: y=a1*x1+a2*x2+a3*x3+a4*x4+a5*x5+b, where a1-a5 and b are the obtained coefficients.

[0100] For example, taking the case where both the preset and target indicators are financial indicators, after clustering the time series, the sequences within the same category can be considered to represent the performance of financial indicators under similar business strategies. Therefore, by clustering the sequences within the same category, the relationship between the preset and target indicators under similar business strategies can be obtained, which facilitates the subsequent analysis of the performance of the target indicators under specific business strategies. In other words, the sequences within the same category can be considered to correspond to similar business strategies. Therefore, based on the time windows included in each category, the various time periods in which companies have similar business strategies can be determined. For example, if the same category includes the series for Company 1 from 2006 to 2009 and the series for Company 2 from 2009 to 2012, it can be considered that Company 1's business strategy from 2006 to 2009 is similar to Company 2's business strategy from 2009 to 2012.

[0101] In step 204 , the preset values ​​of the plurality of preset indicators are input into a target model among the plurality of regression models to obtain the value of the target indicator output by the target model.

[0102] After obtaining the regression model for each category, the relationship between the preset indicator and the target indicator for each category can be determined, and this relationship can be used to obtain the specific value of the target indicator when the preset indicator takes a specific value. Specifically, after selecting the target model corresponding to a specific category from multiple regression models, the preset values ​​of the multiple preset indicators can be input into the target model to obtain the value of the target indicator output by the target model. In this way, based on multiple regression models, it is possible to evaluate the impact of various values ​​of the preset indicator on the value of the target indicator, thereby realizing analysis between indicators and facilitating subsequent decision-making and planning by field personnel.

[0103] Taking the financial sector as an example, the target model described above can correspond to a target category within multiple categories, with the time window within the target category belonging to one or more target companies. Specifically, based on the multiple categories obtained by categorization, a time window that includes a specific target company (e.g., a company to be analyzed or referenced) is first found within the multiple categories, and then the target category is determined within the multiple categories. Based on the target category, a target model can be further determined within multiple regression models. This target model can then be considered to reflect the relationship between indicators for one or more specific target companies when adopting a specific business strategy.

[0104] That is to say, after determining the regression model under a certain business strategy, based on the regression model, the specific values ​​of the target indicators under various preset values ​​of multiple preset indicators can be obtained, which makes it easier for business personnel in the financial field to implement the company's business strategy or preset indicator planning.

[0105] In summary, this embodiment clusters time series and then constructs a regression model based on the resulting clustered sequences within each category. This model can reveal the relationship between the preset indicator and the target indicator under specific temporal dependencies. This regression model can then be used to determine how the value of the preset indicator affects the value of the target indicator under various circumstances, thereby enabling efficient analysis of indicator data without relying on manual analysis.

[0106] The above describes the process of clustering time series and building a regression model to implement time series analysis in this embodiment. For ease of understanding, the following details the process of clustering multiple groups of time series in this embodiment.

[0107] In this embodiment, when clustering multiple time series, reference can be made to the existing Toeplitz Inverse Covariance-Based Clustering (TICC) method. However, TICC is designed for clustering a single set of time series (e.g., data from various sensors in a car at various time points) and cannot be directly used to cluster multiple time series in different dimensions as described in this embodiment.

[0108] Specifically, TICC treats all input sequences as a whole and divides them into multiple continuous time windows by dividing them into time windows. Therefore, when processing multiple time series in different dimensions based on TICC, the sequences in different dimensions will be divided into the same time window, which can easily lead to incorrect clustering results or even failure to achieve clustering.

[0109] For example, see Figure 5 , Figure 5 This is a schematic diagram of processing multiple time series based on TICC provided in an embodiment of the present application. Figure 5 As shown in the figure, when the input is two sets of time series with different dimensions, assuming that TICC divides the time window with a time window length of 4, then TICC will divide the series from 2006 to 2009 in the first set of time series into the same time window, divide the series from 2007 to 2010 in the first set of time series into the same time window, and divide the series from 2008 to 2010 in the first set of time series and the series from 2006 in the second set of time series into the same time window (i.e., the wrong time window). Since the series in the first set of time series and the series in the second set of time series are actually of different dimensions (for example, financial indicator data from different companies), dividing the series in different dimensions into the same time window will not be able to extract the temporal dependencies between the indicators in the series, which can easily lead to clustering failure.

[0110] Based on this, this embodiment does not directly apply TICC, but makes adaptive adjustments based on the input of this embodiment to achieve clustering of multiple groups of time series in different dimensions.

[0111] For example, in this embodiment, when clustering multiple groups of time series, a time window division may be performed on each group of time series based on a preset time window length to obtain multiple time windows. The time window length is used to indicate the number of sequences included in the time window. That is, in this embodiment, the time window division is performed on each group of time series, taking each group of time series as a unit, thereby ensuring that each divided time window only includes sequences within one group of time series, and does not include sequences within different time series. Furthermore, the time window length may be user-specified or set by default, such as a value of 3, 5, or 6. The specific length may be determined based on actual conditions, and this embodiment does not impose any specific limitation on this.

[0112] Then, clustering is performed on the multiple time windows to obtain multiple categories, where the time windows included in each category are part of the multiple time windows. In other words, after dividing the multiple time series into multiple time windows, clustering is performed on the multiple time windows, so that time windows with similar temporal dependencies are grouped into the same category.

[0113] In this solution, each group of time series is divided into multiple time windows independently in advance, and then all the divided time windows are clustered. This ensures that time windows including sequences in different time series will not be clustered during the clustering process, thereby ensuring the accuracy of clustering.

[0114] Optionally, when dividing the time windows, adjacent time windows divided under the same group of time series include at least one repeated sequence. That is to say, when dividing the time windows for each group of time series, there may be partially overlapping sequences and partially non-overlapping sequences between the previous time window and the next time window. For example, for two adjacent time windows, except for the first sequence of the previous time window and the last sequence of the next time window, the others are repeated sequences. The number of non-overlapping sequences can be one or more, which can be adjusted according to actual conditions. For example, in Figure 5 In , the number of sequences that do not overlap between the previous time window and the next time window is one.

[0115] It should be noted that only when two adjacent time windows belong to the same group of time series will the two adjacent time windows contain repeated sequences. In this way, when dividing the time windows, if the previous time window belongs to one group of time series and the next time window belongs to another group of time series, the two adjacent time windows will no longer have repeated sequences.

[0116] In this scheme, by setting the adjacent time windows obtained by division to include repeated sequences, the same sequence can appear in different time windows, thereby ensuring that a sufficient number of sequence combinations can be obtained, which is conducive to discovering the temporal dependency between sequences in the subsequent clustering process.

[0117] Optionally, adjacent time windows divided from the same set of time series may also include non-duplicate sequences. For example, for two adjacent time windows, the first time window includes the 1st to 5th sequences in the set of time series, while the second time window includes the 6th to 10th sequences in the same set of time series. For another example, for two adjacent time windows, the first time window includes the 1st to 3rd sequences in the set of time series, while the second time window includes the 5th to 7th sequences in the same set of time series.

[0118] In addition to the above-mentioned division of multiple time series into groups of time windows, when applying TICC, it is also necessary to adjust the objective function used by TICC.

[0119] Specifically, TICC sets an objective function when implementing time window clustering, and optimizes the clustering of time windows by solving and minimizing the objective function. The objective function set by TICC includes a penalty term related to adjacent time windows, which forces adjacent time windows to be grouped into the same category. However, in this embodiment, when multiple time windows obtained from multiple groups of time series are used as TICC input to implement clustering, the last time window in the previous group of time series and the first time window in the next group of time series, although located adjacently, should not be forced to be clustered into the same category because the two time windows are completely unrelated.

[0120] For example, in this embodiment, when clustering the multiple time windows obtained by the division, an improved TICC may be used to cluster the multiple time windows to obtain multiple categories. In the process of solving the objective function to optimize the time window clustering, the improved TICC cancels the penalty term between adjacent time windows in different time series, while retaining the penalty term between adjacent time windows in the same time series. The retained penalty term is used to encourage adjacent retained time windows in the same time series to be classified into the same category.

[0121] That is, when performing clustering on multiple time windows, multiple time windows belonging to different groups of time series are arranged together for input. Therefore, two adjacent time windows may belong to the same group of time series, or they may belong to different groups of time series (for example, two time windows at the junction of time series). Therefore, when setting the objective function, if two adjacent time windows belong to the same group of time series, the penalty term between the two adjacent time windows can be retained; if the two adjacent time windows do not belong to the same group of time series, the penalty term between the two adjacent time windows can be cancelled.

[0122] In this scheme, based on the characteristics of multiple groups of time series as clustering objects, TICC is improved to adjust the penalty term of the objective function in TICC. This can avoid clustering time windows of different groups of time series into the same category, thereby ensuring the accuracy of clustering.

[0123] To facilitate understanding, the time series analysis method provided in this embodiment will be described in detail below using the financial field as a specific example.

[0124] For example, see Figure 6 , Figure 6 A schematic diagram of a process for processing multiple time series of different companies provided in an embodiment of the present application. Figure 6 In the example shown, taking the decision-making related to the company's business strategy as an example, the processing flow mentioned in this solution can be used to find the time period with similar business strategies in the reference company and simulate the performance of the main indicators of a specific company under a certain business strategy. Figure 6 As shown, the process of processing multiple time series of different companies includes the following steps 601-604.

[0125] Step 601: Acquire multiple time series of different companies.

[0126] First, based on the companies you want to research, collect publicly released operating efficiency data from multiple companies (such as accounts payable, accounts receivable, inventory-to-revenue ratio, contract assets, and contract liabilities) and calculate these indicators periodically (e.g., annually). This way, you can calculate a value for each indicator every year.

[0127] Then, a primary indicator (i.e., the target instruction in the above embodiment) and multiple secondary indicators (i.e., the preset indicators in the above embodiment) are determined. The primary indicator is the value that needs to be simulated and obtained using the regression model, while the secondary indicators are used to input the regression model to obtain the value of the primary indicator.

[0128] Finally, the data of secondary indicators are integrated with the company as the dimension. For example, the indicator data of each company can be combined into an m*n matrix as a set of time series, where m represents the year of the indicator data and n represents the number of different indicator data. For example, if there are 5 secondary indicators and a total of 15 years of data, then each company can be integrated into a 15*5 matrix. Figure 7 As shown, Figure 7 A schematic diagram of an integrated company's indicator data provided in an embodiment of the present application. Figure 7 In the example, each company's secondary indicator data are integrated into a 15*5 matrix. Each row of the matrix contains the data of the five secondary indicators in the year represented by the current row, and the matrix contains a total of 15 rows of data.

[0129] Step 602 , clustering multiple groups of time series from different companies to obtain sequences in multiple categories. Sequences in the same category have similar business strategies.

[0130] Before clustering multiple time series, we first concatenate the indicator data of multiple companies (i.e., the matrix of each company) into a larger matrix in the row direction, that is, concatenate multiple time series of different companies. For example, assuming that the size of each company's matrix is ​​15*5, concatenating the matrices of three companies will result in a 45*5 matrix. Figure 8 As shown, Figure 8 This is a schematic diagram of the splicing of indicator data of multiple companies provided in the embodiment of this application. Figure 8 In the data, the secondary indicator data of each company are integrated into a 15*5 matrix. By splicing the secondary indicator data of each company, a 45*5 matrix can be obtained.

[0131] After concatenating the indicator data of multiple companies to create a concatenated time series, the concatenated time series is divided into time windows. When dividing the time windows, it is necessary to skip adjacent portions of indicator data between two companies. In other words, indicator data from different companies cannot be grouped into the same time window. For example, if the time window length is 5 and each company has 15 years of indicator data, each company can be divided into 11 timestamps. For example, if a company has indicator data from 2001 to 2015, the indicator data from 2001 to 2005 constitutes the first time window, the indicator data from 2002 to 2006 constitutes the second time window, the indicator data from 2003 to 2007 constitutes the third time window, and so on, the indicator data from 2011 to 2015 constitutes the 11th time window.

[0132] In this way, the improved TICC can be used to cluster the spliced ​​time series after dividing the time windows, so that time windows with similar temporal dependencies can be classified into the same category, and finally each time window is assigned a corresponding category. The goal of using the improved TICC to cluster timestamps is to solve an objective function and continuously update the category to which each time window belongs during the solution process until the value of the objective function converges. When the value of the objective function converges, the time windows under each category are the output results of the algorithm. Specifically, the objective function used by the improved TICC is obtained by improving the objective function used by the traditional TICC. The objective function used by the traditional TICC can be expressed by the following formula.

[0133]

[0134] Where K is the input cluster number, representing the total number of categories (one time window belongs to one category); P is the category matrix, used to store the category to which each time window currently belongs; λ is the regularization parameter, entered by the user; θ is the Toeplitz matrix, which contains some characteristics of the time window (each category corresponds to a Toeplitz matrix); X represents the time window; ll represents the likelihood function; β is the penalty term coefficient, entered by the user. The larger β is, the greater the probability that adjacent timestamps will be classified into the same category.

[0135] In the improved TICC, the penalty term in the objective function corresponding to the traditional TICC is modified. In the traditional TICC, the penalty term of the objective function is That is, in the time window X t Classified into category but the previous time window X t-1 Not classified into category P i When , the penalty term exists, which leads to an increase in the value of the objective function, and ultimately causes the two adjacent time windows X to t and X t-1 are classified into the same category. In the improved TICC, the penalty term in the objective function is given by Change to Here, l represents the number of time windows for each company, and mod represents the remainder. Therefore, in the improved TICC, when a time window belongs to the first time window of each company (i.e., the first time window), the penalty term for that time window is canceled, thus no longer forcing two time windows belonging to different companies but adjacent in the matrix to be classified into the same category. However, the penalty term is still retained for other time windows.

[0136] In general, the objective function in the improved TICC consists of three parts: the first is a regularization term, used to control model complexity and avoid overfitting; the second is the log-likelihood term, which measures the degree of fit of a time window to a specific category. A higher log-likelihood indicates that the data point's performance in that category is more consistent with the characteristics of that category; and the third is a penalty term, which encourages temporally adjacent time windows to be assigned to the same category, thereby maintaining the continuity and consistency of the time series data. The specific process for obtaining the objective function is shown in S6021-S6024 below.

[0137] S6021: Initialize the Toeplitz matrix θ and the category matrix P.

[0138] S6022: Fix the Toeplitz matrix θ and update the category matrix P. This step uses a dynamic programming algorithm to assign each time window to a category so as to minimize the likelihood and penalty terms of the objective function.

[0139] S6023: Fix the category matrix P and update the Toeplitz matrix θ. The previous step assigned a new category to each time window. This step requires keeping the category unchanged to find the optimal Toeplitz matrix for the current classification.

[0140] S6024: Repeat S6022 and S6023 continuously, iterate the weight matrix and the category matrix until the objective function converges, and clustering of the time window is achieved.

[0141] Finally, after completing the clustering of the time windows, we can get the time windows under each category. Since each time window includes multiple series, each category actually includes multiple series, and each series includes the values ​​of multiple secondary indicators of a company in a certain year. For example, you can refer to Figure 9 , Figure 9 This is a schematic diagram of a time window classification provided in an embodiment of the present application. Figure 9 As shown, for the indicator data of the above three companies from 2001 to 2015, the time window including the 5-year indicator data of each company can be divided into three categories (i.e., category one to category three), so that each category can have corresponding multiple sequences, where each sequence includes the values ​​of multiple secondary indicators of the company in 1 year.

[0142] Step 603: construct a regression model based on the sequence under each category.

[0143] In this step, it can be assumed that sequences within the same category correspond to the same or similar business strategies. Therefore, a corresponding regression model can be constructed based on the sequences within each category. When constructing the regression model, the primary indicator corresponding to the sequence within each category serves as the dependent variable, and the secondary indicators in the sequence serve as the independent variables, thereby achieving the construction of the regression model.

[0144] After obtaining the constructed regression model, it can be considered that this regression model contains the indicator relationship under the corresponding business strategy. Therefore, the indicators can be calculated based on this regression model, thereby realizing the simulation of the company's operations.

[0145] Step 604: Substitute the indicator data of the specific company into the regression model to obtain the simulation results output by the regression model.

[0146] Because different regression models represent different business strategies, in practical applications, by inputting a specific company's secondary indicator data (for example, the values ​​of the company's planned secondary indicators) into different regression models, the values ​​of the primary indicators under various business strategies can be obtained, thereby obtaining simulation results. In this way, the simulation results output by the regression model can be combined to make relevant business strategy decisions.

[0147] In summary, this solution, based on the improved TICC, clusters multiple time series, taking into account the temporal characteristics of each indicator and the temporal correlations between them. Furthermore, by dividing a company's time series into multiple time windows for clustering, different time windows are assigned to different Lei Xias. This accounts for changes in a company's business strategy over time, making it easier to cluster and obtain categories that reflect the actual situation.

[0148] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.

[0149] See also Figure 10 , Figure 10 This is a schematic diagram of the structure of a time series analysis device provided in an embodiment of the present application. Figure 10As shown, the time series analysis device provided by the embodiment of the present application includes: an acquisition module 1001, which is used to obtain multiple groups of time series under different dimensions, each group of time series includes multiple sequences arranged in chronological order, each sequence includes multiple values ​​of preset indicators, and each sequence also corresponds to the value of a target indicator, and the value of the target indicator is related to the values ​​of multiple preset indicators; a processing module 1002, which is used to cluster the multiple groups of time series to obtain multiple categories, each category includes at least one time window, and each time window includes adjacent partial sequences in a group of time series; the processing module 1002 is also used to construct a regression model based on the sequence under each category and the value of the target indicator corresponding to the sequence, to obtain multiple regression models corresponding to multiple categories, and the multiple regression models are all used to indicate the relationship between multiple preset indicators and the target indicator; the processing module 1002 is also used to input the preset values ​​of the multiple preset indicators into the target model in the multiple regression models to obtain the value of the target indicator output by the target model.

[0150] In a possible implementation, the plurality of preset indicators and the target indicator are any one of the following indicators: a financial indicator, a climate indicator, a health indicator, or a financial indicator.

[0151] In a possible implementation, when the plurality of preset indicators and the target indicator are all financial indicators, different groups of time series in the plurality of groups of time series correspond to different companies.

[0152] In a possible implementation, the target model corresponds to a target category among the multiple categories, and the time windows included in the target category belong to one or more target companies.

[0153] In one possible implementation, when multiple preset indicators and target indicators are all financial indicators, the multiple preset indicators are accounts payable, accounts receivable, inventory-to-revenue ratio, contract assets and contract liabilities, and the target indicator is accounts payable turnover rate.

[0154] In one possible implementation, the processing module 1002 is further used to: perform time window division on each group of time series based on a preset time window length to obtain multiple time windows, wherein the time window length is used to indicate the number of sequences included in the time window; perform clustering on the multiple time windows to obtain multiple categories, and the time windows included in each category are part of the multiple time windows.

[0155] In a possible implementation, adjacent time windows obtained by dividing the same group of time series include at least one repeated sequence.

[0156] In one possible implementation, the processing module 1002 is further configured to: perform clustering on the multiple time windows using an improved TICC to obtain multiple categories; wherein, in the process of solving the objective function to optimize the time window clustering, the improved TICC cancels the penalty terms between adjacent time windows in different time series, and retains the penalty terms between adjacent time windows in the same time series. The retained penalty terms are used to encourage adjacent retained time windows in the same time series to be classified into the same category.

[0157] The acquisition module 1001 and the processing module 1002 can be implemented in software or hardware. For example, the implementation of the processing module 1002 will be described below using the processing module 1002 as an example. Similarly, the implementation of the acquisition module 1001 can refer to the implementation of the processing module 1002.

[0158] The processing module 1002 is taken as an example of a software functional unit. The processing module 1002 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the processing module 1002 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Generally, a region may include multiple AZs.

[0159] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0160] As an example of a hardware functional unit, processing module 1002 may include at least one computing device, such as a server. Alternatively, processing module 1002 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0161] The multiple computing devices included in processing module 1002 can be distributed in the same region or in different regions. The multiple computing devices included in processing module 1002 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in processing module 1002 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0162] It should be noted that the information interaction, implementation process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the embodiment of the present application, and no further details will be given here.

[0163] This embodiment of the application also provides a computing device 1100. Figure 11 , Figure 11 This is a schematic diagram of the structure of a computing device 1100 provided in an embodiment of the present application. Figure 11 As shown, computing device 1100 includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. Processor 1104, memory 1106, and communication interface 1108 communicate with each other via bus 1102. Computing device 1100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1100.

[0164] The bus 1102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 The bus 1102 may include a path for transmitting information between various components of the computing device 1100 (eg, memory 1106, processor 1104, communication interface 1108).

[0165] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0166] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0167] The memory 1106 stores executable program code, which the processor 1104 executes to implement the functions of the acquisition module and the processing module, thereby implementing the time series analysis method. In other words, the memory 1106 stores instructions for executing the time series analysis method.

[0168] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.

[0169] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0170] See also Figure 12 , Figure 12 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. Figure 12 As shown, the computing device cluster includes at least one computing device 1100. The memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the time series analysis method.

[0171] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the time series analysis method. In other words, the combination of one or more computing devices 1100 can jointly execute the instructions for executing the time series analysis method.

[0172] It should be noted that the memory 1106 in different computing devices 1100 in the computing device cluster can store different instructions, each for executing a portion of the functions of the data processing apparatus. In other words, the instructions stored in the memory 1106 in different computing devices 1100 can implement the functions of one or more of the aforementioned acquisition module and processing module.

[0173] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 13 A possible implementation is shown. Figure 13 This is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. Figure 13 As shown, in computing device cluster 1300, two computing devices 1100A and 1100B are connected via a network. Specifically, the connection to the network is achieved through a communication interface within each computing device. In this possible implementation, memory 1106 within computing device 1100A stores instructions for executing the functions of an acquisition module. Simultaneously, memory 1106 within computing device 1100B stores instructions for executing the functions of a processing module.

[0174] It should be understood that Figure 13The functionality of the computing device 1100A shown in FIG. 1 may also be implemented by multiple computing devices 1100. Similarly, the functionality of the computing device 1100B may also be implemented by multiple computing devices 1100.

[0175] See Figure 14 , Figure 14 This is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. This application also provides a computer-readable storage medium. In some embodiments, the workflow executed by the above-mentioned database system can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or products.

[0176] Figure 14 Schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0177] In one embodiment, the computer-readable storage medium 1400 is provided using a signal-bearing medium 1401. The signal-bearing medium 1401 may include one or more program instructions 1402, which when executed by one or more processors may provide the functions or part of the functions described above for the database system.

[0178] In some examples, signal bearing medium 1401 may include computer readable medium 1403 such as, but not limited to, a hard drive, compact disk (CD), digital video disk (DVD), digital tape, memory, ROM or RAM, and the like.

[0179] In some embodiments, the signal-bearing medium 1401 may include a computer-recordable medium 1404, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1401 may include a communication medium 1405, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1401 may be communicated via a wireless form of the communication medium 1405 (e.g., a wireless communication medium conforming to the IEEE 802.X standard or other transmission protocol).

[0180] The one or more program instructions 1402 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1402 communicated to the computing device via one or more of computer-readable media 1403, computer-recordable media 1404, and / or communication media 1405.

[0181] The present application also provides a computer program product containing instructions. This computer program product can be software or a program product containing instructions that can be executed on a computing device or stored on any available medium. When the computer program product is executed on at least one computing device, it causes the at least one computing device to perform the time series analysis method described in the above embodiments.

[0182] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods of each embodiment of the present application.

[0183] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

Claims

1. A time series analysis method, characterized in that: include: Acquire multiple groups of time series in different dimensions, each group of time series including multiple sequences arranged in chronological order, each sequence including multiple preset indicator values, and each sequence also corresponding to a target indicator value, the target indicator value being related to the multiple preset indicator values; Clustering the multiple groups of time series to obtain multiple categories, each category including at least one time window, and each time window including adjacent partial sequences in a group of time series; Constructing a regression model based on the sequence under each category and the value of the target indicator corresponding to the sequence to obtain multiple regression models corresponding to the multiple categories, wherein the multiple regression models are used to indicate the relationship between the multiple preset indicators and the target indicator; The preset values ​​of the plurality of preset indicators are input into a target model among the plurality of regression models to obtain the value of the target indicator output by the target model.

2. The method according to claim 1, characterized in that The plurality of preset indicators and the target indicator are any one of the following indicators: a financial indicator, a climate indicator, a health indicator or a financial indicator.

3. The method according to claim 1 or 2, characterized in that In a case where the plurality of preset indicators and the target indicator are all financial indicators, different groups of time series in the plurality of groups of time series correspond to different companies.

4. The method according to claim 3, characterized in that The target model corresponds to a target category among the multiple categories, and the time windows included in the target category belong to one or more target companies.

5. The method according to claims 1-4, characterized in that When the multiple preset indicators and the target indicator are all financial indicators, the multiple preset indicators are accounts payable, accounts receivable, inventory-to-revenue ratio, contract assets and contract liabilities, and the target indicator is accounts payable turnover rate.

6. The method according to any one of claims 1 to 5, characterized in that The clustering of the multiple groups of time series to obtain multiple categories includes: Based on a preset time window length, each group of time series is divided into time windows to obtain multiple time windows, wherein the time window length is used to indicate the number of sequences included in the time window; Clustering is performed on the multiple time windows to obtain the multiple categories, and the time windows included in each category are part of the multiple time windows.

7. The method according to claim 6, characterized in that Adjacent time windows obtained by dividing the same group of time series include at least one repeated sequence.

8. The method according to claim 6 or 7, characterized in that The performing clustering on the multiple time windows to obtain the multiple categories includes: Using an improved Toeplitz inverse covariance matrix-based clustering method TICC, clustering the multiple time windows is performed to obtain the multiple categories; In the process of solving the objective function to optimize time window clustering, the improved TICC cancels the penalty terms between adjacent time windows in different groups of time series, and retains the penalty terms between adjacent time windows in the same group of time series. The retained penalty terms are used to encourage adjacent retained time windows in the same group of time series to be classified into the same category.

9. A time series analysis device, characterized in that: include: An acquisition module, configured to acquire multiple groups of time series in different dimensions, each group of time series including multiple sequences arranged in chronological order, each sequence including multiple preset indicator values, and each sequence also corresponding to a target indicator value, the target indicator value being related to the multiple preset indicator values; a processing module, configured to cluster the multiple groups of time series to obtain multiple categories, each category including at least one time window, and each time window including adjacent partial sequences in a group of time series; The processing module is further configured to construct a regression model based on the sequence under each category and the value of the target indicator corresponding to the sequence, to obtain a plurality of regression models corresponding to the plurality of categories, wherein the plurality of regression models are each configured to indicate a relationship between the plurality of preset indicators and the target indicator; The processing module is further configured to input the preset values ​​of the plurality of preset indicators into a target model among the plurality of regression models to obtain the value of the target indicator output by the target model.

10. The device according to claim 9, characterized in that The plurality of preset indicators and the target indicator are any one of the following indicators: a financial indicator, a climate indicator, a health indicator or a financial indicator.

11. The device according to claim 9 or 10, characterized in that In a case where the plurality of preset indicators and the target indicator are all financial indicators, different groups of time series in the plurality of groups of time series correspond to different companies.

12. The device according to claim 11, characterized in that The target model corresponds to a target category among the multiple categories, and the time windows included in the target category belong to one or more target companies.

13. The device according to claims 9-12, characterized in that When the multiple preset indicators and the target indicator are all financial indicators, the multiple preset indicators are accounts payable, accounts receivable, inventory-to-revenue ratio, contract assets and contract liabilities, and the target indicator is accounts payable turnover rate.

14. The device according to any one of claims 9 to 13, characterized in that The processing module is further configured to: Based on a preset time window length, each group of time series is divided into time windows to obtain multiple time windows, wherein the time window length is used to indicate the number of sequences included in the time window; Clustering is performed on the multiple time windows to obtain the multiple categories, and the time windows included in each category are part of the multiple time windows.

15. The device according to claim 14, characterized in that Adjacent time windows obtained by dividing the same group of time series include at least one repeated sequence.

16. The device according to claim 14 or 15, characterized in that The processing module is further configured to: Using an improved TICC, clustering is performed on the multiple time windows to obtain the multiple categories; In the process of solving the objective function to optimize time window clustering, the improved TICC cancels the penalty terms between adjacent time windows in different groups of time series, and retains the penalty terms between adjacent time windows in the same group of time series. The retained penalty terms are used to encourage adjacent retained time windows in the same group of time series to be classified into the same category.

17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster performs the operating steps of the method according to any one of claims 1 to 8.

18. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 8.

19. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 8.