Data processing method and related device
By combining panel regression analysis and significance testing with tree search, the problem of weakened dissimilarity of distance metrics in high-dimensional data clustering was solved, and higher clustering accuracy was achieved.
Patent Information
- Application Number
- CN202410608068.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to effectively cluster high-dimensional data because the distances between data points tend to be similar in high-dimensional space, which weakens the differences based on distance metrics, making accurate clustering difficult.
Panel regression analysis is used to transform high-dimensional data into a panel data structure. The difference in intercept of high-dimensional data is tested for significance, and the data is classified by combining the tree search method to achieve clustering of high-dimensional data.
It improves the accuracy of high-dimensional data clustering, overcomes the curse of dimensionality by measuring the differences in individual or time dimensions, and achieves more accurate data classification.
Smart Images

Figure CN120973990A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data clustering, and in particular, to a data processing method and related device. BACKGROUND
[0002] With the rapid development of information technology, various applications have generated massive and multi-dimensional data, such as text, image, bioinformatics data, etc. These data usually exist in the form of high dimension. Clustering of high-dimensional data has become an important research direction in the field of data mining and machine learning, aiming to group data sets with multiple attributes or characteristics, so that data objects in the same group exhibit maximum similarity to some extent.
[0003] Currently, existing schemes mainly cluster data based on distance, that is, by calculating the distance or similarity between each data point in the data set, and then dividing the data points into different clusters according to the obtained distance information, so that the distance between data points in the same cluster is as close as possible (the similarity is as high as possible), and the distance between data points in different clusters is as far as possible (the similarity is as low as possible).
[0004] However, for high-dimensional data, due to the dimension disaster phenomenon in high-dimensional space, the distance between all data points in the data set tends to be close, which greatly weakens the difference between data based on distance measurement, and further leads to difficulty in clustering high-dimensional data based on distance. SUMMARY
[0005] The present application provides a data processing method and related device for performing panel regression analysis on multiple high-dimensional data in a tree search manner to realize clustering of high-dimensional data and improve the accuracy of clustering of high-dimensional data.
[0006] Therefore, in a first aspect, the present application provides a data processing method, which comprises: first, obtaining a data set comprising multiple high-dimensional data; then, performing panel regression analysis on the multiple high-dimensional data to obtain a panel data regression model of the multiple high-dimensional data; performing significance test on multiple first intercepts in the panel data regression model to obtain multiple first significance test results, wherein the first intercepts can represent the difference between the multiple high-dimensional data; and dividing the multiple high-dimensional data into multiple data categories according to the multiple first significance test results, wherein each data category comprises at least one high-dimensional data.
[0007] In the embodiments of the present application, the panel regression analysis can be used to analyze the differences between different individuals or different times corresponding to the plurality of high-dimensional data by constructing a panel data regression model, and the intercepts in the panel data regression model can represent the differences between different individuals or the differences between different times, so that the plurality of high-dimensional data can be clustered according to the individual or time differences by combining the significant test results obtained by performing the significant test on the intercepts, thereby avoiding clustering the high-dimensional data based on the distance, and the difference between the high-dimensional data measured based on the intercepts obtained by performing the panel regression analysis on the high-dimensional data is much stronger than the difference between the high-dimensional data measured based on the distance, so that the accuracy of clustering the high-dimensional data can be improved.
[0008] In a possible implementation, the panel regression analysis on the plurality of high-dimensional data to obtain the panel data regression model of the plurality of high-dimensional data can include: converting the plurality of high-dimensional data into a panel data structure to obtain panel data, the panel data including the plurality of high-dimensional data; and performing panel regression analysis on the panel data to obtain the panel data regression model.
[0009] In the embodiments of the present application, since the panel regression analysis is a regression method based on panel data, the plurality of panel data can be converted into a panel data structure before performing the panel regression analysis on the plurality of high-dimensional data, so that each high-dimensional data corresponds to an individual identifier and a time identifier, thereby facilitating the subsequent panel regression analysis of the high-dimensional data according to the individual or time dimensions.
[0010] In a possible implementation, the panel regression analysis on the panel data to obtain the panel data regression model can include: determining the category of the panel data regression model according to the panel data; and estimating the parameters in the panel data regression model by using a preset estimation method to obtain the panel data regression model, the parameters including regression coefficients and a plurality of first intercepts.
[0011] In a possible implementation, the significant test on the plurality of first intercepts in the panel data regression model to obtain a plurality of first significant test results can include: performing the significant test on the plurality of first intercepts according to a preset significant test method and a preset significant test index to obtain the plurality of first significant test results.
[0012] In a possible implementation, the division of the plurality of high-dimensional data into a plurality of data categories according to the plurality of first significant test results can include: dividing the plurality of high-dimensional data into the plurality of data categories by using a tree search method according to the plurality of first significant test results.
[0013] In a possible implementation, the aforementioned dividing the plurality of high-dimensional data into a plurality of data categories according to the plurality of first significance test results can include: dividing the plurality of high-dimensional data into a first up-to-standard data set and a first not-up-to-standard data set according to the plurality of first significance test results, the first up-to-standard data set including one or more high-dimensional data, and the first not-up-to-standard data set including one or more high-dimensional data; performing panel regression analysis on the high-dimensional data in the first not-up-to-standard data set to obtain a plurality of second intercepts of a panel data regression model of the first not-up-to-standard data set, the second intercepts representing differences between the high-dimensional data in the first not-up-to-standard data set; performing significance test on the plurality of second intercepts to obtain a plurality of second significance test results; and dividing the plurality of high-dimensional data in the first not-up-to-standard data set into a second up-to-standard data set and a second not-up-to-standard data set according to the plurality of second significance test results, until there is no high-dimensional data in the second not-up-to-standard data set that passes the significance test, and the dividing is stopped.
[0014] In the embodiments of the present application, the obtained panel data can be subjected to panel regression cycle analysis in a tree search manner. By performing panel regression analysis on the panel data layer by layer, the plurality of high-dimensional data can be divided into a plurality of data categories, thereby realizing clustering of the high-dimensional data.
[0015] In a possible implementation, the panel data regression model is a fixed effect model or a random effect model.
[0016] In a second aspect, the present application provides a data processing apparatus, comprising:
[0017] an acquisition module configured to acquire a data set, the data set including a plurality of high-dimensional data;
[0018] an analysis module configured to perform panel regression analysis on the plurality of high-dimensional data to obtain a panel data regression model of the plurality of high-dimensional data;
[0019] a test module configured to perform significance test on a plurality of first intercepts in the panel data regression model to obtain a plurality of first significance test results, the first intercepts representing differences between the plurality of high-dimensional data;
[0020] a division module configured to divide the plurality of high-dimensional data into a plurality of data categories according to the plurality of first significance test results, one data category including at least one high-dimensional data.
[0021] In a possible implementation, the aforementioned analysis module is specifically configured to: convert the plurality of high-dimensional data into a panel data structure to obtain panel data, the panel data including the plurality of high-dimensional data; and perform panel regression analysis on the panel data to obtain the panel data regression model.
[0022] In a possible implementation, the analysis module is specifically configured to: determine a category of the panel data regression model according to the panel data; and estimate parameters in the panel data regression model by using a preset estimation method, to obtain the panel data regression model, the parameters including regression coefficients and the plurality of first intercepts.
[0023] In a possible implementation, the inspection module is specifically configured to: perform significance inspection on the plurality of first intercepts according to a preset significance inspection method and a preset significance inspection index, to obtain a plurality of first significance inspection results.
[0024] In a possible implementation, the division module is specifically configured to: divide the plurality of high-dimensional data into a plurality of data categories by using a tree search method according to the plurality of first significance inspection results.
[0025] In a possible implementation, the division module is specifically configured to: divide the plurality of high-dimensional data into a first up-to-standard data set and a first substandard data set according to the plurality of first significance inspection results, the first up-to-standard data set including one or more high-dimensional data, and the first substandard data set including one or more high-dimensional data; perform panel regression analysis on the high-dimensional data in the first substandard data set, to obtain a plurality of second intercepts of a panel data regression model of the first substandard data set, the second intercepts representing differences between the high-dimensional data in the first substandard data set; perform significance inspection on the plurality of second intercepts, to obtain a plurality of second significance inspection results; and divide the plurality of high-dimensional data in the first substandard data set into a second up-to-standard data set and a second substandard data set according to the plurality of second significance inspection results, until there is no high-dimensional data in the second substandard data set that passes the significance inspection, and the division is stopped.
[0026] In a possible implementation, the panel data regression model is a fixed effect model or a random effect model.
[0027] In a third aspect, an embodiment of the present application provides a computing device, including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device performs the method in the first aspect or any possible implementation manner of the first aspect.
[0028] In a fourth aspect, an embodiment of the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method in the first aspect or any possible implementation manner of the first aspect.
[0029] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, including computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method according to the first aspect or any possible implementation manner of the first aspect.
[0030] In a sixth aspect, an embodiment of the present application provides a computer program product including instructions, when the instructions are run by a computing device cluster, the computing device cluster executes the method according to the first aspect or any possible implementation manner of the first aspect.
[0031] The technical effects brought by the second aspect to the sixth aspect or any possible implementation manner thereof can refer to the technical effects brought by the first aspect or the related possible implementation manner of the first aspect, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 A system framework schematic diagram is provided for an embodiment of the present application;
[0033] Figure 2 A flowchart of a data classification process is provided for an embodiment of the present application;
[0034] Figure 3 A flowchart of a data processing method is provided for an embodiment of the present application;
[0035] Figure 4 A flowchart of clustering multiple high-dimensional data in an individual data set is provided for an embodiment of the present application;
[0036] Figure 5 A structure schematic diagram of a data processing apparatus is provided for an embodiment of the present application;
[0037] Figure 6 A structure schematic diagram of a computing device is provided for an embodiment of the present application;
[0038] Figure 7 A structure schematic diagram of a computing device cluster is provided for an embodiment of the present application;
[0039] Figure 8 Another structure schematic diagram of a computing device cluster is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0041] In this application, "at least one" means one or more, "multiple" means two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. The terms "first", "second", etc. in the specification and claims of this application and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a way of distinguishing the objects with the same properties in the description of the embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the process, method, system, product or equipment containing a series of units does not have to be limited to those units, but can include other units not clearly listed or inherent to these processes, methods, products or equipment.
[0042] First, some concepts related to the embodiments of the present application are introduced.
[0043] 1. High-dimensional data
[0044] High-dimensional data generally refers to data points with multiple dimensions. In data and computer science, the dimension of a data point refers to the number of independent parameters that can be used to describe the data point.
[0045] 2. Curse of dimensionality
[0046] The curse of dimensionality phenomenon refers to the fact that in high-dimensional space, the distance between data points becomes very scattered, resulting in sparse, redundant and difficult-to-analyze data. As the dimension of the feature increases, the distance between data points becomes larger and larger, which makes it difficult to accurately estimate and represent the distribution of data in high-dimensional space, while also increasing the difficulty of training for limited samples. The curse of dimensionality involves digital analysis, sampling, combination, machine learning, data mining and databases, and many other fields. In these fields, when the dimension increases, the volume of the space increases rapidly, making the available data sparse.
[0047] 3. Panel data
[0048] Panel data, also known as parallel data or TS-CS data (Time Series-Cross Section), refers to the index data of different individuals at different times, including individual dimensions and time dimensions.
[0049] 4. Panel data regression
[0050] Panel data regression is a regression analysis method suitable for panel data, aiming to explore the relationship between different individuals and time changes. Specifically, panel data regression is a method of modeling and analyzing data across time and across individuals. Panel data consists of observations of multiple individuals at multiple time points, unlike cross-sectional data and time series data. This method can be used to study the differences between individuals and the impact of time trends on them, and has important application value in economics and social sciences.
[0051] 5. Tree search
[0052] Tree search (also known as tree retrieval or tree search algorithm) is a widely used search strategy in computer science, which searches based on tree data structure. This search method is commonly used in various algorithms and applications, especially in the fields of databases, file systems and computer graphics.
[0053] In tree search, data is usually organized in a tree structure, with each node representing a data item and the relationship between nodes represented by links or pointers. The search starts from the root node and moves down the tree path, checking each node according to the specific search conditions until the target node is found or it is determined that the target node does not exist in the tree.
[0054] The specific implementation of tree search algorithm depends on the tree structure used and the search conditions. One common tree search algorithm is Depth-First Search (DFS), which traverses the nodes of the tree according to the depth of the tree, searching as deeply as possible for branches of the tree. Another common algorithm is Breadth-First Search (BFS), which traverses the nodes of the tree according to the level of the tree, first visiting all neighboring nodes and then layer by layer down.
[0055] With the development of artificial intelligence technology, high-dimensional data has appeared in more and more fields. For example, in the financial field, portfolio analysis, credit default analysis, customer behavior analysis and other scenarios often require the use of high-dimensional data for analysis. For example, customer behavior analysis can be used to cluster similar preferences by clustering high-dimensional data.
[0056] Currently, the method for clustering high-dimensional data usually adopts distance-based method, such as K-means clustering algorithm, etc. However, for high-dimensional data, there is a dimension disaster phenomenon. With the increase of the dimension of data, the distance between data points in high-dimensional space tends to be the same, which makes it difficult for traditional distance-based clustering algorithms to cluster high-dimensional data.
[0057] Based on this, the present application provides a data processing method, which can perform panel regression analysis on multiple high-dimensional data, then perform significance analysis on the intercept of the panel data regression model obtained by the panel regression analysis, and divide the multiple high-dimensional data into multiple data categories based on the significance analysis result.
[0058] The system architecture provided by the embodiments of the present application is introduced below.
[0059] Referring to Figure 1 , a system architecture 100 is provided by the present application. As Figure 1 shown, the system architecture 100 can include a computing device cluster 110, which includes at least one computing device. The computing device can be a server, such as a cloud server, a central server, an edge server, or a local server of a local data center. In some possible embodiments, the computing device can also be a desktop computer, a notebook computer, or other terminal device.
[0060] Exemplarily, the computing device cluster can include one or more servers, which call related microservices to implement the data processing method of the embodiments of the present application after obtaining multiple high-dimensional data. Alternatively, the computing device cluster can also be a client device, which implements the data processing method of the embodiments of the present application through an installed application program or plug-in, etc.
[0061] For example, as Figure 2 shown in the flowchart of data classification, the computing device cluster 110 can first obtain initial data, perform panel regression cycle analysis on the initial data, obtain substandard data and standard data A, perform panel regression cycle analysis on the substandard data, obtain substandard data and standard data B, repeat the foregoing steps to obtain substandard data and standard data C, and so on, until there is no standard data in the substandard data, stop dividing, and thus obtain multiple data categories. Each data in the finally obtained substandard data is a data category, and the foregoing standard data A, standard data B and standard data C are respectively a data category.
[0062] It is worth noting that Figure 1The system architecture shown is only as an example, and is not used to limit the specific implementation of the example. For example, in other possible system architectures, the system architecture 100 further includes a database for storing high-dimensional data, and can send the high-dimensional data to be processed to the computing device.
[0063] The method flow provided by the present application is introduced below in combination with the foregoing system architecture.
[0064] Referring to Figure 3 The present application provides a flow diagram of a data processing method, as follows.
[0065] 301, obtaining a data set;
[0066] In the embodiment of the present application, the data set includes a plurality of high-dimensional data, wherein the high-dimensional data is a data point with multiple dimensions. For example, an image data set includes a plurality of images, and each image can be regarded as a high-dimensional data, because an image has thousands of pixel points, and each pixel point can represent a dimension (such as color intensity). For another example, in a text data set, a document can be regarded as a high-dimensional data, and each word or phrase in the document can represent a dimension. For another example, in an enterprise data set, the data set includes the financial status of a plurality of enterprises, and the financial status of an enterprise can be regarded as a high-dimensional data, and a plurality of indicators affecting the financial status of the enterprise represent a dimension, for example, a plurality of financial indicators such as operating income, net profit, profit margin, or asset return rate.
[0067] 302, performing panel regression analysis on the plurality of high-dimensional data in the data set to obtain a panel data regression model of the plurality of high-dimensional data;
[0068] The panel regression analysis is a regression analysis method suitable for panel data, and the panel data includes observation data of a plurality of individuals at a plurality of time points, and the panel data can also be regarded as an m*n data matrix, recording a data indicator of m individuals at n time nodes. This method can be used to analyze the differences between different individuals and the influence of time changes. The individuals can be individuals, enterprises, or industries, and the time points can be units of years, months, days, hours, minutes, or seconds, and the specific units are not limited here.
[0069] Therefore, in the embodiment of the present application, the plurality of high-dimensional data can be unfolded according to the individual dimension or the time dimension of the panel data, and the panel regression analysis is performed, so that the plurality of high-dimensional data can be analyzed according to the differences between different individuals or the influence of time changes, to realize the clustering of the high-dimensional data.
[0070] Optionally, the plurality of high-dimensional data can be converted into a panel data structure to obtain panel data, the panel data comprising the plurality of high-dimensional data. Since the panel data comprises individual dimensions and time dimensions, the plurality of high-dimensional data can be arranged according to the individual dimensions and the time dimensions, one high-dimensional data corresponding to one individual identifier and one time identifier. Subsequently, the panel data is subjected to panel regression analysis to obtain a panel data regression model.
[0071] Specifically, according to the panel data, the category of the panel data regression model can be determined first. The panel data regression model can comprise a fixed effect model or a random effect model, etc. The fixed effect model comprises an individual fixed effect model, a time fixed effect model, and an individual time fixed effect model. Subsequently, a preset estimation method can be used to estimate the parameters of the determined panel data regression model to obtain the panel data regression model.
[0072] The preset estimation method can be ordinary least squares (OLS) or maximum likelihood estimation (MLE), which is not limited here.
[0073] Since the intercept terms of different individuals in the fixed effect model are fixed, but the intercept terms of each individual are different. Therefore, in the embodiments of the present application, the panel data regression model can use the individual fixed effect model or the time fixed effect model in the fixed effect model. When the panel data regression model is the individual fixed effect model, different individuals correspond to different intercepts, which can represent the differences between different individuals, so that subsequent analysis can be performed according to the intercepts, thereby realizing clustering of high-dimensional data.
[0074] For example, according to the financial status of enterprises, a plurality of enterprises can be classified, and enterprises with similar financial status can be divided into a category. A plurality of financial indicators related to different enterprises can be obtained first, such as operating income, net profit, profit margin, or asset return rate, etc. The plurality of financial indicators collectively constitute high-dimensional data for analyzing the financial status of enterprises. Subsequently, the individual fixed effect model can be used to perform panel regression analysis on the plurality of obtained high-dimensional data, and the aforementioned estimation method can be used to estimate the parameters of the individual fixed effect model, thereby obtaining the individual fixed effect model. Different enterprises in the individual fixed effect model correspond to different intercepts, which represent the differences between different enterprises.
[0075] Specifically, the expression of the individual fixed effect model can be:
[0076]
[0077] wherein y itLet x represent the value of the dependent variable for the i-th individual at time t. kit Let λ represent the value of the k-th independent variable for the i-th individual at time t. i Let u represent the fixed effect for the i-th individual. This can be understood as each individual having a separate intercept term or constant term, with different individuals having different fixed effects. it Let represent the random perturbation term of the i-th individual at time t.
[0078] 303. Perform significance tests on multiple first intercepts in the panel data regression model and obtain multiple first significance test results;
[0079] After obtaining the panel data regression model, the significance of multiple first intercepts in the model can be tested. By using the preset significance index, it can be determined whether multiple first intercepts pass the significance test, and multiple first significance test results can be obtained. Among them, the first intercept can represent the difference between multiple high-dimensional data.
[0080] Optionally, multiple first intercepts can be tested for significance based on preset significance test methods and preset significance test indices to obtain multiple first significance test results. The preset significance test methods include commonly used test methods such as t-test and F-test, and the preset test indices are usually values such as 0.1, 0.05, or 0.02, etc., which are not limited here.
[0081] Specifically, this can be based on the estimated value λ of each of the multiple first intercepts. i and the standard error SE(λ) corresponding to this estimate. i Assuming each intercept is 0, the t-statistic corresponding to the first intercept is calculated, where the formula for calculating the t-statistic satisfies the formula t = (λ / 2) / 2. i -0) / SE(λ i Subsequently, we can determine whether the absolute value of the t-statistic is greater than the critical value, or obtain the corresponding p-value based on the t-statistic and determine whether the p-value is less than the preset significance test index. If the absolute value of the t-statistic is greater than the critical value or the p-value is less than the preset significance test index, then the first intercept is significant; otherwise, the first intercept is not significant, thus obtaining the first significance test result for the first intercept. The critical value and p-value can be obtained from the t-distribution table.
[0082] 304. Based on the results of multiple first significance tests, divide multiple high-dimensional data into multiple data categories.
[0083] After obtaining the plurality of first significance test results, the plurality of high-dimensional data can be divided into a plurality of data categories by using a tree search method according to the significance test results, so as to divide the high-dimensional data corresponding to the individuals or time with consistent significance results into the same category.
[0084] Optionally, after obtaining the plurality of first significance test results, the plurality of high-dimensional data can be first divided into a first qualified data set and a first unqualified data set according to the test results, that is, the individuals corresponding to the first intercepts with significant first significance test results are divided into a category (referred to as the first qualified data set), and the high-dimensional data corresponding to the individuals are divided into the category. Similarly, the high-dimensional data corresponding to the first intercepts with insignificant first significance test results can be divided into the first unqualified data set, so as to preliminarily divide the plurality of high-dimensional data into two categories, and the first qualified data set can be marked as category A.
[0085] Subsequently, the high-dimensional data in the first unqualified data set can be subjected to panel regression analysis again to obtain a plurality of second intercepts of the panel data regression model of the first unqualified data set, and the plurality of second intercepts can be subjected to significance test to obtain a plurality of second significance test results; and then, according to the plurality of second significance test results, the plurality of high-dimensional data in the first unqualified data set can be divided into a second qualified data set and a second unqualified data set, until there is no high-dimensional data in the second unqualified data set passing the significance test, or all the high-dimensional data in the second unqualified data set do not pass the significance test, and the division is stopped, that is, by performing panel regression analysis on the high-dimensional data in the unqualified data set round by round, and performing significance test on the obtained intercepts, until there is no intercept corresponding to the high-dimensional data passing the significance test, the division of the high-dimensional data is stopped.
[0086] Among them, the second qualified data set can be marked as category B, if there is an intercept corresponding to the high-dimensional data in the second unqualified data set satisfying the preset significance test index, the high-dimensional data in the second unqualified data set can be divided into a third qualified data set and a third unqualified data set, at this time, the third qualified data set can be marked as category C, if there is no high-dimensional data passing the significance test in the third unqualified data set at this time, the division is stopped, at this time, each high-dimensional data in the third unqualified data set is a category, and each category includes at least one high-dimensional data. According to the analysis of the differences between each individual, similar individuals are divided into a category, and then the high-dimensional data corresponding to the individuals in the same category can be clustered into a category, so as to realize the clustering of the plurality of high-dimensional data.
[0087] In this embodiment, a tree-structured search method is used to classify multiple high-dimensional data points layer by layer. Panel regression analysis and significance tests are performed on the high-dimensional data in each round, grouping individuals with high similarity into the same category until the remaining high-dimensional data is dispersed, thus obtaining as many data categories as possible to achieve clustering of multiple high-dimensional data. Furthermore, measuring the differences between high-dimensional data based on the differences between different individuals or between different time periods is far more effective than measuring differences based on distance, thereby improving the accuracy of clustering high-dimensional data.
[0088] Furthermore, since the high-dimensional data is divided into compliant and non-compliant datasets based on the significance test results after each round, and panel regression analysis is then performed only on the high-dimensional data in the non-compliant dataset, the parameters (slope and intercept) of the panel data regression model obtained in each round are different. Therefore, the preset significance index set when performing significance testing on the intercept in each round can be different. For example, the preset significance index in the first round can be 0.05, and the preset significance index in the second round can be 0.02. No specific limitation is made here.
[0089] For example, such as Figure 4 As shown, multiple high-dimensional data points in the acquired individual dataset can be expanded from the individual dimension, that is, multiple high-dimensional data points can be arranged according to their corresponding individuals. For example, the individuals here can be enterprises. There are 27 initial individuals (enterprises). Now, it is necessary to classify enterprises with similar financial conditions into the same category, and then classify the high-dimensional data corresponding to enterprises in the same category into the same category. First, panel regression analysis is performed on all high-dimensional data corresponding to the initial enterprises, and the significance of the intercepts of different enterprises is tested. According to the test results, the initial enterprises are divided into qualified enterprise category A (11 enterprises in total) and non-qualified enterprises. Subsequently, panel regression analysis is used step by step to search on the tree structure, and finally 10 enterprise categories are obtained. There are 11 enterprises in qualified enterprise category A, 3 enterprises in qualified enterprise category B, and 3 enterprises in qualified enterprise category C. Finally, the remaining 7 enterprises are all non-qualified, and the 7 non-qualified enterprises are in a separate category.
[0090] In this embodiment, multiple high-dimensional data can be converted into panel data, and panel regression analysis can be performed on the panel data to obtain a panel data regression model. Subsequently, significance tests can be performed on multiple intercepts in the panel data regression model to obtain significance test results. Different intercepts can represent differences between different individuals or different times in the panel data. Based on the significance test results, multiple high-dimensional data can be classified from an individual dimension or a time dimension, thereby achieving clustering of high-dimensional data.
[0091] The foregoing introduces the method flow provided by the present application. Based on the foregoing method flow, the device provided by the present application is introduced as follows.
[0092] Referring to Figure 5 The data processing device provided by the present application has the structure as described below.
[0093] The acquisition module 501 is configured to acquire a data set, the data set including a plurality of high-dimensional data.
[0094] The analysis module 502 is configured to perform panel regression analysis on the plurality of high-dimensional data to obtain a panel data regression model of the plurality of high-dimensional data.
[0095] The test module 503 is configured to perform significance test on a plurality of first intercepts in the panel data regression model to obtain a plurality of first significance test results, the first intercepts representing differences between the plurality of high-dimensional data.
[0096] The division module 504 is configured to divide the plurality of high-dimensional data into a plurality of data categories according to the plurality of first significance test results, one data category including at least one high-dimensional data.
[0097] In a possible implementation, the analysis module 502 is specifically configured to: convert the plurality of high-dimensional data into a panel data structure to obtain panel data, the panel data including the plurality of high-dimensional data; and perform panel regression analysis on the panel data to obtain the panel data regression model.
[0098] In a possible implementation, the analysis module 502 is specifically configured to: determine a category of the panel data regression model according to the panel data; and estimate parameters in the panel data regression model by using a preset estimation method to obtain the panel data regression model, the parameters including regression coefficients and the plurality of first intercepts.
[0099] In a possible implementation, the test module 503 is specifically configured to: perform significance test on the plurality of first intercepts according to a preset significance test method and a preset significance test index to obtain the plurality of first significance test results.
[0100] In a possible implementation, the division module 504 is specifically configured to: divide the plurality of high-dimensional data into the plurality of data categories by using a tree search method according to the plurality of first significance test results.
[0101] In a possible implementation, the division module 504 is specifically configured to: divide the plurality of high-dimensional data into a first up-to-standard data set and a first not-up-to-standard data set according to the plurality of first significance test results, the first up-to-standard data set including one or more high-dimensional data, and the first not-up-to-standard data set including one or more high-dimensional data; perform panel regression analysis on the high-dimensional data in the first not-up-to-standard data set to obtain a plurality of second intercepts of a panel data regression model of the first not-up-to-standard data set, the second intercepts representing differences between the high-dimensional data in the first not-up-to-standard data set; perform significance test on the plurality of second intercepts to obtain a plurality of second significance test results; and divide the plurality of high-dimensional data in the first not-up-to-standard data set into a second up-to-standard data set and a second not-up-to-standard data set according to the plurality of second significance test results, until there is no high-dimensional data in the second not-up-to-standard data set that passes the significance test, and the division is stopped.
[0102] In a possible implementation, the panel data regression model is a fixed effects model or a random effects model.
[0103] The acquisition module, the analysis module, the test module, and the division module can be implemented by software or by hardware. For example, the implementation of the acquisition module is described below. Similarly, the implementation of the analysis module, the test module, and the division module can refer to the implementation of the acquisition module.
[0104] As an example of a software functional unit, the acquisition module can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the acquisition module can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region) or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple data centers in a similar geographical location. Generally, one region can include multiple AZs.
[0105] Likewise, the plurality of hosts / virtual machines / containers for running the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, usually one VPC is set in one region, and communication between two VPCs in the same region or between VPCs in different regions needs to set a communication gateway in each VPC to realize the interconnection between VPCs through the communication gateway.
[0106] As an example of a hardware functional unit, the obtaining module can include at least one computing device, such as a server, etc. Alternatively, the obtaining module can also be a device implemented by a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or a programmable logic device (PLD), etc. Among them, the above-mentioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system on chip (SoC), an offload card, an acceleration card, or any combination thereof.
[0107] The plurality of computing devices included in the obtaining module can be distributed in the same region or in different regions. The plurality of computing devices included in the obtaining module can be distributed in the same AZ or in different AZs. Likewise, the plurality of computing devices included in the obtaining module can be distributed in the same VPC or in multiple VPCs. Among them, the plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offload cards, acceleration cards, etc.
[0108] It should be noted that in other embodiments, the obtaining module can be configured to perform any of the steps of the data processing method, the analyzing module can be configured to perform any of the steps of the data processing method, the verifying module can be configured to perform any of the steps of the data processing method, the dividing module can be configured to perform any of the steps of the data processing method, and the steps implemented by the obtaining module, the analyzing module, the verifying module, and the dividing module can be specified as needed, and the overall function of the data processing apparatus can be implemented by the obtaining module, the analyzing module, the verifying module, and the dividing module implementing different steps of the data processing method.
[0109] The present application also provides a computing device 600. As shown in Figure 6 The computing device 600 includes a bus 602, a processor 604, a memory 606, and a communication interface 608. The processor 604, the memory 606, and the communication interface 608 communicate with each other through the bus 602. The computing device 600 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 600.
[0110] The bus 602 can be a peripheral component interconnect Express (PCIe) bus or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. Among them, the unified bus is also called a flexible bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only one line is used, but it does not mean that there is only one bus or one type of bus. The bus 604 can include a path for transmitting information between various components of the computing device 600 (e.g., the memory 606, the processor 604, the communication interface 608). Among them, the unified bus can also be called a flexible bus.
[0111] The processor 604 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), an ASIC, an FPGA, a CPLD, an NPU, a SoC, an offload card, an acceleration card, or other computing device.
[0112] The memory 606 can include volatile memory, such as random access memory (RAM). The processor 604 can further include non-volatile memory, such as read-only memory (ROM), Flash memory, a hard disk drive (HDD), or a solid-state drive (SSD). In addition, the memory 606 can also be implemented by a storage class memory (SCM), a phase change memory (PCM), or other types of storage media.
[0113] It is worth noting that the same type of storage medium can be configured to implement the function of the memory 606 in the same computing device, or two or more types of storage media can be configured to implement the function of the memory 606, which is not limited in the present application.
[0114] The memory 606 stores executable program code, and the processor 604 executes the executable program code to respectively implement the functions of the aforementioned acquisition module, analysis module, verification module, and division module, thereby implementing the data processing method. That is, the memory 606 stores instructions for executing the data processing method.
[0115] The communication interface 608 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 600 and other devices or communication networks.
[0116] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.
[0117] AsFigure 7 As shown, the computing device cluster includes at least one computing device 600. The memory 606 in one or more computing devices 600 in the computing device cluster can have the same instructions for performing the data processing method.
[0118] In some possible implementation manners, the memory 606 in one or more computing devices 600 in the computing device cluster can also respectively have partial instructions for performing the data processing method. In other words, the combination of one or more computing devices 600 can collectively perform the instructions for performing the data processing method.
[0119] It should be noted that the memories 606 in different computing devices 600 in the computing device cluster can store different instructions, respectively for performing partial functions of the data processing apparatus. That is, the instructions stored in the memories 606 in different computing devices 600 can implement the functions of one or more of the obtaining module, the analyzing module, the verifying module, and the dividing module.
[0120] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network, a local area network, or the like. Figure 8 A possible implementation manner is shown. As shown in Figure 8 The two computing devices 600A and 600B are connected through a network. Specifically, the communication interfaces in the respective computing devices are connected to the network. In this type of possible implementation manner, the memory 606 in the computing device 600A has instructions for performing the functions of the obtaining module. Meanwhile, the memory 606 in the computing device 600B has instructions for performing the functions of the analyzing module, the verifying module, and the dividing module.
[0121] It should be understood that Figure 8 The functions of the computing device 600A shown in the foregoing
[0122] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection relationship between the computing devices in the computing device cluster shown in Figure 7 and Figure 8 The connection manner of the computing device cluster. The difference is that the memory 606 in one or more computing devices 600 in the computing device cluster can have the same instructions for performing the data processing method.
[0123] In some possible implementations, partial instructions for performing the data processing method can also be respectively stored in the memory 606 of one or more computing devices 600 in the computing device cluster. In other words, the combination of one or more computing devices 600 can collectively execute the instructions for performing the data processing method.
[0124] The embodiments of the present application further provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, the at least one computing device is caused to perform the method provided by the present application.
[0125] The embodiments of the present application further provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, the at least one computing device is caused to perform the method provided by the present application.
[0126] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, characterized in that, include: Obtain a dataset, which includes multiple high-dimensional data; Panel regression analysis was performed on the multiple high-dimensional data to obtain a panel data regression model for the multiple high-dimensional data. Significance tests are performed on multiple first intercepts in the panel data regression model to obtain multiple first significance test results, where the first intercept represents the difference between the multiple high-dimensional data. Based on the results of the multiple first significance tests, the multiple high-dimensional data are divided into multiple data categories, and each data category includes at least one of the high-dimensional data.
2. The method according to claim 1, characterized in that, The panel regression analysis of the multiple high-dimensional data to obtain the panel data regression model for the multiple high-dimensional data includes: The multiple high-dimensional data are transformed into a panel data structure to obtain panel data, which includes the multiple high-dimensional data. Panel regression analysis was performed on the panel data to obtain the panel data regression model.
3. The method according to claim 2, characterized in that, The panel regression analysis of the panel data to obtain the panel data regression model includes: Based on the panel data, determine the category of the panel data regression model; The parameters in the panel data regression model are estimated using a preset estimation method to obtain the panel data regression model, wherein the parameters include regression coefficients and the plurality of first intercepts.
4. The method according to any one of claims 1 to 3, characterized in that, The significance test is performed on multiple first intercepts in the panel data regression model to obtain multiple first significance test results, including: The multiple first intercepts are subjected to significance tests according to the preset significance test method and preset significance test index to obtain the multiple first significance test results.
5. The method according to any one of claims 1 to 4, characterized in that, Based on the results of the multiple first significance tests, the multiple high-dimensional data are divided into multiple data categories, including: Based on the results of the multiple first significance tests, the multiple high-dimensional data are divided into multiple data categories using a tree search method.
6. The method according to claim 5, characterized in that, Based on the results of the multiple first significance tests, a tree-based search method is used to divide the multiple high-dimensional data into multiple data categories, including: Based on the results of the multiple first significance tests, the multiple high-dimensional data are divided into a first qualified dataset and a first unqualified dataset. The first qualified dataset includes one or more high-dimensional data, and the first unqualified dataset includes one or more high-dimensional data. The panel regression analysis is performed on the high-dimensional data in the first substandard dataset to obtain multiple second intercepts of the panel data regression model of the first substandard dataset. The second intercepts represent the differences between the high-dimensional data in the first substandard dataset. The significance test is performed on the plurality of second intercepts to obtain a plurality of second significance test results; Based on the results of the multiple second significance tests, the multiple high-dimensional data in the first substandard dataset are divided into a second compliant dataset and a second substandard dataset, until no high-dimensional data in the second substandard dataset passes the significance test, at which point the division stops.
7. The method according to any one of claims 1 to 6, characterized in that, The panel data regression model is either a fixed effects model or a random effects model.
8. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire a dataset, which includes multiple high-dimensional data. The analysis module is used to perform panel regression analysis on the multiple high-dimensional data to obtain a panel data regression model for the multiple high-dimensional data. The testing module is used to perform significance tests on multiple first intercepts in the panel data regression model to obtain multiple first significance test results, where the first intercept represents the difference between the multiple high-dimensional data. The partitioning module is used to partition the multiple high-dimensional data into multiple data categories based on the multiple first significance test results, wherein each data category includes at least one of the high-dimensional data.
9. The apparatus according to claim 8, characterized in that, The analysis module is specifically used for: The multiple high-dimensional data are transformed into a panel data structure to obtain panel data, which includes the multiple high-dimensional data. Panel regression analysis was performed on the panel data to obtain the panel data regression model.
10. The apparatus according to claim 9, characterized in that, The analysis module is specifically used for: Based on the panel data, determine the category of the panel data regression model; The parameters in the panel data regression model are estimated using a preset estimation method to obtain the panel data regression model, wherein the parameters include regression coefficients and the plurality of first intercepts.
11. The apparatus according to any one of claims 8 to 10, characterized in that, The inspection module is specifically used for: The multiple first intercepts are subjected to significance tests according to the preset significance test method and preset significance test index to obtain the multiple first significance test results.
12. The apparatus according to any one of claims 8 to 11, characterized in that, The partitioning module is specifically used for: Based on the results of the multiple first significance tests, the multiple high-dimensional data are divided into multiple data categories using a tree search method.
13. The apparatus according to claim 12, characterized in that, The partitioning module is specifically used for: Based on the results of the multiple first significance tests, the multiple high-dimensional data are divided into a first qualified dataset and a first unqualified dataset. The first qualified dataset includes one or more high-dimensional data, and the first unqualified dataset includes one or more high-dimensional data. The panel regression analysis is performed on the high-dimensional data in the first substandard dataset to obtain multiple second intercepts of the panel data regression model of the first substandard dataset. The second intercepts represent the differences between the high-dimensional data in the first substandard dataset. The significance test is performed on the plurality of second intercepts to obtain a plurality of second significance test results; Based on the results of the multiple second significance tests, the multiple high-dimensional data in the first substandard dataset are divided into a second compliant dataset and a second substandard dataset, until no high-dimensional data in the second substandard dataset passes the significance test, at which point the division stops.
14. The apparatus according to any one of claims 8 to 13, characterized in that, The panel data regression model is either a fixed effects model or a random effects model.
15. A computing device, characterized in that, The computing device includes a processor and memory; The processor is configured to execute instructions stored in the memory to cause the computing device to perform the method as described in any one of claims 1 to 7.
16. A computing device cluster, characterized in that, It includes at least one computing device, said at least one computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the operational steps of the method as described in any one of claims 1 to 7.
17. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the operational steps of the method as described in any one of claims 1 to 7.
18. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the operation steps of the method as described in any one of claims 1 to 7.