Multi-dimensional big data analysis method and system based on public space pattern

By employing a multidimensional big data analysis method based on public space patterns and using CSP and CNN to optimize features, the problem of data correlation being ignored in existing technologies is solved, achieving efficient and rigorous multidimensional data analysis and improving data processing efficiency and recognition rate.

CN120832373APending Publication Date: 2025-10-24SHENZHEN ACAD OF AEROSPACE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410477370.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-19
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing data spatial relationship mining methods ignore the correlation between different types of data, resulting in insufficient rigor in data mining, low processing efficiency, and poor applicability.

Method used

We employ a multidimensional big data analysis method based on common spatial patterns, using CSP to extract spatial features and analyze multidimensional common spatial patterns. We then combine a one-vs-rest model and CNN to optimize the features, extending the analysis to multiple types of data.

Benefits of technology

It improves the rigor and efficiency of data analysis, expands the types of data that can be extracted, enhances the applicability of algorithms, simplifies the amount of computation, and improves the recognition rate and data processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832373A_ABST
    Figure CN120832373A_ABST
Patent Text Reader

Abstract

According to the multi-dimensional big data analysis method and system based on the public spatial pattern, the CSP is used for analyzing and processing the data, spatial information of the big data is mined, the purpose of information collaborative analysis is achieved, data analysis is simpler and more convenient, and meanwhile the processing efficiency and preciseness of the data can be improved; data analysis is expanded to multi-dimensional data analysis, and two classes are expanded to multiple classes, so that the applicability of the algorithm is improved, meanwhile, the types of extractable data are expanded, and the benefits of the algorithm are improved; in addition, CNN is added to optimize features, redundant data of multi-channel big data is effectively removed, the reliability of the data is improved, the calculation amount is further simplified, and the recognition rate, the calculation rate and the data processing performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to big data processing technology, and relates to a multi-dimensional big data analysis method and system based on a public space mode. BACKGROUND

[0002] In many fields, there is a need to process various data, and the data is various and miscellaneous. For example, in the process of judicial supervision, a large amount of big data is generated, including various voice, image, video, text, physiological information, hyperspectrum and the like. The legal information is often massive, various and unstructured. When the information needs to be manually screened, sorted, analyzed and extracted, a large amount of time and energy is consumed.

[0003] The common methods for existing data space relationship mining include statistical method, clustering method, association rule mining method, neural network method, fuzzy set theory, genetic algorithm and the like. The statistical method is a common method for analyzing data, and focuses on the analysis of non-spatial characteristics. Generally, the spatial characteristics of data are not considered as a limiting factor. Moreover, experts with domain knowledge, statistical knowledge and statistical experience are needed to complete the analysis. The clustering method divides data into a series of mutually distinguished groups according to the division standard of minimum difference within the class and maximum difference between the classes, and according to a certain distance or similarity coefficient. The association rule mining method counts the frequency of the co-occurrence of various characteristics, and then converts the collocation with high frequency into an association rule. The neural network method is a self-adaptive nonlinear dynamic system composed of a large number of neurons through extremely rich and perfect connections. The system learns the patterns in the data to be analyzed by a large number of neurons through training, forms a nonlinear function describing a complex nonlinear system, and is suitable for mining classification knowledge from a nonlinear spatial system. The biggest disadvantage of the existing analysis method is to assume that the spatial distribution data has statistical irrelevance, and ignores the possibility of correlation between different types of data. Only one type of data is subjected to spatial relationship mining, but in reality, different data are often correlated. This may lead to insufficient rigor in data mining. SUMMARY

[0004] The present application provides a multi-dimensional big data analysis method and system based on a public function space mode. The method uses CSP to analyze and process data, mines the spatial information of big data, solves the technical problems of uncoordinated analysis of information and low data processing efficiency, and is extended to multi-dimensional data analysis. The method is expanded from two categories to multiple categories, improves the applicability of the algorithm, expands the types of extractable data, improves the efficiency of the algorithm, and solves the technical problems of few types of data extraction, low algorithm efficiency and poor applicability.

[0005] The present application provides a multi-dimensional big data analysis method based on a public space, which includes the following steps.

[0006] The spatial feature extraction of data: using CSP to project the original data into a new space, so that the data of different categories has the maximum difference in the new space, and the most discriminative projection direction is found;

[0007] Performing multi-dimensional common space mode: for multi-class data, using one-vs-rest model to perform multiple spatial feature extraction, obtaining multiple data.

[0008] By using CSP to extract spatial features of data, that is, presenting data in another way, so that the maximum difference of different signals can be reflected, and the desired data can be extracted according to the maximum difference. This method considers the spatial distribution of data and analyzes and mines the spatial information of big data, achieves the purpose of information collaborative analysis, and makes data analysis more convenient while improving the processing efficiency and rigor.

[0009] Further, the spatial feature extraction of data includes:

[0010] Data preprocessing: filtering and down-sampling the original data, dividing the original data into different categories, and segmenting according to the needs;

[0011] Covariance calculation: calculating the signal covariance matrix of each category, and the spatial covariance matrix can be represented as:

[0012] Wherein, E represents the data matrix of each category, T represents the number of sampling points, C represents the covariance matrix, and the spatial covariance of two data sets can be represented as , Indicates the mixed spatial covariance matrix;

[0013] Orthogonal whitening transformation: performing eigenvalue decomposition on the covariance matrix, usually using singular value decomposition; the mixed spatial covariance matrix Eigenvalue decomposition is performed:

[0014] Wherein, U represents the eigenvector matrix, and Lambda is a diagonal matrix composed of corresponding eigenvalues; the eigenvalues are arranged in descending order, and the whitening value matrix can be obtained after the whitening transformation U:

[0015] Diagonalization calculation: P is applied to And , obtaining , . And have common eigenvectors and can be decomposed as , , and has where I represents a unit matrix, and B represents an eigenvector matrix;

[0016] Projection transformation: according to the eigenvector, a projection matrix is constructed. Multiplying the original signal by the projection matrix can obtain the CSP feature representation in the new space. The projection matrix is:

[0017] The original data E is projected by the projection matrix W to obtain E', and the variance of the front and rear rows can be taken as the feature of the electroencephalogram signal.

[0018] Further, the multi-dimensional common space mode includes:

[0019] Multi-class covariance calculation: take one class of data as a class, and take all the remaining dimensional data as another class, calculate the corresponding CSP, and calculate the corresponding CSP for each class of data in turn;

[0020]

[0021] where i represents a class, Covariance matrix of the i-th class data, and the mixed space covariance matrix is represented as ;

[0022] One-versus-all common space mode: using the above CSP algorithm, calculate each space filter; for dimension i as a class, the remaining dimensions as another class:

[0023] Further, the method further includes space feature optimization, and the space feature can be optimized and extracted using a CNN.

[0024] Further, the space feature optimization includes:

[0025] CNN network construction: including 5 layers, wherein the first layer is an input layer, the second layer is a convolution layer, the third layer is a convolution layer, the fourth layer is a full connection layer, and the fifth layer is a full connection layer;

[0026] Calculate the feature matrix: let the input feature map of the full connection layer in the CNN be s, and the dimension be , the full connection layer is represented by . Then the weight value of the k-th node of the t-th full connection layer is obtained by multiplying the weight value:

[0027] Calculate the weight value: the matrix pair can be regarded as s weight value images. The t-th image is operated as follows:

[0028] Calculate the feature contribution: ​

[0029] is a column vector, which represents the deviation degree of each row feature in the feature matrix, and the greater the deviation degree is, the greater the contribution of the row feature is

[0030] The feature contribution is sorted.

[0031] The application also provides a public space-based multi-dimensional big data analysis system, which is characterized by being used for executing the public space-based multi-dimensional big data analysis method according to claim 1 and comprising:

[0032] The space feature extraction module is used for projecting original data into a new space so that data of different categories has the greatest difference in the new space and the most discriminative projection direction is found.

[0033] The multi-dimensional common space mode running module is used for extracting features of the multi-class data multiple times by using the one-vs-rest model to obtain multiple data.

[0034] Further, the system further comprises a space feature optimization module which is used for optimizing and extracting the space features.

[0035] The application has the following beneficial effects: the application uses the CSP to analyze and process data, mines the spatial information of big data, achieves the purpose of information collaborative analysis, makes data analysis more convenient, improves the data processing efficiency and rigor, extends to multi-dimensional data analysis, improves the applicability of the algorithm, expands the types of extractable data, improves the efficiency of the algorithm, optimizes the features by using the CNN, effectively removes the redundant data of multi-channel big data, improves the reliability of the data, further simplifies the calculation amount, improves the recognition rate, calculation rate and data processing performance. BRIEF DESCRIPTION OF DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments, and it should be understood that the following drawings only show some embodiments of the application, and therefore should not be regarded as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0037] Figure 1 is a multi-dimensional big data analysis process schematic diagram provided by one of the embodiments of the application;

[0038] Figure 2A spatial feature extraction flowchart provided for one of the embodiments of the present application;

[0039] Figure 3 A multi-dimensional common spatial pattern flowchart provided for one of the embodiments of the present application;

[0040] Figure 4 A multi-dimensional big data analysis flowchart provided for another embodiment of the present application;

[0041] Figure 5 A CNN network structure diagram provided for one of the embodiments of the present application;

[0042] Figure 6 A data optimization flowchart based on CNN provided for one of the embodiments of the present application. DETAILED DESCRIPTION

[0043] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0044] A multi-dimensional big data analysis method based on common space, comprising:

[0045] Spatial feature extraction of data: using common spatial patterns (CSP) to project original data into a new space, so that data of different categories have the greatest difference in the new space, and the most discriminative projection direction is found;

[0046] Multi-dimensional common spatial pattern: for multi-class data, using one-vs-rest model to extract features in multiple spaces to obtain multiple data.

[0047] By using CSP to extract spatial features of data, i.e. presenting data in another way, the maximum difference of different signals can be reflected, and the desired data can be extracted according to the maximum difference. Thus, the analysis and statistics of data are performed on the basis of considering the spatial distribution of data, the spatial information of big data is mined, the purpose of information collaborative analysis is achieved, and the data analysis is more convenient, and the processing efficiency of data is improved.

[0048] Further, the spatial feature extraction of data comprises:

[0049] Data pre-processing: filtering, down-sampling, segmenting, etc. are performed on the original data, i.e. the original data is divided into different categories, each of which represents a different state or activity, and the data is segmented according to the needs. In one embodiment, for example, 50 Hz power frequency interference is used to extract the frequency characteristics of interest. In another embodiment, 100 Hz power frequency interference can also be used to extract other frequency characteristics of interest. Different signals may require different power frequencies for interference, which can be determined by those skilled in the art according to actual conditions.

[0050] Calculate the covariance: calculate the signal covariance matrix of each category; that is, take one category of data as a category, and take all the remaining dimensions of data as another category; let each category of data be represented as a matrix E, and T represents the number of sampling points. The spatial covariance matrix can be represented as:

[0051]

[0052] where C represents the covariance matrix, and the spatial covariance of the two data sets can be represented as , represents the mixed spatial covariance matrix, i.e. the sum of the spatial covariances of the two categories of data.

[0053] Perform orthogonal whitening transformation: perform eigenvalue decomposition on the covariance matrix, usually using singular value decomposition; the mixed spatial covariance matrix Perform eigenvalue decomposition:

[0054]

[0055] where U represents the eigenvector matrix, and Λ is a diagonal matrix composed of corresponding eigenvalues. The eigenvalues are arranged in descending order, and the whitening value matrix U can be obtained after the whitening transformation U:

[0056]

[0057] Diagonalization calculation: the data obtained above is processed by diagonalization to obtain the eigenvector matrix; that is, P acts on and to obtain , . and have common eigenvectors and can be decomposed as , , and where I represents the unit matrix, and B represents the eigenvector matrix.

[0058] Projection transformation: according to the eigenvector, the projection matrix is constructed. That is, the original signal is multiplied by the projection matrix, and the CSP feature representation in the new space can be obtained. The projection matrix is:

[0059]

[0060] The original data E is projected by the projection matrix W to obtain E', and the variance of the front and rear rows is taken as the feature of the electroencephalogram signal.

[0061] After CSP transformation, the data is closer to the Gaussian distribution, and the spatial information feature of the data is extracted. CSP is a two-class spatial analysis method, which can maximize the variance between different classes by simultaneously diagonalizing two covariance matrices, that is, the maximum difference of the data is more obvious after CSP transformation, so that the part with the maximum difference can be extracted, that is, the spatial information feature of the data.

[0062] Multi-dimensional common space mode includes:

[0063] Multi-class covariance calculation: take one class of data as a class, and take all the remaining dimensional data as another class, and then perform CSP transformation, then take another class of data as a class, and take all the remaining dimensional data as another class, and then perform CSP transformation, and repeat the above steps, and calculate the corresponding CSP for each class of data.

[0064]

[0065] where i represents the class, the covariance matrix of the ith class data, and the mixed space covariance matrix is represented as .

[0066] One-versus-all common space mode: use the above CSP algorithm to calculate each spatial filter. For dimension i as a class, the remaining dimensions as another class.

[0067]

[0068] Thus, a two-class CSP is formed, and the corresponding CSP is calculated for each class mode in turn. The eigenvalue of the original covariance matrix after transformation satisfies and equals 1, that is, in the case of maximum variance of the first class signal, the variance of all other mode signals is minimum. Thus, more desired information is extracted according to different maximum variance values.

[0069] The two classes are expanded to multiple classes, and the data is analyzed in multiple dimensions, which improves the number and types of effective data that can be obtained, and the algorithm applicability.

[0070] The multi-channel data contains a large amount of redundant information, which leads to increased computational complexity and reduced recognition rate. To solve this problem, the multi-dimensional big data analysis method further includes spatial feature optimization, that is, using a convolutional neural network (CNN) to optimize and extract spatial features.

[0071] For the feature matrix obtained by the CSP algorithm, important information is mainly concentrated in the head and tail. Therefore, the key is how to select the number of head and tail feature rows. If the selected value is too small, the feature information is insufficient; if the selected value is too large, there will be redundant information. The CNN is a feedforward neural network containing convolution calculation and having a deep structure, which learns the characteristics of cell pooling, and each layer of features is the result of performing convolution on the next layer of features. Therefore, the CNN can extract and optimize features inside the network.

[0072] As shown in FIG. 1, under one embodiment, the spatial feature optimization specifically includes: Figures 5-6

[0073] CNN network construction; in the learning process of the feature matrix, the convolutional neural network structure is composed of 5 layers, wherein the first layer is an input layer, the second layer is a convolutional layer, the third layer is a convolutional layer, the fourth layer is a fully connected layer, and the fifth layer is a fully connected layer.

[0074] Calculate the feature matrix; let the input feature map of the fully connected layer in the CNN be s, the dimension of which is , and the fully connected layer is denoted by . Then the weight of the kth node of the tth fully connected layer is obtained by multiplying the weights:

[0075]

[0076] Calculate the weight; the matrix pair can be regarded as s images of weight. The tth image is operated as follows:

[0077]

[0078] Calculate the feature contribution; construct a feature contribution matrix, that is, obtain T by averaging the full job of each feature dimension, and then obtain the standard deviation of T.

[0079]

[0080]

[0081] is a column vector, which represents the deviation degree of each row feature in the feature matrix. The greater the deviation degree, the greater the contribution of the row feature.

[0082] ​The feature contributions are sorted, and a feature selection threshold is set, so that the most effective feature set in the feature matrix is obtained.

[0083] Through the transformation of the CNN, the redundant information of the multi-channel large data can be effectively removed, the reliability of the spatial features is improved, the calculation amount is simplified, the recognition rate and the calculation rate are improved, and the data processing performance of the multi-channel large data is improved.

[0084] The application further provides a multi-dimensional large data analysis system based on a common space, which is used for executing the multi-dimensional large data analysis method, and comprises:

[0085] The spatial feature extraction module is used for projecting the original data into a new space, so that the data of different categories has the maximum difference in the new space, and the most discriminative projection direction is found.

[0086] The multi-dimensional common space mode running module is used for performing multiple spatial feature extraction on multi-class data by using a one-vs-rest model, and obtaining multiple data.

[0087] Further, the system further comprises a spatial feature optimization module, which is used for optimizing and extracting the spatial features.

[0088] The above only describes the preferred embodiments of the application and is not used to limit the application. For those skilled in the art, the application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.

Claims

1. A method for multi-dimensional big data analysis based on public space, characterized in that, The method comprises the following steps: The spatial feature extraction of data: the original data is projected into a new space using CSP, so that the data of different categories has the maximum difference in the new space, and the most discriminative projection direction is found; Carrying out multi-dimensional common space mode: for multi-class data, using one-vs-rest model to carry out multiple spatial extraction features, and obtaining multiple data.

2. The method of claim 1, wherein, The spatial feature extraction of data comprises: Data preprocessing: filtering and downsampling the original data, dividing the original data into different categories, and segmenting according to the needs; the different categories represent different states or activities; Covariance calculation: Calculate the signal covariance matrix for each class, the spatial covariance matrix can be represented as: wherein E represents each type of data matrix, T represents the number of sampling points, C represents the covariance matrix, and the spatial covariance of the two types of data sets can be represented as , represents a mixed spatial covariance matrix; Performing a whitening transform: Eigen decomposition of the covariance matrix, usually using singular value decomposition; mixing space covariance matrix Performing an eigenvalue decomposition: where U represents the eigenvector matrix, and Λ is a diagonal matrix composed of corresponding eigenvalues; the eigenvalues are arranged in descending order, and the whitening conversion U is performed to obtain a whitening value matrix: Diagonalization is performed: P is applied to and to obtain , . and have a common eigenvector and can be decomposed as , and have where I denotes the identity matrix and B denotes the eigenvector matrix. Projection transformation: According to the eigenvectors, a projection matrix is constructed. Multiplying the original signal by the projection matrix can obtain the CSP feature representation in the new space. The projection matrix is: The original data E is projected through the projection matrix W to obtain E`, and the variance of the front and rear rows can be taken as the feature of the electroencephalogram signal.

3. The method of claim 1, wherein, The multi-dimensional common space mode comprises: Multi-class covariance calculation: taking one class of data as a class, and taking all the remaining dimensional data as another class, calculating the corresponding CSP, and calculating the corresponding CSP for each class of data in turn; where i denotes a class, C i denotes the covariance matrix of the ith class of data, and the mixed space covariance matrix is denoted as ; One-to-many public space pattern: using the CSP algorithm described above, compute the individual space filters; for dimension i is one class, the rest of the dimensions is another class:

4. The method of claim 1-3, wherein, It also includes spatial feature optimization, which can optimize and extract spatial features using CNN.

5. The method of claim 4, wherein, The spatial feature optimization comprises: CNN network construction: including 5 layers, wherein the first layer is an input layer, the second layer is a convolutional layer, the third layer is a convolutional layer, the fourth layer is a fully connected layer, and the fifth layer is a fully connected layer; Calculate the feature matrix: let the input feature map of the full connection layer in the CNN be s, the dimension be , and the full connection layer be denoted by . Then the weight of the kth node of the tth full connection layer is obtained by multiplying the weights: Compute weight: matrix pair This can be considered as s weight images. For the t-th image, the following operations are performed: Computing feature contribution: is a column vector, which represents the deviation degree of each row feature in the feature matrix. The greater the deviation degree, the greater the contribution of the row feature. By sorting the feature contribution and setting a feature selection threshold, the most effective feature set in the feature matrix can be obtained.

6. A public space based multi-dimensional big data analytics system, characterized in that, The method for carrying out the multi-dimensional big data analysis method based on common space comprises: A spatial feature extraction module is used to project the original data into a new space, so that the data of different categories has the maximum difference in the new space, and the most discriminative projection direction is found; A multi-dimensional common space mode running module is used to use one-vs-rest model for multi-class data, to carry out multiple spatial extraction features, and to obtain multiple data.

7. The multi-dimensional big data analytics system of claim 6, wherein, It also includes a spatial feature optimization module for optimizing and extracting spatial features.