Data Processing Method, Apparatus, Device, Medium, and Program Product

The high-dimensional single-view data is dimensionally split through principal component analysis and converted into multi-view data, solving the problems of high redundancy rate and Hughes phenomenon in high-dimensional data modeling, and achieving more comprehensive data correlation mining and model applicability.

CN115687907BActive Publication Date: 2025-07-18INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211043758.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-07-18
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

In the process of high-dimensional single-view data modeling, there is a high feature redundancy rate and Hughes phenomenon, and traditional dimensionality reduction methods will lose information, making it difficult to fully mine data correlation.

Method used

The principal component analysis method is used to dimensionally split the high-dimensional single-view data, and transform it into multi-view data. The multi-view algorithm is used to build a model, retain data dimension information, and reduce redundancy rate and Hughes phenomenon.

Benefits of technology

Without losing data information, the scope of application of data modeling has been expanded, the high redundancy rate and Hughes phenomenon brought about by high-dimensional data has been alleviated, and the data correlation mining capability has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115687907B_ABST
    Figure CN115687907B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method, which can be applied to the field of big data technology. The method includes: obtaining original data, preprocessing the original data to obtain view data to be split, where the view data to be split is single-view data; performing view splitting on the view data to be split according to dimensions based on the principal component analysis method, and converting the view data to be split into multi-view data; and establishing a model based on the multi-view data using a single-view algorithm or a multi-view algorithm, where the single-view data is high-dimensional data, and the dimensions in the high-dimensional data correspond to entity features; the multi-view data is data allocated to multiple views, where the number of dimensions of the data in each view is the same or different, and the aggregation of the number of dimensions of the data in each view is equal to the number of dimensions of the single-view data. The present disclosure also provides a data processing device, equipment, storage medium and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of big data technology, and particularly to a data processing method, apparatus, device, medium and program product. Background Art

[0002] With the rapid development of big data technology, big data modeling has been widely applied in many technical fields such as finance, commerce, and government affairs. The essence of big data modeling is to extract useful information from data in the data feature space for purposes such as classification and regression. From the perspective of the type of data samples, data modeling models can be divided into two categories: single-view and multi-view methods. Single-view means that all samples are sampled from the same data distribution, which is the basis for the research of modeling problems. Multi-view means that samples are sampled from two or more different sources. The distribution information of the samples is more complex, but the information contained is more abundant, and the structure between information is more obvious.

[0003] When using single-view data for modeling, due to the continuous increase in the complexity of application scenarios, a large amount of single-view data used for modeling is high-dimensional data. In the process of using high-dimensional data for single-view modeling, due to the high dimension, problems such as high redundancy rate between features and the Hughes phenomenon are likely to occur. And when the data dimension is relatively high, it is difficult to mine the data structure information. Traditional dimensionality reduction processing methods will lose information after reducing the dimension. How to reduce the feature redundancy rate, reduce the occurrence of the Hughes phenomenon, and at the same time not lose the original data information and more fully mine the association between data when using single-view data for modeling is an urgent problem to be solved. Summary of the Invention

[0004] In view of the above problems, embodiments of the present disclosure provide a data processing method, apparatus, device, medium and program product that can improve the data utilization rate of single-view data modeling, reduce the feature redundancy rate, and reduce the Hughes phenomenon.

[0005] According to a first aspect of the present disclosure, there is provided a data processing method, including: obtaining original data, preprocessing the original data to obtain view data to be split, where the view data to be split is single-view data; performing view splitting on the view data to be split according to dimensions based on the principal component analysis method, and converting the view data to be split into multi-view data; and establishing a model based on the multi-view data using a single-view algorithm or a multi-view algorithm, where the single-view data is high-dimensional data, the dimensions in the high-dimensional data correspond to entity features; the multi-view data is data allocated to multiple views, where the number of dimensions of the data in each view is the same or different, and the aggregation of the number of dimensions of the data in each view is equal to the dimensions of the single-view data.

[0006] According to an embodiment of the present disclosure, the step of splitting the view data to be split into multi-view data according to dimensions based on the principal component analysis method includes a dimension splitting step, wherein the dimension splitting step includes: constructing a principal component analysis classification indicator function, wherein the principal component analysis classification indicator function is constructed based on a feature classification indicator matrix, and the feature classification indicator matrix is an n×k matrix, where n is the dimension of the data to be split and k is the preset number of view splits; solving the optimization problem of the principal component analysis classification indicator function to obtain a solution result, wherein the solution result includes the weight assignment result of the feature classification indicator matrix, and the weight assignment result of the feature classification indicator matrix is a matrix composed of eigenvectors corresponding to the minimum eigenvalue corresponding to the number of view splits; and splitting the dimensions of the view data to be split based on the weight assignment result of the feature classification indicator matrix to obtain a data dimension splitting result, wherein the dimension of the view data to be split is m, and splitting the i-th data dimension includes: and splitting the i-th data dimension into the view corresponding to the maximum eigenvalue in the eigenvector corresponding to the i-th data dimension in the feature classification indicator matrix, where i ∈ [1, m].

[0007] According to an embodiment of the present disclosure, after obtaining the dimension splitting result of the m-dimensional data, the method further includes: allocating the view data to be split according to the dimension splitting result to obtain the multi-view data.

[0008] According to an embodiment of the present disclosure, the preset number of view splits is obtained by processing the dimensions of the data to be split based on one of an automatic clustering algorithm or a similarity algorithm.

[0009] According to an embodiment of the present disclosure, wherein the automatic clustering algorithm includes one of a density-based spatial clustering of applications with noise algorithm, a fuzzy clustering algorithm, and a K-means clustering algorithm; and / or, the similarity algorithm includes one of a cosine similarity algorithm, a distance similarity algorithm, and a Pearson correlation coefficient.

[0010] According to an embodiment of the present disclosure, the preprocessing of the original data includes: normalizing the original data.

[0011] According to an embodiment of the present disclosure, solving the optimization problem of the principal component analysis classification indicator function includes: solving the optimization problem of the principal component analysis classification indicator function based on a generalized eigenvalue solving method.

[0012] According to an embodiment of the present disclosure, the multi-view algorithm includes one of canonical correlation analysis, multiple canonical correlation analysis, kernel canonical correlation analysis, locally preserving canonical correlation analysis, discriminant canonical correlation analysis, generalized multi-view analysis, multi-view discriminant analysis, and multi-view dimensionality reduction model.

[0013] According to an embodiment of the present disclosure, the original data includes user feature data, and the model is used to construct a user portrait.

[0014] A second aspect of the present disclosure provides a data processing device, including: an acquisition module configured to acquire original data, preprocess the original data, and acquire data of a view to be split, where the data of the view to be split is single-view data, the single-view data is high-dimensional data, and the dimensions in the high-dimensional data correspond to entity features; a data splitting module configured to perform view splitting on the data of the view to be split according to dimensions based on the principal component analysis method, and convert the data of the view to be split into multi-view data, where the multi-view data is data distributed in multiple views, and the number of dimensions of the data in each view is the same or different, and the aggregation of the number of dimensions of the data in each view is equal to the number of dimensions of the single-view data; and a model establishment module configured to establish a model based on the multi-view data using a single-view algorithm or a multi-view algorithm.

[0015] According to an embodiment of the present disclosure, the data splitting module may include a construction sub-module, a solution sub-module, and a dimension splitting sub-module. Among them, the construction sub-module is configured to construct a principal component analysis classification indicator function, where the principal component analysis classification indicator function is constructed based on a feature classification indicator matrix, and the feature classification indicator matrix is an n×k matrix, where n is the dimension of the data to be split, and k is the preset number of view splits. The solution sub-module is configured to solve the optimization problem of the principal component analysis classification indicator function and obtain a solution result, where the solution result includes a weight allocation result of the feature classification indicator matrix, and the weight allocation result of the feature classification indicator matrix is a matrix composed of eigenvectors corresponding to the minimum eigenvalue corresponding to the number of view splits. The dimension splitting sub-module is configured to perform view splitting on the dimensions of the data of the view to be split based on the weight allocation result of the feature classification indicator matrix to obtain a data dimension splitting result, where the dimension of the data of the view to be split is m, and splitting the i-th data dimension includes: and splitting the i-th data dimension into the view corresponding to the maximum eigenvalue of the eigenvector corresponding to the i-th data dimension in the feature classification indicator matrix, where i ∈ [1, m].

[0016] According to an embodiment of the present disclosure, the data splitting module may include a construction sub-module, a solution sub-module, a dimension splitting sub-module, and a data allocation sub-module. Among them, the data allocation sub-module is configured to allocate the view data to be split according to the dimension splitting result to obtain the multi-view data.

[0017] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above data processing method.

[0018] A fourth aspect of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above data processing method.

[0019] A fifth aspect of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above data processing method is implemented.

[0020] The method provided by the embodiment of the present disclosure draws on the idea of multi-view modeling, uses the principal component analysis method to split the dimensions of single-view data, transforms the single-view data into multi-view data without losing data dimension information, increases the structural information of the data, expands the range of optional models in the data modeling process, and alleviates to a certain extent problems such as high redundancy rate between features and the Hughes phenomenon caused by too high data dimensions. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0022] Figure 1 Schematically shows an application scenario diagram of the data processing method, device, equipment, medium, and program product according to the embodiment of the present disclosure.

[0023] Figure 2 Schematically shows a flowchart of the data processing method according to the embodiment of the present disclosure.

[0024] Figure 3 Schematically shows a flowchart of the dimension splitting method according to the embodiment of the present disclosure.

[0025] Figure 4 Schematically shows a flowchart of the method for splitting the i-th data dimension according to the embodiment of the present disclosure.

[0026] Figure 5 Schematically shows a flowchart of the method for view data allocation according to the embodiment of the present disclosure.

[0027] Figure 6 Schematically shows a structural block diagram of a data processing apparatus according to an embodiment of the present disclosure.

[0028] Figure 7 Schematically shows a structural block diagram of a data splitting module according to an embodiment of the present disclosure.

[0029] Figure 8 Schematically shows a structural block diagram of a data splitting module according to an embodiment of the present disclosure.

[0030] Figure 9 Schematically shows a block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure. Detailed implementation manners

[0031] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0032] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0033] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0034] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0035] With the rapid development of big data technology, big data modeling has been widely applied in many technical fields such as finance, commerce, and government affairs. The essence of big data modeling is to extract useful information from the data feature space and use it for purposes such as classification and regression. From the perspective of the type of data samples, data modeling models can be divided into two categories: single-view and multi-view methods. Single-view means that all samples are sampled from the same data distribution, which is the basis for the study of modeling problems. Multi-view means that samples are sampled from two or more different sources, and the distribution information of the samples is more complex, but the information contained is more abundant, and the structure between information is more obvious. For example, if a person is recognized from a visual aspect, it can be considered as single-view data; if a person is recognized from multiple perspectives such as appearance, voice, and movement, it can be considered as multi-view data. Typical examples of multi-view data can be: multimedia videos can be represented by video signals and audio signals at the same time; another example is that web pages can be represented by hyperlinks and text information. The implicit features of multi-view data are consistent, but due to the different statistical features of the data, each perspective contains unique information about a certain aspect of the object, and there are also differences. Using multi-view data to find the consistent and complementary information between data to more comprehensively describe the data, so as to achieve in-depth understanding and analysis of the target, has been proven to have good application effects in practical applications.

[0036] In many scenarios, the data used for modeling is still single-view data. When using single-view data for modeling, due to the continuous increase in the complexity of the application scenario, a large amount of single-view data used for modeling is high-dimensional data. In the process of using high-dimensional data for single-view modeling, due to the high dimension, problems such as high redundancy rate between features and the Hughes phenomenon are likely to occur. And when the data dimension is relatively high, it is difficult to mine the data structure information. Traditional dimensionality reduction methods will lose information after reducing the dimension. How to reduce the feature redundancy rate when using single-view data for modeling, reduce the occurrence of the Hughes phenomenon, and at the same time not lose the original data information, and more fully mine the association between data is an urgent problem to be solved.

[0037] In view of the above problems existing in the prior art, embodiments of the present disclosure provide a data processing method, including: obtaining original data, preprocessing the original data to obtain data of views to be split, wherein the data of views to be split is single-view data; based on the principal component analysis method, splitting the data of views to be split according to dimensions to convert the data of views to be split into multi-view data; and establishing a model based on the multi-view data using a single-view algorithm or a multi-view algorithm, wherein the single-view data is high-dimensional data, and the dimensions in the high-dimensional data correspond to entity features; the multi-view data is data distributed in multiple views, wherein the number of dimensions of the data in each view is the same or different, and the aggregation of the number of dimensions of the data in each view is equal to the dimension of the single-view data.

[0038] The method provided by the embodiments of the present disclosure draws on the idea of multi-view modeling, uses the principal component analysis method to split the dimensions of single-view data, and transforms the single-view data into multi-view data without losing data dimension information, increasing the structural information of the data. The split multi-view data can be modeled using multi-view algorithms, or the multi-view data can be merged and restored to single-view data and modeled using single-view algorithms. This can expand the range of optional models in the data modeling process and, to a certain extent, alleviate problems such as high redundancy rate between features and the Hughes phenomenon caused by high data dimensions.

[0039] It should be noted that the data processing methods, devices, equipment, media, and program products provided by the embodiments of the present disclosure can be used in aspects related to data preprocessing before modeling in big data technology, and can also be used in various fields other than big data technology, such as the financial field. The application fields of the data processing methods, devices, equipment, media, and program products provided by the embodiments of the present disclosure are not limited.

[0040] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information involved all comply with the provisions of relevant laws and regulations, adopt necessary confidentiality measures, and do not violate public order and good customs.

[0041] In the technical solution of the present disclosure, before obtaining or collecting user personal information, the authorization or consent of the user is obtained.

[0042] The following will elaborate on the above operations around achieving at least one object of the present disclosure in conjunction with the accompanying drawings and their explanatory text.

[0043] Figure 1 Schematically shows an application scenario diagram of the data processing method, device, equipment, medium, and program product according to the embodiments of the present disclosure.

[0044] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0045] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only). For example, users can use terminal devices 101, 102, and 103 to send raw data to server 105 via network 104 for modeling.

[0046] Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.

[0047] Server 105 can be a server that provides various services, such as a background management server that supports the websites browsed by users using terminal devices 101, 102, and 103 (for example only). The background management server can analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. For example, after receiving the raw data sent by terminal devices 101, 102, and 103 via network 104, server 105 can complete view data splitting and model establishment based on the data processing method of the embodiments of the present disclosure, and send the model processing results to terminal devices 101, 102, and 103 via network 104.

[0048] It should be noted that the data processing method provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the data processing device provided by the embodiments of the present disclosure can generally be set in server 105. The data processing method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, and 103 and / or server 105. Correspondingly, the data processing device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, and 103 and / or server 105.

[0049] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0050] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. Figure 1 The following will be based on Figures 2 to 5 the scenarios described below, and will describe in detail the data processing methods, devices, equipment, media, and program products of the embodiments of the present disclosure through

[0051] Figure 2 Schematically shows a flowchart of a data processing method according to an embodiment of the present disclosure.

[0052] As Figure 2 shown, the data processing method of this embodiment includes operations S210 to S230. This data processing method can be executed by a processor or any electronic device including a processor.

[0053] In operation S210, obtain the original data, preprocess the original data, and obtain the view data to be split, where the view data to be split is single-view data.

[0054] According to an embodiment of the present disclosure, the original data is single-view data. The single-view data is high-dimensional data, and the dimensions in the high-dimensional data correspond to entity features, that is, each dimension in the high-dimensional data can be used to describe an attribute feature of a certain aspect of the entity. In an embodiment of the present disclosure, the high-dimensional data can be data with dozens of dimensions or more. In some application scenarios, the high-dimensional data can be data with more than 50 dimensions. Before splitting the original data into views, the original data can be preprocessed to optimize the data samples. Typical preprocessing methods can include data cleaning, data standardization, etc. For example, removing unique attributes, filling in missing values, removing outliers, and giving meaning to feature values. In some embodiments, the preprocessing method of the original data preferably includes normalizing the original data. Specifically, common normalization methods such as the maximum-minimum method, zero-mean normalization method, logarithmic function transformation method, arctangent function transformation, or Z-Score method can be used to scale all dimensions to the range [0, 1], eliminate the influence of the dimension of the features, make the features comparable, and improve the scientificity of data splitting. After preprocessing, the view data to be split for view splitting can be obtained.

[0055] In operation S220, perform view splitting on the view data to be split according to dimensions based on the principal component analysis method, and convert the view data to be split into multi-view data.

[0056] In operation S230, establish a model based on the multi-view data using a single-view algorithm or a multi-view algorithm.

[0057] In the embodiments of the present disclosure, the idea of principal component analysis is used to split the view data to be split according to dimensions. The traditional principal component analysis method is a typical data analysis method. By linearly transforming the original data into a set of linearly independent representations in each dimension, it can be used to extract the main feature components of the data and is often used for dimensionality reduction of high-dimensional data. However, in the process of dimensionality reduction, since only the main feature components are extracted or the data in multiple dimensions are transformed, dimension information will inevitably be lost. To avoid the loss of dimension information, the embodiments of the present disclosure draw on the idea of principal component analysis, obtain a feature classification indication matrix by maximizing the dispersion degree between samples, and then allocate the dimension features to the view with the largest projection weight through the feature classification indication matrix, further realizing data splitting, and allocating the original single-view data to multiple views to form multi-view data. After splitting, the data between each view has internal consistency and complementarity, so as to realize in-depth understanding and analysis of the target. Among them, consistency means that there is an internal connection between the data in different views; complementarity means that the difference between the data in different views makes the data in each view contain unique information about a certain aspect of the target object described by the original data. These complementary information can describe the data more comprehensively, and using such information can deeply mine the association between features to build a better-performing model. After obtaining the multi-view data, a multi-view algorithm can be used to model the split multi-view data. It can be understood that the multi-view data is the original single-view data after aggregation, so a single-view algorithm can also be used for modeling. The data processing method of the embodiments of the present disclosure expands the applicable range of the data and the optional range of the model by splitting the data dimensions, so as to be applicable to modeling in more scenarios. On the other hand, the number of dimensions of the data in each view is the same or different, and the aggregation of the number of dimensions of the data in each view is equal to the number of dimensions of the single-view data. Thus, the information of all dimensions of the original data is retained. And because the data dimensions in each view are reduced, it can alleviate problems such as high redundancy rate between features and Hughes phenomenon caused by too high dimensions of the original single-view data to a certain extent. And it saves computing resources to a certain extent when modeling based on view data.

[0058] In the embodiments of the present disclosure, the key to the step of splitting the view data to be split according to dimensions based on the principal component analysis method and converting the view data to be split into multi-view data lies in dimension splitting.

[0059] Figure 3 The flowchart of the dimension splitting method according to the embodiments of the present disclosure is schematically shown.

[0060] As Figure 3 shown, the dimension splitting method of this embodiment includes operation S310 to operation S330.

[0061] In operation S310, a principal component analysis classification indicator function is constructed, where the principal component analysis classification indicator function is constructed based on a feature classification indicator matrix, and the feature classification indicator matrix is an n×k matrix, where n is the dimension of the data to be split, and k is the preset number of view splits.

[0062] In some embodiments, the preset number of view splits can be determined based on expert experience.

[0063] In some embodiments, solving the optimization problem of the principal component analysis classification indicator function includes: solving the optimization problem of the principal component analysis classification indicator function based on a generalized eigenvalue solving method. When solving the optimization problem of the principal component analysis classification indicator function, different methods including generalized eigenvalue solving and singular value decomposition can be applied. By using the generalized eigenvalue solving method of the embodiments of the present disclosure to solve the optimization problem of the principal component analysis classification indicator function, the classification of features can be realized to complete view splitting.

[0064] Specifically, the optimization problem of the principal component analysis classification indicator function can be shown as formula (1):

[0065]

[0066] In formula (1), W∈R n×k is the feature classification indicator matrix, n is the dimension of the original data, k is the number of views, and tr(W T XX T W) is the principal component term, which is used to maximize the divergence between variables after feature aggregation. Solving the optimization problem based on formula (1) for the principal component analysis classification indicator function means that: by maximizing the divergence between variables after feature aggregation (principal component analysis), the original n-dimensional feature data is projected onto a set of k-dimensional orthogonal bases, and the projection weights of each feature on the k mutually orthogonal dimensions can be obtained, that is, W.

[0067] In operation S320, the optimization problem of the principal component analysis classification indicator function is solved to obtain a solution result, where the solution result includes the weight assignment result of the feature classification indicator matrix, and the weight assignment result of the feature classification indicator matrix is a matrix composed of eigenvectors corresponding to the minimum eigenvalues corresponding to the number of view splits. Specifically, XX T W = (-α)W. The solution of W is a matrix composed of the eigenvectors corresponding to the first k minimum eigenvalues, that is, the weight assignment result of the feature classification indicator matrix, which can be used to split features.

[0068] In operation S330, based on the weight assignment result of the feature classification indicator matrix, the dimension of the view data to be split is split into views to obtain a data dimension split result, where the dimension of the view data to be split is m.

[0069] Figure 4 A flowchart schematically showing a method for splitting the i-th data dimension according to an embodiment of the present disclosure is shown.

[0070] As Figure 4 shown, the method for splitting the i-th data dimension in this embodiment includes operation S410.

[0071] In operation S410, the i-th data dimension is split into the view corresponding to the maximum eigenvalue in the eigenvector corresponding to the i-th data dimension in the feature classification indication matrix, where i ∈ [1, m].

[0072] According to the solution result of the feature classification indication matrix, each row thereof represents the projection weight of this feature on different views. Among them, the column where the maximum eigenvalue is located is the view with the largest projection weight after aggregation of this feature, that is, the view containing the most information of this feature after projection. Therefore, the feature of the corresponding dimension should be split into the view corresponding to the column where the maximum eigenvalue is located. It should be understood that each of the m data dimensions can be split into the corresponding view according to the method for splitting the i-th data dimension to complete the data dimension splitting.

[0073] In the embodiment of the present disclosure, the optimization problem solving method of principal component analysis is skillfully applied, the principal component analysis optimization function is constructed based on the number of views, and the feature classification indication matrix is solved. Finally, based on the solution result of the feature classification indication matrix, the feature with the largest projection weight on each view during the aggregation process is found, and the corresponding feature dimension is extracted as the splitting dimension of the current view to complete the dimension splitting.

[0074] According to an embodiment of the present disclosure, after obtaining the dimension splitting result of the m-dimensional data, the method further includes the step of view data allocation.

[0075] Figure 5 A flowchart schematically showing a method for view data allocation according to an embodiment of the present disclosure is shown.

[0076] As Figure 5 shown, the method for view data allocation in this embodiment includes operation S510.

[0077] In operation S510, the view data to be split is distributed according to the dimension splitting result to obtain the multi-view data. After dimension splitting is completed, data can be reorganized based on the dimension splitting rules. The original single-view data with n features is split into data with k views, and the sum of the features of all views is n, thus completing data splitting. This data splitting method can ensure the consistency of features within the view, that is, the current view contains the most information corresponding to the features. At the same time, there is consistency and complementarity between views, thereby mining data correlation information to a greater extent and improving the utilization rate of data.

[0078] In some preferred embodiments, the preset number of view splits is obtained by processing the dimension of the data to be split based on one of an automatic clustering algorithm or a similarity algorithm. Specifically, by automatically clustering the feature dimensions through an automatic clustering algorithm, or calculating the similarity of the feature dimensions through a similarity algorithm, the association between feature dimensions can be initially mined. Using the clustering result or the similarity result as the basis for setting the number of view splits is conducive to optimizing the view splitting result, further improving the scientificity and accuracy of the data processing method of the embodiments of the present disclosure, and enhancing the model performance of subsequent modeling.

[0079] In some specific embodiments, the automatic clustering algorithm includes one of a density-based spatial clustering of applications with noise (DBSCAN) algorithm, a fuzzy clustering algorithm, and a K-means clustering algorithm; and / or, the similarity algorithm includes one of a cosine similarity algorithm, a distance similarity algorithm, and a Pearson correlation coefficient. Among them, the automatic clustering algorithm preferably can use the density-based spatial clustering of applications with noise (DBSCAN) algorithm. Density-based spatial clustering of applications with noise (DBSCAN) is an unsupervised ML clustering algorithm. It does not use pre-labeled targets to cluster data points. DBSCAN does not require specifying the number of clusters, avoids outliers, and has a better clustering effect in clusters of any shape and size. DBSCAN has no centroid, and the clustering clusters are formed through a process of connecting adjacent points. In the specific embodiments of the present disclosure, applying the DBSCAN clustering algorithm or the cosine similarity algorithm can automatically conduct an initial mining of feature correlation to preset a more accurate number of views.

[0080] In some embodiments, the multi-view algorithm includes one of canonical correlation analysis (CCA), local preserving canonical correlation analysis, multiple canonical correlation analysis (MCCA), kernel canonical correlation analysis, discriminant canonical correlation analysis, generalized multi-view analysis, multi-view discriminant analysis, and multi-view dimensionality reduction model. Among them, canonical correlation analysis (CCA) is the most representative two-view data dimensionality reduction method. Its idea is to reduce the dimensionality of the de-meaned two-view data, extract multiple pairs of canonical uncorrelated variables during the dimensionality reduction process, and maximize the correlation between variables. Local preserving canonical correlation analysis (LPCCA) deletes non-neighboring points when calculating the correlation matrix using canonical correlation analysis to protect the local manifold structure of sample points, thereby improving the discriminant ability of the model to a certain extent. In addition, when calculating the overall correlation, not only the correlation of the same points under different views is considered, but also the correlation of neighboring points under different views is considered, which can form another local preserving canonical correlation analysis (ALPCCA). By adding the local neighboring information of the data to the model in this way, the local manifold structure of the data can be maintained during dimensionality reduction. Multiple canonical correlation analysis (MCCA) can reduce the dimensionality of data from more views to achieve the expansion from two views to more views. Kernel canonical correlation analysis (KCCA) introduces the idea of kernel function on the basis of CCA. First, the sample points are non-linearly mapped to a high-dimensional kernel function space, so that the non-linearly distributed sample points are linearly distributed in the kernel function space. Then, the traditional canonical correlation analysis is used to reduce the dimensionality of the two-view data in the kernel function space. The goal of discriminant canonical correlation analysis (DCCA) is to find a new subspace in which the same-class points between different perspectives have the maximum correlation, and the different-class points have a smaller correlation. Generalized multi-view analysis (GMA) considers both the single view itself and the relationship between different views. Its goal is to find a dimensionality reduction direction for each view. After dimensionality reduction in the single view, the distance between different-class points should be far, and there should be a large correlation between different views. The goal of multi-view discriminant analysis (MvDA) is to find a common subspace that simultaneously maintains the compactness within the class and the separability between classes. The multi-view dimensionality reduction model (MDcR) maximizes the similarity of the data from different views in the kernel space and the correlation of the single view itself, and no longer requires the data of any view to be reduced to the same space. Based on the data characteristics of the corresponding application scenarios in practical applications, selecting and applying the above multi-view models can maximize the correlation between different views, fully mine data information, and improve the performance of the model.

[0081] In some embodiments, the original data includes user feature data, and the model is used to construct a user profile. It can be understood that the data processing method of the embodiments of the present disclosure can be used for view splitting of user feature data, and further use the multi-view algorithm to model the user feature data to construct a user profile. When the dimension of user feature data is relatively high, using the traditional single-view modeling method may cause modeling problems such as high redundancy rate between features and the Hughes phenomenon due to the high data dimension. By applying the data processing method of the embodiments of the present disclosure, splitting the dimension of user feature data into views and constructing a model based on the split multi-views can better mine the correlation between user feature data, effectively improve the accuracy and efficiency of subsequent model learning, avoid overfitting, and thus construct a more accurate user profile.

[0082] In the embodiments of the present disclosure, before obtaining user feature data, consent or authorization of the user may be obtained. For example, a request to obtain user feature data may be sent to the user. When the user consents or authorizes to obtain user feature data, the user feature data is obtained.

[0083] It should be understood that the above construction of the user profile is only an exemplary applicable specific scenario of the data processing method of the embodiments of the present disclosure, and does not constitute a limitation on the actual application scenario of the data processing method of the embodiments of the present disclosure.

[0084] In a specific example, taking a bank constructing a user profile as an example. Suppose there are currently 30,000 pieces of customer data, and for each piece of data, there are 500-dimensional features such as customer age, occupation, assets, deposits, and consumption transactions. Applying the data processing of the embodiments of the present disclosure, the original data is split into 3-view data and modeled to construct a user profile. Specifically, first, normalization processing is performed. The maximum-minimum method is used to scale all dimensions to the range of [0, 1] to eliminate the influence of the dimension of the features. Then, a classification indication function based on principal component analysis is constructed and solved, where the feature classification indication matrix W ∈ R 500×3 is used to split the feature dimensions. Solve the optimization problem to obtain the weight allocation result of the matrix W. Suppose the first row vector of the weight allocation result matrix is [0.7, 0.2, 0.1], which means that the first view contains the most information of this feature. Therefore, the features corresponding to this row are assigned to the first view. All feature dimensions are assigned in the same way to complete the dimension splitting. Further, the original data can be assigned to each view based on the rule of dimension splitting to complete the view splitting. Further, a model is established using the multi-view algorithm, and the user profile can be constructed.

[0085] Based on the above data processing method, the embodiments of the present disclosure also provide a data processing device. The following will be combined with Figure 6 to describe this device in detail.

[0086] Figure 6 Schematically shows a structural block diagram of a data processing device according to an embodiment of the present disclosure.

[0087] As Figure 6 shown, the data processing device 600 of this embodiment includes an acquisition module 610, a data splitting module 620, and a model building module 630.

[0088] The acquisition module 610 is configured to acquire original data, preprocess the original data, and acquire data of a view to be split. Among them, the data of the view to be split is single-view data. Among them, the single-view data is high-dimensional data, and the dimensions in the high-dimensional data correspond to entity features. In one embodiment, the acquisition module 610 can be used to perform the operation S210 described above, which will not be elaborated here.

[0089] The data splitting module 620 is configured to perform view splitting on the data of the view to be split according to dimensions based on the principal component analysis method, and convert the data of the view to be split into multi-view data. The multi-view data is data distributed in multiple views. Among them, the number of dimensions of the data in each view may be the same or different, and the aggregation of the number of dimensions of the data in each view is equal to the number of dimensions of the single-view data. In one embodiment, the data splitting module 620 can be used to perform the operation S220 described above, which will not be elaborated here.

[0090] The model building module 630 is configured to build a model based on the multi-view data using a single-view algorithm or a multi-view algorithm. In one embodiment, the model building module 630 can be used to perform the operation S230 described above, which will not be elaborated here.

[0091] According to an embodiment of the present disclosure, the data splitting module may further include a construction sub-module, a solution sub-module, and a dimension splitting sub-module.

[0092] Figure 7 Schematically shows a structural block diagram of a data splitting module according to an embodiment of the present disclosure.

[0093] As Figure 7 shown, the data splitting module 620 of this embodiment includes a construction sub-module 6201, a solution sub-module 6202, and a dimension splitting sub-module 6203.

[0094] Among them, the construction sub-module 6201 is configured to construct a principal component analysis classification indication function. Among them, the principal component analysis classification indication function is constructed based on a feature classification indication matrix. The feature classification indication matrix is an n×k matrix, where n is the dimension of the data to be split, and k is the preset number of view splits.

[0095] The solving sub-module 6202 is configured to solve the optimization problem of the principal component analysis classification indication function and obtain a solving result, where the solving result includes a weight allocation result of the feature classification indication matrix, and the weight allocation result of the feature classification indication matrix is a matrix composed of eigenvectors corresponding to the minimum eigenvalue corresponding to the number of view splits.

[0096] The dimension splitting sub-module 6203 is configured to perform view splitting on the dimension of the view data to be split based on the weight allocation result of the feature classification indication matrix, and obtain a data dimension splitting result, where the dimension of the view data to be split is m, and splitting the i-th data dimension includes: and splitting the i-th data dimension into the view corresponding to the maximum eigenvalue in the eigenvector corresponding to the i-th data dimension in the feature classification indication matrix, where i ∈ [1, m].

[0097] According to an embodiment of the present disclosure, the data splitting module may further include a data allocation sub-module.

[0098] Figure 8 Schematically shows a structural block diagram of a data splitting module according to an embodiment of the present disclosure.

[0099] As Figure 8 shown, in addition to including a construction sub-module 6201, a solving sub-module 6202, and a dimension splitting sub-module 6203, the data splitting module 620 of this embodiment may further include a data allocation sub-module 6204.

[0100] Among them, the data allocation sub-module 6204 is configured to perform view data allocation on the view data to be split according to the dimension splitting result, and obtain the multi-view data.

[0101] According to an embodiment of the present disclosure, any plurality of modules among the acquisition module 610, the data splitting module 620, the model establishment module 630, the construction sub-module 6201, the solution sub-module 6202, the dimension splitting sub-module 6203, and the data allocation sub-module 6204 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the acquisition module 610, the data splitting module 620, the model establishment module 630, the construction sub-module 6201, the solution sub-module 6202, the dimension splitting sub-module 6203, and the data allocation sub-module 6204 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the acquisition module 610, the data splitting module 620, the model establishment module 630, the construction sub-module 6201, the solution sub-module 6202, the dimension splitting sub-module 6203, and the data allocation sub-module 6204 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0102] Figure 9 FIG. schematically shows a block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure.

[0103] As Figure 9 shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 901 may also include on-board memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0104] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.

[0105] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, and the input / output (I / O) interface 905 is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the I / O interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage portion 908 as needed.

[0106] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.

[0107] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or apparatus. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903.

[0108] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present disclosure.

[0109] When the computer program is executed by the processor 901, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0110] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 909, and / or be installed from the removable medium 911. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0111] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or be installed from the removable medium 911. When the computer program is executed by the processor 901, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0112] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0114] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0115] The embodiments of the present disclosure have been described above. However, these embodiments are merely for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present disclosure.

Claims

1. A data processing method, characterized in that, Including: Obtain the original data, preprocess the original data, and obtain the view data to be split, where the view data to be split is single-view data; Based on the principal component analysis method, split the view data to be split according to dimensions, and convert the view data to be split into multi-view data; and Based on the multi-view data, establish a model using a multi-view algorithm, wherein the single-view data is high-dimensional data, and the dimensions in the high-dimensional data correspond to entity features; the multi-view data is data allocated to multiple views, where the number of dimensions of the data in each view is the same or different, and the aggregation of the number of dimensions of the data in each view is equal to the dimensions of the single-view data, wherein the step of splitting the view data to be split according to dimensions based on the principal component analysis method and converting the view data to be split into multi-view data includes the step of dimension splitting, and the step of dimension splitting includes: Construct a principal component analysis classification indicator function, where the principal component analysis classification indicator function is constructed based on a feature classification indicator matrix, and the feature classification indicator matrix is an n×k matrix, where n is the dimension of the view data to be split, and k is the preset number of view splits; Solve the optimization problem of the principal component analysis classification indicator function to obtain a solution result, where the solution result includes the weight allocation result of the feature classification indicator matrix, and the weight allocation result of the feature classification indicator matrix is a matrix composed of eigenvectors corresponding to the minimum eigenvalue corresponding to the number of view splits; and Based on the weight allocation result of the feature classification indicator matrix, split the dimensions of the view data to be split to obtain a data dimension split result, where the dimension of the view data to be split is m, and splitting the i-th data dimension includes: Split the i-th data dimension into the view corresponding to the maximum eigenvalue in the eigenvector corresponding to the i-th data dimension in the feature classification indicator matrix, where i ∈ [1, m].

2. The method according to claim 1, wherein After obtaining the dimension split result of the m-dimensional data, the method further includes: Allocate the view data to be split according to the dimension split result to obtain the multi-view data.

3. A method according to claim 1, wherein, The preset number of view splits is obtained by processing the dimensions of the view data to be split based on one of the automatic clustering algorithm or the similarity algorithm.

4. A method according to claim 3, wherein, The automatic clustering algorithm includes one of the density-based spatial clustering of applications with noise algorithm, fuzzy clustering algorithm, and K-means clustering algorithm; and / or, the similarity algorithm includes one of the cosine similarity algorithm, distance similarity algorithm, and Pearson correlation coefficient.

5. A method according to claim 1, wherein, The preprocessing of the original data includes: Normalize the original data.

6. A method according to claim 1, wherein, The solving of the optimization problem of the principal component analysis classification indicator function includes: Solve the optimization problem of the principal component analysis classification indicator function based on the generalized eigenvalue solving method.

7. A method according to claim 1, wherein, The multi-view algorithm includes one of canonical correlation analysis, multi-canonical correlation analysis, kernel canonical correlation analysis, local preserving canonical correlation analysis, discriminant canonical correlation analysis, generalized multi-view analysis, multi-view discriminant analysis, and multi-view dimensionality reduction model.

8. A method according to claim 1, wherein, The original data includes user feature data, and the model is used to construct a user profile.

9. A data processing device, comprising: An acquisition module, configured to acquire original data, preprocess the original data, and acquire data of views to be split, where the data of views to be split is single-view data, and the single-view data is high-dimensional data, and the dimensions in the high-dimensional data correspond to entity features; A data splitting module, configured to split the data of views to be split by dimension based on the principal component analysis method, and convert the data of views to be split into multi-view data, where the multi-view data is data allocated to multiple views, and the number of dimensions of the data in each view is the same or different, and the aggregation of the number of dimensions of the data in each view is equal to the number of dimensions of the single-view data; and A model establishment module, configured to establish a model based on the multi-view data by using a multi-view algorithm, where the data splitting module includes: A construction sub-module: configured to construct a principal component analysis classification indication function, where the principal component analysis classification indication function is constructed based on a feature classification indication matrix, and the feature classification indication matrix is an n×k matrix, where n is the number of dimensions of the data of views to be split, and k is the preset number of view splits; A solution sub-module: configured to solve the optimization problem of the principal component analysis classification indication function, and obtain a solution result, where the solution result includes a weight allocation result of the feature classification indication matrix, and the weight allocation result of the feature classification indication matrix is a matrix composed of eigenvectors corresponding to the minimum eigenvalue corresponding to the number of view splits; and A dimension splitting sub-module: configured to split the dimensions of the data of views to be split based on the weight allocation result of the feature classification indication matrix, and obtain a data dimension splitting result, where the number of dimensions of the data of views to be split is m, and splitting the i-th data dimension includes: Splitting the i-th data dimension into the view corresponding to the maximum eigenvalue in the eigenvector corresponding to the i-th data dimension in the feature classification indication matrix, where i∈[1,m].

10. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 8.

12. A computer program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Eigenvector dimension reduction method and medical image recognition method, device and storage medium

    CN109409416A

  • Distributed similarity learning for high-dimensional image features

    US20150146973A1