A Dimensionality Reduction Visualization Method, Device and Medium for High-Dimensional Scientific and Technological Innovation Data

Through cleaning, standardizing, automated dimension reduction and intelligent interactive visualization of high-dimensional science and technology innovation data, the efficiency and visualization problems in high-dimensional data processing are solved, and efficient data simplification and user-friendly interactive analysis are achieved.

CN119416178BActive Publication Date: 2025-08-05WUHAN STARLINK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510035459.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-08-05
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Traditional data processing methods and visualization technologies are difficult to cope with the "dimensional disaster" of high-dimensional science and technology innovation data, resulting in inefficient data processing and poor visualization effects. The existing dimensionality reduction algorithms are limited in scope and have high computational complexity, which cannot support users' in-depth exploration and analysis of data characteristics.

Method used

By obtaining high-dimensional science and technology data for cleaning and standardization, the dimensionality reduction algorithm is automatically selected and distributed dimensionality reduction processing is performed, combined with intelligent interactive visualization tools for three-dimensional visual display, and dynamic optimization is performed in combination with user interaction behavior.

Benefits of technology

It realizes efficient dimensionality reduction processing, improves the accuracy and speed of data processing, supports users' in-depth exploration and analysis of data, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416178B_ABST
    Figure CN119416178B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing and visualization, and specifically to a method, device, and medium for dimensionality reduction and visualization of high-dimensional scientific and technological innovation data. The method includes the following steps: obtaining high-dimensional scientific and technological innovation data, performing data cleaning and standardization processing to obtain a preprocessed data set; extracting features from the preprocessed data set and automatically selecting a dimensionality reduction algorithm to generate dimensionality reduction parameters; performing distributed dimensionality reduction processing based on the dimensionality reduction parameters to obtain a low-dimensional data representation; inputting the low-dimensional data into an intelligent interactive visualization tool for three-dimensional visualization display, and dynamically optimizing it in combination with the user's interaction behavior. The present invention improves the processing efficiency of high-dimensional data and supports users' in-depth interaction and analysis of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing and visualization, and particularly to a method, device and medium for dimensionality reduction visualization of high-dimensional scientific and technological innovation data. Background Art

[0002] With the continuous development of science and technology, the amount of data and the data dimension in the scientific research field are increasing day by day. Especially in the process of scientific innovation and technological R & D, the high-dimensional characteristics of data bring great challenges to data analysis and visualization. Traditional data processing methods and visualization technologies are difficult to cope with this "curse of dimensionality" problem, resulting in low data processing efficiency and poor visualization effects.

[0003] Dimensionality reduction technology is an effective means to deal with high-dimensional data processing problems. It simplifies the data representation form by reducing the data dimension while retaining the main features of the data. However, existing dimensionality reduction algorithms often have problems such as limited applicable range, high computational complexity, and poor adaptability to different data sets. Especially for large-scale high-dimensional scientific and technological innovation data, how to select a suitable dimensionality reduction algorithm and achieve efficient computational processing is still an urgent technical problem to be solved.

[0004] In addition, traditional visualization tools usually lack intelligent interaction functions when facing the dimensionality-reduced data, and cannot effectively support users to deeply explore and analyze data features. When users interact with data, it is difficult to dynamically adjust the display content according to their own needs, which limits the data exploration process and cannot fully exploit the potential value of data. Therefore, developing a visualization system that can automate dimensionality reduction processing and support user interaction has important practical significance for improving the efficiency of high-dimensional data analysis and user experience. Summary of the Invention

[0005] The present invention provides a method, device and medium for dimensionality reduction visualization of high-dimensional scientific and technological innovation data to solve the problem of how to efficiently perform dimensionality reduction on complex high-dimensional scientific and technological innovation data, so as to maintain the core features during the visualization process and support users to dynamically explore and deeply analyze the data through interactive behaviors.

[0006] To solve the above technical problems, the present invention provides a method for dimensionality reduction visualization of high-dimensional scientific and technological innovation data, including:

[0007] Obtain high-dimensional scientific and technological innovation data, clean and standardize the data to obtain a preprocessed data set;

[0008] Extract and analyze features from the preprocessed data set, automatically select a dimensionality reduction algorithm, and obtain dimensionality reduction parameters;

[0009] Based on the dimensionality reduction parameters, perform distributed dimensionality reduction processing on the data set to obtain a low-dimensional data representation after dimensionality reduction;

[0010] Input the low-dimensional data into an intelligent interactive visualization tool, perform three-dimensional visualization on the data, and dynamically optimize it in combination with user interaction behaviors to generate an interactive visualization display result.

[0011] Further, the method further includes:

[0012] Obtain the high-dimensional scientific and technological innovation data, and extract a high-dimensional data set including multi-dimensional features, indicators, and experimental data from a specified data source;

[0013] Perform data cleaning on the data set, remove missing values, outliers, and duplicates to obtain a cleaned data set;

[0014] Perform standardization processing on the cleaned data to unify the dimensions and scales of different features and generate a standardized data set.

[0015] Further, the method further includes:

[0016] Extract the main features of the data from the standardized data set;

[0017] Perform statistical analysis on the extracted features, evaluate the complexity of the data, and determine the dimensionality reduction requirements;

[0018] Automatically select a dimensionality reduction algorithm based on the analysis results and generate dimensionality reduction parameters.

[0019] Further, the dimensionality reduction algorithm includes an improved PCA algorithm and a t-SNE algorithm, and the dimensionality reduction parameters include the target number of dimensions, the number of iterations, and the learning rate.

[0020] Further, the method further includes:

[0021] Determine a dimensionality reduction strategy based on the dimensionality reduction parameters and initialize the dimensionality reduction calculation process in the data set;

[0022] Use distributed computing technology to perform distributed dimensionality reduction processing on the data set and map the data into a low-dimensional space;

[0023] Merge the calculation results of distributed computing nodes to obtain a complete low-dimensional data representation.

[0024] Further, the distributed dimensionality reduction processing includes:

[0025] Execute the dimensionality reduction algorithm on multiple distributed nodes to perform dimensionality reduction calculation on the sub-data sets held by each;

[0026] Each node calculates a data projection matrix or a low-dimensional embedding matrix according to the target number of dimensions of the dimensionality reduction algorithm.

[0027] Further, the method further includes:

[0028] Inputting the low-dimensional data representation into an intelligent interactive visualization tool to generate an initial three-dimensional data visualization display;

[0029] Performing three-dimensional visualization rendering on the data to support users in exploring the data features through interaction methods such as rotation, scaling, and selection.

[0030] Further, the system includes:

[0031] A data preprocessing module for obtaining high-dimensional science and technology innovation data from multiple data sources, and performing data cleaning and standardization processing on the data to generate a preprocessed data set;

[0032] A feature extraction and dimensionality reduction algorithm selection module for extracting features from the preprocessed data set and automatically selecting a dimensionality reduction algorithm according to the data complexity to generate dimensionality reduction parameters;

[0033] A distributed dimensionality reduction calculation module for performing distributed dimensionality reduction processing on the preprocessed data set based on the dimensionality reduction parameters to generate a low-dimensional data representation;

[0034] An intelligent interactive visualization module for inputting the low-dimensional data representation into an intelligent interactive visualization tool and supporting users in exploring data features through interaction methods such as rotation, scaling, and selection to generate an interactive visualization display result.

[0035] Further, the device includes:

[0036] A data acquisition unit for obtaining high-dimensional science and technology innovation data from multiple data sources, where the data includes experimental data, sensor data, and historical data;

[0037] A data processing unit for performing data cleaning and standardization processing on the acquired high-dimensional data to obtain a preprocessed data set;

[0038] A feature extraction and dimensionality reduction unit for extracting effective features from the preprocessed data set, automatically selecting a dimensionality reduction algorithm, and generating dimensionality reduction parameters;

[0039] A visualization rendering unit for inputting the dimensionally reduced low-dimensional data into a three-dimensional visualization engine to generate a visualization display of the three-dimensional data, supporting users in exploring the data through interaction methods.

[0040] Further, a computer program is stored thereon, and when the program is executed by a processor, it executes the dimensionality reduction visualization method for high-dimensional science and technology innovation data described in any one of the above.

[0041] The key innovation points of the present invention include:

[0042] (1) Automatically select a dimensionality reduction algorithm and combine it with data complexity evaluation to achieve a more intelligent and accurate dimensionality reduction process;

[0043] (2) A distributed dimensionality reduction calculation method that improves the processing efficiency of large-scale data;

[0044] (3) An intelligent interactive visualization tool that provides optimized functions for user dynamic interaction and data display, enhancing the intuitiveness and convenience of data exploration.

[0045] The following are its main beneficial effects:

[0046] (1) Efficient data processing: Through data cleaning, standardization, and feature extraction, the present invention effectively simplifies complex high-dimensional data, automatically selects an appropriate dimensionality reduction algorithm, improves the accuracy and speed of dimensionality reduction processing. Especially when processing multi-dimensional science and innovation data, it can significantly reduce the computational complexity and solve the bottleneck of traditional dimensionality reduction methods in high-dimensional data processing.

[0047] (2) Distributed dimensionality reduction calculation: Adopting a distributed architecture, the data set is allocated to multiple nodes for parallel dimensionality reduction processing, ensuring that large-scale data can be quickly dimensionally reduced and improving the processing efficiency. Compared with the dimensionality reduction method of a single node, distributed computing can significantly shorten the calculation time and improve the adaptability to large-scale data sets.

[0048] (3) Intelligent interactive visualization: Through a three-dimensional visualization tool, the present invention displays the dimensionally reduced data in three dimensions and combines the user's interaction behavior to dynamically optimize the display effect in real time, supporting the user to conduct in-depth exploration and analysis of the data. Compared with traditional static displays, the intelligent interactive function of the present invention improves the user experience, enabling the user to flexibly view and analyze data features. Brief Description of the Drawings

[0049] Figure 1 It is a schematic flowchart of a dimensionality reduction visualization method for high-dimensional science and innovation data provided by an embodiment of the present application;

[0050] Figure 2 It is a system structure block diagram of a dimensionality reduction visualization method for high-dimensional science and innovation data provided by an embodiment of the present application;

[0051] Figure 3 It is a structure block diagram of a computer device provided by an embodiment of the present application. Detailed Embodiments

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "comprising" and "having" and any variations thereof in the specification, claims and drawings of this application are intended to cover non-exclusive inclusion. The terms "first" and "second" in the specification, claims or drawings of this application are used to distinguish different objects and not to describe a specific order.

[0053] Reference herein to "embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0054] Embodiment 1:

[0055] Refer to Figure 1 , which is a schematic flowchart of a method for dimensionality reduction visualization of high-dimensional scientific and technological innovation data provided by an embodiment of the present invention. This process can at least include steps S100 - S400:

[0056] S100. Obtain high-dimensional scientific and technological innovation data, clean and standardize the data to obtain a preprocessed data set;

[0057] S200. Extract and analyze features from the preprocessed data set, automatically select an appropriate dimensionality reduction algorithm, and obtain dimensionality reduction parameters;

[0058] S300. Based on the dimensionality reduction parameters, perform distributed dimensionality reduction processing on the data set to obtain a low-dimensional data representation after dimensionality reduction;

[0059] S400. Input the low-dimensional data into an intelligent interactive visualization tool, perform three-dimensional visualization on the data, and perform dynamic optimization in combination with user interaction behavior to generate an interactive visualization display result.

[0060] Step S100 at least includes S110 - S130:

[0061] S110. Obtain high-dimensional scientific and technological innovation data, and extract a high-dimensional data set from a specified data source, including various multi-dimensional features, indicators, and experimental data.

[0062] In step S110, the system extracts a high-dimensional scientific and technological innovation dataset from a specified data source. The dataset includes multi-dimensional features, metrics, and experimental data, specifically including data tables obtained from a scientific and technological innovation data platform and experimental data collected by sensors. Define this dataset as , where = { , , …, }, and each is a sample with multi-dimensional features, and the number of features is m. These features include scientific and technological innovation metrics, project parameters, and research result data. For the feature set = { , , …, } of each sample, the data sources of each dimension are different and the dimensions are different, laying a foundation for subsequent unified processing.

[0063] S120: Clean the data to remove missing values, outliers, and duplicates.

[0064] Check the integrity of the dataset , process the features with missing values, and define the processed dataset as . For the missing part in each feature dimension , use the mean filling strategy and interpolation method to complete it. Specifically, if a certain dimension is missing, calculate the mean of this dimension in all samples:

[0065] , if is missing, then =

[0066] Furthermore, use the z-score normalization method to detect outliers for each feature , and calculate the standard deviation and mean of each feature:

[0067]

[0068] When | | > 3, determine this data point as an outlier and process it, for example, replace it with the mean of adjacent normal values. The processed dataset continues to be .

[0069] Check the dataset for duplicates. If it is found that the sample is exactly the same as other samples, delete it. The dataset obtained after removing duplicates is defined as , for standardization processing.

[0070] S130: Perform standardization processing on the cleaned data to unify the dimensions and scales of different features and generate a standardized data set.

[0071] After completing the data cleaning, enter step S130 to perform standardization processing on the cleaned data set to obtain a standardized data set .

[0072] Specifically, for each feature , use the Min - Max normalization method to scale it to the interval [0,1]. The formula is as follows:

[0073]

[0074] where and are respectively the minimum and maximum values of the feature in the data set. Through this standardization step, the obtained standardized data set is ={ , ,…, }, and each contains the standardized feature.

[0075] Furthermore, after the standardization processing in step S130, the obtained data set will be further input into the feature extraction module (S200) of the next stage for automatically selecting a suitable dimensionality reduction algorithm.

[0076] Step S200 at least includes steps S210 - S230:

[0077] S210: Extract the main features of the data from the pre - processed data set for use in the analysis of S220.

[0078] In step S210, based on the data set ={ , ,…, } that has been standardized in S130, extract the main features of the data set. Specifically, by analyzing the data distribution of each feature dimension and the correlation between features, extract the core features of the data.

[0079] For each standardized feature dimension in the data set calculate its statistics, including the mean , variance , and the skewness and kurtosis of the data distribution.

[0080]

[0081]

[0082] Furthermore, based on the linear correlations among all features calculate the correlation coefficient matrix R between features. The elements in the matrix represent the Pearson correlation coefficient between feature and , and are defined as follows:

[0083]

[0084] Through the correlation coefficient matrix R, the linear relationships between features can be further analyzed to obtain highly correlated feature pairs for reference in subsequent dimensionality reduction.

[0085] According to the feature distribution and correlation analysis results, extract the main features of the dataset and generate a feature set ={ , … } for statistical analysis in the subsequent step S220.

[0086] S220: Conduct statistical analysis on the extracted features, and evaluate the complexity of the data and the dimensionality reduction requirements according to the nature of the features.

[0087] Evaluate the linear and non - linear properties of each feature ∈ and calculate its linearity index and non - linearity index . Specifically, represents the degree of linear relationship between the feature and other features, while measures the non - linear distribution of the feature. They are defined as follows:

[0088]

[0089] Furthermore, for the feature set, calculate the complexity of the overall dataset to determine the difficulty of dimensionality reduction. The complexity is defined as the weighted average of the non - linearity indices of all features:

[0090]

[0091] According to the complexity of the data, classify and evaluate the dimensionality reduction requirements. When > When the data has a high complexity, a non-linear dimensionality reduction method is preferably selected; when ≤ a linear dimensionality reduction method can meet the requirements, where is a preset complexity threshold.

[0092] S230: Based on the analysis results, automatically select an appropriate dimensionality reduction algorithm and generate dimensionality reduction parameters.

[0093] After completing the statistical analysis and complexity evaluation of the features, step S230 is entered. Based on the analysis results, an appropriate dimensionality reduction algorithm is automatically selected and dimensionality reduction parameters are generated to guide the subsequent dimensionality reduction process.

[0094] According to the complexity classification results obtained in step S220, select a suitable data dimensionality reduction algorithm. When > a non-linear dimensionality reduction method is selected; otherwise, a linear dimensionality reduction method, such as an improved PCA algorithm, is selected.

[0095] For the selected dimensionality reduction algorithm, further determine the dimensionality reduction parameter set , including the target dimension number d, the number of iterations , and the learning rate α. Specifically, the target dimension number d can be set according to the cumulative contribution rate of the features, satisfying a contribution rate greater than 90%.

[0096] Output the generated dimensionality reduction parameter set to guide the subsequent dimensionality reduction processing (S300 module), including the selected dimensionality reduction algorithm type and specific parameter configuration.

[0097] Step S300 at least includes steps S310 - S330:

[0098] S310: Use the selected dimensionality reduction parameters to determine a dimensionality reduction strategy and initialize the dimensionality reduction calculation process in the dataset.

[0099] Based on the dimensionality reduction parameter set , determine the type of dimensionality reduction algorithm to be used. If the improved PCA algorithm is selected, use the target dimension d in the parameter set and the sample matrix ={ , ,…, } to initialize the PCA dimensionality reduction matrix calculation; if the improved t-SNE algorithm is selected, initialize the t-SNE similarity metric matrix calculation.

[0100] Furthermore, set the initial parameters of the dimensionality reduction calculation according to the selected dimensionality reduction algorithm, such as the target dimension number d, the number of iterations and the learning rate α. For PCA, calculate the feature covariance matrix C:

[0101]

[0102] where is the mean vector of the dataset. For t-SNE, calculate the similarity probability based on the Euclidean distance :

[0103]

[0104] Initialize the distributed computing environment and divide the dataset into multiple sub-datasets ([[]] , ,… ), where p is the number of distributed nodes, ensuring that each node is assigned a dimensionality reduction task at the initial stage.

[0105] S320: Perform distributed dimensionality reduction on the dataset based on the improved PCA or t-SNE algorithm and map the data into a low-dimensional space.

[0106] On each distributed node, perform dimensionality reduction calculation on the respective sub-dataset where (q = 1, 2, …, p). For the nodes using the PCA algorithm, calculate the data projection matrix :

[0107]

[0108] where is the eigenvector matrix of the first d principal components obtained by eigenvalue decomposition. For the nodes using the t-SNE algorithm, update the low-dimensional embedding by gradient descent:

[0109]

[0110] where is the cost function of t-SNE, and ∇ is its gradient.

[0111] Perform multiple iterative calculations on each distributed node to make the low-dimensional representation converge to a reasonable range. For the calculation of t-SNE, use the locally calculated similarity probability to map the low-dimensional points until the cost function converges:

[0112]

[0113] Each distributed node will output the calculated low-dimensional representation matrix Keep it local for subsequent merging in step S330.

[0114] S330: Perform dimensionality reduction calculation in a distributed environment, merge the calculation results of each distributed node, and obtain the low-dimensional data representation after dimensionality reduction.

[0115] Collect the low-dimensional representation matrices of each sub-dataset from all distributed nodes ( , …, } Further, merge all sub-matrices according to the sample index to generate a complete low-dimensional dataset representation .

[0116] Perform smoothing on the merged low-dimensional dataset to remove possible inconsistencies caused by distributed computing and ensure that the overall data structure in the low-dimensional space remains consistent.

[0117] Output the final low-dimensional data representation , which is used in the subsequent intelligent interaction visualization tool (S400 module) for users to conduct in-depth exploration and interactive operations on the data.

[0118] Step S400 at least includes steps S410 - S430:

[0119] S410: Input the low-dimensional data representation after dimensionality reduction into the intelligent interaction visualization tool to generate an initial three-dimensional data visualization display.

[0120] Map the low-dimensional data representation = { , , …, } to the three-dimensional coordinate space, where each data point represents the three-dimensional coordinate position of a sample, defined as = ( , , ). Determine the range and distribution of the three-dimensional coordinate axes according to the feature range.

[0121] Further, use a three-dimensional rendering engine to draw all data points in the three-dimensional space. For each point , mark it with a corresponding color or shape in the three-dimensional space to distinguish different types of samples or categories. The feature label information is extracted from the feature set in the previous feature extraction stage (S210) to display the corresponding feature names in the visualization.

[0122] Generate an initial three-dimensional data graph , as the basis for subsequent interactive display. The graph is displayed on the user interface and awaits the user's interactive input.

[0123] S420: Perform three-dimensional visualization rendering on the data, supporting the user to explore different features of the data through rotation, scaling, and selection interaction methods.

[0124] After completing the initial three-dimensional data display, step S420 is entered to perform three-dimensional visualization rendering on the data, supporting the user to explore different features of the data through rotation, scaling, and selection interaction methods.

[0125] Understandably, receiving the user's interactive operation instructions, including rotation, scaling, and selection operations. When the user requests to rotate the three-dimensional view, the system is based on the angle θ input by the user, for the three-dimensional graph to perform a rotation transformation. The rotation matrix R(θ, ) acts on each data point , and the rotated coordinates are calculated:

[0126]

[0127] where R(θ, ) is a three-dimensional rotation matrix constructed according to the rotation angle specified by the user.

[0128] Furthermore, the user can adjust the visualization scale through the scaling operation. The system performs a scaling transformation on all data point coordinates in the three-dimensional graph according to the scaling factor s input by the user:

[0129]

[0130] where s > 1 represents a magnification operation, and s < 1 represents a reduction operation.

[0131] When the user performs a data point selection operation, the system captures the selected point , and displays the feature information of this point based on the previous feature extraction results (S210). The feature information of the selected data point is displayed by querying the feature set for the user to deeply understand the specific feature content of this data.

[0132] S430: Combine the user's interaction behavior to dynamically optimize the visualization display result, adjust the display angle and data details to generate an interactive visualization display result.

[0133] Based on the user's interaction behavior record, the system analyzes the user's operation frequency and operation type to judge the user's focus and interaction preference during the data exploration process. Define the user behavior matrix B = { ), where Indicates the frequency of the j-th type of operation (such as rotation, scaling, selection) performed by the user on the data point For example, if the user selects a certain type of data point multiple times, the system automatically magnifies the display of that area and adds detailed markings to the data points. At the same time, for the direction of frequent user interaction, the default display angle of the 3D graph is adjusted to better conform to the user's operation habits.

[0134] According to the user behavior matrix B, dynamically optimize the display angle and details of the 3D visualization.

[0135] Generate the final interactive visualization display result and present it in the user interface. The optimized visualization results include the adjusted positions of the data points, automatically labeled feature information, and optimized 3D perspectives, supporting users to conduct more in-depth data exploration.

[0136] The key innovations of this invention include:

[0137] (1) Automatically select the dimensionality reduction algorithm and combine it with data complexity evaluation to achieve a more intelligent and accurate dimensionality reduction process;

[0138] (2) Distributed dimensionality reduction calculation method, which improves the processing efficiency of large-scale data;

[0139] (3) Intelligent interactive visualization tool, which provides optimization functions for user dynamic interaction and data display, enhancing the intuitiveness and convenience of data exploration.

[0140] The following are its main beneficial effects:

[0141] (1) Efficient data processing: Through data cleaning, standardization, and feature extraction, this invention effectively simplifies complex high-dimensional data, automatically selects appropriate dimensionality reduction algorithms, improves the accuracy and speed of dimensionality reduction processing. Especially when dealing with multi-dimensional scientific and technological innovation data, it can significantly reduce the computational complexity and solve the bottleneck of traditional dimensionality reduction methods in high-dimensional data processing.

[0142] (2) Distributed dimensionality reduction calculation: Adopting a distributed architecture, the data set is allocated to multiple nodes for parallel dimensionality reduction processing, ensuring that large-scale data can be quickly dimensionally reduced, improving the processing efficiency. Compared with the dimensionality reduction method of a single node, distributed computing can significantly shorten the calculation time and improve the adaptability to large-scale data sets.

[0143] (3) Intelligent interaction visualization: Through three-dimensional visualization tools, the present invention performs three-dimensional display on the data after dimensionality reduction, and combines the interactive behavior of users to dynamically optimize the display effect in real time, supporting users to conduct in-depth exploration and analysis of the data. Compared with traditional static display, the intelligent interaction function of the present invention improves the user experience, enabling users to flexibly view and analyze data features.

[0144] Embodiment 2:

[0145] Figure 2 It is a structural block diagram of a dimensionality reduction visualization system for high-dimensional science and technology innovation data provided by an embodiment of the present invention. As Figure 2 shown, the system may include:

[0146] The data preprocessing module 11 is used to obtain high-dimensional science and technology innovation data from multiple data sources and perform preprocessing on it, including data cleaning, missing value filling, and data standardization operations. Specifically, the data preprocessing module 10 first obtains a data set containing multi-dimensional features from multiple sources (such as experimental data, sensor-collected data, historical literature data), performs cleaning processing on it to remove outliers and missing values, and then standardizes different features to unify the data scale, providing a basis for subsequent data analysis.

[0147] The feature extraction and dimensionality reduction algorithm selection module 12 is used to extract effective features from the data set output by the data preprocessing module 10 and automatically select a suitable dimensionality reduction algorithm according to the nature of the features and the complexity of the data. Specifically, this module performs feature extraction on the preprocessed data set, determines the dimensionality reduction requirements by analyzing the correlation between features and the complexity of the data, and automatically selects an appropriate dimensionality reduction algorithm, such as an improved PCA or t-SNE algorithm, to ensure the efficiency and effect of dimensionality reduction.

[0148] The distributed dimensionality reduction calculation module 13 is used to perform dimensionality reduction calculation and map high-dimensional data into a low-dimensional space. Specifically, based on the dimensionality reduction parameters generated in the feature extraction and dimensionality reduction algorithm selection module 20, this module uses distributed computing technology to perform dimensionality reduction processing on the data to improve the calculation efficiency for large-scale data sets. Adopting a distributed architecture, the data is distributed to multiple nodes for parallel dimensionality reduction calculation, and finally the calculation results of all nodes are merged to obtain the low-dimensional data representation after dimensionality reduction.

[0149] The intelligent interaction visualization module 14 is used to perform three-dimensional visualization display on the data after dimensionality reduction and support dynamic interaction operations of users. Specifically, the intelligent interaction visualization module 40 inputs the low-dimensional data output by the distributed dimensionality reduction calculation module 30 into a three-dimensional visualization engine to generate an initial visualization display. Users can perform rotation, scaling, and selection operations on the data through the interactive interface, and the system dynamically adjusts the display effect according to the interactive behavior of users, enabling users to more intuitively explore and understand data features.

[0150] Through the integration of multifunctional modules for data preprocessing, feature extraction, distributed dimensionality reduction, and intelligent interactive visualization, the present invention can efficiently perform dimensionality reduction processing and visualization of high-dimensional scientific and technological innovation data, solving the problems of computational complexity and visualization difficulties caused by the curse of dimensionality in traditional methods. This system can significantly improve the data processing efficiency and provide users with an intuitive and dynamic visualization interaction interface, improving the user experience and enabling users to better mine useful information from high-dimensional data.

[0151] The high-dimensional scientific and technological innovation data dimensionality reduction and visualization system provided by the present invention can efficiently complete the dimensionality reduction processing and visualization of high-dimensional data through the integration of a data preprocessing module, a feature extraction and dimensionality reduction module, a distributed dimensionality reduction calculation module, and an intelligent interactive visualization module. This system achieves efficient processing of large-scale data sets through a distributed architecture and helps users intuitively understand the features and internal relationships of data in an interactive manner through visualization, significantly improving the user experience.

[0152] Embodiment III:

[0153] Figure 3 The structural block diagram of a computer device according to an embodiment of the present application is shown. As Figure 3 shown, the device may include:

[0154] The data acquisition unit 21 is used to acquire high-dimensional scientific and technological innovation data from multiple data sources, and the data includes experimental data, sensor data, and historical data; further, the data acquisition unit 21 is also used to acquire multi-dimensional feature information in the high-dimensional data for subsequent processing.

[0155] The data processing unit 22 is used to clean and standardize the acquired high-dimensional scientific and technological innovation data, including removing missing values, outliers, and duplicates to ensure the consistency and accuracy of the data; the data processed by the data processing unit 22 will be used for subsequent feature extraction and dimensionality reduction analysis.

[0156] The feature extraction and dimensionality reduction unit 23 is used to extract effective features from the data output by the data processing unit 22 and automatically select a dimensionality reduction algorithm; specifically, the feature extraction and dimensionality reduction unit 23 selects a suitable dimensionality reduction method based on the correlation and complexity analysis of features, including improved PCA and t-SNE algorithms, to generate low-dimensional data after dimensionality reduction.

[0157] The visualization presentation unit 24 is used to input the low-dimensional data after dimensionality reduction into a three-dimensional visualization engine to generate a visualization display of the three-dimensional data; further, it supports users to rotate, zoom, and select the visualization content interactively to more intuitively explore and understand the features of the high-dimensional data.

[0158] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for dimensionality reduction and visualization of high-dimensional scientific and technological innovation data can be realized.

[0159] This computer-readable storage medium can be a non-volatile storage medium, such as a read-only memory (ROM), a flash memory, or other types of memories, or it can also be a volatile memory, such as a random access memory (RAM).

[0160] Furthermore, an embodiment of the present application also provides a chip, which includes a processor and a memory. The processor is used to load and run the instructions stored in the memory, so that a computing device installed with the chip executes the above-mentioned method for dimensionality reduction and visualization of high-dimensional scientific and technological innovation data. Through the integration of data processing, dimensionality reduction calculation, and visualization display, the present invention can significantly improve the analysis efficiency of high-dimensional data and the user interaction experience.

[0161] The computer device structure provided by the present invention combines units such as data acquisition, processing, feature extraction and dimensionality reduction, and visualization presentation, and can realize the comprehensive processing and display of high-dimensional data. Through effective data preprocessing and feature extraction, the device automatically selects an appropriate dimensionality reduction algorithm, and combines with the dynamic interactive display of the visualization presentation unit, greatly improving the analysis efficiency of high-dimensional data and the user interaction experience. By combining the calculation process with user interaction, users can quickly gain insights into the data, supporting decision-making and scientific research work.

[0162] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure made directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is equally within the scope of the patent protection of the present application.

Claims

1. A dimensionality reduction and visualization method for high-dimensional scientific and technological data, characterized in that: The following steps are involved: S100, obtaining high-dimensional scientific and technological innovation data, cleaning and standardizing the data to obtain a preprocessed data set; S200, extracting and analyzing features from the preprocessed data set, automatically selecting an appropriate dimensionality reduction algorithm, and obtaining dimensionality reduction parameters; S300, performing distributed dimensionality reduction processing on the data set based on the dimensionality reduction parameters to obtain a low-dimensional data representation after dimensionality reduction; S400: Input the low-dimensional data into an intelligent interactive visualization tool, perform three-dimensional visualization on the low-dimensional data, and dynamically optimize it in combination with user interaction behavior to generate an interactive visualization display result; Step S200 at least includes steps S210-S230: S210: extracting the main features of the data from the preprocessed data set for analysis in S220; S220: Perform statistical analysis on the extracted features and evaluate the complexity of the data and the need for dimensionality reduction based on the nature of the features; complexity It is defined as the weighted average of all characteristic nonlinearity indices: ; According to the complexity of the data , classify and evaluate the dimensionality reduction requirements; when When , nonlinear dimensionality reduction method is selected; when When , the linear dimensionality reduction method is selected, where is a pre-set complexity threshold; m is the number of features, and j is an integer greater than or equal to 1; S230: Based on the analysis results, automatically select a suitable dimensionality reduction algorithm and generate dimensionality reduction parameters; The method further comprises: Based on the dimensionality reduction parameters, determine a dimensionality reduction strategy and initialize a dimensionality reduction calculation process in the data set; Using distributed computing technology, performing distributed dimensionality reduction processing on the data set to map the data into a low-dimensional space; Merge the computation results of distributed computing nodes to obtain a complete low-dimensional data representation.

2. The dimensionality reduction and visualization method of high-dimensional scientific and technological innovation data according to claim 1 is characterized in that: The method further comprises: Smoothing is performed on the merged low-dimensional dataset.

3. The dimensionality reduction and visualization method for high-dimensional scientific and technological data according to claim 1, characterized in that: The dimensionality reduction algorithm includes an improved PCA algorithm and / or a t-SNE algorithm, and the dimensionality reduction parameters include a target number of dimensions, a number of iterations, and a learning rate.

4. The dimensionality reduction and visualization method for high-dimensional scientific and technological innovation data according to claim 3 is characterized in that: The distributed dimensionality reduction process includes: Execute the dimensionality reduction algorithm on multiple distributed nodes to perform dimensionality reduction calculations on the sub-data sets held by each node; Each node calculates the data projection matrix and low-dimensional embedding matrix according to the target dimension of the dimensionality reduction algorithm.

5. The dimensionality reduction and visualization method of high-dimensional scientific and technological data according to claim 1 is characterized in that: The method further comprises: Inputting the low-dimensional data representation into an intelligent interactive visualization tool to generate an initial three-dimensional data visualization display; The low-dimensional data is visualized in three dimensions, allowing users to explore the characteristics of the data through interactive methods such as rotation, scaling, and selection.

6. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, executes the dimensionality reduction and visualization method for high-dimensional scientific and technological data according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • High-dimensional data processing method and device

    CN111950651A

  • Visual agricultural big data analysis interaction system and method

    CN118277745A