Multi-dimensional analysis-based medical and invasive big data model generation method and system

The method for generating big data models for science and technology innovation through multidimensional analysis solves the problem of incomplete data collection, improves the scientificity and reliability of the generated data, enhances the accuracy and applicability of the model, and ensures the high-quality conduct of science and technology innovation activities.

CN122020588APending Publication Date: 2026-05-12HUBEI KEHUITONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUBEI KEHUITONG TECH CO LTD
Filing Date
2026-01-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for generating big data models for science and technology innovation suffer from incomplete data collection and imprecise model evaluation, leading to a decrease in the quality and reliability of science and technology innovation activities.

Method used

A multidimensional analysis-based approach is adopted to divide the data into different batches. The model input and feedback data acquisition unit collects technical basic data, resource correlation data, correlation feature data, innovative theoretical achievement data, and model stability data in real time. Multidimensional data analysis nodes are used to calculate the model input complexity, resource adaptability, feature effectiveness, model innovation, model stability, and model adaptability coefficients for comprehensive evaluation and early warning feedback.

Benefits of technology

It improves the scientific rigor and reliability of the data generated by the model, enhances the accuracy and applicability of the model, enables timely detection of model anomalies and adjustments and optimizations, and ensures high-quality generation and decision-making of scientific and technological innovation big data models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020588A_ABST
    Figure CN122020588A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-dimensional analysis-based big data model generation method and system, and particularly relates to the field of big data model generation, and the method comprises a data batch division step, a data collection step, a data analysis step, a data comprehensive evaluation step and an early warning feedback step. The data acquisition step comprises a model input data acquisition unit and a model feedback data acquisition unit and is used for acquiring target data in real time and transmitting the acquired data to the data analysis step. The model input data acquisition unit is used for acquiring technical basic data, resource associated data and associated feature data; the feedback data acquisition unit is used for acquiring innovation theory result data, model stability data and model adaptability data; according to the method, the data generated by the large-model-based medical and invasive big data model is divided into different batches, so that the scientificity and the reliability of the source of the model generated data are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of science and technology innovation big data model generation technology, and more specifically, to a method and system for generating science and technology innovation big data models based on multidimensional analysis. Background Technology

[0002] With the vigorous development of technological innovation and the continuous advancement of digital transformation, the position of science and technology innovation in the global economy is becoming increasingly prominent, and the demand for the mining, analysis, and application of science and technology innovation data has reached an unprecedented level. The technology for generating science and technology innovation big data models based on multidimensional analysis, due to its key role in science and technology innovation decision-making and development, brings new challenges and opportunities to the field. Accurate big data model generation can ensure in-depth insights and effective utilization of science and technology innovation data, meeting the needs of high-quality development in science and technology innovation.

[0003] Existing methods for generating big data models for science and technology innovation mainly include a data acquisition module, a data preprocessing module, a model building module, and a model evaluation module. The data acquisition module deploys intelligent acquisition tools at key locations across various science and technology innovation-related data sources, such as research institution databases, innovative enterprise information systems, and patent platforms, to collect data including research project information, technological innovation achievements, market trend data, and talent mobility, thereby achieving the collection of science and technology innovation data. The data preprocessing module uses data cleaning algorithms to denoise and standardize the collected data, ensuring its quality and usability before entering the model building stage. The model building module, based on specific algorithms and architectures, combines the preprocessed data to construct a big data model for science and technology innovation, uncovering the inherent correlations and potential patterns between data points, providing model support for strategic planning for science and technology innovation development. The model evaluation module allows professionals to intuitively understand the model's performance, adjust and optimize model parameters, promptly correct the model structure, and improve the accuracy and stability of the big data model for science and technology innovation.

[0004] However, this method still has some shortcomings in practical applications. For example, in terms of data collection, due to the diversity and complexity of data sources and the compatibility issues of collection tools, some scientific and technological innovation data may be missed, resulting in incomplete data collection, which will adversely affect subsequent model construction. In terms of model evaluation, facing the complex scientific and technological innovation environment and diverse evaluation indicators, the evaluation results are prone to being inaccurate and incomplete. This may lead to the inability to discover potential defects in the model in a timely manner, reducing the ability of the scientific and technological innovation big data model to guide scientific and technological innovation practices.

[0005] Therefore, there is an urgent need to provide a method and system for generating big data models for science and technology innovation based on multidimensional analysis, in order to solve the problems of insufficient data collection and inaccurate evaluation results in existing methods and systems for generating big data models for science and technology innovation, and to further improve the quality and reliability of science and technology innovation activities. Summary of the Invention

[0006] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a method and system for generating science and technology innovation big data models based on multidimensional analysis, which solves the problems mentioned in the background art through the following solutions.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for generating a science and technology innovation big data model based on multidimensional analysis, comprising: S1. Data batch division steps: This step is used to determine the data to be collected as the target data, divide the target data into different batches according to the equal time division method, and label them as 1, 2, ..., n in sequence; S2. Data Acquisition Steps: This includes a model input data acquisition unit and a model feedback data acquisition unit, used to acquire target data in real time and transmit the acquired data to the data analysis steps. The model input data acquisition unit is used to acquire basic technical data, resource-related data, and related feature data; the feedback data acquisition unit is used to acquire data on innovative theoretical achievements, model stability data, and model adaptability data. S3. Data Analysis Steps: This includes a model input data analysis unit and a model feedback data analysis unit, used to collect target data in real time and transmit the collected data to the data comprehensive evaluation step. The model input data analysis unit includes a technical foundation data analysis node, a resource association data analysis node, and an association feature data analysis node; the feedback data analysis unit includes an innovative theoretical achievement data analysis node, a model stability data analysis node, and a model adaptability data analysis node. S4. Data Comprehensive Evaluation Step: This includes a data analysis unit generated by the science and technology innovation big data model, which is used to comprehensively analyze the analysis results transmitted from the data analysis step and transmit the analysis results to the early warning feedback step. S5. Early Warning Feedback Steps: Used to establish a preset value for the comprehensive evaluation index generated by the science and technology innovation big data model, to judge the value of the comprehensive evaluation index generated by the science and technology innovation big data model based on the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model, and to issue corresponding signals based on the judgment results.

[0008] Preferably, the technical foundation data includes algorithm complexity Ac, data dimension Ds, and data sparsity Sd; the resource correlation data includes energy consumption rate Ec and data transmission bandwidth requirement Dt; the correlation feature data includes feature correlation coefficient Fc and feature entropy value Er; the innovative theoretical achievement data includes theoretical logical correlation strength St, theoretical novelty Nt, and theoretical extension dimension Te; the model stability data includes model convergence speed Ms and variance inflation factor Vf; and the model adaptability data includes new data fitness rate Ar and data scale scalability Dc.

[0009] Preferably, the technical foundation data analysis node is used to establish a technical foundation data calculation model, importing the technical foundation data transmitted in the data acquisition step into the technical foundation data calculation model to obtain the model input complexity coefficient value. The technical foundation data calculation model is specifically represented as follows: , in, Ac represents the model input complexity coefficient value for the i-th computation. i Ds represents the algorithm complexity of the i-th data collection. i Sd represents the dimension of the data collected in the i-th iteration. i This represents the sparsity of the data collected in the i-th iteration.

[0010] Preferably, the resource association data analysis node is used to establish a resource association data calculation model, import the resource association data transmitted in the data acquisition step into the resource association data calculation model, and obtain the resource suitability coefficient value. The resource association data calculation model is specifically represented as follows: , in, Ec represents the resource fit coefficient value calculated in the i-th iteration. i Dt represents the energy consumption rate of the i-th data collection. i Let σ represent the data transmission bandwidth requirement for the i-th data acquisition, σ represent the standard deviation of the data transmission bandwidth requirement, and μ represent the mean of the data transmission bandwidth requirement.

[0011] Preferably, the associated feature data analysis node is used to establish an associated feature data calculation model, importing the associated feature data transmitted in the data acquisition step into the associated feature data calculation model to obtain the feature validity coefficient value. The associated feature data calculation model is specifically represented as follows: , in, Fc represents the feature validity coefficient value calculated in the i-th iteration. i Er represents the feature correlation coefficient of the i-th acquisition. iThis represents the feature entropy value of the i-th acquisition.

[0012] Preferably, the innovation theory achievement data analysis node is used to establish an innovation theory achievement data calculation model, importing the innovation theory achievement data transmitted in the data acquisition step into the innovation theory achievement data calculation model to obtain the model's innovation coefficient value. The innovation theory achievement data calculation model is specifically represented as follows: , in, St represents the model innovation coefficient value calculated in the i-th iteration. i Nt represents the theoretical logical correlation strength of the i-th data collection. i Te represents the theoretical novelty of the i-th acquisition. i This represents the theoretical extension dimension of the i-th data collection.

[0013] Preferably, the model stability data analysis node is used to establish a model stability data calculation model, importing the model stability data transmitted in the data acquisition step into the model stability data calculation model to obtain model stability coefficient values. The model stability data calculation model is specifically represented as follows: , in, Ms represents the model stability coefficient value calculated in the i-th iteration. i Vf represents the model convergence rate during the i-th acquisition. i Let represent the variance inflation factor of the i-th data collection.

[0014] Preferably, the model adaptability data analysis node is used to establish a model adaptability data calculation model, import the model adaptability data transmitted in the data acquisition step into the model adaptability data calculation model, and obtain the model adaptability coefficient value. The model adaptability data calculation model is specifically represented as follows: , in, Ar represents the model fitness coefficient value calculated in the i-th iteration. i Dc represents the fitness rate of the new data collected in the i-th iteration. i This indicates the scalability of the data size collected in the i-th iteration.

[0015] Preferably, the data analysis unit for generating the science and technology innovation big data model is used to establish a data calculation model for generating the science and technology innovation big data model. It imports the model input complexity coefficient, resource adaptability coefficient, feature effectiveness coefficient, model innovation coefficient, model stability coefficient, and model adaptability coefficient transmitted from the data analysis steps into the data calculation model for generating the science and technology innovation big data model, thereby obtaining a comprehensive evaluation index value for generating the science and technology innovation big data model. Specifically, the data calculation model for generating the science and technology innovation big data model is represented as follows: , Where A represents the comprehensive evaluation index value generated by the calculated science and technology innovation big data model. This represents the model input complexity coefficient value for the i-th computation. This represents the resource fit coefficient value calculated in the i-th iteration. This represents the feature validity coefficient value calculated in the i-th iteration. This represents the model innovation coefficient value calculated in the i-th iteration. This represents the minimum value of the calculated model's innovativeness coefficient. This represents the maximum value of the calculated model innovation coefficient. This represents the model stability coefficient value calculated in the i-th iteration. This represents the minimum calculated model stability coefficient. This represents the maximum calculated model stability coefficient. This represents the model fitness coefficient value calculated in the i-th iteration. This represents the minimum calculated model fitness coefficient. This represents the maximum value of the calculated model fitness coefficient, where i indicates starting from the i-th number and n indicates ending at the n-th number.

[0016] Preferably, a science and technology innovation big data model generation system based on multidimensional analysis includes: Data batch segmentation module: This module is used to identify the data to be collected as target data, divide the target data into different batches according to the equal time interval, and label them sequentially as 1, 2, ..., n; The data acquisition module includes a model input data acquisition unit and a model feedback data acquisition unit, used to acquire target data in real time and transmit the acquired data to the data analysis module. The model input data acquisition unit is used to collect basic technical data, resource-related data, and related feature data; the feedback data acquisition unit is used to collect data on innovative theoretical achievements, model stability data, and model adaptability data. The data analysis module includes a model input data analysis unit and a model feedback data analysis unit, used to collect target data in real time and transmit the collected data to the data comprehensive evaluation module. The model input data analysis unit includes a technical foundation data analysis node, a resource-related data analysis node, and a related feature data analysis node; the feedback data analysis unit includes an innovative theoretical achievement data analysis node, a model stability data analysis node, and a model adaptability data analysis node. Data comprehensive evaluation module: including the data analysis unit generated by the science and technology innovation big data model, which is used to comprehensively analyze the analysis results transmitted by the data analysis module and transmit the analysis results to the early warning feedback module; Early warning feedback module: used to establish a preset value for the comprehensive evaluation index generated by the science and technology innovation big data model, judge the value of the comprehensive evaluation index generated by the science and technology innovation big data model based on the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model, and issue corresponding signals based on the judgment results.

[0017] The technical effects and advantages of this invention are as follows: This invention effectively improves the scientific rigor and reliability of the data sources generated by the science and technology innovation big data model by dividing the data into batches based on multidimensional analysis. By collecting technical foundation data, resource-related data, related feature data, innovative theoretical achievement data, model stability data, and model adaptability data from multiple dimensions, this multi-faceted data collection mode greatly overcomes the limitations of traditional, one-sided model data sources, providing rich and authoritative information support for subsequent model construction. This invention performs in-depth data mining through data analysis, and then calculates the model input complexity coefficient, resource suitability coefficient, feature effectiveness coefficient, model innovation coefficient, model stability coefficient, and model adaptability coefficient, accurately identifying factors that may affect model quality. This invention significantly enhances the accuracy and applicability of the model by comprehensively considering and deeply integrating the data analysis results. Through dynamic monitoring and continuous tracking, it immediately issues alerts upon detecting model anomalies. Relevant researchers and decision-makers can quickly understand the status and potential problems of the science and technology innovation big data model, and take timely action to adjust and optimize the model parameters and architecture. This provides a solid guarantee for achieving high-quality generation of science and technology innovation big data models based on large models and for efficient and accurate decision-making. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the method structure of the present invention.

[0019] Figure 2 This is a schematic diagram of the method structure of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] As attached Figure 1 The method for generating a science and technology innovation big data model based on multidimensional analysis is shown, which specifically includes: data batch division step, data collection step, data analysis step, data comprehensive evaluation step, and early warning feedback step.

[0022] S1. Data batch division steps: This step is used to determine the data to be collected as the target data, divide the target data into different batches according to the equal time division method, and mark them as 1, 2, ..., n in sequence.

[0023] In this embodiment, it should be specifically explained that: the method of dividing according to equal time is based on the characteristics of each target data. For the real-time requirements of each type of target data, the database is run by a method to collect data once at equal intervals and divide it into different collection times.

[0024] S2. Data Acquisition Steps: This includes a model input data acquisition unit and a model feedback data acquisition unit, used to acquire target data in real time and transmit the acquired data to the data analysis steps. The model input data acquisition unit is used to acquire basic technical data, resource-related data, and related feature data; the feedback data acquisition unit is used to acquire data on innovative theoretical achievements, model stability data, and model adaptability data.

[0025] In this embodiment, it should be specifically noted that: the technical foundation data includes algorithm complexity Ac, data dimension Ds, and data sparsity Sd; the resource association data includes energy consumption rate Ec and data transmission bandwidth requirement Dt; the association feature data includes feature correlation coefficient Fc and feature entropy value Er; the innovative theoretical achievement data includes theoretical logical association strength St, theoretical novelty Nt, and theoretical extension dimension Te; the model stability data includes model convergence speed Ms and variance inflation factor Vf; and the model adaptability data includes new data fitness rate Ar and data scale scalability Dc.

[0026] The algorithm complexity accurately reflects the growth trend of computational load under different input scales. It is crucial for resource planning and time estimation. When optimizing model training and inference processes, the complexity can guide the selection of more efficient algorithm improvement directions, such as adopting parallel computing strategies or optimizing internal algorithm steps. Higher complexity may require more powerful computing hardware or a more reasonable task allocation mechanism. The data collection method involves starting with the algorithm's code logic and using mathematical analysis to determine the relationship between the number of executions of its basic operations and the input scale. For loop structures, the dependency between the number of loops and the input scale is analyzed; for recursive algorithms, recursive equations are established to solve for the complexity. Simultaneously, the algorithm can be run on simulated datasets of different scales, and the computation time recorded. Data analysis of time and input scale can then be used to verify the theoretical complexity.

[0027] The dimensionality of the data directly determines the number of model parameters and the computational cost. High-dimensional data may require more complex model structures to fully extract information, but it may also lead to overfitting. It affects all aspects of model design, from the number of layers and nodes per layer in neural networks to the feature combination methods in traditional machine learning models. Furthermore, dimensionality is also related to data storage and processing methods. High-dimensional data may require special dimensionality reduction or feature selection methods. The acquisition method is as follows: if the data is in its raw numerical form, the dimensionality can be determined by examining the number of features at each data point. For image data, its dimension may include the number of pixel rows, columns, and color channels. For text data, after vectorization, such as the vectors generated by word vector models, its dimension is the length of the vector. Data processing libraries and tools can be used to automatically obtain this information.

[0028] Data sparsity implies the existence of a large amount of redundant information. Specialized sparse data structures and algorithms can be used during storage and computation to save space and improve efficiency. For example, in matrix operations, sparse matrix multiplication algorithms can significantly reduce unnecessary computations. It also influences the choice of feature selection methods, as sparse features may contribute little to the model. The collection method involves traversing the entire dataset, counting the number of elements with a statistical value of zero or less than a specific threshold (determined based on data characteristics), and then dividing by the total number of elements in the dataset to obtain the data sparsity. For data stored in matrix form, the characteristics of the matrix storage format (such as CSR, CSC, and other sparse matrix storage formats) can be utilized to calculate sparsity more efficiently.

[0029] The energy consumption rate provides insight into the model's energy consumption, which helps optimize model algorithms and hardware selection. Reducing the energy consumption rate can lower operating costs, especially when running a large number of model instances in large-scale data centers. It also provides a basis for researching the application of new energy-efficient computing technologies in models. The data collection method involves using specialized hardware power monitoring equipment, such as a power meter, connected to the power input line of the computing device running the model (e.g., server, GPU cluster). During model operation, the total energy consumption (joules) of the computing device and the number of operations performed by the model are recorded simultaneously (this can be achieved by embedding a counter in the code). The energy consumption rate is obtained by dividing the total energy consumption by the number of operations. To improve accuracy, multiple measurements can be taken and the average value calculated. The data transmission bandwidth requirements can optimize the network architecture and storage system configuration. Insufficient bandwidth may lead to data transmission delays, affecting the model's real-time performance or training efficiency. For distributed training models, reasonable bandwidth planning can ensure timely data synchronization between different computing nodes. The data collection method is as follows: during model runtime, use network monitoring tools (such as network analyzers) or the operating system's built-in network performance monitoring function to monitor the data transmission speed of the network interface during the data input and output phases. For data transmission to storage devices, storage performance testing tools can be used. Multiple measurements should be performed under different data volume input and output scenarios, and the peak value should be used as a reference value for the data transmission bandwidth requirements. Simultaneously, considering sudden traffic spikes in the network and storage, a certain margin can be appropriately added.

[0030] The feature correlation coefficient reveals the linear relationship between different features. Highly correlated features may lead to multicollinearity, affecting the stability and interpretability of the model. By analyzing correlations, feature selection or combination can be performed to improve model performance. In regression models, it helps to understand the interaction between independent variables. The data collection method is as follows: for numerical features, calculate the Pearson correlation coefficient or Spearman correlation coefficient between features (for nonlinear relationships). The correlation coefficient is calculated using statistical formulas by traversing all feature pairs. For large-scale datasets, distributed computing frameworks or specialized data analysis software can be used to improve computational efficiency.

[0031] The feature entropy value measures the uncertainty of a feature's value. High-entropy features contain more information and may be more valuable for model training, but they can also increase model complexity. It is instructive in feature selection and data preprocessing; for example, it can be used to determine whether further processing is needed to reduce uncertainty. The acquisition method is as follows: for discrete features, the information entropy is calculated based on the probability distribution of their values. For continuous features, discretization can be performed before calculating the entropy value, or a continuous entropy calculation method based on the probability density function can be used.

[0032] A strong logical connection between theories signifies high scientific validity and reliability. It ensures that the theory avoids self-contradiction during derivation and application, forming the foundation for its widespread acceptance and application. For example, in mathematical theory, rigorous logical connections are crucial for proving the correctness of theorems. Euclidean geometry, with its axiomatic method establishing strong logical connections, has allowed the entire geometric system to develop steadily. In scientific research and education, theories with strong logical connections are easier to understand and disseminate, helping researchers and students better grasp the theory's meaning and application, improving the efficiency of knowledge transfer. The data collection method involves decomposing the theory to be evaluated into multiple basic propositions, concepts, and reasoning steps to construct a logical structure model. This model can be represented as a directed graph, where nodes represent elements in the theory and edges represent the logical relationships between them. Graph theory algorithms, such as computing graph connectivity and shortest paths, are used to analyze the tightness of the logical structure. For example, by calculating the shortest reasoning path length between any two key propositions, the shorter the path, the more direct and tight the logical connection. Consider the type and weight of logical relations, such as causality and implication, which have different levels of importance in the theory. Based on logical norms and expert opinions within the field, assign appropriate weights to different types of logical relations to further refine the calculation of logical affinity strength.

[0033] The theoretical novelty described here refers to the ability of highly novel theories to open up entirely new research paths and guide researchers to break free from the constraints of traditional thinking. It is a key indicator for measuring the level of scientific innovation, helping research institutions and funders identify theoretical directions with potentially significant impact, allocate resources rationally, and promote the rapid development of theoretical science. For example, the theory of relativity exemplifies extremely high theoretical novelty; it fundamentally changed people's understanding of spacetime and triggered a series of research breakthroughs in related fields. The data collection method involves: First, constructing a large text database containing classic theories, cutting-edge theories, and newly proposed theories within the field. This database needs to cover theoretical content from different periods and schools of thought, ensuring the accuracy and completeness of the text. Then, using advanced natural language processing techniques, such as word vector models and semantic analysis algorithms, the new theory and other theories in the database are segmented and vectorized, transforming the text into a computer-understandable mathematical form. Finally, the distance between the new theory and other theories in the semantic space is calculated, for example, using cosine similarity. The greater the distance, the higher the novelty. Meanwhile, by combining the weighting rules formulated by domain experts and considering the impact of unique concepts and novel perspectives in the theory on novelty, a comprehensive theoretical novelty index is derived.

[0034] The aforementioned theoretical expansion dimension reflects the theory's development potential and vitality. Theories with a high expansion dimension coefficient can be integrated with other theories in multiple directions or applied in different fields, thereby generating more scientific research results. For example, game theory was initially applied to the field of economics, but due to its high expansion dimension, it has later been widely applied in multiple fields such as biology and computer science, promoting the development of interdisciplinary research. The data collection method involves: conducting comprehensive knowledge mining of the theory, analyzing its basic principles, assumptions, mathematical models, and other core contents, and determining the essential characteristics of the theory.

[0035] The simulation and evaluation process considers multiple dimensions, including interdisciplinary applications, the potential for integration with other existing theories, and the ability to solve different types of problems. For example, in terms of interdisciplinary applications, the applicability of the theory in other disciplines is examined by analyzing similar problems and theoretical structures in other disciplines to identify possible points of convergence. Regarding integration with other theories, the compatibility and complementarity of the theory with other related theories in terms of concepts and methods are studied. Based on the evaluation results of each dimension, and combined with the development trends and needs within the field, appropriate weights are assigned to each dimension, and a comprehensive calculation yields the theoretical expansion dimension coefficient.

[0036] The convergence speed of the model reflects how quickly the model reaches a stable state during training. Fast convergence reduces training time and resource consumption, improving training efficiency. Simultaneously, convergence speed is related to model stability; convergence that is too fast or too slow may indicate overfitting, underfitting, or other problems, helping to adjust the training strategy. The data collection method is as follows: during model training, record the loss function value (or other suitable evaluation metric) for each iteration. Observe the trend of the loss function value; when the change is less than a preset threshold (e.g., 0.001) or remains essentially unchanged for multiple consecutive iterations (e.g., 10 times), record the number of iterations at this point as a measure of convergence speed. Different models can determine their own convergence criteria based on their characteristics.

[0037] The variance inflation factor (VIF) is used to detect multicollinearity among independent variables in a model. A high VIF value (typically greater than 5 or 10) indicates a strong linear relationship among the independent variables, which increases the variance of the model parameter estimates, leading to model instability and inaccurate parameter estimation. By monitoring the VIF value, the model can be adjusted, such as by removing collinear variables or using regularization methods. The data collection method is as follows: For linear models (such as linear regression), the regression coefficient estimates of the independent variables are obtained after fitting the model. The correlation coefficient matrix among the independent variables is calculated, and the VIF value of each independent variable is calculated. For nonlinear models, linearization approximation or other collinearity detection methods suitable for nonlinear situations can be used.

[0038] The new data fitness rate describes the model's ability to adapt to new data as it dynamically changes. It measures the model's capacity to adapt to new input data types or distributions. A high fitness rate means the model can maintain performance on new data, reducing the need for retraining and is crucial for long-term, dynamic applications. The data collection method involves gathering new data that differs from the training data (in distribution, feature type, etc.). This new data is then input into the model for prediction. The number of correctly processed new data samples (based on specific evaluation criteria, such as accuracy within a certain range) is counted, divided by the total number of new data samples, and multiplied by 100% to obtain the fitness rate. The selected new data should represent the new situations encountered in real-world applications.

[0039] Data scalability describes the continuous increase in data volume in big data scenarios. This metric reflects the performance trend of a model as the data volume grows. Understanding scalability helps in planning computing resources and model optimization strategies, ensuring the model remains effective when handling large-scale data. The data collection method is as follows: gradually increase the input data volume (each time by a certain percentage, such as 10%). Run the model at each data volume level and record performance metrics (such as runtime and accuracy). Calculate the rate of change of performance metrics (such as accuracy rate of change and runtime rate of change) at each data volume increase stage, divided by the percentage increase in data volume, to obtain the data scalability metric. A data volume-performance metric change curve can be plotted to observe scalability.

[0040] S3. Data Analysis Steps: This includes a model input data analysis unit and a model feedback data analysis unit, used to collect target data in real time and transmit the collected data to the data comprehensive evaluation step. The model input data analysis unit includes a technical foundation data analysis node, a resource association data analysis node, and an association feature data analysis node; the feedback data analysis unit includes an innovative theoretical achievement data analysis node, a model stability data analysis node, and a model adaptability data analysis node.

[0041] In this embodiment, it should be specifically noted that: the technical basic data analysis node is used to establish a technical basic data calculation model, importing the technical basic data transmitted in the data acquisition step into the technical basic data calculation model to obtain the model input complexity coefficient value. The technical basic data calculation model is specifically represented as follows: , in, Ac represents the model input complexity coefficient value for the i-th computation. i Ds represents the algorithm complexity of the i-th data collection. i Sd represents the dimension of the data collected in the i-th iteration. i This represents the sparsity of the data collected in the i-th iteration.

[0042] In this embodiment, it should be specifically noted that: the resource association data analysis node is used to establish a resource association data calculation model, import the resource association data transmitted in the data acquisition step into the resource association data calculation model, and obtain the resource adaptability coefficient value. The resource association data calculation model is specifically represented as follows: , in, Ec represents the resource fit coefficient value calculated in the i-th iteration. i Dt represents the energy consumption rate of the i-th data collection. i Let σ represent the data transmission bandwidth requirement for the i-th data acquisition, σ represent the standard deviation of the data transmission bandwidth requirement, and μ represent the mean of the data transmission bandwidth requirement.

[0043] In this embodiment, it should be specifically noted that: the associated feature data analysis node is used to establish an associated feature data calculation model, importing the associated feature data transmitted in the data acquisition step into the associated feature data calculation model to obtain the feature validity coefficient value. The associated feature data calculation model is specifically represented as follows: , in, Fc represents the feature validity coefficient value calculated in the i-th iteration. i Er represents the feature correlation coefficient of the i-th acquisition. i This represents the feature entropy value of the i-th acquisition.

[0044] In this embodiment, it should be specifically noted that: the innovation theory achievement data analysis node is used to establish an innovation theory achievement data calculation model, importing the innovation theory achievement data transmitted in the data acquisition step into the innovation theory achievement data calculation model to obtain the model's innovation coefficient value. The innovation theory achievement data calculation model is specifically represented as follows: , in, St represents the model innovation coefficient value calculated in the i-th iteration. i Nt represents the theoretical logical correlation strength of the i-th data collection. i Te represents the theoretical novelty of the i-th acquisition. i This represents the theoretical extension dimension of the i-th data collection.

[0045] In this embodiment, it should be specifically noted that: the model stability data analysis node is used to establish a model stability data calculation model, import the model stability data transmitted in the data acquisition step into the model stability data calculation model, and obtain the model stability coefficient value. The model stability data calculation model is specifically represented as follows: , in, Ms represents the model stability coefficient value calculated in the i-th iteration. i Vf represents the model convergence rate during the i-th acquisition. i Let represent the variance inflation factor of the i-th data collection.

[0046] In this embodiment, it should be specifically noted that: the model adaptability data analysis node is used to establish a model adaptability data calculation model, import the model adaptability data transmitted in the data acquisition step into the model adaptability data calculation model, and obtain the model adaptability coefficient value. The model adaptability data calculation model is specifically represented as follows: , in, Ar represents the model fitness coefficient value calculated in the i-th iteration. i Dc represents the fitness rate of the new data collected in the i-th iteration. i This indicates the scalability of the data size collected in the i-th iteration.

[0047] S4. Data Comprehensive Evaluation Step: This includes a data analysis unit generated by the science and technology innovation big data model, which is used to comprehensively analyze the analysis results transmitted from the data analysis step and transmit the analysis results to the early warning feedback step.

[0048] In this embodiment, it should be specifically explained that: the data analysis unit for generating the science and technology innovation big data model is used to establish a data calculation model for generating the science and technology innovation big data model. It imports the model input complexity coefficient, resource adaptability coefficient, feature effectiveness coefficient, model innovation coefficient, model stability coefficient, and model adaptability coefficient transmitted from the data analysis steps into the data calculation model for generating the science and technology innovation big data model, thereby obtaining a comprehensive evaluation index value for generating the science and technology innovation big data model. The data calculation model for generating the science and technology innovation big data model is specifically represented as follows: , Where A represents the comprehensive evaluation index value generated by the calculated science and technology innovation big data model. This represents the model input complexity coefficient value for the i-th computation. This represents the resource fit coefficient value calculated in the i-th iteration. This represents the feature validity coefficient value calculated in the i-th iteration. This represents the model innovation coefficient value calculated in the i-th iteration. This represents the minimum value of the calculated model's innovativeness coefficient. This represents the maximum value of the calculated model innovation coefficient. This represents the model stability coefficient value calculated in the i-th iteration. This represents the minimum calculated model stability coefficient. This represents the maximum calculated model stability coefficient. represents the value of the model adaptation ability coefficient for the i-th calculation, represents the minimum value of the calculated model adaptation ability coefficient, represents the maximum value of the calculated model adaptation ability coefficient. i represents starting from the i-th number, and n represents ending at the n-th number.

[0049] S5. Early warning feedback step: used to establish a preset value of the comprehensive evaluation index generated by the science and technology innovation big data model, judge the value of the comprehensive evaluation index generated by the science and technology innovation big data model according to the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model, and send corresponding signals according to the judgment results.

[0050] In this embodiment, specifically, it should be noted that: the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model is marked as A def , when A def = <A, a normal signal is sent. This signal indicates that the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model is less than or equal to the value of the comprehensive evaluation index generated by the science and technology innovation big data model, indicating that the generation state of the science and technology innovation big data model is normal; when A def > A, an early warning signal is sent to relevant technical management personnel. This signal indicates that the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model is greater than the value of the comprehensive evaluation index generated by the science and technology innovation big data model, indicating that the generation state of the science and technology innovation big data model is poor.

[0051] Based on the above solution and appendix Figure 2 , this application embodiment also discloses a science and technology innovation big data model generation system based on multi-dimensional analysis, including: Data batch division module: used to determine the to-be-collected data as target data, divide the target data into different batches in an equal-time division manner, and sequentially mark them as 1, 2,..., n; Data collection module: includes a model input data collection unit and a model feedback data collection unit, used to collect the target data in real time and transmit the collected data to the data analysis module. The model input data collection unit is used to collect technical basic data, resource association data, and associated feature data; the feedback data collection unit is used to collect innovation theory achievement data, model stability data, and model adaptability data; Data analysis module: includes a model input data analysis unit and a model feedback data analysis unit, used to collect the target data in real time and transmit the collected data to the data comprehensive evaluation module. The model input data analysis unit includes a technical basic data analysis node, a resource association data analysis node, and an associated feature data analysis node; the feedback data analysis unit includes an innovation theory achievement data analysis node, a model stability data analysis node, and a model adaptability data analysis node; Data comprehensive evaluation module: including the data analysis unit generated by the science and technology innovation big data model, which is used to comprehensively analyze the analysis results transmitted by the data analysis module and transmit the analysis results to the early warning feedback module; Early warning feedback module: used to establish a preset value for the comprehensive evaluation index generated by the science and technology innovation big data model, judge the value of the comprehensive evaluation index generated by the science and technology innovation big data model based on the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model, and issue corresponding signals based on the judgment results.

[0052] This invention effectively improves the scientific rigor and reliability of data sources for generating science and technology innovation big data models by dividing the generated data into batches. By collecting data from multiple dimensions—including technical foundation data, resource-related data, related feature data, innovative theoretical achievements data, model stability data, and model adaptability data—this multi-dimensional data collection model significantly overcomes the limitations of traditional, one-sided data sources, providing rich and authoritative information support for subsequent model construction. Through in-depth data analysis, the collected data is mined to calculate coefficients for model input complexity, resource suitability, feature effectiveness, model innovation, model stability, and model adaptability, accurately identifying factors that may affect model quality. Comprehensive consideration and deep integration of the data analysis results significantly enhance the model's accuracy and applicability. Dynamic monitoring and continuous tracking ensure that any model anomalies are immediately detected, allowing relevant researchers and decision-makers to quickly understand the status and potential problems of the science and technology innovation big data model and take timely action to adjust and optimize model parameters and architecture. This provides a solid guarantee for achieving high-quality generation of science and technology innovation big data models based on large models and for efficient and accurate decision-making.

[0053] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other. In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating a science and technology innovation big data model based on multidimensional analysis, characterized in that, include: S1. Data batch division steps: This step is used to determine the data to be collected as the target data, divide the target data into different batches according to the equal time division method, and label them as 1, 2, ..., n in sequence; S2. Data Acquisition Step: This step includes a model input data acquisition unit and a model feedback data acquisition unit, used to acquire target data in real time and transmit the acquired data to the data analysis step. The model input data acquisition unit is used to acquire basic technical data, resource-related data, and related feature data. The feedback data acquisition unit is used to collect data on innovative theoretical achievements, model stability, and model adaptability. S3. Data Analysis Step: This step includes a model input data analysis unit and a model feedback data analysis unit, used to collect target data in real time and transmit the collected data to the data comprehensive evaluation step. The model input data analysis unit includes a technical foundation data analysis node, a resource-related data analysis node, and a related feature data analysis node. The feedback data analysis unit includes an innovative theoretical achievement data analysis node, a model stability data analysis node, and a model adaptability data analysis node; S4. Data Comprehensive Evaluation Step: This includes a data analysis unit generated by the science and technology innovation big data model, which is used to comprehensively analyze the analysis results transmitted from the data analysis step and transmit the analysis results to the early warning feedback step. S5. Early Warning Feedback Steps: Used to establish a preset value for the comprehensive evaluation index generated by the science and technology innovation big data model, to judge the value of the comprehensive evaluation index generated by the science and technology innovation big data model based on the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model, and to issue corresponding signals based on the judgment results.

2. The method for generating a science and technology innovation big data model based on multidimensional analysis according to claim 1, characterized in that: The technical foundation data includes algorithm complexity Ac, data dimension Ds, and data sparsity Sd; the resource correlation data includes energy consumption rate Ec and data transmission bandwidth requirement Dt; the correlation feature data includes feature correlation coefficient Fc and feature entropy value Er; the innovative theoretical achievement data includes theoretical logical correlation strength St, theoretical novelty Nt, and theoretical extension dimension Te; the model stability data includes model convergence speed Ms and variance inflation factor Vf; and the model adaptability data includes new data fitness rate Ar and data scale scalability Dc.

3. The method for generating a science and technology innovation big data model based on multidimensional analysis according to claim 1, characterized in that: The technical foundation data analysis node is used to establish a technical foundation data calculation model. It imports the technical foundation data transmitted during the data acquisition steps into the technical foundation data calculation model to obtain the model input complexity coefficient value. The technical foundation data calculation model is specifically represented as follows: , in, Ac represents the model input complexity coefficient value for the i-th computation. i Let Ds represent the algorithm complexity of the i-th data collection. i Sd represents the dimension of the data collected in the i-th iteration. i This represents the sparsity of the data collected in the i-th iteration.

4. The method for generating a science and technology innovation big data model based on multidimensional analysis according to claim 1, characterized in that: The resource association data analysis node is used to establish a resource association data calculation model. It imports the resource association data transmitted during the data acquisition step into the resource association data calculation model to obtain the resource suitability coefficient value. The resource association data calculation model is specifically represented as follows: , in, Ec represents the resource fit coefficient value calculated in the i-th iteration. i Dt represents the energy consumption rate of the i-th data collection. i Let σ represent the data transmission bandwidth requirement for the i-th data acquisition, σ represent the standard deviation of the data transmission bandwidth requirement, and μ represent the mean of the data transmission bandwidth requirement.

5. The method for generating a science and technology innovation big data model based on multidimensional analysis according to claim 1, characterized in that: The associated feature data analysis node is used to establish an associated feature data calculation model. It imports the associated feature data transmitted during the data acquisition step into the associated feature data calculation model to obtain feature validity coefficient values. The associated feature data calculation model is specifically represented as follows: , in, Fc represents the feature validity coefficient value calculated in the i-th iteration. i Er represents the feature correlation coefficient of the i-th acquisition. i This represents the feature entropy value of the i-th acquisition.

6. The method for generating a science and technology innovation big data model based on multidimensional analysis according to claim 1, characterized in that: The innovation theory achievement data analysis node is used to establish an innovation theory achievement data calculation model. It imports the innovation theory achievement data transmitted during the data acquisition step into the innovation theory achievement data calculation model to obtain the model's innovation coefficient value. The innovation theory achievement data calculation model is specifically represented as follows: , in, St represents the model innovation coefficient value calculated in the i-th iteration. i Nt represents the theoretical logical correlation strength of the i-th data collection. i Te represents the theoretical novelty of the i-th acquisition. i This represents the theoretical extension dimension of the i-th data collection.

7. The method for generating a science and technology innovation big data model based on multidimensional analysis according to claim 1, characterized in that: The model stability data analysis node is used to establish a model stability data calculation model. It imports the model stability data transmitted during the data acquisition step into the model stability data calculation model to obtain the model stability coefficient values. The model stability data calculation model is specifically represented as follows: , in, Ms represents the model stability coefficient value calculated in the i-th iteration. i Vf represents the model convergence rate during the i-th acquisition. i Let represent the variance inflation factor of the i-th data collection.

8. The method for generating a science and technology innovation big data model based on multidimensional analysis according to claim 1, characterized in that: The model adaptability data analysis node is used to establish a model adaptability data calculation model. It imports the model adaptability data transmitted during the data acquisition step into the model adaptability data calculation model to obtain the model adaptability coefficient value. The model adaptability data calculation model is specifically represented as follows: , in, Ar represents the model fitness coefficient value calculated in the i-th iteration. i Dc represents the fitness rate of the new data collected in the i-th iteration. i This indicates the scalability of the data size collected in the i-th iteration.

9. The method for generating a science and technology innovation big data model based on multidimensional analysis according to claim 1, characterized in that: The data analysis unit for generating the science and technology innovation big data model is used to establish a data calculation model for generating the science and technology innovation big data model. It imports the model input complexity coefficient, resource adaptability coefficient, feature effectiveness coefficient, model innovation coefficient, model stability coefficient, and model adaptability coefficient transmitted from the data analysis steps into the data calculation model for generating the science and technology innovation big data model, thereby obtaining a comprehensive evaluation index value for generating the science and technology innovation big data model. The specific representation of the data calculation model for generating the science and technology innovation big data model is as follows: , Where A represents the comprehensive evaluation index value generated by the calculated science and technology innovation big data model. This represents the model input complexity coefficient value for the i-th computation. This represents the resource fit coefficient value calculated in the i-th iteration. This represents the feature validity coefficient value calculated in the i-th iteration. This represents the model innovation coefficient value calculated in the i-th iteration. This represents the minimum value of the calculated model's innovativeness coefficient. This represents the maximum value of the calculated model innovation coefficient. This represents the model stability coefficient value calculated in the i-th iteration. This represents the minimum calculated model stability coefficient. This represents the maximum calculated model stability coefficient. This represents the model fitness coefficient value calculated in the i-th iteration. This represents the minimum calculated model fitness coefficient. This represents the maximum value of the calculated model fitness coefficient, where i indicates starting from the i-th number and n indicates ending at the n-th number.

10. A science and technology innovation big data model generation system based on multidimensional analysis, used to implement the science and technology innovation big data model generation method based on multidimensional analysis as described in any one of claims 1 and 9, characterized in that, include: Data batch segmentation module: This module is used to identify the data to be collected as target data, divide the target data into different batches according to the equal time interval, and label them sequentially as 1, 2, ..., n; The data acquisition module includes a model input data acquisition unit and a model feedback data acquisition unit, used to acquire target data in real time and transmit the acquired data to the data analysis module. The model input data acquisition unit is used to collect basic technical data, resource-related data, and related feature data. The feedback data acquisition unit is used to collect data on innovative theoretical achievements, model stability, and model adaptability. The data analysis module includes a model input data analysis unit and a model feedback data analysis unit, used to collect target data in real time and transmit the collected data to the data comprehensive evaluation module. The model input data analysis unit includes a technical foundation data analysis node, a resource-related data analysis node, and a related feature data analysis node. The feedback data analysis unit includes an innovative theoretical achievement data analysis node, a model stability data analysis node, and a model adaptability data analysis node; Data comprehensive evaluation module: including the data analysis unit generated by the science and technology innovation big data model, which is used to comprehensively analyze the analysis results transmitted by the data analysis module and transmit the analysis results to the early warning feedback module; Early warning feedback module: used to establish a preset value for the comprehensive evaluation index generated by the science and technology innovation big data model, judge the value of the comprehensive evaluation index generated by the science and technology innovation big data model based on the preset value of the comprehensive evaluation index generated by the science and technology innovation big data model, and issue corresponding signals based on the judgment results.