A Construction Method of an Online Industrial Analysis Big Data Platform
By building an online industry analysis big data platform, the problem of insufficient real-time and accuracy of the existing industry analysis platform is solved, and enterprise information and traffic data are used for data desensitization and timing prediction, combined with graph neural network analysis, structured data is generated and knowledge graphs are constructed, which realizes efficient multi-grained, cross-regional industrial analysis, and improves the accuracy of analysis and decision-making efficiency.
Patent Information
- Application Number
- CN202510483948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing industrial analysis platform is slow to respond, lacks real-time, and cannot quickly respond to changes in multiple heterogeneous data sources. The traditional personalized recommendation and prediction methods are insufficient in accuracy. The computing resources and model complexity bottlenecks of graph neural networks in industrial data analysis cannot provide multi-grained and cross-region comprehensive analysis, which limits the accurate judgment of the overall trend of the industry and regional competition.
By obtaining enterprise information and real-time traffic information, UGC data desensitization, building an online industry preprocessing data set, performing text semantic analysis and coordinate deviation correction, building a personalized analysis model, using graph neural network to perform enterprise association network analysis, determining model optimization hyperparameters, generating structured data, and building an analysis knowledge graph, and uploading it to cloud platform storage.
It improves data security, enhances the time and space perception ability of analysis, accurately captures the timing laws of enterprise activities, reduces the information island effect, enhances the depth and breadth of analysis, improves the accuracy and decision-making efficiency of industrial analysis, and realizes information sharing and intelligent decision-making support.
Smart Images

Figure CN120012140B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a method for constructing an online industrial analysis big data platform. Background Art
[0002] When processing dynamic industrial data, many platforms are slow to respond, lack real-time performance, and cannot quickly cope with changes in multiple heterogeneous data sources. Secondly, traditional personalized recommendation and prediction methods often cannot deeply explore the complex associations in the industrial chain, resulting in limited accuracy and effectiveness of the analysis results. At the same time, the application of graph neural networks in industrial data analysis faces bottlenecks in computing resources and model complexity, especially in the processing of spatio-temporal data with low efficiency. In addition, existing platforms usually rely on single-level analysis and cannot provide comprehensive analysis of multiple granularities and across regions, restricting the accurate judgment of the overall industrial trend and regional competition. These deficiencies urge the industrial analysis platform to urgently adopt more advanced technologies to improve the processing ability and analysis depth. Summary of the Invention
[0003] Based on this, it is necessary to provide a method for constructing an online industrial analysis big data platform to solve at least one of the above technical problems.
[0004] To achieve the above object, a method for constructing an online industrial analysis big data platform, the method includes the following steps:
[0005] Step S1: Obtain enterprise information and real-time traffic information; perform UGC data desensitization on the enterprise information and real-time traffic information, and construct an online industrial preprocessing data set;
[0006] Step S2: Generate text semantic parsing data according to the online industrial preprocessing data set; correct coordinate deviation according to the online industrial preprocessing data set to obtain coordinate-corrected data; perform industrial time-series prediction deduction on the text semantic parsing data and the coordinate-corrected data, and construct a personalized model to obtain an online industrial personalized preliminary analysis model;
[0007] Step S3: Output industrial personalized results through the online industrial personalized preliminary analysis model; perform enterprise association network analysis using the industrial personalized results to obtain industrial graph neural analysis data; determine industrial model optimization hyperparameters using the industrial graph neural analysis data; perform model hyperparameter iteration on the online industrial personalized preliminary analysis model based on the industrial model optimization hyperparameters to obtain an online industrial personalized analysis model;
[0008] Step S4: Generate industrial structured data according to the online industrial personalized analysis model; construct an analysis knowledge graph using the industrial structured data and upload it to the cloud platform for storage.
[0009] The beneficial effect of the present invention is that, by acquiring enterprise information and real-time traffic information and desensitizing UGC data, an online industry preprocessing data set is generated. At the data level, the risk of enterprise privacy leakage is reduced by data desensitization technology, while retaining the effective information of the data, which is helpful for subsequent analysis and modeling. Subsequently, text semantic analysis is used to generate industry-related semantic data, and coordinate deviation correction is performed based on real-time traffic data to improve the accuracy of geographic information. In the process of time series prediction and deduction, the industry personalized preliminary analysis model is further constructed by combining semantic analysis data and corrected geographic data. Through this process, the model has a stronger time and space perception ability and can accurately capture the time series law of enterprise activities. In the stage of industry personalized analysis, enterprise association network analysis is performed through industry personalized results, and graph neural network (GNN) is used to mine potential associations between enterprises. This method effectively reduces the information island effect in industrial network analysis, and enhances the depth and breadth of analysis through network structure optimization. Further using graph neural analysis data to determine the optimized hyperparameters of the industrial model, and implementing hyperparameter iterative optimization, not only improves the stability of the model, but also significantly enhances the accuracy and generalization ability of industrial analysis. Finally, the present invention generates structured data through an industry personalized analysis model to construct a complete analysis knowledge graph. After being stored on the cloud platform, the graph can provide real-time updated industry relationship information for upstream and downstream enterprises in the industrial chain, realizing information sharing and intelligent decision-making support. The dynamic update capability of the knowledge graph ensures the timeliness of the data and enhances the resilience of enterprises in market competition. Therefore, the present invention solves the problems of low data security, insufficient analysis accuracy, and poor model generalization ability in traditional industrial analysis through steps such as industrial data desensitization, time series deduction, graph neural network analysis, and hyperparameter optimization, thereby improving the accuracy of industrial analysis and decision-making efficiency.
[0010] Preferably, step S1 comprises the following steps:
[0011] Step S11: Obtaining basic enterprise information and real-time traffic information;
[0012] Step S12: screening the upstream and downstream relationships of the industrial chain for the basic information of the enterprise, and eliminating blank data to obtain the enterprise integrated data; generating commuting time circle analysis data for the real-time traffic information;
[0013] Step S13: Desensitize the UGC data based on the enterprise integration data and commuting time circle analysis data, and build an online industry preprocessing data set.
[0014] The present invention forms a basic data source by obtaining enterprise authoritative information and real-time traffic information. The enterprise authoritative information provides the operating conditions of enterprises, industrial chain relationships, and upstream and downstream collaboration data, ensuring the authenticity and integrity of the data; the real-time traffic information reflects traffic flow, travel patterns, and regional commuting characteristics, providing spatio-temporal dynamic information for subsequent regional industrial analysis. At the data level, by screening the upstream and downstream relationships of the industrial chain for enterprise authoritative information, redundant and invalid information is effectively removed, data noise is reduced, and at the same time, key data with high value for analyzing industrial chain relationships is retained. The screening process uses industrial chain relationship mapping and data screening algorithms to ensure the high accuracy and relevance of the enterprise authoritative integrated data by identifying transaction chains, supply chain networks, and collaboration relationships between enterprises. In the process of processing real-time traffic information, the commuting time circle analysis method is used to convert traffic data into associated data such as the flow of people and logistics between enterprises and the surrounding areas. By dividing the commuting circle and analyzing the traffic load and travel behavior characteristics at different time periods, commuting time circle analysis data is generated. This data not only reflects the traffic accessibility of the area where the enterprise is located but also reveals the dynamic changes in regional economic activities, providing chronological support for subsequent regional industrial vitality analysis. In the UGC data desensitization link, privacy protection of user-generated content (UGC) is carried out based on enterprise authoritative integrated data and commuting time circle analysis data. By applying technologies such as differential privacy, hash encryption, and data de-identification, the risk of data leakage is effectively avoided, while the statistical characteristics and behavior patterns of the data are retained. The desensitized data set retains core information such as enterprise location distribution, traffic association characteristics, and industrial chain relationships, while meeting the requirements of data security and compliance. The finally constructed online industrial preprocessing data set realizes the fusion and standardized processing of multi-source data, providing reliable basic data support for subsequent personalized analysis, industrial prediction, and optimization.
[0015] Preferably, the UGC data desensitization in step S1 includes:
[0016] Extracting enterprise user data according to the enterprise integrated data;
[0017] Extracting user trajectory information data according to the commuting time circle analysis data;
[0018] Generating and anonymizing the identifiers of enterprise user data, where the anonymization key range is 128-256 bits, to obtain online industrial identifier anonymized data;
[0019] Generalizing the regional scope of the user trajectory information data to obtain online regional generalized data;
[0020] Hashing the sensitive fields of the enterprise user data generation and the user trajectory information data to obtain online industrial sensitive field data;
[0021] Fuse multi-source data including anonymized online industry identifier data, generalized online geographical data, and sensitive field data of the online industry, and construct a fused dataset to obtain a preprocessed dataset for the online industry.
[0022] Through the two-way extraction of enterprise authoritative integration data and commuting time circle analysis data, the present invention forms enterprise user data and user trajectory information data. The enterprise user data includes the basic information, business status, and upstream and downstream relationships in the industrial chain of the enterprise, which helps to depict the industrial ecological network of the enterprise; the user trajectory information data is based on traffic travel characteristics, shows the dynamic activities of users within a specific area, and reflects the economic vitality and traffic flow distribution within the area. At the data level, the enterprise user data is protected through identifier anonymization technology, and a 128-256-bit anonymization key is used to ensure the security of the data after encryption. At the same time, the direct identity identifiers in the data are removed, reducing the risk of privacy leakage. The anonymized online industry identifier data still retains the association characteristics between enterprises, enabling subsequent industrial chain analysis and business relationship mining to continue. In the processing of user trajectory information data, the user location data is blurred through regional range generalization technology, converting the precise geographical coordinates into a location representation at the regional level to form generalized online geographical data. This technology effectively protects user privacy, avoids the exposure of precise locations, and still supports travel pattern analysis and traffic behavior research based on macro regions. In addition, sensitive fields (such as user identity information, enterprise-specific business data, etc.) in the enterprise user data and user trajectory information data are hashed to form sensitive field data of the online industry. Hashing converts sensitive data through an irreversible encryption algorithm to ensure that even in the case of data leakage, attackers cannot restore the original data. Finally, through the fusion of multi-source data including anonymized online industry identifier data, generalized online geographical data, and sensitive field data of the online industry, a unified fused dataset, namely the preprocessed dataset for the online industry, is formed. The multi-source data fusion uses feature alignment and data standardization technologies to solve the problems of data heterogeneity and inconsistent formats, ensuring that information from different data sources can be associated and analyzed in the same data space. The fused dataset provides a complete and secure data foundation for subsequent industrial analysis, trend prediction, and business strategy formulation, and at the same time achieves a good balance among data security, privacy protection, and data integrity.
[0023] Preferably, step S2 includes the following steps:
[0024] Step S21: Generate text semantic parsing data according to the preprocessed dataset for the online industry;
[0025] Step S22: Correct the coordinate deviation according to the preprocessed dataset for the online industry. When the GPS coordinate deviation is greater than or equal to 50m, perform POI space correction to obtain coordinate correction data;
[0026] Step S23: Extract the regional week-on-week traffic spatio-temporal features from the text semantic parsing data and the coordinate correction data, and construct the time series prediction and deduction data;
[0027] Step S24: Construct a personalized model based on the time series prediction and deduction data to obtain a preliminary online industry personalized analysis model.
[0028] The present invention gradually forms a preliminary online industry personalized analysis model by performing text semantic parsing, coordinate deviation correction, regional traffic spatio-temporal feature extraction, and time series prediction and deduction based on the online industry preprocessing data set. At the data level, through the text semantic parsing data generation step, the implicit relationships in text data such as enterprise information, user evaluations, and industrial policies are effectively mined. Through natural language processing technology, complex industry terms and semantic structures are parsed to extract entity relationships, sentiment tendencies, and topic information, forming text semantic parsing data with industrial attributes. This process ensures the structured transformation of text data, facilitating subsequent correlation analysis and model construction. In the coordinate deviation correction stage, the present invention adopts spatial correction technology for the errors in GPS coordinate data to ensure the accuracy of location data. When the GPS coordinate deviation is greater than or equal to 50m, through the POI (Point of Interest) spatial correction method, historical positioning data, map reference data, and traffic node information are compared to dynamically correct the offset position and generate high-precision coordinate correction data. This process improves the reliability of spatial data and helps to accurately locate the geographical locations of enterprises and users in subsequent spatio-temporal feature extraction. In the process of extracting regional week-on-week traffic spatio-temporal features, the text semantic parsing data and the coordinate correction data are used for spatial position correlation analysis and time series feature extraction. By performing week-on-week analysis on traffic flow, travel demand, and commuting patterns at different times within the region, traffic anomalies and commuting rules are identified, generating comprehensive spatio-temporal feature data. Combining time series modeling methods, time series prediction and deduction data are further constructed to effectively depict the traffic evolution trend and industrial activity changes, providing dynamic time series support for subsequent personalized analysis. Finally, a preliminary online industry personalized analysis model is constructed based on the time series prediction and deduction data. The model forms a personalized industrial analysis perspective through multi-dimensional feature embedding and dynamic parameter adjustment on the basis of making full use of spatial, temporal, and text data. The model can accurately deduce the industrial development trend of a specific enterprise or region, providing customized decision-making support for enterprises.
[0029] Preferably, step S2 includes the following steps:
[0030] Step S21: Generate text semantic parsing data according to the online industry preprocessing data set;
[0031] Step S22: Perform coordinate deviation correction based on the online industrial preprocessing dataset. When the GPS coordinate deviation is greater than or equal to 50 m, perform POI spatial correction to obtain coordinate correction data;
[0032] Step S23: Extract regional week-on-week traffic spatio-temporal features from the text semantic parsing data and the coordinate correction data, and construct time series prediction and deduction data;
[0033] Step S24: Build a personalized model based on the time series prediction and deduction data to obtain a preliminary online industrial personalized analysis model.
[0034] The present invention gradually forms a preliminary online industry personalized analysis model through text semantic parsing, coordinate deviation correction, traffic spatio-temporal feature extraction, and time series prediction and deduction based on an online industry preprocessed dataset. At the data level, in the process of generating text semantic parsing data, natural language processing technology is adopted to deeply analyze text information such as enterprise data, user evaluations, and industry news, and extract key information and semantic features related to the industry. Through technical means such as word segmentation, entity recognition, dependency analysis, and sentiment classification, text semantic parsing data with context relationships is generated. This process converts unstructured text data into structured information, providing a data basis for subsequent spatio-temporal feature extraction and prediction and deduction. In the coordinate deviation correction link, aiming at the positioning error problem of GPS data, the present invention sets a deviation threshold of 50m and introduces POI (Point of Interest) spatial correction technology. By comparing spatial feature information such as enterprise addresses, road nodes, and transportation hubs, coordinates with significant offsets are corrected to generate coordinate correction data. This process improves the accuracy of geographical location information and ensures the reliability of spatio-temporal feature extraction. At the same time, the accuracy optimization of the coordinate correction data provides a solid data basis for the analysis of regional traffic spatio-temporal features. In the process of extracting regional week-on-week traffic spatio-temporal features, the text semantic parsing data and the coordinate correction data are further integrated. By constructing a time series model, traffic flow, travel demand, and enterprise commuting data in a specific region at different times are analyzed. The week-on-week analysis method is used to identify short-term trends and periodic fluctuations in regional traffic changes, and extract abnormal patterns and key influencing factors in spatio-temporal features. In addition, combining historical traffic data and enterprise operation data, a time series prediction and deduction model is established to speculate on future traffic conditions and industrial activity trends. The deduction data not only provides quantitative information on enterprise development trends but also provides a reference for urban traffic planning and regional industrial policies. Finally, a preliminary online industry personalized analysis model is constructed based on the time series prediction and deduction data. Based on the characteristics of multi-source data, the model uses adaptive parameter optimization and multi-dimensional feature fusion technology to form personalized analysis results for different industrial scenarios. The model can dynamically adjust the analysis strategy according to the actual situations of different regions and enterprises, and realize refined industrial trend prediction and business decision support.
[0035] Preferably, the construction of the personalized model described in step S2 includes:
[0036] Setting personalized target indexes according to the time series prediction and deduction data, where the personalized target indexes include a macroscopic congestion index and a passenger source distribution index;
[0037] Setting initial parameters of the model according to the time series prediction and deduction data;
[0038] Use a preset spatio-temporal graph convolutional network to perform regional-industry correlation analysis on the macro congestion index, and use the passenger source distribution index to analyze the passenger source POI trajectory to obtain the initial model construction data;
[0039] Construct the model framework with the initial model construction data and assign a loss function to obtain a preliminary online industrial personalization analysis model.
[0040] The present invention sets personalized target indexes, initial parameters, conducts regional-industry correlation analysis and model construction through time series prediction and deduction data, and forms a preliminary online industrial personalization analysis model. At the data level, first, time series features such as traffic flow, travel patterns, and business district activity are extracted based on time series prediction and deduction data, and personalized target indexes such as the macro congestion index and the passenger source distribution index are set. The macro congestion index quantifies the regional traffic pressure by analyzing the traffic congestion conditions in different regions and reflects the load level of the urban traffic network; the passenger source distribution index, based on user travel trajectories and consumption behavior data, depicts the passenger source flow characteristics of the target region and provides data support for commercial site selection, customer group positioning, and market demand prediction. In the initial model parameter setting stage, the trend characteristics of time series prediction and deduction data are used to dynamically assign values to the model parameters. Specifically, it includes the learning rate, regularization parameter, convolution kernel size, etc., to ensure that the initial state of the model has good convergence and generalization capabilities. The parameter setting process adopts an adaptive adjustment strategy to optimize the initial parameters according to the fluctuations of the target index, thereby enhancing the adaptability of the model to different regions and industry scenarios. The regional-industry correlation analysis is realized through a spatio-temporal graph convolutional network (ST-GCN). As a representative quantity of regional traffic pressure, the macro congestion index forms a spatial correlation graph of regional traffic states through the spatial dependence modeling of the graph convolutional network. The passenger source distribution index is further used for POI trajectory analysis, combined with time series features, to track the activity trajectories of users in specific regions. Through cross-regional spatio-temporal correlation analysis, the core regions of industrial activities and the correlation relationships between the upstream and downstream of the industrial chain are identified to obtain the initial model construction data. This graph network-based analysis method effectively captures the non-linear correlations in the spatial structure, reduces noise interference, and improves the accuracy of regional industry relationship recognition. In the model framework construction link, the initial model construction data is used to build a multi-layer network architecture including a feature extraction layer, a graph convolutional layer, and a time series prediction layer. In the process of assigning the loss function, a weighted loss function is adopted by combining the passenger source distribution index and the macro congestion index to balance the optimization requirements of different targets. This method optimizes the convergence speed and prediction accuracy of the model while taking into account the regional traffic pressure and market demand distribution.
[0041] Preferably, step S3 includes the following steps:
[0042] Step S31: Output the industrial personalization result through the online industrial personalization preliminary analysis model; use the industrial personalization result to conduct enterprise association network analysis to obtain industrial graph neural analysis data;
[0043] Step S32: Use the industrial graph neural analysis data to verify the ratio of user retention rate to task completion rate, and determine the hyperparameters for optimizing the industrial model;
[0044] Step S33: Perform model hyperparameter iteration on the online industrial personalization preliminary analysis model based on the hyperparameters for optimizing the industrial model to obtain the online industrial personalization analysis model.
[0045] The present invention outputs the industrial personalization result through the online industrial personalization preliminary analysis model, and on this basis, conducts enterprise association network analysis, verification of hyperparameters for model optimization, and hyperparameter iteration, so as to construct the online industrial personalization analysis model. At the data level, first, multi-dimensional data such as the business characteristics, market performance, and user behavior patterns of enterprises are extracted through the industrial personalization result. This result can reflect the market positioning, competitive relationships, and potential cooperation opportunities of enterprises within the region. By inputting this data into the enterprise association network analysis module, the graph neural network (Graph Neural Network, GNN) is used to model the business associations, supply chain relationships, and industry competition patterns among enterprises, forming industrial graph neural analysis data. Such data effectively captures the non-linear association relationships among enterprises, making up for the deficiencies of traditional linear analysis. In the verification link of the ratio of user retention rate to task completion rate, by tracking and analyzing user behavior data, the long-term activity and task completion of users are quantified. The user retention rate reflects the market attractiveness of the enterprise's products or services, while the ratio of task completion rate measures the interaction depth and goal achievement of users on the platform. Based on the industrial graph neural analysis data, further analyze the performance of enterprises in the user conversion path, and identify the key factors affecting user retention and task completion from the data level. Through cross-validation of multi-dimensional features, ensure that the selection of optimized hyperparameters has stronger robustness and adaptability. In the model hyperparameter iteration stage, use the verification results to dynamically adjust key hyperparameters such as the learning rate, regularization parameter, node feature dimension, and network layer number of the model. During the iteration process, methods such as Bayesian optimization, grid search, or adaptive gradient adjustment are used to ensure the optimal performance of the model in different scenarios. By continuously optimizing the parameter configuration of the model, improve the model's representation ability and prediction accuracy for complex industrial relationships. The finally formed online industrial personalization analysis model has higher accuracy and stability. This model can not only provide personalized market insights for enterprises, but also perform real-time optimization and adjustment in a dynamic industrial environment. Its data-driven optimization method effectively avoids the limitations of traditional models relying on static parameter settings, ensuring that the analysis results have high applicability and decision-making reference value in different scenarios.
[0046] Preferably, the enterprise association network analysis described in step S3 includes:
[0047] Using the industrial personalized output results to generate comprehensive data through multi-granularity result rendering, where the comprehensive data of multi-granularity result rendering includes macro-level rendering data and meso-level rendering data;
[0048] Extracting traffic flow-enterprise density based on the industrial personalized output results, and performing regional rasterization rendering with the raster size ranging from 100m to 1000m to obtain macro-level rendering data;
[0049] Performing cluster coupling based on the industrial personalized output results to obtain an enterprise cluster coupling index; using the enterprise cluster coupling index to perform coupling association on the industrial personalized output results and performing road network rendering to obtain meso-level rendering data.
[0050] The present invention generates comprehensive data by using industry-specific output results for multi-granularity result rendering, forming macro-level rendering data and meso-level rendering data. At the data level, traffic flow and enterprise density features are first extracted based on the industry-specific output results. Traffic flow data includes information such as road traffic conditions, vehicle driving speeds, and travel volume distributions within a region, while enterprise density data is calculated based on the spatial distribution of enterprises, industry types, and market activity intensities. By mapping this data into a spatial coordinate system, a traffic flow-enterprise density feature matrix is formed, and a regional rasterization rendering method is used to visually represent the data in space. The raster size range is set between 100m and 1000m to adapt to analysis requirements at different scales. Small-scale rasters are used for refined urban traffic management and enterprise layout optimization, while large-scale rasters are suitable for regional economic macro-analysis and industrial policy evaluation. Rasterization rendering not only intuitively shows the spatial relationship between traffic flow and enterprise distribution but also reveals the traffic pressure and industrial agglomeration degree of urban functional areas, providing a decision-making basis for traffic planning and industrial layout. During the generation of meso-level rendering data, cluster coupling analysis is performed on enterprise clusters to calculate the enterprise cluster coupling index. This index quantifies the intensity of the internal connections of enterprise clusters by measuring the spatial proximity, industry relevance, and collaborative cooperation degree among enterprises. At the data level, the calculation of the cluster coupling index uses multi-dimensional feature fusion, including enterprise geographical locations, upstream and downstream relationships in the industrial chain, and supply-demand matching degrees. Further, the enterprise cluster coupling index is used to perform coupling correlation analysis on the industry-specific output results to construct an enterprise cluster relationship network. On this basis, through road network rendering technology, the spatial connections and traffic networks between enterprise clusters are visualized. Road network rendering not only reflects the actual commuting and logistics paths between enterprises but also reveals the dependence of industrial clusters on transportation infrastructure, providing data support for optimizing traffic networks and promoting industrial coordinated development. The finally generated macro-level and meso-level rendering data have multi-granularity spatial expression capabilities. At the macro level, the government and planning agencies can intuitively understand the traffic pressure, enterprise distribution, and economic activity intensity within a region to assist in formulating regional economic development strategies. At the meso level, enterprises and investors can identify potential industrial cooperation opportunities, evaluate the competitive advantages and synergistic effects of enterprise clusters, and optimize resource allocation and business strategies.
[0051] Preferably, step S32 includes the following steps:
[0052] Step S321: Analyze the user retention rate using industrial graph neural analysis data to obtain enterprise user retention rate data;
[0053] Step S322: Detect the enterprise supply chain dependence based on the enterprise user retention rate data to obtain enterprise user-supply chain analysis data;
[0054] Step S323: Analyze the completeness of the enterprise user - supply chain analysis data and conduct ratio verification to obtain the enterprise user - supply chain - completeness data;
[0055] Step S324: Generate hyperparameters for the enterprise user - supply chain - completeness data based on the online industrial personalization preliminary analysis model to obtain the optimized hyperparameters for the industrial model.
[0056] The present invention forms the optimized hyperparameters for the industrial model by performing user retention rate analysis, supply chain dependency detection, completeness analysis, and hyperparameter generation based on industrial graph neural analysis data. At the data level, first, the industrial graph neural analysis data is used to analyze the user retention rate, and key features such as the active behavior, consumption frequency, and transaction records of enterprise users are extracted. By calculating the continuous activity rate and repurchase rate of users within a specific time period, the enterprise user retention rate data is generated. This data effectively reflects the market recognition and user loyalty of the enterprise's products or services, providing a reliable quantitative basis for the enterprise to evaluate user stickiness and market competitiveness. In the enterprise supply chain dependency detection stage, the enterprise user retention rate data is used to analyze the position and dependency relationship of the enterprise in the supply chain. At the data level, the supply chain dependency detection is based on multi - dimensional data features, including the raw material procurement volume, order fulfillment rate, logistics time of the enterprise, and the transaction intensity of upstream and downstream enterprises. Through the association analysis method, the degree of dependence of the enterprise on specific supply chain links is quantified, and the existing supply chain vulnerabilities are identified. The generation of the enterprise user - supply chain analysis data not only reveals the role of the enterprise in the supply chain network but also provides data support for supply chain optimization and risk warning. In the completeness analysis and ratio verification stage, the enterprise user - supply chain analysis data is further used to measure the task completion degree of the enterprise in supply chain collaboration. By analyzing indicators such as the order execution rate, delivery timeliness rate, and production plan achievement rate of the enterprise, the enterprise user - supply chain - completeness data is formed. The ratio verification process detects the performance of the enterprise in market demand response and supply chain coordination by calculating the ratio of the user retention rate to the task completion degree. If the ratio fluctuates abnormally, it indicates that there are problems in the enterprise's supply chain management or customer relationship maintenance. This process provides intuitive performance evaluation data for the enterprise, facilitating timely adjustment of business strategies. In the hyperparameter generation stage, the enterprise user - supply chain - completeness data is modeled based on the online industrial personalization preliminary analysis model. Through multiple rounds of parameter optimization and adaptive learning, the optimized hyperparameters for the industrial model are generated. These hyperparameters include the learning rate, regularization parameter, number of neural network layers, etc., which can dynamically adjust the training strategy and prediction accuracy of the model. Through the hyperparameter generation process, the adaptability and generalization ability of the model are significantly improved, enabling it to maintain a high analysis accuracy in different industrial scenarios.
[0057] Preferably, step S4 includes the following steps:
[0058] Step S41: Extract industrial structured data according to the online industrial personalized analysis model;
[0059] Step S42: Use the industrial structured data to construct an analysis knowledge graph and set industrial confidence parameters to obtain comprehensive industrial analysis data;
[0060] Step S43: Upload the comprehensive industrial analysis data to the cloud platform for storage.
[0061] The present invention extracts industrial structured data, constructs an analysis knowledge graph, and sets industrial confidence parameters based on the online industrial personalized analysis model, and finally forms comprehensive industrial analysis data and uploads it to the cloud platform for storage. At the data level, first, through the process of extracting industrial structured data, the unstructured data output by the model is converted and sorted to form a standardized data format convenient for analysis. The extracted data includes multi-dimensional information such as enterprise operation data, user behavior data, industrial chain relationship data, and market dynamic data. Through technical means such as feature mapping, entity relationship extraction, and anomaly detection, the integrity and accuracy of the data are ensured. The generation of structured data not only provides high-quality input for subsequent analysis but also effectively solves the problems of information redundancy and excessive noise in traditional unstructured data. In the stage of constructing the knowledge graph, the extracted structured data is used to generate multiple types of nodes such as enterprise entities, industrial chain nodes, and market trend nodes, and an industrial analysis knowledge graph is constructed based on graph relationship modeling. At the data level, through relationship reasoning and multi-hop path analysis, the cooperation relationships, supply chain dependence relationships, and market competition patterns among enterprises are identified. Further, a graph neural network (GNN) is used for feature embedding and node classification to realize the dynamic tracking of the upstream and downstream relationships of the industrial chain. The construction of the knowledge graph not only provides an intuitive view of the enterprise and the industrial environment but also provides data support for the causal analysis and trend prediction of complex industrial problems. In the process of setting industrial confidence parameters, the confidence of the nodes and edges in the knowledge graph is calculated to measure the credibility of the data and the reliability of the model output results. At the data level, the calculation of the confidence parameters comprehensively considers factors such as the authority of the data source, the data update frequency, and the historical verification results. Methods such as Bayesian inference and confidence interval estimation are used to quantify the uncertainty of the industrial analysis results. The setting of the confidence parameters ensures the reliability of the comprehensive industrial analysis data in decision-making support and effectively reduces the risk of data noise and misjudgment. The finally generated comprehensive industrial analysis data is uploaded to the cloud platform for storage. The cloud platform has a distributed storage architecture and an efficient retrieval mechanism, which can ensure the security and availability of the data. At the data level, the storage process includes links such as data deduplication, partition storage, and multi-copy backup, which further improves the efficiency and reliability of data storage. At the same time, through the computing resources of the cloud platform, users can perform real-time queries, dynamic analysis, and multi-dimensional mining to realize the wide application of industrial analysis results.
[0062] Preferably, step S42 includes the following steps:
[0063] Step S421: Construct an analysis knowledge graph using industrial structured data;
[0064] Step S422: Set industrial confidence parameters based on the analysis knowledge graph;
[0065] Step S423: Use the industrial confidence parameters to confirm the heat map of the analysis knowledge graph, and conduct comprehensive heat map analysis to obtain comprehensive industrial analysis data.
[0066] The present invention constructs an analysis knowledge graph based on industrial structured data, sets industrial confidence parameters, and confirms the heat map, and finally forms comprehensive industrial analysis data. At the data level, first, an analysis knowledge graph is constructed using industrial structured data, and information such as enterprises, industrial chain links, and market behaviors from different sources and types is connected through an entity relationship graph. This process involves modeling the cooperation relationships, upstream and downstream supply chain relationships, market trends, and competition patterns among enterprises. Through graph algorithms and natural language processing technologies, unstructured information is further transformed into analyzable structured data, enabling various nodes and edges in the industrial chain to be clearly marked in the knowledge graph, thus laying a foundation for subsequent data analysis and decision support. This knowledge graph not only intuitively shows the mutual relationships among various links in the industrial chain but also provides in-depth multi-dimensional analysis for complex industrial problems. Next, industrial confidence parameters are set based on the constructed knowledge graph. The setting of confidence parameters involves evaluating the credibility of each node and edge in the industrial graph at the data level. By weighted calculation of factors such as the weight of the data source, historical verification, and the accuracy of the data source, the confidence value of each node is generated. At the data level, methods such as Bayesian inference and weighted average are used in the calculation of confidence, thus ensuring the reliability of the data and the accuracy of the model output results. The setting of industrial confidence parameters not only provides a reliability assessment of the data for subsequent analysis but also effectively reduces the error in the model inference process and improves the accuracy of the prediction results. In the heat map confirmation stage, the industrial confidence parameters are used to generate the heat map of the analysis knowledge graph. Through spatial distribution analysis, the heat map visualizes the key features of different regions, enterprises, and market performances in the industrial chain, making the changes in industrial behaviors and trends more intuitive. Through comprehensive heat map analysis, key nodes, bottleneck regions, and potential risk points in the industrial chain can be identified. At the data level, the generation of the heat map depends on the application of clustering analysis, numerical processing, and heat mapping algorithms for industrial structured data, thus representing complex data through color gradients to help decision-makers quickly identify high-risk regions and priority development areas. Finally, the obtained comprehensive industrial analysis data provides data-driven decision support for the optimization, adjustment, and prediction of the industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 FIG. 1 is a schematic flow chart of the steps of a method for constructing an online industrial analysis big data platform;
[0068] Figure 2 FIG. 2 Figure 1 is a detailed implementation step flow chart of step S3 in FIG. 1;
[0069] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work belong to the scope of protection of the present invention.
[0071] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0072] It should be understood that although terms such as "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0073] To achieve the above object, please refer to Figures 1 to 2 , a method for constructing an online industrial analysis big data platform, the method comprising the following steps:
[0074] Step S1: Obtain enterprise information and real-time traffic information; perform UGC data desensitization on the enterprise information and real-time traffic information, and construct an online industrial preprocessing data set;
[0075] Step S2: Generate text semantic parsing data based on the online industry preprocessed dataset; correct coordinate deviations according to the online industry preprocessed dataset to obtain coordinate-corrected data; perform industrial time-series prediction and deduction on the text semantic parsing data and the coordinate-corrected data, and construct a personalized model to obtain a preliminary online industry personalized analysis model;
[0076] Step S3: Output industrial personalized results through the preliminary online industry personalized analysis model; perform enterprise association network analysis using the industrial personalized results to obtain industrial graph neural analysis data; determine the hyperparameters for optimizing the industrial model using the industrial graph neural analysis data; perform model hyperparameter iteration on the preliminary online industry personalized analysis model based on the hyperparameters for optimizing the industrial model to obtain an online industry personalized analysis model;
[0077] Step S4: Generate industrial structured data according to the online industry personalized analysis model; construct an analysis knowledge graph using the industrial structured data and upload it to the cloud platform for storage.
[0078] In the embodiment of the present invention, with reference to Figure 1 As shown, it is a schematic diagram of the step flow of a method for constructing an online industry analysis big data platform of the present invention. In this example, the method for constructing the online industry analysis big data platform includes the following steps:
[0079] In the embodiments of the present invention, enterprise information and real-time traffic information are obtained. This process requires real-time data collection through multiple data sources, and the business data, location data, and real-time traffic flow information related to the enterprise are obtained through data interfaces and web crawler technology. These data usually come from public data platforms, government statistical bureaus, traffic management departments, and APIs released by enterprises, etc. The data obtained are mostly unstructured or semi-structured data, and need to be transformed into a unified format through cleaning, formatting, and structuring processes to facilitate subsequent data analysis and mining. Secondly, UGC data desensitization of enterprise information and real-time traffic information is to protect personal privacy and sensitive enterprise information and avoid data leakage. Specific technical means include desensitizing sensitive information such as personal identifiers, contact information, and addresses in user-generated content (UGC). Desensitization methods include data encryption, anonymization, pseudonymization technology, etc. By replacing or encrypting sensitive fields in the data, it cannot be directly identified or traced back to a specific individual or enterprise. These desensitization processes ensure that the data still retains a certain value for use while complying with privacy protection regulations. After the data processing is completed, an online industrial preprocessing data set is constructed. This process depends on standardizing, normalizing, and deduplicating different types of data to ensure data quality. Data preprocessing also involves steps such as timestamp alignment, missing value processing, and noise removal, and finally generates a structured and cleaned online industrial preprocessing data set. This data set provides high-quality data input for subsequent industrial analysis, model construction, and decision support, ensuring the integrity, accuracy, and timeliness of the data.
[0080] Step S2: Generate text semantic analysis data according to the online industrial preprocessing data set; correct coordinate deviations according to the online industrial preprocessing data set to obtain coordinate correction data; perform industrial time series prediction deduction on the text semantic analysis data and the coordinate correction data, and construct a personalized model to obtain an initial online industrial personalized analysis model;
[0081] In the embodiments of the present invention, text semantic parsing data is generated. In this process, natural language processing (NLP) technology is used to perform semantic analysis on the text data in the preprocessed dataset of the online industry. Specifically, by performing word segmentation, part-of-speech tagging, named entity recognition (NER), etc. on the text data such as enterprise information, industry dynamics, and market descriptions, key information in the text is extracted, such as industry terms, company names, product categories, etc., and deep semantic relationships in the text are obtained through semantic analysis models (such as BERT, Word2Vec, etc.). This technical means can convert unstructured text data into structured information and provide rich feature data support for subsequent industry trend analysis, enterprise relationship modeling, etc. Secondly, coordinate deviation correction is performed according to the preprocessed dataset of the online industry to obtain coordinate-corrected data. The real-time traffic information contains deviations of GPS coordinates. Especially in urban environments, GPS positioning errors will be caused by factors such as buildings and high-density signal interference. To ensure the accuracy of the data, coordinate deviation correction technology, such as the POI (Point of Interest) correction algorithm for spatial data, is used to correct the position deviation in the GPS data. This process is usually based on existing geographic information system (GIS) data, matches with accurate map data, and applies a certain error correction model (such as Kalman filtering or least squares method) to adjust the GPS coordinates to ensure the accuracy of the data in the spatial dimension, thereby generating coordinate-corrected data. Finally, the text semantic parsing data and the coordinate-corrected data are used for industrial time-series prediction and deduction, and a personalized model is constructed to obtain a preliminary online industry personalized analysis model. In the data fusion stage, first, different types of data need to be merged to form a multi-dimensional dataset including time dimension and space dimension. Then, through time-series prediction models (such as long short-term memory network LSTM or autoregressive model ARIMA), time-series predictions are made on industry development trends, traffic flow changes, market demand fluctuations, etc. This process learns from historical data to predict industrial changes, enterprise dynamics, etc. that will occur within a certain future period. At the same time, combined with personalized analysis requirements, machine learning methods (such as decision trees, support vector machines, or neural networks) are used to perform personalized modeling on the operation data and market behaviors of different enterprises, thereby generating a personalized analysis model adapted to the characteristics of specific enterprises or industries.
[0082] Step S3: Output industrial personalization results through the preliminary online industry personalized analysis model; perform enterprise association network analysis using the industrial personalization results to obtain industrial graph neural analysis data; determine the optimization hyperparameters of the industrial model using the industrial graph neural analysis data; perform model hyperparameter iteration on the preliminary online industry personalized analysis model based on the optimization hyperparameters of the industrial model to obtain an online industry personalized analysis model;
[0083] In the embodiments of the present invention, industry personalized results are output through an online industry personalized preliminary analysis model. This process uses a previously constructed personalized analysis model to analyze enterprise and industry data, thereby generating personalized conclusions and prediction results for specific industries or enterprises. The analysis model is usually based on machine learning or deep learning methods, such as support vector machines (SVMs), random forests (RFs), or neural networks (NNs). It outputs multi-dimensional personalized results such as industry development trends, enterprise performance, and market demands by learning patterns from industry data. Next, enterprise association network analysis is performed using the industry personalized results to obtain industry graph neural analysis data. This process uses graph neural network (GNN) technology to model the association relationships between industries and enterprises. Graph neural network is a deep learning method that can effectively process graph-structured data. By treating each enterprise in the industry as a node in the graph and the cooperation, competition, or supply-demand relationships between enterprises as edges between nodes, graph neural network can capture the complex dependencies between enterprises and their dynamic changes. By learning these relationships, graph neural network can deduce a more accurate industry network structure and then generate industry graph neural analysis data, which reflect information such as the interactions within the industry, cooperation patterns, and risk propagation paths between enterprises. Then, the industry model optimization hyperparameters are determined using the industry graph neural analysis data. In the model optimization stage, by analyzing the industry graph neural analysis data, hyperparameter optimization algorithms (such as Bayesian optimization, grid search, or genetic algorithms) are used to adjust the hyperparameters of the industry analysis model. These hyperparameters are crucial for the performance of the model, and reasonable hyperparameter settings can significantly improve the prediction accuracy and generalization ability of the model. In this process, the optimization objective (such as accuracy, F1 value, etc.) is first defined, and based on the features extracted from the graph neural network and the prediction results of the model, the model parameters are automatically adjusted to achieve the best performance. Finally, model hyperparameter iteration is performed on the online industry personalized preliminary analysis model based on the industry model optimization hyperparameters to obtain the online industry personalized analysis model. This step relies on optimization techniques such as backpropagation algorithm and gradient descent. By continuously adjusting the model hyperparameters and performing iterative training until the output of the model reaches the expected performance level. Through this iterative process, the accuracy, robustness, and adaptability of the model are improved, thereby providing more refined and personalized prediction results for industry analysis.
[0084] Step S4: Generate industry structured data according to the online industry personalized analysis model; construct an analysis knowledge graph using the industry structured data and upload it to the cloud platform for storage.
[0085] In an embodiment of the present invention, industry structured data is generated according to an online industry personalized analysis model. The core of this process is to convert the output data of the personalized analysis model into a structured data format, usually by feature extraction and data collation methods, to convert the original industry analysis results into a more standardized and easy-to-analyze data structure. At this time, the structured data not only contains the basic information of each enterprise and industry, but also includes multi-dimensional numerical or categorical data such as market trends, industrial chain nodes, product demand, consumer behavior, etc. These data are pre-processed by means of data cleaning, data normalization and standardization to ensure that they are suitable for subsequent deep analysis and graph construction. Next, the industry structured data is used to construct an analytical knowledge graph. The technical means of knowledge graph construction usually rely on graph database technology and natural language processing (NLP) technology, and establish a relationship network between entities such as enterprises, products, services, and markets by processing the industry data with semantic parsing, entity recognition, and relationship extraction. In the graph, each entity (such as an enterprise, product, etc.) is represented as a node in the graph, and the relationship between entities (such as supply chain relationships, cooperative relationships, etc.) is represented as an edge, thereby forming a complex network containing multiple information. Through the knowledge graph, various knowledge, relationships and laws in the industry field can be effectively captured, so that the industrial data not only has a structured form of expression, but also can provide semantic support and reasoning capabilities for complex business problems. Finally, the constructed industry analysis knowledge graph is uploaded to the cloud platform for storage. This step relies on cloud storage technology and distributed computing frameworks. It usually uses large-scale data storage solutions provided by cloud service platforms (such as AWS, Azure or Google Cloud) to store knowledge graphs in the cloud in an efficient and scalable manner. The cloud platform not only supports large-scale data storage, but also provides convenient data access, query and sharing services. Through cloud storage, enterprises or research institutions can access and use knowledge graphs at any time, and combine cloud computing resources for further computational analysis and model training.
[0086] Preferably, step S1 comprises the following steps:
[0087] Step S11: Obtaining basic enterprise information and real-time traffic information;
[0088] Step S12: screening the upstream and downstream relationships of the industrial chain for the basic information of the enterprise, and eliminating blank data to obtain the enterprise integrated data; generating commuting time circle analysis data for the real-time traffic information;
[0089] Step S13: Desensitize the UGC data based on the enterprise integration data and commuting time circle analysis data, and build an online industry preprocessing data set.
[0090] In the embodiments of the present invention, the collection of enterprise data and traffic data is achieved by obtaining enterprise authoritative information and real-time traffic information. Enterprise authoritative information is usually obtained through official databases, enterprise official websites or industry reports, and these data include basic information of enterprises, industrial chain roles, market behaviors, etc.; real-time traffic information is obtained through means such as traffic monitoring systems, GPS data, traffic flow monitoring, etc., which reflects the dynamic changes of the traffic system, such as traffic flow, road conditions information, commuting patterns, etc. In the data collection stage, it is necessary to ensure the comprehensiveness, timeliness and accuracy of the data for subsequent processing and analysis. Screening the upstream and downstream relationships of the industrial chain for enterprise authoritative information is to identify enterprises related to the target analysis from a large amount of enterprise data and eliminate information that is irrelevant or invalid to the analysis. The screening of the upstream and downstream relationships of the industrial chain usually adopts graph analysis techniques in graph theory. By analyzing the positions and relationships of enterprises in the supply chain and value chain, their upstream and downstream roles in the industrial chain are determined. Through this screening, blank data is removed, ensuring the accuracy and integrity of the data set, and thus obtaining more accurate integrated enterprise authoritative data. At the same time, for the real-time traffic information, commuting time circle analysis data is generated. Using geographic information system (GIS) and spatio-temporal data analysis techniques, the traffic flow patterns in a specific area are analyzed, and the commuting time circle is drawn, so as to provide spatio-temporal background information for subsequent traffic pattern analysis and industrial decision-making. UGC (user-generated content) data desensitization is performed according to the integrated enterprise authoritative data and commuting time circle analysis data. The main purpose of UGC data desensitization is to protect personal privacy, and usually encryption techniques, anonymization techniques or data perturbation techniques are used to process sensitive information. Specifically, sensitive information of enterprises and individuals, such as addresses, contact information, etc., can be processed by data replacement or masking to make it untraceable. Then, the desensitized data is used to construct an online industrial preprocessing data set, which will integrate enterprise information and traffic data to provide data support for subsequent industrial analysis. This process includes data cleaning, noise removal, standardization processing, etc., to ensure that the data can still be used for effective analysis while maintaining data privacy.
[0091] Preferably, the UGC data desensitization described in step S1 includes:
[0092] Extracting enterprise user data according to the integrated enterprise data;
[0093] Extracting user trajectory information data according to the commuting time circle analysis data;
[0094] Generating and anonymizing the identifiers of the enterprise user data, where the range of the anonymization key is 128 - 256 bits, to obtain online industrial identifier anonymized data;
[0095] Generalizing the regional scope of the user trajectory information data to obtain online regional generalized data;
[0096] Hash the sensitive fields of enterprise user data generation and user trajectory information data to obtain online industry sensitive field data;
[0097] Fuse the multi-source data of online industry identifier anonymized data, online geographical generalization data, and online industry sensitive field data, and construct a fusion dataset to obtain an online industry preprocessing dataset.
[0098] In the embodiments of the present invention, enterprise user data extraction is performed based on the authoritative integration of enterprise data. First, key features related to enterprise users need to be extracted from the original data through data screening and feature extraction methods. Usually, feature selection algorithms or data cleaning methods are used to screen the data to ensure the accuracy and relevance of the data. This process relies on traditional data processing technologies, such as data integration, deduplication, and missing value filling, to ensure that the extracted enterprise user data has high quality and high credibility. Secondly, for the commuting time circle analysis data, the geographical information system (GIS) technology is used to analyze the user's movement trajectory. The spatio-temporal analysis method is adopted to extract the user's movement path and activity range from the traffic data, and the user trajectory information data is constructed. In this process, by analyzing the geographical location data, timestamp, and traffic flow data, a commuting time circle can be established, and then spatio-temporal data support for user behavior can be provided for subsequent analysis. To protect user privacy and comply with data protection regulations, anonymization techniques are adopted. After the enterprise user data is generated, a 128-256-bit key is used for identifier anonymization. By encrypting the identifier, it is ensured that the data can be analyzed without revealing the user's identity. Symmetric encryption or hash encryption algorithms are usually used for identifier anonymization, and the selection of the key range depends on the requirements of data privacy protection and the specific implementation environment. Next, for the user trajectory information data, regional range generalization is performed. Generalization technology is a common technology in data privacy protection. It mainly converts precise location coordinates into regional information with a larger range through fuzzy operations, thereby reducing the risk of location privacy leakage. This operation usually realizes through clustering algorithms or block division technologies in GIS. Further, sensitive fields (such as location, behavior data, etc.) in the enterprise user data and user trajectory information data are processed using hashing technology. Hashing technology encrypts and hashes the sensitive fields to generate a hash value with a fixed length, thereby ensuring the privacy and integrity of the data. Hashing technology has one-wayness, that is, the original data is irreversibly converted into a hash value, which increases the security of the data. Through these methods, the anonymity and non-traceability of the sensitive field data in the online industry are ensured. Finally, through multi-source data fusion technology, the anonymized data of the online industry identifier, the generalized data of the online region, and the sensitive field data of the online industry are integrated. In this process, data fusion technology involves the merging of multi-dimensional data. Usually, through technical means such as data alignment, matching, and standardization, the association between different data sources is realized, and a unified data set is generated. The fused data set is the preprocessed data set of the online industry, which provides a secure, privacy-protected, and well-structured data basis for subsequent analysis.
[0099] Preferably, step S2 includes the following steps:
[0100] Step S21: Generate text semantic parsing data according to the preprocessed data set of the online industry;
[0101] Step S22: Perform coordinate deviation correction based on the online industrial preprocessing dataset. When the GPS coordinate deviation is greater than or equal to 50 m, perform POI spatial correction to obtain coordinate correction data;
[0102] Step S23: Extract regional week-on-week traffic spatio-temporal features from the text semantic parsing data and the coordinate correction data, and construct time-series prediction deduction data;
[0103] Step S24: Construct a personalized model based on the time-series prediction deduction data to obtain a preliminary online industrial personalized analysis model.
[0104] In the embodiments of the present invention, text data is split into independent words or phrases through a word segmentation technique; then, each word is assigned a grammatical role (such as noun, verb, etc.) through part-of-speech tagging (POS Tagging); next, text is transformed into a vector form using word vector techniques (such as Word2Vec, GloVe, or BERT), enabling the machine to understand the semantic relationships of the text. These processed vector representations can capture the grammar, semantics, and context information of the text and provide basic data for subsequent analysis tasks. For example, a text classification model will classify different types of industrial texts based on these vectors, providing guidance for the further processing of industrial data. The mentioned "coordinate deviation correction" and "POI spatial correction" techniques involve spatial data processing in a geographic information system (GIS). GPS coordinates are usually affected by errors, especially in urban environments where buildings and other obstacles can cause signal deviations. To improve the position accuracy, spatial correction is first required by comparing real-time GPS data with known point-of-interest (POI) data. During the spatial correction process, methods such as the nearest neighbor interpolation algorithm and the k-nearest neighbor algorithm (KNN) are usually adopted to adjust the GPS coordinates by finding the closest known location data in space. Further, the coordinate accuracy can be improved through multi-source data fusion (such as mobile base station information, Wi-Fi hotspots, etc.) to obtain more reliable geographical location information, which is crucial for accurately analyzing user behavior and traffic patterns. When performing "regional week-on-week traffic spatio-temporal feature extraction" and "temporal sequence prediction and deduction data construction", spatio-temporal data analysis and time series prediction techniques are involved. First, the core of traffic spatio-temporal feature extraction is to extract meaningful spatio-temporal features from large-scale traffic flow data. This can be achieved by using spatio-temporal clustering analysis (such as the DBSCAN algorithm) to identify traffic hotspots and patterns in different regions at different time periods, or by spatio-temporal association rule mining (such as the Apriori algorithm) to find traffic flow rules in specific regions and time periods. These extracted features provide a data basis for the temporal sequence prediction model. Temporal sequence prediction uses algorithms such as ARIMA, LSTM (long short-term memory network), Prophet, etc., to predict future traffic flow or congestion conditions based on the time series characteristics of historical traffic data. Deep learning models such as LSTM can process long temporal data and capture long-term dependencies by memorizing historical states, thus providing more accurate future trend predictions. "Personalized model construction" is based on the temporal sequence prediction and deduction data, and models are constructed for individuals or specific industries through machine learning and deep learning methods. At this stage, the model construction techniques can include traditional machine learning methods such as support vector machines (SVM), decision trees, and random forests, or deep learning methods such as convolutional neural networks (CNN) and recurrent neural networks (RNN).Through training data, these models can automatically learn personalized demand patterns from time-series data and predict future change trends based on historical data. The construction of personalized models aims to optimize resource allocation, traffic prediction, etc. according to the specific needs of industries or users, enabling the models to be dynamically adjusted according to actual situations, and then outputting an initial online industry personalized analysis model to provide support for further analysis.
[0105] Preferably, the construction of the personalized model described in step S2 includes:
[0106] Setting personalized target indexes according to time-series prediction and deduction data, where the personalized target indexes include a macro congestion index and a passenger source distribution index;
[0107] Setting initial model parameters according to time-series prediction and deduction data;
[0108] Using a preset spatio-temporal graph convolutional network to conduct regional-industry correlation analysis on the macro congestion index and parsing the passenger source POI trajectory with the passenger source distribution index to obtain initial model construction data;
[0109] Conducting model framework construction on the initial model construction data and assigning a loss function to obtain an initial online industry personalized analysis model.
[0110] In the embodiments of the present invention, the setting of personalized target indices is based on time-series prediction and deduction data. The personalized target indices include a macro congestion index and a passenger source distribution index. To generate these target indices, time-series analysis techniques such as the ARIMA (Autoregressive Integrated Moving Average) model or the LSTM (Long Short-Term Memory Network) are required to model the historical data of traffic and passenger source flows, capturing their long-term and short-term trend changes. The macro congestion index reflects the traffic congestion level within a certain time period through time-series modeling of data such as traffic flow, road capacity, and vehicle speed; while the passenger source distribution index obtains the spatial distribution characteristics of passenger sources through the analysis of data such as passenger flow volume and user activity trajectories at different times and in different regions, thereby providing personalized prediction results for subsequent analysis. In the initial parameter setting stage of the model, the data used mainly comes from time-series prediction and deduction data, involving model tuning techniques. The initial parameters of the model can be set through methods such as Bayesian optimization or grid search. These methods can automatically adjust the hyperparameters of the model according to the goal of minimizing the prediction error, enabling the model to better adapt to the actual industrial and traffic environment. The application of the Spatio-Temporal Graph Convolutional Network (ST-GCN) technology further enhances the spatio-temporal modeling ability of the model. ST-GCN can take spatio-temporal data as input, use the Graph Convolutional Network (GCN) to process graph-structured data, and combine the time dimension to model the evolution of nodes in space, and then conduct regional-industry correlation analysis. Through the spatio-temporal graph convolutional network, the model can deeply explore the relationship between traffic congestion and industrial development in different regions, and based on this, conduct dynamic traffic flow prediction and optimization. In addition, the passenger source distribution index conducts detailed spatial analysis through POI (Point of Interest) trajectory parsing, using trajectory data analysis techniques such as Dynamic Time Warping (DTW) and clustering analysis to identify the activity patterns of users at different times and spatial positions. Through these analyses, the accurate trajectories of passenger sources can be obtained, further promoting the in-depth development of personalized analysis. Finally, in the model framework construction stage, all the preliminarily constructed data is iteratively optimized through the construction of the model framework and the assignment of the loss function. The setting of the loss function usually depends on the objective to be optimized, such as the Mean Squared Error (MSE) or cross-entropy, etc., which is used to evaluate the error between the model and the actual data during the training process, thereby guiding the adjustment of model parameters. Through this process, an online industrial personalized preliminary analysis model is finally formed, which can provide data support for further online industrial analysis and decision-making.
[0111] As an example of the present invention, refer to Figure 2 As shown, in this example, step S3 includes:
[0112] Step S31: Output the industrial personalization result through the online industrial personalization preliminary analysis model; use the industrial personalization result to conduct enterprise association network analysis to obtain industrial graph neural analysis data;
[0113] Step S32: Use the industrial graph neural analysis data to verify the ratio of user retention rate to task completion rate, and determine the hyperparameters for optimizing the industrial model;
[0114] Step S33: Iterate the model hyperparameters of the online industrial personalized preliminary analysis model based on the hyperparameters for optimizing the industrial model to obtain the online industrial personalized analysis model.
[0115] In the embodiment of the present invention, the output of the industrial personalized result by the online industrial personalized preliminary analysis model means that data prediction and analysis are performed according to the previously trained personalized model to generate personalized features for a specific industry. In this process, the data completes the task of extracting key indicators from the large dataset through steps such as spatio-temporal prediction, feature extraction, and target index calculation of the model. Then, the industrial personalized result is used for enterprise association network analysis to construct a graph structure based on the mutual connections and interactions between enterprises. At this time, the relationships between enterprises are represented as nodes and edges in the graph, and the graph neural network (GNN), as an effective graph structure data processing technology, uses the information of nodes and edges for in-depth analysis and reasoning. Through the propagation of neighborhood information, GNN enables the model to capture the complex relationship patterns between enterprises, thereby obtaining the industrial graph neural analysis data. Use the industrial graph neural analysis data to verify the ratio of user retention rate to task completion rate. User retention rate analysis usually involves tracking the continuous participation behavior of users on the platform, using methods such as time series analysis, regression analysis, or survival analysis to evaluate the activity level and churn risk of users over a period of time. The verification of the task completion rate ratio focuses on the situation of users completing the specified tasks. Combining the task completion records in the dataset, calculate the task completion rate of users in a certain period, and compare the completion situations of different tasks to evaluate the service efficiency and user experience of the platform. Through these verifications, the key factors affecting the model accuracy can be identified and a basis for optimizing the hyperparameters can be provided. The mentioned hyperparameters for optimizing the industrial model are to adjust the hyperparameters of the existing model to improve the accuracy and generalization ability of the model. Through optimization techniques such as grid search, Bayesian optimization, or genetic algorithms, the optimal combination of hyperparameters is found. The optimization of hyperparameters can ensure the optimal performance of the model in practical applications by maximizing the performance of the model on the validation set. This process includes adjusting parameters such as the learning rate, regularization term, and hidden layer size. After completing the hyperparameter optimization, based on the optimized hyperparameters, the model enters the iteration stage, that is, further training and adjustment are performed, and finally the optimized online industrial personalized analysis model is generated to ensure that it can perform efficient and accurate prediction and analysis according to the new data environment.
[0116] Preferably, the enterprise association network analysis described in step S3 includes:
[0117] Generate comprehensive data by rendering multi-granularity results using industry-specific output results, where the multi-granularity result-rendered comprehensive data includes macro-layer rendering data and meso-layer rendering data;
[0118] Extract traffic flow-enterprise density based on industry-specific output results, and perform regional rasterization rendering with the raster size ranging from 100m to 1000m to obtain macro-layer rendering data;
[0119] Perform cluster coupling based on industry-specific output results to obtain the enterprise cluster coupling index; use the enterprise cluster coupling index to perform coupling association on industry-specific output results and perform road network rendering to obtain meso-layer rendering data.
[0120] In the embodiments of the present invention, multi-granularity result rendering comprehensive data is generated using the industry personalized output results, aiming to refine the presentation of industry characteristics through data analysis and rendering at different levels. The macro-layer rendering data and the meso-layer rendering data respectively correspond to the spatial and spatio-temporal characteristics at large and medium scales. These rendering data provide in-depth analysis of industry operation, traffic flow, enterprise distribution, etc. at the macro and meso dimensions respectively, and can comprehensively present multi-level information of the industrial ecosystem. Next, traffic flow-enterprise density extraction is performed based on the industry personalized output results. Using spatial analysis techniques, by analyzing Geographic Information System (GIS) data, the traffic flow and enterprise density are quantitatively extracted. In this process, first, the enterprise density and traffic flow per unit area need to be calculated through traffic flow data and the spatial distribution data of enterprises to reflect the industrial activity level and traffic pressure in different regions. These data will then enter the regional rasterization rendering step, that is, the entire region is divided into small grids by the rasterization method, and the corresponding flow and density information are calculated in each grid. The selection of the raster size between 100m and 1000m can be flexibly adjusted at different scales according to the required analysis accuracy, so as to generate the spatial distribution map of traffic flow and enterprise density at the macro layer. In further processing, through cluster coupling analysis, combining the spatial and operational relationships between enterprises, an enterprise cluster coupling index is generated. This index reflects the degree of aggregation of enterprises in the geographical space and their interaction relationships in the industrial chain. The generation of the enterprise cluster coupling index involves clustering enterprises according to their geographical locations and business relationships, evaluating the closeness of their interaction with surrounding enterprises, and quantifying this interaction relationship. Through this index, the formation and development trend of industrial clusters can be effectively depicted. Finally, the enterprise cluster coupling index is used to perform coupling correlation analysis on the industry personalized output results and perform road network rendering. This step combines road network data and industrial distribution data for spatial analysis and rendering to reveal the relationship between industrial activities and the road network. Road network rendering can be carried out in different ways, such as line width change based on traffic flow, traffic density heat map, etc., to show the traffic conditions of different road sections and their relationship with enterprise clusters.
[0121] Preferably, step S32 includes the following steps:
[0122] Step S321: Analyze the user retention rate using industrial graph neural analysis data to obtain enterprise user retention rate data;
[0123] Step S322: Detect the enterprise supply chain dependence according to the enterprise user retention rate data to obtain enterprise user-supply chain analysis data;
[0124] Step S323: Analyze the completeness of the enterprise user-supply chain analysis data and perform ratio verification to obtain enterprise user-supply chain-completeness data;
[0125] Step S324: Generate hyperparameters for the enterprise user - supply chain - completion data based on the online industry personalized preliminary analysis model to obtain the hyperparameters for optimizing the industry model.
[0126] In the embodiments of the present invention, the "analysis of user retention rate using industrial graph neural analysis data" mainly uses the graph neural network (GNN) technology to analyze the relationship between enterprises and their users within the industrial chain, and then calculates the retention situation of users in different time periods. The graph neural network can effectively capture the relationship and interaction pattern between users. By modeling nodes (enterprises or users) and edges (interaction relationships), it analyzes the user retention rate within a given period, reflecting the ability of different enterprises to attract and retain users in the industrial chain. "Detection of enterprise supply chain dependence based on enterprise user retention rate data" further analyzes the supply chain dependence relationship between enterprises on the basis of the user retention rate. Here, network analysis and supply chain models are mainly used to reflect the demand and supply relationship between enterprises through user retention rate data, generating interactive analysis data between enterprise users and the supply chain. These data help to identify which enterprises have higher dependence in the supply chain and which enterprises are the key nodes in the chain. "Conduct completion analysis on the enterprise user - supply chain analysis data and perform ratio verification to obtain the enterprise user - supply chain - completion data" involves in - depth processing of the analysis results. Statistical methods are used for completion analysis to examine the performance and task completion degree of enterprises in the supply chain. Through ratio verification, the accuracy and rationality of the data are verified to ensure that the obtained data can reflect the actual business situation and provide reliable data for the next - step optimization. "Generate hyperparameters for the enterprise user - supply chain - completion data based on the online industry personalized preliminary analysis model" optimizes the data obtained from the previous analysis through regression models, machine learning, or deep - learning algorithms. Hyperparameter generation determines the optimal parameter combination of the model through experiments and tuning, using methods such as cross - validation, so as to improve the accuracy and generalization ability of the model prediction, and then form the hyperparameters for optimizing the industry model, ensuring that the online industry personalized analysis model can achieve optimal performance in practical applications.
[0127] Preferably, step S4 includes the following steps:
[0128] Step S41: Extract industrial structured data according to the online industry personalized analysis model;
[0129] Step S42: Use the industrial structured data to construct an analysis knowledge graph and set industrial confidence parameters to obtain comprehensive industrial analysis data;
[0130] Step S43: Upload the comprehensive industrial analysis data to the cloud platform for storage.
[0131] In the embodiments of the present invention, "extracting industrial structured data according to the online industry personalized analysis model" involves using the industry personalized analysis model to perform structured processing on the original data. This process converts the unstructured data related to enterprises into structured data through multi-level data mining and information processing methods, such as natural language processing (NLP) and data cleaning techniques. This process is not just about converting the data format. More importantly, it identifies the key features in the data and maps them to a predefined data structure for further analysis and processing. "Constructing an analysis knowledge graph using the industrial structured data and setting industrial confidence parameters to obtain comprehensive industrial analysis data". This process uses graph database technology and knowledge graph construction methods. A knowledge graph represents industrial data by constructing a graph structure of entities (such as enterprises, products, industries, etc.) and the relationships between them. Using the industrial structured data, a multi-dimensional and multi-level knowledge graph is constructed based on the preset relationship rules and semantic models, thereby revealing the mutual relationships and influence paths between different entities in the industry. At the same time, the setting of industrial confidence parameters is a process of evaluating the credibility of the constructed knowledge graph, aiming to provide a reliable basis for subsequent analysis and ensure high credibility of the data in the reasoning and decision-making processes. This process generally requires evaluating and adjusting the parameters through machine learning or statistical models to make the information in the graph more accurate. "Uploading the comprehensive industrial analysis data to the cloud platform for storage". This step mainly uses cloud computing technology to upload the data processed through industrial structured data and knowledge graphs to the cloud platform for long-term storage and management. The cloud platform provides efficient storage and computing capabilities, capable of supporting the processing and analysis of large-scale data. Through the storage of the cloud platform, the comprehensive industrial analysis data can be efficiently backed up and accessed, and can also be shared and integrated with other systems, providing basic data support for subsequent data analysis, decision support, and model iteration.
[0132] Preferably, step S42 includes the following steps:
[0133] Step S421: Constructing an analysis knowledge graph using the industrial structured data;
[0134] Step S422: Setting industrial confidence parameters based on the analysis knowledge graph;
[0135] Step S423: Using the industrial confidence parameters to confirm the heat map of the analysis knowledge graph and performing comprehensive heat map analysis to obtain comprehensive industrial analysis data.
[0136] In the embodiments of the present invention, an analytical knowledge graph is constructed using industrial structured data. This process relies on graph databases and knowledge graph technologies. Through in-depth mining of industry-related data, multi-dimensional entities such as enterprises, products, services, markets, etc. and their mutual relationships (such as cooperation, dependence, competition, etc.) are transformed into a graph form for expression. During the data structuring process, raw data (such as text, numerical values, geographical information, etc.) is processed through technologies such as natural language processing (NLP), entity recognition, and relationship extraction, and mapped into nodes and edges with clear hierarchies and correlations. This knowledge graph can reveal the complex connections between various elements in the industry, providing visual basic data for further analysis. In the next step S422, industrial confidence parameters are set based on the analytical knowledge graph. This operation means quantitatively evaluating the information quality in the knowledge graph. The setting of confidence parameters usually relies on machine learning algorithms, such as Bayesian inference or other statistical models, to determine the credibility of each node and edge in the graph. This process weights the relationships between nodes in the knowledge graph by evaluating historical data, model errors, or the feedback of domain experts to improve the accuracy and effectiveness of the graph. Finally, in step S423, the analytical knowledge graph is confirmed by using industrial confidence parameters, and a comprehensive heat map analysis is performed to obtain comprehensive industrial analysis data. This technical means combines heat map generation and analysis methods. Using confidence parameters, the intensity of each node and relationship in the industrial knowledge graph is visualized in the form of a heat map. During this process, hot spots and highly confident connections will be marked as darker parts, reflecting the most critical links and influential factors in the industry. By comprehensively analyzing the heat map data, the industrial structure, development trend, and potential risks can be evaluated as a whole, further forming actionable comprehensive industrial analysis data and providing in-depth insights for the decision support system.
[0137] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.
[0138] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing an online industrial analysis big data platform, characterized in that, The following steps are involved: Step S1: Obtaining basic enterprise information and real-time traffic information; Desensitize UGC data of enterprise basic information and real-time traffic information, and build an online industry pre-processing data set; Step S2: performing text semantic analysis based on the online industry preprocessing data set to generate text semantic data; performing coordinate deviation correction based on the online industry preprocessing data set to obtain coordinate correction data; The text semantic data and coordinate correction data are used to perform industry time series forecasting and deduction, and a personalized model is constructed to obtain an online industry personalized preliminary analysis model; Step S3: outputting industry personalized results through the online industry personalized preliminary analysis model; Use the industry personalized results to conduct enterprise association network analysis and obtain industry graph neural analysis data; Use industry graph neural analysis data to determine industry model optimization hyperparameters; Based on the industry model optimization hyperparameters, the model hyperparameters of the online industry personalized preliminary analysis model are iterated to obtain the online industry personalized analysis model; Step S4: Generate industry structured data based on the online industry personalized analysis model; Use industry structured data to build analytical knowledge graphs and upload them to the cloud platform for storage.
2. The construction method of the online industrial analysis big data platform according to claim 1, wherein, Step S1 includes the following steps: Step S11: Obtaining basic enterprise information and real-time traffic information; Step S12: screening the upstream and downstream relationships of the industrial chain for the basic information of the enterprise, and eliminating blank data to obtain the enterprise integrated data; generating commuting time circle analysis data for the real-time traffic information; Step S13: Desensitize the UGC data based on the enterprise integration data and commuting time circle analysis data, and build an online industry preprocessing data set.
3. The construction method of the online industrial analysis big data platform according to claim 2, characterized in that, The UGC data desensitization described in step S1 includes: Extract enterprise user data based on enterprise integrated data; Extract user trajectory information data based on commuting time circle analysis data; The enterprise user data is generated to anonymize the identifier, where the anonymization key ranges from 128 to 256 bits, to obtain the online industry identifier anonymized data; Generalize the user trajectory information data to a regional scope to obtain online regional generalization data; Hash sensitive fields of enterprise user data generation and user trajectory information data to obtain online industry sensitive field data; The online industry identifier anonymized data, online regional generalized data and online industry sensitive field data are fused, and a fused data set is constructed to obtain an online industry preprocessed data set.
4. The construction method of the online industrial analysis big data platform according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Generate text semantic analysis data based on the online industry preprocessing data set; Step S22: Correct the coordinate deviation according to the online industry preprocessing data set. When the GPS coordinate deviation is greater than or equal to 50m, perform POI space correction to obtain coordinate correction data; Step S23: extracting the spatiotemporal features of regional weekly traffic from the text semantic analysis data and the coordinate correction data, and constructing the time series prediction deduction data; Step S24: construct a personalized model based on the time series prediction and deduction data to obtain an online industry personalized preliminary analysis model.
5. The construction method of the online industrial analysis big data platform according to claim 4, characterized in that, The personalized model building described in step S2 includes: Set personalized target indices based on time - series prediction and deduction data, where the personalized target indices include a macro congestion index and a passenger source distribution index; Set the initial parameters of the model according to the time - series prediction and deduction data; Use a preset spatio - temporal graph convolutional network to conduct regional - industry correlation analysis on the macro congestion index, and use the passenger source distribution index to analyze the passenger source POI trajectory to obtain the initial model construction data; Construct the model framework from the initial model construction data and assign a loss function to obtain a preliminary online industrial personalized analysis model.
6. The construction method of the online industrial analysis big data platform according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Output the industrial personalization result through the preliminary online industrial personalized analysis model; use the industrial personalization result to conduct enterprise correlation network analysis to obtain industrial graph neural analysis data; Step S32: Use the industrial graph neural analysis data to verify the ratio of user retention rate to task completion rate and determine the hyperparameters for optimizing the industrial model; Step S33: Iterate the model hyperparameters of the preliminary online industrial personalized analysis model based on the hyperparameters for optimizing the industrial model to obtain an online industrial personalized analysis model.
7. The construction method of the online industrial analysis big data platform according to claim 6, characterized in that The enterprise correlation network analysis described in Step S3 includes: Use the output result of industrial personalization to generate comprehensive data for multi - granularity result rendering, where the comprehensive data for multi - granularity result rendering includes macro - layer rendering data and meso - layer rendering data; Extract traffic flow - enterprise density according to the output result of industrial personalization and conduct regional rasterization rendering with the raster size ranging from 100m to 1000m to obtain the macro - layer rendering data; Conduct cluster coupling according to the output result of industrial personalization to obtain an enterprise cluster coupling index; use the enterprise cluster coupling index to conduct coupling correlation on the output result of industrial personalization and conduct road network rendering to obtain the meso - layer rendering data.
8. The construction method of the online industrial analysis big data platform according to claim 6, characterized in that, Step S32 includes the following steps: Step S321: Use the industrial graph neural analysis data to conduct user retention rate analysis to obtain enterprise user retention rate data; Step S322: Conduct enterprise supply - chain dependence detection based on the enterprise user retention rate data to obtain enterprise user - supply - chain analysis data; Step S323: Conduct completion analysis on the enterprise user - supply - chain analysis data and conduct ratio verification to obtain enterprise user - supply - chain - completion data; Step S324: Generate hyperparameters for the industrial model based on the enterprise user - supply - chain - completion data of the preliminary online industrial personalized analysis model to obtain the hyperparameters for optimizing the industrial model.
9. The method for constructing an online industrial analysis big data platform according to claim 1, wherein Step S4 includes the following steps: Step S41: Extract industrial structured data according to the online industrial personalized analysis model; Step S42: Use the industrial structured data to construct an analysis knowledge graph and set industrial confidence parameters to obtain comprehensive industrial analysis data; Step S43: Upload the comprehensive industrial analysis data to the cloud platform for storage.
10. The method for constructing an online industrial analysis big data platform according to claim 9, wherein Step S42 includes the following steps: Step S421: Use the industrial structured data to construct an analysis knowledge graph; Step S422: Set industrial confidence parameters based on the analysis knowledge graph; Step S423: Use the industrial confidence parameter to confirm the heat map of the analysis knowledge graph and conduct a comprehensive heat map analysis to obtain comprehensive industrial analysis data.
Citation Information
Patent Citations
Time series data event prediction method and system based on graph convolutional neural network and application thereof
CN111367961A
Intelligent traffic flow analysis method based on Beidou data
CN117935561A