Construction method of online industry analysis big data platform
By building an online industry analysis big data platform, using technologies such as data desensitization, text semantic analysis, timing prediction and graph neural network, the existing platform's shortcomings in data processing and analysis depth have been solved, and efficient and real-time industrial analysis and decision-making support have been achieved.
Patent Information
- Application Number
- CN202510483948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing industrial analysis platform responds slowly when processing dynamic industrial data, lacks real-timeness, and cannot deeply explore complex industrial chain relationships. The graph neural network is inefficient in spatiotemporal data processing, and cannot provide multi-grained and cross-regional comprehensive analysis.
By obtaining enterprise information and real-time traffic information, UGC data desensitization is carried out and an online industry preprocessing data set is built. Then, text semantic analysis and coordinate deviation correction are carried out, industrial timing prediction and deduction are carried out, personalized models are constructed, enterprise correlation network analysis is carried out, industrial model optimization hyperparameters are determined, industrial structured data is generated, and analytical knowledge graph is constructed.
It improves the real-time and analysis depth of data processing, enhances the mining ability of industrial chain associations, improves the stability and prediction accuracy of the model, realizes multi-grained and cross-region comprehensive analysis, and supports real-time updated industrial relationship information sharing and intelligent decision-making.
Smart Images

Figure CN120012140A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and in particular to a method for constructing an online industry analysis big data platform. Background Art
[0002] Many platforms are slow to respond to dynamic industrial data, lack real-time performance, and are unable to quickly respond to changes in multiple heterogeneous data sources. Secondly, traditional personalized recommendation and prediction methods often cannot deeply explore the complex connections in the industrial chain, resulting in limited accuracy and effectiveness of analysis results. At the same time, the application of graph neural networks in industrial data analysis faces bottlenecks in computing resources and model complexity, especially in the processing of spatiotemporal data. In addition, existing platforms usually rely on single-level analysis and cannot provide multi-granular, cross-regional comprehensive analysis, which limits the accurate judgment of the overall industry trend and regional competition. These shortcomings have prompted industrial analysis platforms to urgently adopt more advanced technologies to improve processing capabilities and analysis depth. Summary of the invention
[0003] Based on this, it is necessary to provide a method for constructing an online industry analysis big data platform to solve at least one of the above technical problems.
[0004] To achieve the above purpose, a method for constructing an online industry analysis big data platform is provided, the method comprising the following steps: Step S1: Obtain enterprise information and real-time traffic information; desensitize the enterprise information and real-time traffic information to UGC data, and build an online industry preprocessing data set; Step S2: Generate text semantic analysis data based on the online industry preprocessing data set; perform coordinate deviation correction based on the online industry preprocessing data set to obtain coordinate correction data; perform industry time series prediction and deduction on the text semantic analysis data and the coordinate correction data, and build a personalized model to obtain an online industry personalized preliminary analysis model; Step S3: outputting industry personalized results through the online industry personalized preliminary analysis model; using the industry personalized results to perform enterprise association network analysis to obtain industry graph neural analysis data; using the industry graph neural analysis data to determine the industry model optimization hyperparameters; iterating the model hyperparameters of the online industry personalized preliminary analysis model based on the industry model optimization hyperparameters to obtain the online industry personalized analysis model; Step S4: Generate industry structured data based on the online industry personalized analysis model; use the industry structured data to build an analysis knowledge graph and upload it to the cloud platform for storage.
[0005] The beneficial effect of the present invention is that, by acquiring enterprise information and real-time traffic information and desensitizing UGC data, an online industry preprocessing data set is generated. At the data level, the risk of enterprise privacy leakage is reduced by data desensitization technology, while retaining the effective information of the data, which is helpful for subsequent analysis and modeling. Subsequently, text semantic analysis is used to generate industry-related semantic data, and coordinate deviation correction is performed based on real-time traffic data to improve the accuracy of geographic information. In the process of time series prediction and deduction, the industry personalized preliminary analysis model is further constructed by combining semantic analysis data and corrected geographic data. Through this process, the model has a stronger time and space perception ability and can accurately capture the time series law of enterprise activities. In the stage of industry personalized analysis, enterprise association network analysis is performed through industry personalized results, and graph neural network (GNN) is used to mine potential associations between enterprises. This method effectively reduces the information island effect in industrial network analysis, and enhances the depth and breadth of analysis through network structure optimization. Further using graph neural analysis data to determine the optimized hyperparameters of the industrial model, and implementing hyperparameter iterative optimization, not only improves the stability of the model, but also significantly enhances the accuracy and generalization ability of industrial analysis. Finally, the present invention generates structured data through an industry personalized analysis model to construct a complete analysis knowledge graph. After being stored on the cloud platform, the graph can provide real-time updated industry relationship information for upstream and downstream enterprises in the industrial chain, realizing information sharing and intelligent decision-making support. The dynamic update capability of the knowledge graph ensures the timeliness of the data and enhances the resilience of enterprises in market competition. Therefore, the present invention solves the problems of low data security, insufficient analysis accuracy, and poor model generalization ability in traditional industrial analysis through steps such as industrial data desensitization, time series deduction, graph neural network analysis, and hyperparameter optimization, thereby improving the accuracy of industrial analysis and decision-making efficiency.
[0006] Preferably, step S1 comprises the following steps: Step S11: Obtaining basic enterprise information and real-time traffic information; Step S12: screening the upstream and downstream relationships of the industrial chain for the basic information of the enterprise, and eliminating blank data to obtain the enterprise integrated data; generating commuting time circle analysis data for the real-time traffic information; Step S13: Desensitize the UGC data based on the enterprise integration data and commuting time circle analysis data, and build an online industry preprocessing data set.
[0007] The present invention forms a basic data source by acquiring authoritative information of enterprises and real-time traffic information. The authoritative information of enterprises provides the operating conditions, industrial chain relations and upstream and downstream collaboration data of enterprises to ensure the authenticity and integrity of the data; the real-time traffic information reflects the traffic flow, travel mode and regional commuting characteristics, and provides spatiotemporal dynamic information for subsequent regional industrial analysis. At the data level, by screening the upstream and downstream relations of the industrial chain of the authoritative information of enterprises, redundant and invalid information is effectively eliminated, data noise is reduced, and key data with high value for the analysis of industrial chain relations is retained. The screening process adopts industrial chain relationship mapping and data screening algorithms to ensure that the authoritative integrated data of enterprises has high accuracy and relevance by identifying the transaction chain, supply chain network and collaborative relationship between enterprises. In the process of processing real-time traffic information, the commuting time circle analysis method is used to convert traffic data into related data such as human flow and logistics between enterprises and surrounding areas. By dividing the commuting circle, analyzing the traffic load and travel behavior characteristics in different time periods, and generating commuting time circle analysis data. This data not only reflects the traffic accessibility of the area where the enterprise is located, but also reveals the dynamic changes of regional economic activities, providing temporal support for subsequent regional industrial vitality analysis. In the UGC data desensitization stage, the privacy of user-generated content (UGC) is protected based on the authoritative integrated data of the enterprise and the commuting time circle analysis data. By applying technologies such as differential privacy, hash encryption and data de-identification, the risk of data leakage is effectively avoided while retaining the statistical characteristics and behavior patterns of the data. The desensitized data set retains core information such as the location distribution of enterprises, traffic-related characteristics and industrial chain relationships, while meeting data security and compliance requirements. The online industry pre-processing data set finally constructed realizes the fusion and standardized processing of multi-source data, providing reliable basic data support for subsequent personalized analysis, industry forecasting and optimization.
[0008] Preferably, the UGC data desensitization in step S1 includes: Extract enterprise user data based on enterprise integrated data; Extract user trajectory information data based on commuting time circle analysis data; Generate and anonymize the enterprise user data, where the anonymization key ranges from 128 to 256 bits, to obtain online industry identifier anonymized data; Generalize the user trajectory information data to a regional scope to obtain online regional generalization data; Hash sensitive fields of enterprise user data generation and user trajectory information data to obtain online industry sensitive field data; The online industry identifier anonymized data, online regional generalized data and online industry sensitive field data are fused, and a fused data set is constructed to obtain an online industry preprocessed data set.
[0009] The present invention forms enterprise user data and user trajectory information data through the bidirectional extraction of enterprise authoritative integration data and commuting time circle analysis data. Enterprise user data contains the basic information, business status and upstream and downstream relationships of the enterprise industry chain, which helps to depict the enterprise's industrial ecological network; user trajectory information data is based on traffic travel characteristics, showing the dynamic activities of users in a specific area, reflecting the economic vitality and traffic flow distribution in the area. At the data level, the enterprise user data is protected by identifier anonymization technology, and a 128-256-bit anonymization key is used to ensure the security of the data after encryption. At the same time, the direct identity identification in the data is removed, reducing the risk of privacy leakage. The anonymized online industry identifier data still retains the association characteristics between enterprises, allowing subsequent industrial chain analysis and business relationship mining to continue. In the processing of user trajectory information data, the user location data is fuzzy processed by regional scope generalization technology, and the precise geographic coordinates are converted into regional level location representations to form online regional generalization data. This technology effectively protects user privacy and avoids the exposure of precise locations, while still supporting travel mode analysis and traffic behavior research based on macro-regions. In addition, sensitive fields in enterprise user data and user trajectory information data (such as user identity information, enterprise-specific business data, etc.) are hashed to form online industry sensitive field data. Hashing converts sensitive data through an irreversible encryption algorithm to ensure that even in the event of a data leak, attackers cannot restore the original data. Finally, a unified fusion data set is formed by multi-source data fusion of online industry identifier anonymization data, online regional generalization data, and online industry sensitive field data, namely, the online industry preprocessing data set. Multi-source data fusion uses feature alignment and data standardization technology to solve the problems of data heterogeneity and inconsistent formats, ensuring that information from different data sources can be associated and analyzed in the same data space. The fusion data set provides a complete and secure data foundation for subsequent industry analysis, trend forecasting, and business strategy formulation, while achieving a good balance between data security, privacy protection, and data integrity.
[0010] Preferably, step S2 comprises the following steps: Step S21: Generate text semantic analysis data based on the online industry preprocessing data set; Step S22: Correct the coordinate deviation according to the online industry preprocessing data set. When the GPS coordinate deviation is greater than or equal to 50m, perform POI space correction to obtain coordinate correction data; Step S23: extracting the spatiotemporal features of regional weekly traffic from the text semantic analysis data and the coordinate correction data, and constructing the time series prediction deduction data; Step S24: construct a personalized model based on the time series prediction and deduction data to obtain an online industry personalized preliminary analysis model.
[0011] The present invention gradually forms an online industry personalized preliminary analysis model by performing text semantic analysis, coordinate deviation correction, regional traffic spatiotemporal feature extraction and time series prediction deduction based on the online industry preprocessing data set. At the data level, through the text semantic analysis data generation step, the implicit relationship in text data such as enterprise information, user evaluation, and industrial policy is effectively mined. Through natural language processing technology, complex industry terms and semantic structures are parsed, entity relationships, emotional tendencies and subject information are extracted, and text semantic analysis data with industrial attributes are formed. This process ensures the structured conversion of text data and facilitates subsequent association analysis and model construction. In the coordinate deviation correction stage, the present invention adopts spatial correction technology to ensure the accuracy of location data for errors in GPS coordinate data. When the GPS coordinate deviation is greater than or equal to 50m, the POI (Point of Interest) spatial correction method is used to compare historical positioning data, map benchmark data and traffic node information, and dynamically correct the offset position to generate high-precision coordinate correction data. This process improves the reliability of spatial data and helps to accurately locate the geographical location of enterprises and users in subsequent spatiotemporal feature extraction. In the process of extracting the spatiotemporal features of regional weekly traffic, text semantic parsing data and coordinate correction data are used to perform spatial location association analysis and time series feature extraction. By performing a month-on-month analysis of traffic flow, travel demand and commuting patterns in different periods in the region, abnormal traffic fluctuations and commuting patterns are identified, and comprehensive spatiotemporal feature data are generated. Combined with the time series modeling method, time series prediction and deduction data are further constructed to effectively depict traffic evolution trends and changes in industrial activities, providing dynamic time series support for subsequent personalized analysis. Finally, an online personalized preliminary analysis model for the industry is constructed based on the time series prediction and deduction data. Based on the full use of space, time and text data, the model forms a personalized industry analysis perspective through multi-dimensional feature embedding and dynamic parameter adjustment. The model can accurately deduce the industrial development trends of specific enterprises or regions, providing customized decision support for enterprises.
[0012] Preferably, step S2 comprises the following steps: Step S21: Generate text semantic analysis data based on the online industry preprocessing data set; Step S22: Correct the coordinate deviation according to the online industry preprocessing data set. When the GPS coordinate deviation is greater than or equal to 50m, perform POI space correction to obtain coordinate correction data; Step S23: extracting the spatiotemporal features of regional weekly traffic from the text semantic analysis data and the coordinate correction data, and constructing the time series prediction deduction data; Step S24: construct a personalized model based on the time series prediction and deduction data to obtain an online industry personalized preliminary analysis model.
[0013] The present invention gradually forms an online industry personalized preliminary analysis model through text semantic analysis, coordinate deviation correction, traffic spatiotemporal feature extraction and time series prediction and deduction based on online industry preprocessing data sets. At the data level, the text semantic analysis data generation process uses natural language processing technology to deeply analyze text information such as enterprise data, user evaluation, industry information, and extract key information and industry-related semantic features. Through technical means such as word segmentation, entity recognition, dependency analysis and sentiment classification, text semantic analysis data with contextual relationships is generated. This process converts unstructured text data into structured information, providing a data basis for subsequent spatiotemporal feature extraction and prediction and deduction. In the coordinate deviation correction link, the present invention sets a 50m deviation threshold for the positioning error problem of GPS data and introduces POI (Point of Interest) spatial correction technology. By comparing spatial feature information such as enterprise addresses, road nodes, and transportation hubs, coordinates with significant offsets are corrected to generate coordinate correction data. This process improves the accuracy of geographic location information and ensures the reliability of spatial feature extraction. At the same time, the accuracy optimization of coordinate correction data provides a solid data basis for the analysis of regional traffic spatiotemporal characteristics. In the process of extracting the spatiotemporal features of regional weekly month-on-month traffic, text semantic parsing data and coordinate correction data are further integrated. By constructing a time series model, the traffic flow, travel demand and enterprise commuting data of a specific area at different times are analyzed. The month-on-month analysis method is used to identify short-term trends and cyclical fluctuations in regional traffic changes, and to extract abnormal patterns and key influencing factors in spatiotemporal features. In addition, a time series prediction and deduction model is established by combining historical traffic data and enterprise operation data to speculate on future traffic conditions and industrial activity trends. The deduction data not only provides quantitative information on the development trends of enterprises, but also provides a reference for urban traffic planning and regional industrial policies. Finally, an online personalized preliminary analysis model for the industry is constructed based on the time series prediction and deduction data. Based on the characteristics of multi-source data, the model uses adaptive parameter optimization and multi-dimensional feature fusion technology to form personalized analysis results for different industrial scenarios. The model can dynamically adjust the analysis strategy according to the actual situation of different regions and enterprises, and realize refined industrial trend prediction and business decision support.
[0014] Preferably, the personalized model building described in step S2 includes: Personalized target index is set based on time series forecasting data, where the personalized target index includes macro congestion index and passenger source distribution index; Set the initial model parameters based on the time series forecasting and deduction data; The preset spatiotemporal graph convolutional network is used to analyze the regional-industry correlation of the macro congestion index, and the source distribution index is used to analyze the source POI trajectory to obtain the initial model construction data; The model framework is constructed using the initial model construction data, and the loss function is assigned to obtain a preliminary analysis model for online industry personalization.
[0015] The present invention uses time series prediction and deduction data to perform personalized target index setting, initial parameter setting, regional-industry association analysis and model construction to form an online industry personalized preliminary analysis model. At the data level, firstly, based on the time series prediction and deduction data, time series characteristics such as traffic flow, travel rules, and business district activity are extracted, and personalized target indexes such as macro congestion index and customer source distribution index are set. The macro congestion index quantifies regional traffic pressure by analyzing the traffic congestion conditions in different regions and reflects the load level of the urban transportation network; the customer source distribution index is based on user travel trajectories and consumption behavior data to characterize the customer source flow characteristics of the target area, and provide data support for commercial site selection, customer group positioning and market demand forecasting. In the initial parameter setting stage of the model, the trend characteristics of the time series prediction and deduction data are used to dynamically assign model parameters. Specifically, it includes learning rate, regularization parameter, convolution kernel size, etc., to ensure that the initial state of the model has good convergence and generalization ability. The parameter setting process adopts an adaptive adjustment strategy to optimize the initial parameters according to the fluctuation of the target index, thereby enhancing the adaptability of the model to different regions and industry scenarios. The regional-industry association analysis is implemented through the spatiotemporal graph convolutional network (ST-GCN). The macro congestion index is used as a representation of regional traffic pressure. After the spatial dependency modeling of the graph convolutional network, a spatial correlation graph of regional traffic status is formed. The source distribution index is further used for POI trajectory analysis, and combined with time series features, the activity trajectory of users in a specific area is tracked. Through cross-regional spatiotemporal association analysis, the core areas of industrial activities and the association between upstream and downstream of the industrial chain are identified to obtain the initial model construction data. This graph network-based analysis method effectively captures the nonlinear association in the spatial structure, while reducing noise interference and improving the accuracy of regional industrial relationship identification. In the model framework construction stage, the initial model construction data is used to build a multi-layer network architecture including feature extraction layer, graph convolution layer and time series prediction layer. In the process of assigning the loss function, the source distribution index and the macro congestion index are combined, and a weighted loss function is used to balance the optimization needs of different objectives. This method optimizes the convergence speed and prediction accuracy of the model while taking into account the regional traffic pressure and market demand distribution.
[0016] Preferably, step S3 comprises the following steps: Step S31: outputting industry personalized results through an online industry personalized preliminary analysis model; using the industry personalized results to perform enterprise association network analysis to obtain industry graph neural analysis data; Step S32: using the industry graph neural analysis data to verify the user retention rate and task completion ratio, and determine the industry model optimization hyperparameters; Step S33: Based on the industry model optimization hyperparameters, the model hyperparameters of the online industry personalized preliminary analysis model are iterated to obtain the online industry personalized analysis model.
[0017] The present invention outputs industry personalized results through an online industry personalized preliminary analysis model, and on this basis, conducts enterprise association network analysis, model optimization hyperparameter verification and hyperparameter iteration, thereby constructing an online industry personalized analysis model. At the data level, firstly, multi-dimensional data such as the business characteristics, market performance and user behavior patterns of the enterprise are extracted through the industry personalized results. The results can reflect the market positioning, competitive relationship and potential cooperation opportunities of enterprises in the region. By inputting these data into the enterprise association network analysis module, the business association, supply chain relationship and industry competition pattern between enterprises are modeled using the Graph Neural Network (GNN) to form industry graph neural analysis data. Such data effectively captures the nonlinear association relationship between enterprises and makes up for the shortcomings of traditional linear analysis. In the verification link of user retention rate and task completion ratio, the long-term activity and task completion of users are quantified by tracking and analyzing user behavior data. The user retention rate reflects the market attractiveness of the enterprise's products or services, while the task completion ratio measures the user's interaction depth and goal achievement on the platform. Based on the industry graph neural analysis data, the performance of the enterprise in the user conversion path is further analyzed, and the key factors affecting user retention and task completion are identified from the data level. Through cross-validation of multi-dimensional features, the selection of optimized hyperparameters is ensured to be more robust and adaptable. In the model hyperparameter iteration stage, the verification results are used to dynamically adjust the model's key hyperparameters such as learning rate, regularization parameter, node feature dimension, and number of network layers. During the iteration process, methods such as Bayesian optimization, grid search, or adaptive gradient adjustment are used to ensure the optimal performance of the model in different scenarios. By continuously optimizing the model's parameter configuration, the model's ability to represent complex industrial relationships and its prediction accuracy are improved. The resulting online industry personalized analysis model has higher accuracy and stability. This model can not only provide personalized market insights for enterprises, but also perform real-time optimization and adjustment in a dynamic industrial environment. Its data-driven optimization method effectively avoids the limitations of traditional models that rely on static parameter settings, ensuring that the analysis results have high applicability and decision-making reference value in different scenarios.
[0018] Preferably, the enterprise association network analysis in step S3 includes: The industry-specific output results are used to generate multi-granularity result rendering comprehensive data, where the multi-granularity result rendering comprehensive data includes macro-layer rendering data and meso-layer rendering data; According to the personalized output results of the industry, traffic flow-enterprise density is extracted, and regional raster rendering is performed with a grid size range of 100m-1000m to obtain macro-layer rendering data; Cluster coupling is performed based on the personalized output results of the industry to obtain the enterprise cluster coupling index; the enterprise cluster coupling index is used to couple and associate the personalized output results of the industry, and road network rendering is performed to obtain the meso-level rendering data.
[0019] The present invention generates comprehensive data by rendering multi-granularity results using the personalized output results of the industry, forming macro-layer rendering data and meso-layer rendering data. At the data level, the traffic flow and enterprise density characteristics are first extracted based on the personalized output results of the industry. The traffic flow data includes information such as road traffic conditions, vehicle speed and travel volume distribution in the region, and the enterprise density data is calculated based on the spatial distribution of the enterprise, industry type and market activity intensity. By mapping these data into the spatial coordinate system, a traffic flow-enterprise density feature matrix is formed, and the data is spatially visualized using a regional rasterization rendering method. The grid size range is set between 100m and 1000m to adapt to the analysis requirements at different scales. Small-scale grids are used for refined urban traffic management and enterprise layout optimization, while large-scale grids are suitable for regional economic macro-analysis and industrial policy evaluation. Rasterization rendering not only intuitively displays the spatial relationship between traffic flow and enterprise distribution, but also reveals the traffic pressure and industrial agglomeration degree of urban functional areas, providing a decision-making basis for traffic planning and industrial layout. In the process of generating meso-layer rendering data, the enterprise cluster coupling index is calculated by performing cluster coupling analysis on enterprise clusters. The index quantifies the strength of the internal connection of enterprise clusters by measuring the spatial proximity, industry relevance and degree of synergy between enterprises. At the data level, the calculation of the cluster coupling index adopts multi-dimensional feature fusion, including the geographical location of enterprises, upstream and downstream relationships in the industrial chain, and the degree of supply and demand matching. The enterprise cluster coupling index is further used to conduct coupling correlation analysis on the personalized output results of the industry to construct an enterprise cluster relationship network. On this basis, the spatial connection and transportation network between enterprise clusters are visualized through road network rendering technology. Road network rendering not only reflects the actual commuting and logistics paths between enterprises, but also reveals the degree of dependence of industrial clusters on transportation infrastructure, providing data support for optimizing transportation networks and promoting industrial collaborative development. The resulting macro-layer and meso-layer rendering data have multi-granular spatial expression capabilities. At the macro level, governments and planning agencies can intuitively understand the traffic pressure, enterprise distribution and economic activity intensity in the region, and assist in formulating regional economic development strategies. At the meso level, enterprises and investors can identify potential industrial cooperation opportunities, evaluate the competitive advantages and synergy effects of enterprise clusters, and optimize resource allocation and business strategies.
[0020] Preferably, step S32 includes the following steps: Step S321: Perform user retention rate analysis using industry graph neural analysis data to obtain enterprise user retention rate data; Step S322: performing enterprise supply chain dependency detection based on enterprise user retention rate data to obtain enterprise user-supply chain analysis data; Step S323: Performing completion analysis on the enterprise user-supply chain analysis data and performing ratio verification to obtain enterprise user-supply chain-completion data; Step S324: Generate hyperparameters for enterprise user-supply chain-completion data based on the online industry personalized preliminary analysis model to obtain industry model optimization hyperparameters.
[0021] The present invention forms the industrial model optimization hyperparameters by performing user retention rate analysis, supply chain dependency detection, completion analysis and hyperparameter generation based on the industrial graph neural analysis data. At the data level, the user retention rate is first analyzed using the industrial graph neural analysis data to extract key features such as the active behavior, consumption frequency, and transaction records of enterprise users. By calculating the continuous activity rate and repurchase rate of users in a specific time period, the enterprise user retention rate data is generated. This data effectively reflects the market recognition and user loyalty of the enterprise's products or services, and provides a reliable quantitative basis for enterprises to evaluate user stickiness and market competitiveness. In the enterprise supply chain dependency detection stage, the enterprise user retention rate data is used to analyze the position and dependency of the enterprise in the supply chain. At the data level, supply chain dependency detection is based on multidimensional data features, including the enterprise's raw material procurement volume, order fulfillment rate, logistics time, and transaction intensity of upstream and downstream enterprises. Through the association analysis method, the degree of dependence of the enterprise on a specific supply chain link is quantified, and the existing supply chain vulnerability is identified. The generation of enterprise user-supply chain analysis data not only reveals the role of the enterprise in the supply chain network, but also provides data support for supply chain optimization and risk warning. In the completion analysis and ratio verification stage, the enterprise user-supply chain analysis data is further used to measure the task completion of the enterprise in supply chain collaboration. By analyzing the enterprise's order execution rate, delivery timeliness rate, and production plan achievement rate, the enterprise user-supply chain-completion data is formed. The ratio verification process calculates the ratio of user retention rate to task completion to detect the performance of the enterprise in market demand response and supply chain collaboration. If the ratio fluctuates abnormally, it indicates that the enterprise has problems in supply chain management or customer relationship maintenance. This process provides enterprises with intuitive performance evaluation data, which facilitates timely adjustment of business strategies. In the hyperparameter generation stage, the enterprise user-supply chain-completion data is modeled based on the online industry personalized preliminary analysis model. Through multiple rounds of parameter optimization and adaptive learning, the industry model optimization hyperparameters are generated. These hyperparameters include learning rate, regularization parameter, number of neural network layers, etc., which can dynamically adjust the model's training strategy and prediction accuracy. Through the hyperparameter generation process, the adaptability and generalization ability of the model are significantly improved, so that it can maintain a high analysis accuracy in different industry scenarios.
[0022] Preferably, step S4 comprises the following steps: Step S41: extracting industry structured data according to the online industry personalized analysis model; Step S42: construct an analysis knowledge graph using industry structured data, and set industry confidence parameters to obtain comprehensive industry analysis data; Step S43: Upload the comprehensive industry analysis data to the cloud platform for storage.
[0023] The present invention extracts industry structured data, builds analysis knowledge graphs and sets industry confidence parameters based on an online industry personalized analysis model, and finally forms comprehensive industry analysis data and uploads it to a cloud platform for storage. At the data level, the unstructured data output by the model is first converted and sorted through the industry structured data extraction process to form a standardized data format that is easy to analyze. The extracted data includes multi-dimensional information such as business operation data, user behavior data, industrial chain relationship data and market dynamic data. The integrity and accuracy of the data are ensured by technical means such as feature mapping, entity relationship extraction and anomaly detection. The generation of structured data not only provides high-quality input for subsequent analysis, but also effectively solves the problems of information redundancy and excessive noise in traditional unstructured data. In the knowledge graph construction stage, the extracted structured data is used to generate multiple types of nodes such as enterprise entities, industrial chain nodes, market trend nodes, and an industry analysis knowledge graph is constructed based on graph relationship modeling. At the data level, through relationship reasoning and multi-hop path analysis, the cooperative relationship between enterprises, the supply chain dependency relationship and the market competition pattern are identified. The graph neural network (GNN) is further used for feature embedding and node classification to achieve dynamic tracking of upstream and downstream relationships in the industrial chain. The construction of the knowledge graph not only provides an intuitive view of the enterprise and industrial environment, but also provides data support for causal analysis and trend prediction of complex industrial problems. In the process of setting the industry confidence parameters, the credibility of the data and the reliability of the model output results are measured by calculating the confidence of the nodes and edges in the knowledge graph. At the data level, the calculation of the confidence parameters comprehensively considers factors such as the authority of the data source, the frequency of data updates, and historical verification results. The uncertainty of the industrial analysis results is quantified by using methods such as Bayesian inference and confidence interval estimation. The setting of confidence parameters ensures the reliability of the comprehensive industrial analysis data in decision support and effectively reduces the risk of data noise and misjudgment. The final generated comprehensive industrial analysis data is uploaded to the cloud platform for storage. The cloud platform has a distributed storage architecture and an efficient retrieval mechanism to ensure the security and availability of data. At the data level, the storage process includes data deduplication, partition storage, and multi-copy backup, which further improves the efficiency and reliability of data storage. At the same time, through the computing resources of the cloud platform, users can perform real-time queries, dynamic analysis, and multi-dimensional mining to achieve the wide application of industrial analysis results.
[0024] Preferably, step S42 includes the following steps: Step S421: construct an analytical knowledge graph using industry structured data; Step S422: Setting industry confidence parameters based on the analysis of the knowledge graph; Step S423: Use the industry confidence parameter to perform heat map confirmation on the analysis knowledge graph, and perform comprehensive heat map analysis to obtain comprehensive industry analysis data.
[0025] The present invention constructs an analytical knowledge graph based on industrial structured data, sets industrial confidence parameters and confirms heat maps, and finally forms comprehensive industrial analysis data. At the data level, first, the analytical knowledge graph is constructed using industrial structured data, and information such as enterprises of different sources and types, industrial chain links, and market behaviors are connected through entity relationship diagrams. This process involves modeling the cooperative relationship between enterprises, upstream and downstream supply chain relationships, market trends, and competitive landscapes. Through graph algorithms and natural language processing technology, unstructured information is further converted into analyzable structured data, so that various nodes and edges in the industrial chain are clearly marked in the graph, thereby laying the foundation for subsequent data analysis and decision support. The knowledge graph not only intuitively shows the relationship between each link in the industrial chain, but also provides in-depth multi-dimensional analysis for complex industrial problems. Next, the industry confidence parameters are set based on the constructed knowledge graph. The confidence parameter setting involves credibility assessment of each node and edge in the industrial graph at the data level. The confidence value of each node is generated by weighted calculation of factors such as the weight of the data source, historical verification, and the accuracy of the data source. At the data level, the calculation of confidence uses methods such as Bayesian inference and weighted average to ensure the reliability of the data and the accuracy of the model output results. The setting of industry confidence parameters not only provides data reliability assessment for subsequent analysis, but also effectively reduces errors in the model reasoning process and improves the accuracy of prediction results. In the heat map confirmation stage, the industry confidence parameters are used to generate heat maps for the analysis knowledge map. The heat map visualizes the key characteristics of different regions, enterprises and market performance in the industrial chain through spatial distribution analysis, making the changes in industrial behavior and trends more intuitive. Through comprehensive analysis of heat maps, key nodes, bottleneck areas and potential risk points in the industrial chain can be identified. At the data level, the generation of heat maps relies on cluster analysis, numerical processing and application of heat mapping algorithms for industrial structured data, so that complex data can be represented by color gradients, helping decision makers quickly identify high-risk areas and priority development areas. Finally, the comprehensive data of industrial analysis obtained provides data-driven decision support for industrial optimization, adjustment and prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A schematic diagram of the steps of a method for building an online industry analysis big data platform; Figure 2 for Figure 1 Detailed implementation steps of step S3 in FIG. The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0027] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.
[0028] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different network and / or processor methods and / or microcontroller methods.
[0029] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0030] To achieve this, please refer to Figure 1 to Figure 2 , a method for constructing an online industry analysis big data platform, the method comprising the following steps: Step S1: Obtain enterprise information and real-time traffic information; desensitize the enterprise information and real-time traffic information to UGC data, and build an online industry preprocessing data set; Step S2: Generate text semantic analysis data based on the online industry preprocessing data set; perform coordinate deviation correction based on the online industry preprocessing data set to obtain coordinate correction data; perform industry time series prediction and deduction on the text semantic analysis data and the coordinate correction data, and build a personalized model to obtain an online industry personalized preliminary analysis model; Step S3: outputting industry personalized results through the online industry personalized preliminary analysis model; using the industry personalized results to perform enterprise association network analysis to obtain industry graph neural analysis data; using the industry graph neural analysis data to determine the industry model optimization hyperparameters; iterating the model hyperparameters of the online industry personalized preliminary analysis model based on the industry model optimization hyperparameters to obtain the online industry personalized analysis model; Step S4: Generate industry structured data based on the online industry personalized analysis model; use the industry structured data to build an analysis knowledge graph and upload it to the cloud platform for storage.
[0031] In the embodiment of the present invention, reference Figure 1 FIG. 1 is a schematic diagram of a step flow of a method for constructing an online industry analysis big data platform of the present invention. In this example, the method for constructing an online industry analysis big data platform includes the following steps: In the embodiment of the present invention, the enterprise information and real-time traffic information are obtained. This process requires real-time data collection through multiple data sources, and the enterprise-related business data, location data and real-time traffic flow information are obtained through data interfaces and crawler technology. These data usually come from public data platforms, government statistics bureaus, traffic management departments, and APIs released by enterprises. The acquired data are mostly unstructured or semi-structured data, which need to be converted into a unified format through cleaning, formatting and structuring to facilitate subsequent data analysis and mining. Secondly, the UGC data desensitization of enterprise information and real-time traffic information is to protect personal privacy and sensitive information of enterprises and avoid data leakage. Specific technical means include desensitizing sensitive information such as personal identifiers, contact information, addresses, etc. in user-generated content (UGC). Desensitization methods include data encryption, anonymization, pseudonymization technology, etc., which replace or encrypt sensitive fields in the data so that they cannot be directly identified or traced back to specific individuals or enterprises. These desensitization processes ensure that the data can still maintain a certain use value under the premise of complying with privacy protection regulations. After the data is processed, an online industry preprocessing data set is constructed. This process relies on standardization, normalization and deduplication of different types of data to ensure data quality. Data preprocessing also involves steps such as timestamp alignment, missing value processing, and noise removal, ultimately generating a structured, cleaned online industry preprocessing dataset. This dataset provides high-quality data input for subsequent industry analysis, model building, and decision support, ensuring data integrity, accuracy, and timeliness.
[0032] Step S2: Generate text semantic analysis data based on the online industry preprocessing data set; perform coordinate deviation correction based on the online industry preprocessing data set to obtain coordinate correction data; perform industry time series prediction and deduction on the text semantic analysis data and the coordinate correction data, and build a personalized model to obtain an online industry personalized preliminary analysis model; In an embodiment of the present invention, text semantic analysis data is generated. This process uses natural language processing (NLP) technology to perform semantic analysis on text data in online industry preprocessing data sets. Specifically, by performing word segmentation, part-of-speech tagging, named entity recognition (NER) and other processing on text data such as enterprise information, industry dynamics, market descriptions, etc., key information in the text, such as industry terms, company names, product categories, etc., is extracted, and deep semantic relationships in the text are obtained through semantic analysis models (such as BERT, Word2Vec, etc.). This technical means can convert unstructured text data into structured information and provide rich feature data support for subsequent industry trend analysis, enterprise relationship modeling, etc. Secondly, coordinate deviation correction is performed according to the online industry preprocessing data set to obtain coordinate correction data. Real-time traffic information contains deviations in GPS coordinates, especially in urban environments, which can cause GPS positioning errors due to factors such as buildings and high-density signal interference. In order to ensure the accuracy of the data, coordinate deviation correction technology, such as the POI (Point of Interest) correction algorithm of spatial data, is used to correct the position deviation in GPS data. This process is usually based on existing geographic information system (GIS) data, matched with accurate map data, and applied certain error correction models (such as Kalman filtering or least squares method) to adjust GPS coordinates to ensure the accuracy of data in spatial dimension, thereby generating coordinate correction data. Finally, the text semantic analysis data and coordinate correction data are used for industry time series forecasting and deduction, and a personalized model is constructed to obtain an online industry personalized preliminary analysis model. In the data fusion stage, different types of data need to be merged to form a multivariate data set containing time and space dimensions. Then, time series forecasting models (such as long short-term memory network LSTM or autoregressive model ARIMA) are used to predict industry development trends, traffic flow changes, market demand fluctuations, etc. This process predicts industry changes and corporate dynamics that will occur in a certain period of time in the future by learning from historical data. At the same time, combined with personalized analysis needs, machine learning methods (such as decision trees, support vector machines or neural networks) are used to model the operation data and market behaviors of different enterprises in a personalized way, thereby generating personalized analysis models that adapt to the characteristics of specific enterprises or industries.
[0033] Step S3: outputting industry personalized results through the online industry personalized preliminary analysis model; using the industry personalized results to perform enterprise association network analysis to obtain industry graph neural analysis data; using the industry graph neural analysis data to determine the industry model optimization hyperparameters; iterating the model hyperparameters of the online industry personalized preliminary analysis model based on the industry model optimization hyperparameters to obtain the online industry personalized analysis model; In an embodiment of the present invention, an industry personalized result is output through an online industry personalized preliminary analysis model. This process uses a previously constructed personalized analysis model to analyze the data of enterprises and industries, thereby generating personalized conclusions and prediction results for specific industries or enterprises. The analysis model is usually based on machine learning or deep learning methods, such as support vector machines (SVM), random forests (RF) or neural networks (NN), which output personalized results in multiple dimensions such as industry development trends, enterprise performance, and market demand through patterns learned from industry data. Next, the industry personalized results are used to perform enterprise association network analysis to obtain industry graph neural analysis data. This process uses graph neural network (GNN) technology to model the association between industries and enterprises. Graph neural network is a deep learning method that can effectively process graph structure data. By treating each enterprise in the industry as a node in the graph, and the cooperation, competition or supply and demand relationship between enterprises as edges between nodes, the graph neural network can capture the complex dependencies between enterprises and their dynamic changes. By learning these relationships, the graph neural network can derive a more accurate industry network structure, and then generate industry graph neural analysis data, which reflects information such as interactions within the industry and between enterprises, cooperation models, and risk propagation paths. Then, the industry graph neural analysis data is used to determine the industry model optimization hyperparameters. In the model optimization stage, the industry graph neural analysis data is analyzed and the hyperparameter optimization algorithms (such as Bayesian optimization, grid search or genetic algorithm) are used to adjust the hyperparameters of the industry analysis model. These hyperparameters are crucial to the performance of the model. Reasonable hyperparameter settings can significantly improve the prediction accuracy and generalization ability of the model. In this process, the optimization objectives (such as accuracy, F1 value, etc.) are first defined, and the model parameters are automatically adjusted according to the features extracted from the graph neural network and the prediction results of the model to achieve the best performance. Finally, the model hyperparameters of the online industry personalized preliminary analysis model are iterated based on the industry model optimization hyperparameters to obtain the online industry personalized analysis model. This step relies on optimization techniques such as backpropagation algorithm and gradient descent. By continuously adjusting the model hyperparameters, iterative training is performed until the output of the model reaches the expected performance level. Through this iterative process, the accuracy, robustness and adaptability of the model are improved, thereby providing more refined and personalized prediction results for industry analysis.
[0034] Step S4: Generate industry structured data based on the online industry personalized analysis model; use the industry structured data to build an analysis knowledge graph and upload it to the cloud platform for storage.
[0035] In an embodiment of the present invention, industry structured data is generated according to an online industry personalized analysis model. The core of this process is to convert the output data of the personalized analysis model into a structured data format, usually by feature extraction and data collation methods, to convert the original industry analysis results into a more standardized and easy-to-analyze data structure. At this time, the structured data not only contains the basic information of each enterprise and industry, but also includes multi-dimensional numerical or categorical data such as market trends, industrial chain nodes, product demand, consumer behavior, etc. These data are pre-processed by means of data cleaning, data normalization and standardization to ensure that they are suitable for subsequent deep analysis and graph construction. Next, the industry structured data is used to construct an analytical knowledge graph. The technical means of knowledge graph construction usually rely on graph database technology and natural language processing (NLP) technology, and establish a relationship network between entities such as enterprises, products, services, and markets by processing the industry data with semantic parsing, entity recognition, and relationship extraction. In the graph, each entity (such as an enterprise, product, etc.) is represented as a node in the graph, and the relationship between entities (such as supply chain relationships, cooperative relationships, etc.) is represented as an edge, thereby forming a complex network containing multiple information. Through the knowledge graph, various knowledge, relationships and laws in the industry field can be effectively captured, so that the industrial data not only has a structured form of expression, but also can provide semantic support and reasoning capabilities for complex business problems. Finally, the constructed industry analysis knowledge graph is uploaded to the cloud platform for storage. This step relies on cloud storage technology and distributed computing frameworks. It usually uses large-scale data storage solutions provided by cloud service platforms (such as AWS, Azure or Google Cloud) to store knowledge graphs in the cloud in an efficient and scalable manner. The cloud platform not only supports large-scale data storage, but also provides convenient data access, query and sharing services. Through cloud storage, enterprises or research institutions can access and use knowledge graphs at any time, and combine cloud computing resources for further computational analysis and model training.
[0036] Preferably, step S1 comprises the following steps: Step S11: Obtaining basic enterprise information and real-time traffic information; Step S12: screening the upstream and downstream relationships of the industrial chain for the basic information of the enterprise, and eliminating blank data to obtain the enterprise integrated data; generating commuting time circle analysis data for the real-time traffic information; Step S13: Desensitize the UGC data based on the enterprise integration data and commuting time circle analysis data, and build an online industry preprocessing data set.
[0037] In the embodiment of the present invention, the collection of enterprise data and traffic data is realized by obtaining authoritative information of enterprises and real-time traffic information. Authoritative information of enterprises is usually obtained through official databases, official websites of enterprises or industry reports. These data include basic information of enterprises, roles in the industrial chain, market behavior, etc.; real-time traffic information is obtained through means such as traffic monitoring systems, GPS data, and traffic flow monitoring, reflecting the dynamic changes of the traffic system, such as traffic flow, road conditions, commuting patterns, etc. The data collection stage must ensure the comprehensiveness, timeliness and accuracy of the data for subsequent processing and analysis. The purpose of screening the upstream and downstream relationships of the industrial chain for authoritative information of enterprises is to identify enterprises related to the target analysis from a large amount of enterprise data and eliminate information that is irrelevant or invalid to the analysis. The screening of upstream and downstream relationships of the industrial chain usually adopts the graph analysis technology in graph theory, and determines its upstream and downstream roles in the industrial chain by analyzing the position and relationship of enterprises in the supply chain and value chain. Through this screening, blank data is removed, the accuracy and completeness of the data set are ensured, and more accurate authoritative integrated data of enterprises are obtained. At the same time, the commuting time circle analysis data of real-time traffic information is generated. The traffic flow pattern in a specific area is analyzed by using geographic information system (GIS) and spatiotemporal data analysis technology, and the commuting time circle is drawn, thereby providing spatiotemporal background information for subsequent traffic pattern analysis and industry decision-making. The UGC (user-generated content) data is desensitized based on the authoritative integrated data of the enterprise and the commuting time circle analysis data. The main purpose of UGC data desensitization is to protect personal privacy. Encryption technology, anonymization technology or data perturbation technology are usually used to process sensitive information. Specifically, sensitive information of enterprises and individuals, such as addresses, contact information, etc., can be replaced or masked to make it untraceable. Then, the desensitized data is used to construct an online industry preprocessing data set, which will integrate enterprise information and traffic data to provide data support for subsequent industry analysis. This process includes data cleaning, noise removal, standardization, etc., to ensure that the data can still be used for effective analysis while maintaining data privacy.
[0038] Preferably, the UGC data desensitization in step S1 includes: Extract enterprise user data based on enterprise integrated data; Extract user trajectory information data based on commuting time circle analysis data; Generate and anonymize the enterprise user data, where the anonymization key ranges from 128 to 256 bits, to obtain online industry identifier anonymized data; Generalize the user trajectory information data to a regional scope to obtain online regional generalization data; Hash sensitive fields of enterprise user data generation and user trajectory information data to obtain online industry sensitive field data; The online industry identifier anonymized data, online regional generalized data and online industry sensitive field data are fused, and a fused data set is constructed to obtain an online industry preprocessed data set.
[0039] In the embodiment of the present invention, the enterprise user data is extracted on the basis of the authoritative integration data of the enterprise. First, the key features related to the enterprise user need to be extracted from the original data through data screening and feature extraction methods. The data is usually screened by feature selection algorithms or data cleaning methods to ensure the accuracy and relevance of the data. This process relies on traditional data processing technologies, such as data integration, deduplication, missing value filling, etc., to ensure that the extracted enterprise user data has high quality and high credibility. Secondly, the commuting time circle analysis data analyzes the user's movement trajectory through geographic information system (GIS) technology, and uses spatiotemporal analysis methods to extract the user's movement path and activity range from traffic data to construct user trajectory information data. In this process, by analyzing geographic location data, timestamps and traffic flow data, a commuting time circle can be established, thereby providing spatiotemporal data support for user behavior for subsequent analysis. In order to protect user privacy and comply with data protection specifications, anonymization technology is adopted. After the enterprise user data is generated, a 128-256-bit key is used to anonymize the identifier. By encrypting the identifier, it is ensured that the data can be analyzed without leaking the user's identity. Symmetric encryption or hash encryption algorithms are often used to anonymize identifiers. The choice of key range depends on the needs of data privacy protection and the specific implementation environment. Next, user trajectory information data is generalized in the regional scope. Generalization technology is a common technology in data privacy protection. It mainly converts precise location coordinates into regional information of a larger range through fuzzification operations, thereby reducing the risk of location privacy leakage. This operation is usually achieved through clustering algorithms or block division technologies in GIS. Furthermore, sensitive fields (such as location, behavior data, etc.) in enterprise user data and user trajectory information data are processed using hashing technology. Hashing technology generates a fixed-length hash value by encrypting and hashing sensitive fields, thereby ensuring the privacy and integrity of the data. Hashing technology is unidirectional, that is, the original data is irreversibly converted into a hash value, which increases the security of the data. Through these methods, the anonymity and non-traceability of online industry sensitive field data are ensured. Finally, through multi-source data fusion technology, the online industry identifier anonymization data, online regional generalization data, and online industry sensitive field data are integrated. In this process, data fusion technology involves the merging of multidimensional data, usually through data alignment, matching and standardization, to achieve the association between different data sources and generate a unified data set. The fused data set is the online industry pre-processing data set, which provides a secure, privacy-protected and well-structured data foundation for subsequent analysis.
[0040] Preferably, step S2 comprises the following steps: Step S21: Generate text semantic analysis data based on the online industry preprocessing data set; Step S22: Correct the coordinate deviation according to the online industry preprocessing data set. When the GPS coordinate deviation is greater than or equal to 50m, perform POI space correction to obtain coordinate correction data; Step S23: extracting the spatiotemporal features of regional weekly traffic from the text semantic analysis data and the coordinate correction data, and constructing the time series prediction deduction data; Step S24: construct a personalized model based on the time series prediction and deduction data to obtain an online industry personalized preliminary analysis model.
[0041] In an embodiment of the present invention, the text data is split into independent words or phrases by word segmentation technology; then, each word is assigned a grammatical role (such as noun, verb, etc.) by part-of-speech tagging (POS Tagging); next, the text is converted into a vector form using word vector technology (such as Word2Vec, GloVe or BERT) so that the machine can understand the semantic relationship of the text. These processed vector representations can capture the grammatical, semantic and contextual information of the text and provide basic data for subsequent analysis tasks. For example, the text classification model will classify different types of industrial texts based on these vectors to provide guidance for further processing of industrial data. The "coordinate deviation correction" and "POI spatial correction" technologies mentioned involve spatial data processing in a geographic information system (GIS). GPS coordinates are often affected by errors, especially in urban environments, where buildings and other obstacles can cause signal deviations. In order to improve the position accuracy, it is first necessary to perform spatial correction by comparing real-time GPS data with known point of interest (POI) data. In the process of spatial correction, the nearest neighbor interpolation algorithm, K-nearest neighbor algorithm (KNN) and other methods are usually used to adjust the GPS coordinates by finding the closest known location data in space. Furthermore, the coordinate accuracy can be improved by fusing multi-source data (such as mobile base station information, Wi-Fi hotspots, etc.), so as to obtain more reliable geographic location information, which is crucial for accurately analyzing user behavior and traffic patterns. When performing "regional weekly traffic spatiotemporal feature extraction" and "time series prediction deduction data construction", spatiotemporal data analysis and time series prediction technology are involved. First of all, the core of traffic spatiotemporal feature extraction is to extract meaningful spatiotemporal features from large-scale traffic flow data. This can be done by using spatiotemporal clustering analysis (such as DBSCAN algorithm) to identify traffic hotspots and patterns in different regions in different time periods, or by mining spatiotemporal association rules (such as Apriori algorithm) to find traffic flow patterns in specific regions and time periods. These extracted features provide a data basis for the time series prediction model. Time series prediction uses algorithms such as ARIMA, LSTM (Long Short-Term Memory Network), Prophet, etc. to predict future traffic flow or congestion based on the time series characteristics of historical traffic data. Deep learning models such as LSTM can process long time series data and capture long-term dependencies by memorizing historical states, thereby providing more accurate future trend predictions. "Personalized model building" is to build models for individuals or specific industries based on time series forecasting and deduction data through machine learning and deep learning methods. At this stage, the model building technology can include traditional machine learning methods such as support vector machine (SVM), decision tree, random forest, or deep learning methods such as convolutional neural network (CNN), recursive neural network (RNN), etc.Through training data, these models can automatically learn personalized demand patterns from time series data and predict future trends based on historical data. The construction of personalized models aims to optimize resource allocation, traffic forecasts, etc. according to the specific needs of industries or users, so that the models can be dynamically adjusted according to actual conditions, and then output online personalized preliminary analysis models to provide support for further analysis.
[0042] Preferably, the personalized model building described in step S2 includes: Personalized target index is set based on time series forecasting data, where the personalized target index includes macro congestion index and passenger source distribution index; Set the initial model parameters based on the time series forecasting and deduction data; The preset spatiotemporal graph convolutional network is used to analyze the regional-industry correlation of the macro congestion index, and the source distribution index is used to analyze the source POI trajectory to obtain the initial model construction data; The model framework is constructed using the initial model construction data, and the loss function is assigned to obtain a preliminary analysis model for online industry personalization.
[0043] In an embodiment of the present invention, the personalized target index is set based on the time series prediction and deduction data, and the personalized target index includes a macro congestion index and a passenger source distribution index. In order to generate these target indexes, it is necessary to model the historical data of traffic and passenger flow through time series analysis techniques, such as ARIMA (autoregressive integrated moving average) model or LSTM (long short-term memory network), to capture its long-term and short-term trend changes. The macro congestion index reflects the degree of traffic congestion in a certain period of time by modeling the time series of data such as traffic flow, road capacity and vehicle speed; while the passenger source distribution index obtains the spatial distribution characteristics of the passenger source by analyzing data such as passenger flow and user activity trajectory in different time periods and different regions, thereby providing personalized prediction results for subsequent analysis. In the initial parameter setting stage of the model, the data used mainly comes from the time series prediction and deduction data, involving model parameter adjustment technology. The initial parameter setting of the model can be performed by methods such as Bayesian optimization or grid search, which can automatically adjust the hyperparameters of the model according to the goal of minimizing the prediction error, so that the model can better adapt to the actual industry and traffic environment. The application of spatiotemporal graph convolutional network (ST-GCN) technology further enhances the model's spatial-temporal modeling capabilities. ST-GCN can take spatiotemporal data as input, use graph convolutional network (GCN) to process graph structure data, and combine the time dimension to model the evolution of nodes in space, and then conduct regional-industry association analysis. Through the spatiotemporal graph convolutional network, the model can deeply explore the relationship between traffic congestion and industrial development between different regions, and on this basis, conduct dynamic traffic flow prediction and optimization. In addition, the customer source distribution index conducts detailed spatial analysis through POI (point of interest) trajectory analysis, and uses trajectory data analysis techniques such as dynamic time warping (DTW) and cluster analysis to identify the activity patterns of users in different time periods and spatial locations. Through these analyses, the precise trajectory of the customer source can be obtained, further promoting the in-depth personalized analysis. Finally, in the model framework construction stage, all the data initially constructed are iteratively optimized through the construction of the model framework and the assignment of loss functions. The setting of the loss function usually depends on the target to be optimized, such as mean square error (MSE) or cross entropy, which is used to evaluate the error between the model and the actual data during the training process, thereby guiding the adjustment of model parameters. Through this process, an online industry personalized preliminary analysis model is finally formed, which can provide data support for further online industry analysis and decision-making.
[0044] As an example of the present invention, refer to Figure 2 As shown, in this example, step S3 includes: Step S31: outputting industry personalized results through an online industry personalized preliminary analysis model; using the industry personalized results to perform enterprise association network analysis to obtain industry graph neural analysis data; Step S32: using the industry graph neural analysis data to verify the user retention rate and task completion ratio, and determine the industry model optimization hyperparameters; Step S33: Based on the industry model optimization hyperparameters, the model hyperparameters of the online industry personalized preliminary analysis model are iterated to obtain the online industry personalized analysis model.
[0045] In the embodiment of the present invention, outputting the industry personalized results through the online industry personalized preliminary analysis model means that data prediction and analysis are performed according to the previously trained personalized model to generate personalized features for a specific industry. In this process, the data completes the task of extracting key indicators from a large data set through steps such as spatiotemporal prediction, feature extraction, and target index calculation of the model. Then, the industry personalized results are used for enterprise association network analysis to build a graph structure based on the mutual connection and interaction between enterprises. At this time, the relationship between enterprises is represented as nodes and edges in the graph, and the graph neural network (GNN) is used as an effective graph structure data processing technology to perform in-depth analysis and reasoning using the information of nodes and edges. GNN enables the model to capture the complex relationship patterns between enterprises through the propagation of neighborhood information, thereby obtaining industry graph neural analysis data. The industry graph neural analysis data is used to verify the user retention rate and task completion ratio. User retention rate analysis usually involves tracking the user's continuous participation behavior on the platform, and using time series analysis, regression analysis, or survival analysis methods to evaluate the user's activity level and loss risk over a period of time. The task completion ratio verification focuses on the completion of the specified task by the user. It combines the task completion records in the data set to calculate the task completion rate of the user in a certain period, and compares the completion of different tasks to evaluate the service efficiency and user experience of the platform. Through these verifications, the key factors affecting the accuracy of the model can be identified and the basis for optimizing the hyperparameters can be provided. The mentioned industrial model optimization hyperparameters are to adjust the hyperparameters of the existing model to improve the accuracy and generalization ability of the model. The optimal hyperparameter combination is found through optimization techniques such as grid search, Bayesian optimization or genetic algorithm. The optimization of hyperparameters can ensure the optimal performance in practical applications by maximizing the performance of the model on the validation set. This process includes the adjustment of parameters such as learning rate, regularization term, hidden layer size, etc. After completing the hyperparameter optimization, based on the optimized hyperparameters, the model enters the iteration stage, that is, further training and adjustment are carried out, and finally an optimized online industry personalized analysis model is generated to ensure that it can perform efficient and accurate prediction and analysis according to the new data environment.
[0046] Preferably, the enterprise association network analysis in step S3 includes: The industry-specific output results are used to generate multi-granularity result rendering comprehensive data, where the multi-granularity result rendering comprehensive data includes macro-layer rendering data and meso-layer rendering data; According to the personalized output results of the industry, traffic flow-enterprise density is extracted, and regional raster rendering is performed with a grid size range of 100m-1000m to obtain macro-layer rendering data; Cluster coupling is performed based on the personalized output results of the industry to obtain the enterprise cluster coupling index; the enterprise cluster coupling index is used to couple and associate the personalized output results of the industry, and road network rendering is performed to obtain the meso-level rendering data.
[0047] In the embodiment of the present invention, the personalized output results of the industry are used to generate multi-granularity results to render comprehensive data, which is intended to present the industry characteristics in detail through data analysis and rendering at different levels. The macro-layer rendering data and the meso-layer rendering data correspond to the large-scale and medium-scale spatial and temporal characteristics, respectively. These rendering data provide in-depth analysis of industry operation, traffic flow, enterprise distribution, etc. in the macro and meso dimensions, respectively, and can fully present multi-level information of the industrial ecology. Next, based on the personalized output results of the industry, traffic flow-enterprise density extraction is performed, and spatial analysis technology is used to analyze the geographic information system (GIS) data to quantitatively extract traffic flow and enterprise density. In this process, it is first necessary to calculate the enterprise density and traffic flow per unit area through traffic flow data and the spatial distribution data of enterprises to reflect the industrial activity and traffic pressure in different regions. These data will then enter the regional rasterization rendering step, that is, the entire area is divided into small grids by the rasterization method, and the corresponding flow and density information is calculated in each grid. The selection of grid size between 100m and 1000m can be flexibly adjusted at different scales according to the required analysis accuracy, thereby generating a spatial distribution map of traffic flow and enterprise density at the macro level. In further processing, the enterprise cluster coupling index is generated by combining the spatial and operational relationships between enterprises through cluster coupling analysis. This index reflects the degree of clustering of enterprises in geographical space and their interactive relationship in the industrial chain. The generation of the enterprise cluster coupling index involves clustering enterprises according to their geographical location and business relationship, evaluating the closeness of their interaction with surrounding enterprises, and quantifying this interactive relationship. Through this index, the formation and development trend of industrial clusters can be effectively depicted. Finally, the enterprise cluster coupling index is used to conduct coupling correlation analysis on the personalized output results of the industry, and road network rendering is performed. This step uses road network data combined with industry distribution data for spatial analysis and rendering to reveal the relationship between industrial activities and traffic road networks. Road network rendering can be performed in different ways, such as line width changes based on flow, traffic density heat maps, etc., to show the traffic conditions of different sections and their relationship with enterprise clusters.
[0048] Preferably, step S32 includes the following steps: Step S321: Perform user retention rate analysis using industry graph neural analysis data to obtain enterprise user retention rate data; Step S322: performing enterprise supply chain dependency detection based on enterprise user retention rate data to obtain enterprise user-supply chain analysis data; Step S323: Performing completion analysis on the enterprise user-supply chain analysis data and performing ratio verification to obtain enterprise user-supply chain-completion data; Step S324: Generate hyperparameters for enterprise user-supply chain-completion data based on the online industry personalized preliminary analysis model to obtain industry model optimization hyperparameters.
[0049] In the embodiment of the present invention, the "user retention rate analysis using industrial graph neural analysis data" mentioned mainly uses the graph neural network (GNN) technology to analyze the relationship between enterprises and their users in the industrial chain, and then calculate the retention of users in different time periods. Graph neural networks can effectively capture the relationship and interaction mode between users, and analyze the user retention rate in a given period through the modeling of nodes (enterprises or users) and edges (interaction relationships), reflecting the ability of different enterprises to attract and retain users in the industrial chain. "Enterprise supply chain dependency detection based on enterprise user retention rate data" is to further analyze the supply chain dependency relationship between enterprises on the basis of user retention rate. Here, network analysis and supply chain models are mainly used to reflect the demand and supply relationship between enterprises through user retention rate data, and generate interactive analysis data between enterprise users and supply chains. These data help to identify which enterprises have higher dependence in the supply chain and which enterprises are key nodes in the chain. "Enterprise user-supply chain analysis data is analyzed for completion, and ratio verification is performed to obtain enterprise user-supply chain-completion data" involves in-depth processing of the analysis results, using statistical methods for completion analysis, and testing the performance and task completion of enterprises in the supply chain. Through ratio verification, the accuracy and rationality of the data are verified to ensure that the obtained data can reflect the actual business situation and provide reliable data for the next step of optimization. "Generate hyperparameters for enterprise user-supply chain-completion data based on the online industry personalized preliminary analysis model" is to optimize the data obtained from the previous analysis through regression models, machine learning or deep learning algorithms. Hyperparameter generation is to determine the best parameter combination of the model through experiments and tuning, using methods such as cross-validation, so as to improve the accuracy and generalization ability of model predictions, and then form industry model optimization hyperparameters to ensure that the online industry personalized analysis model can achieve optimal performance in actual applications.
[0050] Preferably, step S4 comprises the following steps: Step S41: extracting industry structured data according to the online industry personalized analysis model; Step S42: construct an analysis knowledge graph using industry structured data, and set industry confidence parameters to obtain comprehensive industry analysis data; Step S43: Upload the comprehensive industry analysis data to the cloud platform for storage.
[0051] In the embodiment of the present invention, "extracting industry structured data according to the online industry personalized analysis model" involves using the industry personalized analysis model to perform structured processing on the original data. This process converts the enterprise-related unstructured data into structured data through multi-level data mining and information processing methods, such as natural language processing (NLP) and data cleaning technology. This process is not only a format conversion of the data, but more importantly, it identifies the key features in the data and maps them to a predefined data structure for further analysis and processing. "Using industry structured data to build an analysis knowledge graph, and setting industry confidence parameters to obtain comprehensive industry analysis data." This process uses graph database technology and knowledge graph construction methods. The knowledge graph represents industry data by constructing a graph structure of entities (such as enterprises, products, industries, etc.) and their relationships. Using industry structured data, according to preset relationship rules and semantic models, a multi-dimensional and multi-level knowledge graph is constructed, thereby revealing the mutual relationship and influence path between different entities in the industry. At the same time, the setting of industry confidence parameters is the process of evaluating the credibility of the constructed knowledge graph, which aims to provide a reliable basis for judgment for subsequent analysis and ensure that the data has high credibility in the reasoning and decision-making process. This process generally requires machine learning or statistical models to evaluate and adjust parameters to make the information in the graph more accurate. "Upload the comprehensive data of industry analysis to the cloud platform for storage". This step mainly uses cloud computing technology to upload the data processed by industry structured data and knowledge graph to the cloud platform for long-term storage and management. The cloud platform provides efficient storage and computing capabilities, which can support the processing and analysis of large-scale data. Through the storage of the cloud platform, the comprehensive data of industry analysis can be efficiently backed up and accessed, and can also share and integrate data with other systems, providing basic data support for subsequent data analysis, decision support and model iteration.
[0052] Preferably, step S42 includes the following steps: Step S421: construct an analytical knowledge graph using industry structured data; Step S422: Setting industry confidence parameters based on the analysis of the knowledge graph; Step S423: Use the industry confidence parameter to perform heat map confirmation on the analysis knowledge graph, and perform comprehensive heat map analysis to obtain comprehensive industry analysis data.
[0053] In the embodiment of the present invention, the industry structured data is used to construct an analytical knowledge graph. This process relies on graph databases and knowledge graph technology. Through in-depth mining of industry-related data, multi-dimensional entities such as enterprises, products, services, markets and their relationships (such as cooperation, dependence, competition, etc.) are converted into graphs for expression. In the process of data structuring, the original data (such as text, numerical values, geographic information, etc.) are processed by natural language processing (NLP), entity recognition and relationship extraction, and mapped into nodes and edges with clear hierarchy and association. This knowledge graph can reveal the complex connections between various elements in the industry and provide visual basic data for further analysis. In the next step S422, the industry confidence parameter is set based on the analytical knowledge graph. This operation means quantitatively evaluating the information quality in the knowledge graph. The setting of confidence parameters usually relies on machine learning algorithms, such as Bayesian reasoning or other statistical models, to determine the credibility of each node and edge in the graph. This process weights the relationship between nodes in the knowledge graph by evaluating historical data, model errors or feedback from domain experts to improve the accuracy and effectiveness of the graph. Finally, in step S423, the industry confidence parameter is used to confirm the analysis knowledge graph with a heat map, and a comprehensive analysis of the heat map is performed to obtain comprehensive industry analysis data. This technical means combines the heat map generation and analysis methods, and uses confidence parameters to visualize the strength of each node and relationship in the industry knowledge graph in the form of a heat map. In this process, hot spots and high-confidence connections will be marked as darker parts, reflecting the most critical links and influential factors in the industry. By comprehensively analyzing the heat map data, it is possible to evaluate the industry structure, development trends and potential risks as a whole, and further form actionable comprehensive industry analysis data to provide in-depth insights for decision support systems.
[0054] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.
[0055] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.
Claims
1. A method for constructing an online industry analysis big data platform, characterized in that: The following steps are involved: Step S1: Obtaining basic enterprise information and real-time traffic information; Desensitize UGC data of enterprise basic information and real-time traffic information, and build an online industry pre-processing data set; Step S2: performing text semantic analysis based on the online industry preprocessing data set to generate text semantic data; performing coordinate deviation correction based on the online industry preprocessing data set to obtain coordinate correction data; The text semantic data and coordinate correction data are used to perform industry time series forecasting and deduction, and a personalized model is constructed to obtain an online industry personalized preliminary analysis model; Step S3: outputting industry personalized results through the online industry personalized preliminary analysis model; Use the industry personalized results to conduct enterprise association network analysis and obtain industry graph neural analysis data; Use industry graph neural analysis data to determine industry model optimization hyperparameters; Based on the industry model optimization hyperparameters, the model hyperparameters of the online industry personalized preliminary analysis model are iterated to obtain the online industry personalized analysis model; Step S4: Generate industry structured data based on the online industry personalized analysis model; Use industry structured data to build analytical knowledge graphs and upload them to the cloud platform for storage.
2. The method for constructing an online industry analysis big data platform according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: Obtaining basic enterprise information and real-time traffic information; Step S12: screening the upstream and downstream relationships of the industrial chain for the basic information of the enterprise, and eliminating blank data to obtain the enterprise integrated data; generating commuting time circle analysis data for the real-time traffic information; Step S13: Desensitize the UGC data based on the enterprise integration data and commuting time circle analysis data, and build an online industry preprocessing data set.
3. The method for constructing an online industry analysis big data platform according to claim 1, characterized in that: The UGC data desensitization described in step S1 includes: Extract enterprise user data based on enterprise integrated data; Extract user trajectory information data based on commuting time circle analysis data; Generate and anonymize the enterprise user data, where the anonymization key ranges from 128 to 256 bits, to obtain online industry identifier anonymized data; Generalize the user trajectory information data to a regional scope to obtain online regional generalization data; Hash sensitive fields of enterprise user data generation and user trajectory information data to obtain online industry sensitive field data; The online industry identifier anonymized data, online regional generalized data and online industry sensitive field data are fused into multiple sources, and a fused data set is constructed to obtain an online industry preprocessed data set.
4. The method for constructing an online industry analysis big data platform according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: Generate text semantic analysis data based on the online industry preprocessing data set; Step S22: Correct the coordinate deviation according to the online industry preprocessing data set. When the GPS coordinate deviation is greater than or equal to 50m, perform POI space correction to obtain coordinate correction data; Step S23: extracting the spatiotemporal features of regional weekly traffic from the text semantic analysis data and the coordinate correction data, and constructing the time series prediction deduction data; Step S24: construct a personalized model based on the time series prediction and deduction data to obtain an online industry personalized preliminary analysis model.
5. The method for constructing an online industry analysis big data platform according to claim 4, characterized in that: The personalized model construction described in step S2 includes: Personalized target index is set based on time series forecasting data, where the personalized target index includes macro congestion index and passenger source distribution index; Set the initial model parameters based on the time series forecasting and deduction data; The preset spatiotemporal graph convolutional network is used to analyze the regional-industry correlation of the macro congestion index, and the source distribution index is used to analyze the source POI trajectory to obtain the initial model construction data; The model framework is constructed using the initial model construction data, and the loss function is assigned to obtain a preliminary analysis model for online industry personalization.
6. The method for constructing an online industry analysis big data platform according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: outputting industry personalized results through an online industry personalized preliminary analysis model; using the industry personalized results to perform enterprise association network analysis to obtain industry graph neural analysis data; Step S32: using the industry graph neural analysis data to verify the user retention rate and task completion ratio, and determine the industry model optimization hyperparameters; Step S33: Based on the industry model optimization hyperparameters, the model hyperparameters of the online industry personalized preliminary analysis model are iterated to obtain the online industry personalized analysis model.
7. The method for constructing an online industry analysis big data platform according to claim 6, characterized in that: The enterprise association network analysis in step S3 includes: The industry-specific output results are used to generate multi-granularity result rendering comprehensive data, where the multi-granularity result rendering comprehensive data includes macro-layer rendering data and meso-layer rendering data; According to the personalized output results of the industry, traffic flow-enterprise density is extracted, and regional raster rendering is performed with a grid size range of 100m-1000m to obtain macro-layer rendering data; Cluster coupling is performed based on the personalized output results of the industry to obtain the enterprise cluster coupling index; the enterprise cluster coupling index is used to couple and associate the personalized output results of the industry, and road network rendering is performed to obtain the meso-level rendering data.
8. The method for constructing an online industry analysis big data platform according to claim 6, characterized in that: Step S32 includes the following steps: Step S321: Perform user retention rate analysis using industry graph neural analysis data to obtain enterprise user retention rate data; Step S322: performing enterprise supply chain dependency detection based on enterprise user retention rate data to obtain enterprise user-supply chain analysis data; Step S323: Performing completion analysis on the enterprise user-supply chain analysis data and performing ratio verification to obtain enterprise user-supply chain-completion data; Step S324: Generate hyperparameters for enterprise user-supply chain-completion data based on the online industry personalized preliminary analysis model to obtain industry model optimization hyperparameters.
9. The method for constructing an online industry analysis big data platform according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: extracting industry structured data according to the online industry personalized analysis model; Step S42: construct an analysis knowledge graph using industry structured data, and set industry confidence parameters to obtain comprehensive industry analysis data; Step S43: Upload the comprehensive industry analysis data to the cloud platform for storage.
10. The method for constructing an online industry analysis big data platform according to claim 9, characterized in that: Step S42 includes the following steps: Step S421: construct an analytical knowledge graph using industry structured data; Step S422: Setting industry confidence parameters based on the analysis of the knowledge graph; Step S423: Use the industry confidence parameter to perform heat map confirmation on the analysis knowledge graph, and perform comprehensive heat map analysis to obtain comprehensive industry analysis data.
Citation Information
Patent Citations
Time series data event prediction method and system based on graph convolutional neural network and application thereof
CN111367961A
Intelligent traffic flow analysis method based on Beidou data
CN117935561A
Method for constructing chemical-plastic industry chain knowledge graph by using graph convolutional network
CN119250172A