Data relationship mining method, system and equipment

By processing multi-source heterogeneous data streams and generating 3D interactive graphs using hybrid learners, the problems of insufficient model versatility, limited visualization, and incomplete interaction in investment promotion are solved. This enables cross-scenario adaptability assessment and real-time data management, improving the accuracy and efficiency of investment promotion.

CN120852058APending Publication Date: 2025-10-28CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511042175.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies for investment promotion suffer from problems such as insufficient model universality and adaptability, limited visualization effects, incomplete interactive functions, and insufficient data updates and real-time performance, making it difficult to meet the needs of complex and ever-changing investment promotion environments.

Method used

It adopts multi-source heterogeneous data stream acquisition and standardization processing, combines a hybrid learner of tree model and neural network to mine data relationship, generates a 3D interaction graph, performs multi-hop association scoring and path reasoning through graph neural network, and uses WebGL for real-time 3D rendering and blockchain notarization.

Benefits of technology

It enables cross-scenario adaptability assessment and interpretable output, improves the accuracy and efficiency of investment promotion, shortens the project implementation cycle, and forms a reliable closed loop for full life cycle management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852058A_ABST
    Figure CN120852058A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of business data processing, and provides a data relationship mining method, system and equipment, and the method comprises the steps: obtaining an original data set through accessing a multi-source heterogeneous data stream, carrying out the standardization processing of the original data set, and outputting a standardized data set; performing mixed training and relation path reasoning on the standardized data set, and outputting a score tensor and a path vector of multi-hop association between nodes; and inputting the score tensor and the path vector into a graph rendering engine, and generating a three-dimensional interaction graph in real time. According to the data relationship mining method, system and equipment, the precision, efficiency and landing rate of investment attraction can be synchronously increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of commercial data processing technology, and in particular to a data relationship mining method, system, and device. Background Technology

[0002] The integration and structuring of investment attraction data using big data technology is a current hot topic. Existing technologies use past industrial chain development as data tags and employ algorithms such as BP neural networks, decision trees, and random forests to train industrial chain suitability assessment models. By analyzing data characteristics related to various conditions within the industrial chain, including enterprise development, talent development, policy development, and project development, these models assist investment promotion personnel in accurately identifying target companies, addressing the information asymmetry between investment promotion personnel and the companies or projects they are attracting, and improving the conversion rate of investment attraction.

[0003] In practical applications, existing technologies have the following shortcomings: 1. Insufficient Model Universality and Adaptability: Evaluation models are often built for specific scenarios or industries, resulting in poor universality. The investment promotion needs and influencing factors vary greatly across different regions and industries, making it difficult for a single model to adapt to the complex and ever-changing investment promotion environment. For example, if a region focuses on developing the cultural and creative industries, traditional supply chain suitability assessment models based on industrial manufacturing are insufficient to effectively evaluate investment projects in the cultural and creative industries.

[0004] 2. Limited Visualization: Visualizations are limited to basic information presentation, such as simple charts and map annotations, lacking in-depth visualization of complex investment relationships. Multi-level and multi-type relationships between enterprises are difficult to present intuitively, affecting users' understanding of the overall picture of investment relationships.

[0005] 3. Inadequate Interactive Functionality: Although an interactive response mechanism exists, the interactive functions are not rich or smooth enough. For example, in map interaction, zooming and switching map types may result in lag, and the interactive operation is not convenient enough to meet users' needs for quickly obtaining information and flexibly exploring investment resources. Adjusting the visual condition list involves cumbersome steps, negatively impacting the user experience.

[0006] Chinese patent CN115170327A discloses a business investment promotion auxiliary decision-making method and system based on big data. This solution has the following shortcomings: 1. Insufficient relationship mining depth: It only outputs a single "fitness score" and fails to perform path-level reasoning on multi-hop relationships such as investment, supply, and cooperation between entities, thus failing to reveal implicit connections between upstream and downstream sectors of the industrial chain. 2. Weak model versatility and interpretability: It uses traditional shallow models (BP, decision tree, random forest), requiring retraining for different industrial chains, resulting in high migration costs; furthermore, the model's internal weights lack intuitive and interpretable rules, making it difficult for investment promotion personnel to adjust strategies accordingly. 3. Insufficient data updates and real-time performance: It relies on batch offline import and manual reporting, lacking incremental log parsing and streaming processing mechanisms, leading to data lag and an inability to capture the latest enterprise dynamics. 4. Limited visualization dimensions: It uses a two-dimensional knowledge graph, only presenting static node-edge relationships, lacking three-dimensional spatial mapping and real-time interaction, resulting in low information density and difficulty in supporting rapid decision-making in complex scenarios. 5. The incomplete business loop means that the assessment results are only used for initial screening and do not cover the entire lifecycle management of project negotiation, signing, implementation, and policy implementation. The lack of state machine or blockchain evidence storage mechanism makes it difficult to continuously track and quantify the implementation effect.

[0007] Therefore, there is an urgent need for an intelligent technical solution that can automatically collect multi-source heterogeneous data, assess the suitability of the subject and the scenario in real time, deeply mine multi-hop relationships and present them intuitively in a 3D interactive diagram, so as to improve investment promotion efficiency, reduce decision-making risks and shorten the project implementation cycle. Summary of the Invention

[0008] In view of this, in order to overcome the shortcomings of the prior art, the present invention aims to provide a data relationship mining method, system and device.

[0009] According to a first aspect of the present invention, a data relationship mining method is provided, the method comprising: Step S1: Obtain the original dataset by accessing multi-source heterogeneous data streams, standardize the original dataset, and output a standardized dataset; Step S2: Perform hybrid training and relational path reasoning on the standardized dataset, and output the rating tensor and path vector of multi-hop associations between nodes; Step S3: Input the rating tensor and path vector into the graphics rendering engine to generate a 3D interactive graph in real time.

[0010] Optionally, in the data relationship mining method of the present invention, in step S1, a multi-source heterogeneous data stream from an external system is accessed in parallel by a distributed collector to obtain an original dataset including subject information, technical information, transaction information, legal information, rule information, and financial information.

[0011] Optionally, in the data relationship mining method of the present invention, in step S1, data change events are captured based on the log incremental parsing mechanism, and the original dataset is denoised, deduplicated, and uniformly formatted using a fingerprint deduplication algorithm and a semantic entity linking model to output a standardized dataset.

[0012] Optionally, in the data relationship mining method of the present invention, step S2 includes: A hybrid learner consisting of tree model and neural network model is used to train a standardized dataset to generate subject-scene fit evaluation parameters. The generated subject-scene fit evaluation parameters are injected into the graph neural network as the initial weights of the nodes. Perform relational path reasoning on the same standardized dataset, and output the rating tensor and path vector of multi-hop associations between nodes.

[0013] Optionally, in the data relationship mining method of the present invention, in step S2, before relationship path reasoning, nodes are threshold filtered according to the subject and scene fit evaluation parameters, and nodes with fit higher than the set threshold are retained to participate in subsequent relationship path reasoning.

[0014] Optionally, in the data relationship mining method of the present invention, in step S2, during relationship path reasoning, the subject and scene fit evaluation parameters are concatenated with the original feature vector as part of the node feature vector and then input into the graph neural network, so that the subject and scene fit evaluation parameters are propagated layer by layer with the graph convolution process.

[0015] Optionally, in step S2 of the data relationship mining method of the present invention, the following steps are further included: after outputting the scoring tensor, the scoring tensor is weighted and fused using subject and scene fit evaluation parameters, and the fusion weight is dynamically determined by a Bayesian optimization algorithm.

[0016] Optionally, in the data relationship mining method of the present invention, step S3 includes: mapping the rating tensor and path vector to visual attribute values ​​of nodes and edges through linear normalization; constructing spherical nodes and cylindrical edge geometry in parallel on the GPU based on the visual attribute values ​​of nodes and edges; writing the cylindrical edge geometry into a G-Buffer through the WebGL delay pipeline and overlaying lighting and highlighting effects to generate a 3D frame; capturing user events and locating nodes through ray casting, sending back identifiers to trigger recalculation of the learning and inference steps, and updating the next frame with highlighting animation.

[0017] According to a second aspect of the present invention, a data relationship mining system is provided, characterized in that the system includes a relationship mining server, the relationship mining server comprising: Data acquisition and processing module: used to acquire raw datasets by accessing multi-source heterogeneous data streams, standardize the raw datasets, and output standardized datasets; The training and inference modules are used for mixed training and relational path inference on standardized datasets, and output the rating tensor and path vector of multi-hop associations between nodes; The visualization module is used to input the rating tensor and path vector into the graphics rendering engine to generate a 3D interactive graph in real time.

[0018] According to a third aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect of the present invention.

[0019] This invention provides a data relationship mining method, system, and device that unifies real-time and complete data foundation through distributed acquisition and incremental governance; unifies cross-scenario adaptability assessment and interpretable output through a tree-neural hybrid model; unifies deep relationship mining and millisecond-level inference through adaptability pre-filtering and graph neural networks; unifies smooth display and personalized interaction of hundreds of thousands of nodes through WebGL 3D rendering; and unifies the entire lifecycle management and trusted closed loop of investment promotion through a rule engine and blockchain contracts, thereby achieving a simultaneous leap in investment promotion accuracy, efficiency, and implementation rate. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is an example architecture diagram of a data relationship mining system according to an embodiment of the present invention; Figure 2 This is an example architecture diagram of a relationship mining server in a data relationship mining system according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the steps of a data relationship mining method according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the structure of the device provided by the present invention. Detailed Implementation

[0022] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0023] It should be noted that, in the absence of conflict, the following embodiments and features can be combined with each other; and, based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0024] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0025] Figure 1 This is an example architecture diagram of a data relationship mining system according to an embodiment of the present invention, such as... Figure 1 As shown, the system may include a relationship mining server 101, a communication network 102, and / or one or more relationship mining clients 103. Figure 1 The example in the text is a client 103 for multiple relationship mining.

[0026] The relationship mining server 101 can be any suitable server used to store information, data, programs, and / or any other suitable type of content. In some embodiments, the relationship mining server 101 can perform appropriate functions. For example, in some embodiments, the relationship mining server 101 can be used for data relationship mining.

[0027] Figure 2 This is an example architecture diagram of a relationship mining server in a data relationship mining system according to an embodiment of the present invention, as shown below. Figure 2 As shown, in this embodiment, the relationship mining server includes: Data acquisition and processing module: used to acquire raw datasets by accessing multi-source heterogeneous data streams, standardize the raw datasets, and output standardized datasets; The training and inference modules are used for mixed training and relational path inference on standardized datasets, and output the rating tensor and path vector of multi-hop associations between nodes; The visualization module is used to input the rating tensor and path vector into the graphics rendering engine to generate a 3D interactive graph in real time.

[0028] As another example, in some embodiments, the relationship mining server 101 may send data relationship mining methods to the relationship mining client 103 for user use, based on a request from the relationship mining client 103.

[0029] As an optional example, in some embodiments, the relationship mining client 103 is used to provide a visual relationship mining interface, which is used to receive a user's selection input operation for data relationship mining, and to obtain and display the relationship mining interface corresponding to the option selected by the selection input operation from the relationship mining server 101 in response to the selection input operation. The relationship mining interface displays at least data relationship mining information and operation options for the data relationship mining information.

[0030] In some embodiments, communication network 102 can be any suitable combination of one or more wired and / or wireless networks. For example, communication network 102 can include any one or more of the following: the Internet, intranet, wide area network (WAN), local area network (LAN), wireless network, digital subscriber line (DSL) network, frame relay network, asynchronous transfer mode (ATM) network, virtual private network (VPN), and / or any other suitable communication network. Relationship mining client 103 can connect to communication network 102 via one or more communication links (e.g., communication link 104), which can be linked to relationship mining server 101 via one or more communication links (e.g., communication link 105). Communication links can be any communication link suitable for transmitting data between relationship mining client 103 and relationship mining server 101, such as network links, dial-up links, wireless links, hardwired links, any other suitable communication links, or any suitable combination of such links.

[0031] The relationship mining client 103 may include any one or more clients that present an interface related to data relationship mining in an appropriate form for user use and operation. In some embodiments, the relationship mining client 103 may include any suitable type of device. For example, in some embodiments, the relationship mining client 103 may include a mobile device, tablet computer, laptop computer, desktop computer, and / or any other suitable type of client device.

[0032] Although the relationship mining server 101 is illustrated as a single device, in some embodiments, any suitable number of devices may be used to perform the functions performed by the relationship mining server 101. For example, in some embodiments, multiple devices may be used to implement the functions performed by the relationship mining server 101. Alternatively, cloud services may be used to implement the functions of the relationship mining server 101.

[0033] Based on the above system, this invention provides a data relationship mining method, which is described below through the following embodiments.

[0034] Figure 3This is a flowchart illustrating the steps of a data relationship mining method according to an embodiment of the present invention. The data relationship mining method of this embodiment can be executed on a relationship mining server, and includes the following steps: Step S1: Obtain the original dataset by accessing multi-source heterogeneous data streams, standardize the original dataset, and output a standardized dataset.

[0035] In this embodiment, a distributed collector is used to access multi-source heterogeneous data streams from an external system in parallel to obtain a raw dataset including entity information, technical information, transaction information, legal information, rule information, and financial information.

[0036] After obtaining the original dataset, this embodiment captures data change events based on the log incremental parsing mechanism, and uses the fingerprint deduplication algorithm and semantic entity linking model to denoise, deduplicatize and uniformly format the original dataset, outputting a standardized dataset.

[0037] The following describes step S1 of this embodiment in a specific scenario, focusing on the standardization process of the original dataset for the new energy industry chain. In this scenario, a certain investment promotion big data platform needs to aggregate multi-dimensional information on relevant entities within the new energy industry chain in a real-time manner to support subsequent investment relationship mining. The platform processes heterogeneous data streams daily, including business registration changes, patent applications, bidding announcements, court judgments, policy documents, and financing events, with a data volume of approximately 1.2TB / day. The specific implementation process is as follows: 1. Data access (parallel acquisition of multi-source heterogeneous data streams) 1.1 Deploy 8 groups of distributed collectors Three groups are oriented towards structured databases: Directly connect to various information disclosure systems via JDBC (Java Database Connectivity) to incrementally pull change logs.

[0038] Three sets of semi-structured interfaces: Based on the RESTful API (Representational State Transfer Application Programming Interface), poll the new energy industry association and the project publicity website in the region to retrieve project announcements in JSON (JavaScript Object Notation) format.

[0039] Two sets of tools are designed for unstructured web pages: using a Scrapy cluster to deeply crawl policy interpretation pages at all levels and extract the PDF / Word text.

[0040] 1.2 Unified Data Bus All collectors push the raw data stream to the Kafka Topic "raw_new_energy". Each message includes a source identifier, collection timestamp, and original payload to ensure traceability.

[0041] 2. Incremental log parsing (real-time capture of change events) 2.1MySQL Binlog analysis For the two major databases, namely industrial and commercial and patent databases, the platform deploys Maxwell processes to parse the Binlog in real time, capture "insert / update / delete" events in milliseconds, and write the parsing results to the Kafka Topic "binlog_delta" in JSON format.

[0042] 2.2 API Polling Time Window For public interfaces without Binlog, a 5-minute sliding window polling method is used; if a change in ETag or Last-Modified is detected, a full fetch is triggered and the incremental time point is marked.

[0043] 3. Data cleaning and deduplication (fingerprint deduplication + semantic entity linking) 3.1 Fingerprint Deduplication In Flink's real-time computing task, a 64-bit SimHash fingerprint is extracted for each original record. If the Hamming distance between two record fingerprints is ≤3, they are considered duplicates, and the latest version is retained. This step filters out approximately 18% of duplicate records daily.

[0044] 3.2 Semantic Entity Linking Company name disambiguation: A pre-trained BERT-NER model is used to identify company entities, and then vector similarity matching (threshold 0.92) is performed with the internal company knowledge base to solve the problem of "XX New Energy Co., Ltd." and "XX New Energy Co., Ltd." referring to the same entity.

[0045] Standardize field formats: Convert registered capital to "ten thousand yuan"; standardize dates to ISO-8601; standardize addresses to the fourth level of administrative division.

[0046] 3.3 Outlier Removal Abnormal values: If the registered capital is less than 0 or greater than 5 billion yuan, it will be marked as abnormal and manually reviewed.

[0047] Text anomalies: Use regular expressions to remove obvious garbled characters and HTML tags, retaining content with a valid character length of ≥10.

[0048] 3.4 Output Standardized Dataset After cleaning, the data is written to the HDFS (Hadoop Distributed File System) directory " / standard / new_energy / " in Parquet (Parquet Columnar Storage Format) format. The schema includes uniform fields: entity_id, entity_name, reg_capital, patent_cnt, bid_cnt, policy_cnt, court_cnt, funding_amt, and update_time. This directory is partitioned by day, supporting minute-level incremental reads by downstream Spark (Apache Spark) jobs.

[0049] Through the above steps, the platform completes the conversion from heterogeneous stream access to standardized dataset within 30 minutes, outputting approximately 950GB of high-quality data daily, laying a solid foundation for subsequent hybrid training and relational path reasoning.

[0050] This embodiment uses a distributed collector to access six heterogeneous data streams in parallel: subject, technology, transaction, legal, rule, and financial. It achieves millisecond-level change capture based on incremental log parsing. Combined with fingerprint deduplication and semantic entity linking model, the data accuracy can be improved to over 98%, solving the problems of data fragmentation and update lag in traditional platforms, and providing a high-confidence real-time data foundation for subsequent analysis.

[0051] Step S2: Perform hybrid training and relational path reasoning on the standardized dataset, and output the rating tensor and path vector of multi-hop associations between nodes.

[0052] As an optional example, this embodiment uses a hybrid learner composed of a tree model and a neural network model to train a standardized dataset to generate subject-scene fit evaluation parameters; the generated subject-scene fit evaluation parameters are injected into the graph neural network as the initial weights of the nodes; relational path reasoning is performed in the same standardized dataset to output the rating tensor and path vector of multi-hop association between nodes.

[0053] It should be noted that in this embodiment, before relational path reasoning, nodes are threshold-filtered based on the subject and scene fit evaluation parameters, retaining nodes with fit scores higher than the set threshold for subsequent relational path reasoning. During relational path reasoning, the subject and scene fit evaluation parameters are concatenated with the original feature vector as part of the node feature vector and input into the graph neural network, allowing the subject and scene fit evaluation parameters to propagate layer by layer during the graph convolution process.

[0054] In this embodiment, after outputting the scoring tensor, the scoring tensor is weighted and fused using subject and scene fit evaluation parameters, and the fusion weights are dynamically determined using a Bayesian optimization algorithm.

[0055] The following provides a further explanation of step S2 in a specific scenario. This scenario focuses on the hybrid training and graph neural network inference process of "subject-scenario adaptability" in the new energy industry chain. Specifically, this embodiment selects the three-tier industry chain of "new energy vehicles - power batteries - positive and negative electrode materials" as the target scenario, identifies potential investment entities with high adaptability to the above scenario from enterprises registered nationwide, and explores their multi-hop investment, cooperation, and supply chain relationships.

[0056] The method of this embodiment is implemented in this scenario as follows: I. Data Preparation 1. Standardized dataset: Static characteristics of the entity: registered capital, years of establishment, business scope (one-hot vector) Dynamic features of the scenario: installed capacity of power batteries, shipment volume of cathode materials, patent IPC classification number vector, and TF-IDF (Term Frequency-Inverse Document Frequency) vector of bidding announcements in the past three years. Tags: "True fit" of historical projects in this scenario (0-1 continuous value, obtained by expert scoring + regression analysis of project implementation results) 2. Dataset size: 2.1M nodes, 3.8M edges, of which approximately 45K nodes are related to the new energy scenario.

[0057] II. Hybrid Learner Training 1. Model Building Tree Model: Lightweight GBDT (500 regression trees, maximum depth 7) is used to capture interpretable segmented rule features. GBDT stands for Gradient Boosting Decision Tree, which is used in hybrid learners to extract interpretable segmented rule features and provide high-order class inputs to the neural network.

[0058] Neural network: 3 fully connected layers + BatchNorm + PReLU, hidden layer dimension 512→256→128, output layer is single-node regression; BatchNorm is used to stabilize the distribution after each layer of the neural network, accelerate convergence, and improve the training efficiency of the hybrid learner. PreLU (Parametric Rectified Linear Unit) is used as the activation function of the neural network, introducing a learnable negative slope to enhance nonlinear expressive power.

[0059] Fusion: The leaf node indexes output by GBDT are fed into the neural network as categorical features along with the original numerical features to form the joint loss L=α·L_tree+(1-α)·L_nn of "GBDT+NN (neural network)". α is optimized between 0.2 and 0.8 through Bayesian optimization, and finally α=0.35 is locked.

[0060] 2. Training Configuration Training / validation / testing are divided in a 7:1.5:1.5 time sequence; Mini-batch Adam optimizer (Adaptive Moment Estimation), initial learning rate 3e-4, cosine annealing; Early stop patience = 10 epochs (training cycles), validation set R 2 (The coefficient of determination, used to measure the ability of the hybrid learner's predictions to explain the true fitness labels,) stopped increasing at 0.892.

[0061] 3. Output Each subject receives a "subject-scene fit evaluation parameter" score_i in the range of 0-1.

[0062] III. Node Pre-filtering (Threshold Filtering) Setting score_i≥0.65 as valid nodes, the remaining nodes after filtering are 12K (accounting for 26.7% of the original 45K), the computational cost is reduced by 73% while the recall rate remains at 91%.

[0063] IV. Graph Neural Network Inference 1. Graph Construction Node: The aforementioned 12K high-adaptability main body; Edge: Three categories: investment, supply, and cooperation, with rights (amount or contract frequency); Features: The original 128-dimensional vector is combined with score_i to form the 129th dimension.

[0064] 2. Network Structure 3-layer GraphSAGE (Graph Sample and Aggregate) mean aggregation, with the number of neighbor samples in each layer [25, 15, 10]; Attention head 4: The attention output is element-wise multiplied with score_i and then fed into the next layer; Output: Two decoders, generating: The rating tensor S∈R^{N×N} represents the strength of the association between subjects in the scene; Path vector P∈R^{N×K×L}: records the first K=5 maximum multi-hop paths, with each path length L≤5.

[0065] 3. Training and Reasoning The side label is a binary value of "whether there is a real business flow", and Focal Loss (Focal Loss Function) is used to solve the imbalance of 1:20 positive and negative samples. After training for 200 epochs, the AUC (Area Under the Curve, used to evaluate the performance of graph neural networks in edge-level classification tasks) reached 0.943. The inference phase uses a GPU batch of 512, and all 12K nodes can be completed in just 1.8 seconds.

[0066] V. Results Output Scoring Tensor: Provides investment promotion personnel with a full-chain strength heatmap of "vehicle manufacturers - battery manufacturers - material manufacturers"; Path vector: Highlights a 4-hop investment path from a vehicle manufacturer to a third-tier battery manufacturer to a leading lithium iron phosphate material company, which can be directly used for priority ranking in investment negotiations.

[0067] This embodiment employs a hybrid learner combining a tree model and a neural network. GBDT is responsible for extracting interpretable segmentation rules, while the neural network handles high-order feature interactions. The two are trained together, and the weights are dynamically adjusted using Bayesian optimization. This allows the same model to maintain an R-value above 0.89 across multiple industry chains, including new energy, biomedicine, and high-end equipment, without requiring retraining. 2 This significantly reduces migration costs and improves decision-making transparency. Using "subject and scenario fit evaluation parameters" as initial node weights, threshold filtering is performed before graph neural network inference, retaining only highly fit nodes for computation, reducing computational load by over 70%. Simultaneously, the fit vector is concatenated with the original features and input into GraphSAGE+Attention, allowing fit information to propagate layer by layer with convolution. The final output is a scoring tensor and path vector representing multi-hop associations between nodes, achieving "hundreds of millions of edges, millisecond-level" inference, capable of revealing investment, cooperation, and supply chain pathways within 5 hops in a single operation.

[0068] Step S3: Input the rating tensor and path vector into the graphics rendering engine to generate a 3D interactive graph in real time.

[0069] In this embodiment, the scoring tensor and path vector are linearly normalized and mapped to the visual attribute values ​​of nodes and edges. Based on the visual attribute values ​​of nodes and edges, spherical nodes and cylindrical edge geometry are constructed in parallel on the GPU (Graphics Processing Unit). The cylindrical edge geometry is written into the G-Buffer (Geometry Buffer) through the WebGL (Web Graphics Library) delay pipeline and lighting and highlighting effects are superimposed to generate a 3D frame. User events are captured and nodes are located through ray casting. The returned identifier triggers the recalculation of the learning and inference steps, and the next frame is updated with highlight animation.

[0070] For example, this embodiment generates 3D interactive diagrams in real time in the following manner: 1. Data mapping: The rating tensor and path vector are mapped to the [0,1] interval through linear normalization to obtain the node color value, node size value, edge color value, and edge thickness value; 2. Geometry generation: The vertex shader of the graphics rendering engine is called to generate sphere geometry in batches based on node position vectors and cylinder geometry in batches based on path vectors. Vertex coordinates, normals and texture coordinates are calculated in parallel on the GPU. 3. Rendering pipeline: A WebGL-based deferred rendering pipeline is adopted. First, the position, color, and normal information of nodes and edges are written into the G-Buffer. Then, the screen space ambient occlusion algorithm is used to enhance the sense of 3D depth during the lighting stage. Finally, a highlight effect is superimposed on the forward rendering. 4. Interactive Response: Listen for user mouse or gesture events, detect selected nodes through ray casting algorithm, send the node ID back to the learning and inference steps to trigger local recalculation, and present the updated result in the next frame with a highlighted border animation, realizing real-time and interactive visualization of the scoring tensor and path vector into a 3D interactive graph.

[0071] In this embodiment, the rating tensor and path vector are linearly normalized and mapped to node color, size, and edge thickness. By utilizing the WebGL deferred pipeline and GPU parallel geometry construction, smooth rendering at 60fps can be maintained at a scale of 100,000 nodes. User interaction sequences are analyzed in real time through K-means clustering, and the force-directed layout and LOD (Level of Detail) level are dynamically adjusted. Raycasting is also supported to quickly locate nodes, realizing an immersive exploration experience that is "what you see is what you get".

[0072] In practical applications, the system can also incorporate a built-in rule engine and state machine to track each stage of project negotiation, signing, implementation, and policy fulfillment in a node-based manner. When the state transitions to "signing" or "payment", the smart contract deployed on the blockchain is automatically triggered to complete the condition verification and fund storage, further shortening the policy fulfillment cycle, improving the success rate of implementation, and forming a trusted closed loop of "data-model-interaction-business".

[0073] In practical applications, the data relationship mining method and system of this embodiment have the following beneficial technical effects: a unified real-time and complete data foundation is achieved through distributed acquisition and incremental governance; a unified cross-scenario adaptability assessment and interpretable output is achieved through a tree-neural hybrid model; a unified deep relationship mining and millisecond-level inference is achieved through adaptability pre-filtering and graph neural networks; a unified smooth display and personalized interaction of hundreds of thousands of nodes is achieved through WebGL 3D rendering; and a unified investment promotion lifecycle management and trusted closed loop is achieved through a rule engine and blockchain contract, thereby achieving a simultaneous leap in investment promotion accuracy, efficiency, and implementation rate.

[0074] like Figure 4 As shown, the present invention also provides a device including a processor 210, a communication interface 220, a memory 230 for storing processor-executable computer programs, and a communication bus 240. The processor 210, communication interface 220, and memory 230 communicate with each other via the communication bus 240. The processor 210 implements the aforementioned data relationship mining method by running the executable computer program.

[0075] The computer program in memory 230, when implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0076] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected based on actual needs to achieve the purpose of this embodiment. Those skilled in the art can understand and implement this without any creative effort.

[0077] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0078] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A data relationship mining method, characterized in that, The method includes: Step S1: Obtain the original dataset by accessing multi-source heterogeneous data streams, standardize the original dataset, and output a standardized dataset; Step S2: Perform hybrid training and relational path reasoning on the standardized dataset, and output the rating tensor and path vector of multi-hop associations between nodes; Step S3: Input the rating tensor and path vector into the graphics rendering engine to generate a 3D interactive graph in real time.

2. The data relationship mining method according to claim 1, characterized in that, In step S1, a distributed collector is used to access multi-source heterogeneous data streams from external systems in parallel to obtain raw datasets including entity information, technical information, transaction information, legal information, rule information, and financial information.

3. The data relationship mining method according to claim 1, characterized in that, In step S1, data change events are captured based on the log incremental parsing mechanism. The fingerprint deduplication algorithm and semantic entity linking model are used to denoise, deduplicatize and uniformly format the original dataset, and output a standardized dataset.

4. The data relationship mining method according to claim 1, characterized in that, Step S2 includes: A hybrid learner consisting of tree model and neural network model is used to train a standardized dataset to generate subject-scene fit evaluation parameters. The generated subject-scene fit evaluation parameters are injected into the graph neural network as the initial weights of the nodes. Perform relational path reasoning on the same standardized dataset, and output the rating tensor and path vector of multi-hop associations between nodes.

5. The data relationship mining method according to claim 4, characterized in that, In step S2, before relational path reasoning, nodes are threshold-filtered based on subject and scene fit evaluation parameters, and nodes with fit higher than the set threshold are retained to participate in subsequent relational path reasoning.

6. The data relationship mining method according to claim 4, characterized in that, In step S2, during relational path reasoning, the subject and scene fit evaluation parameters are concatenated with the original feature vector as part of the node feature vector and then input into the graph neural network, so that the subject and scene fit evaluation parameters are propagated layer by layer with the graph convolution process.

7. The data relationship mining method according to claim 4, characterized in that, Step S2 further includes: after outputting the score tensor, using the subject and scene fit evaluation parameters to perform weighted fusion of the score tensor, and dynamically determining the fusion weights through a Bayesian optimization algorithm.

8. The data relationship mining method according to claim 1, characterized in that, Step S3 includes: mapping the rating tensor and path vector to visual attribute values ​​of nodes and edges through linear normalization; constructing spherical nodes and cylindrical edge geometry in parallel on the GPU based on the visual attribute values ​​of nodes and edges; writing the cylindrical edge geometry into the G-Buffer through the WebGL deferred pipeline and overlaying lighting and highlighting effects to generate a 3D frame; capturing user events and locating nodes through ray casting, sending back the identifier to trigger the recalculation of the learning and inference steps, and updating the next frame with highlighting animation.

9. A data relationship mining system, characterized in that, The system includes a relationship mining server, which includes: Data acquisition and processing module: used to acquire raw datasets by accessing multi-source heterogeneous data streams, standardize the raw datasets, and output standardized datasets; The training and inference modules are used for mixed training and relational path inference on standardized datasets, and output the rating tensor and path vector of multi-hop associations between nodes; The visualization module is used to input the rating tensor and path vector into the graphics rendering engine to generate a 3D interactive graph in real time.

10. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Business investment attraction auxiliary decision-making method and system based on big data

    CN115170327A