Internet big data extraction method, device and equipment, and storage medium
By using distributed web crawling and heterogeneous data processing, a semantic feature graph is constructed and a hidden knowledge network is mined. This solves the problem of price prediction bias caused by the wide range of data sources and semantic ambiguity in the used car market, and achieves efficient data extraction and intelligent query expansion.
Patent Information
- Application Number
- CN202510809252.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing technologies in the used car market suffer from significant biases in price prediction and maintenance cost analysis due to the wide range of data sources and semantic ambiguity, making it difficult to effectively integrate and reason across different domains.
Data is collected from the Internet through distributed web crawlers, heterogeneous data is structurally transformed, semantic feature maps are constructed and subject segmentation and classification are performed, implicit knowledge networks are mined, semantic decomposition and expansion are performed based on implicit knowledge networks, similarity calculation and multi-factor comprehensive ranking are performed, and target extraction data is obtained.
It improves the accuracy of price prediction and maintenance cost analysis in the used car market, realizes the semantic understanding and interactive capabilities of intelligent queries, and significantly improves the system's query accuracy and business relevance.
Smart Images

Figure CN120632187B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data information technology, and in particular to a method and related equipment for extracting big data from the Internet. Background Technology
[0002] With the rapid development of internet technology, massive amounts of used car transaction information, vehicle repair records, and parts price data are constantly emerging. This data contains enormous commercial value and decision-making support potential. However, traditional data processing methods are inadequate when faced with such vast and heterogeneous online data, struggling to efficiently extract information with practical application value. Especially in the used car market, the wide range of data sources, diverse formats, and semantic ambiguity lead to significant biases and uncertainties in existing systems when conducting vehicle evaluation, price prediction, and repair cost analysis.
[0003] Current big data analytics methods for the used car market largely rely on structured databases or simple keyword matching techniques, lacking the ability to deeply understand the semantics and mine knowledge from multi-source, heterogeneous data. For example, when assessing the actual value of a used car, traditional methods often rely solely on basic fields such as model, year, and mileage, neglecting implicit but crucial influencing factors such as accident records, repair history, regional differences, and parts replacement frequency. This information gap not only reduces the accuracy of the assessment results but also increases the risk for consumers and the trust costs in the market.
[0004] Furthermore, although some studies have attempted to introduce graph techniques and association rule mining to enhance the depth of data analysis, they still face many challenges in practical applications, such as insufficient semantic parsing capabilities, ambiguous topic classification, and low efficiency in knowledge network construction. Especially in complex scenarios involving depreciation of used parts prices and repair strategy recommendations, existing systems struggle to effectively integrate and reason across domains, thus limiting their application in refined services and intelligent decision-making. Therefore, there is an urgent need for a big data extraction method that can integrate multi-source data, mine deep semantic relationships, and support intelligent query expansion to address core issues in the current used car and related fields, such as information fragmentation, coarse evaluation, and non-standardized services. Summary of the Invention
[0005] The main objective of this invention is to provide a method for extracting big data from the Internet, which solves the technical problem that existing systems have significant deviations in price prediction and maintenance cost analysis in the used car market due to the wide range of data sources and semantic ambiguity.
[0006] To achieve the above objectives, the present invention provides a method for extracting big data from the Internet, comprising the following steps:
[0007] Distributed crawling is performed on the Internet data source to obtain the original network dataset, and the original network dataset is subjected to heterogeneous data structuring transformation to obtain a structured data matrix;
[0008] Multi-level semantic parsing is performed on the structured data matrix to obtain a semantic feature map;
[0009] The semantic feature graph is segmented and classified into topics to obtain a topic domain knowledge tree, and association rules are mined from the topic domain knowledge tree to obtain a hidden knowledge network.
[0010] Based on the implicit knowledge network, the preset query conditions are semantically decomposed and expanded to obtain extended query data;
[0011] The similarity between the query extended data and the implicit knowledge network is calculated to obtain a candidate dataset;
[0012] The candidate dataset is sorted and extracted based on multiple factors to obtain the target extraction data.
[0013] Furthermore, the heterogeneous data structuring transformation of the original network dataset to obtain a structured data matrix includes:
[0014] Heterogeneous data parsing and classification labeling are performed on the original network dataset to obtain type identification results. Based on the type identification results, the original network dataset is cleaned and format unified to obtain a standardized data stream.
[0015] The standardized data stream is subjected to feature extraction and dimension mapping using a deep semantic vectorization method to obtain a high-dimensional feature tensor. The high-dimensional feature tensor is then subjected to sparsity optimization and dimensionality reduction to obtain a compact feature representation.
[0016] Based on the compact feature representation, a heterogeneous data association network is constructed to obtain a data element relationship graph. The data element relationship graph is then subjected to a structured transformation to obtain a structured data matrix.
[0017] Furthermore, the step of performing topic segmentation and classification on the semantic feature map through unsupervised hierarchical clustering to obtain a topic domain knowledge tree includes:
[0018] Local neighborhood topology construction is performed on the high-dimensional semantic nodes in the semantic feature map to obtain the adjacency similarity matrix, and the adjacency similarity matrix is decomposed into feature subspace based on spectral analysis to obtain a low-rank embedding vector field.
[0019] Based on the low-rank embedding vector field, the high-dimensional semantic nodes are dynamically density-awarely partitioned to obtain an initial clustering seed cluster. Then, an adaptive diffusion fusion operation is performed on the initial clustering seed cluster to obtain a multi-level semantic condensation structure.
[0020] The topic discrimination function is calculated for the semantic regions in the multi-level semantic cohesion structure to obtain the topic attribution probability distribution. Based on the topic attribution probability distribution, the semantic regions are propagated in a hierarchical manner to obtain a preliminary topic domain hierarchical mapping. The preliminary topic domain hierarchical mapping includes a top-level abstract topic, a middle-level classification topic, and a bottom-level fine-grained topic label.
[0021] Based on the preliminary topic domain hierarchical mapping, the semantic feature graph is optimized for topic structure to obtain a topic domain knowledge tree.
[0022] Furthermore, the step of performing association rule mining on the topic domain knowledge tree to obtain the implicit knowledge network includes:
[0023] The branch topic nodes in the topic domain knowledge tree are extracted based on the information entropy increase of the context co-occurrence pattern to obtain the high-frequency semantic co-occurrence sequence, and the high-frequency semantic co-occurrence sequence is probabilistically dependent modeled based on the conditional random field to obtain the semantic dependency probability map.
[0024] Based on the semantic dependency probability graph, multi-granularity path sampling is performed on leaf-level specific topic instances in the topic domain knowledge tree to obtain a set of topic evolution paths. Knowledge flow propagation modeling is then performed on the set of topic evolution paths to obtain the causal transmission weight matrix between topics.
[0025] The causal transmission weight matrix is calculated using the maximum likelihood estimation method to obtain a set of potential association rules. Based on the set of potential association rules, a bidirectional reasoning chain is constructed for the root topic node in the topic domain knowledge tree to obtain a cross-domain knowledge reasoning path.
[0026] Based on the cross-domain knowledge reasoning path, the topic domain knowledge tree is extended by knowledge entity mapping to obtain the implicit knowledge network.
[0027] Furthermore, the semantic decomposition and expansion of the preset query conditions based on the implicit knowledge network to obtain extended query data includes:
[0028] Deep syntactic dependency analysis is performed on the preset query conditions to obtain a query semantic structure tree. Key entities are extracted and relationships are identified on the query semantic structure tree to obtain a query intent skeleton. The query intent skeleton includes core query entities, query constraints, inter-entity relationship links, and implicit query intent.
[0029] Based on the query intent skeleton, semantic neighborhood activation is performed on the implicit knowledge network to obtain a related knowledge subgraph. Then, semantic path traversal and aggregation are performed on the related knowledge subgraph to obtain an extended semantic set, which includes synonymous concept groups, hierarchical concept chains, cross-domain related concepts, and temporal evolution variants.
[0030] The extended semantic set is transformed into a vector representation by distributed semantic coding to obtain a semantic embedding matrix. The semantic embedding matrix is then fused with attention weighting to obtain a context-enhanced query representation, which includes an explicit query vector, an implicit intent vector, a relational constraint vector, and a timeliness weight distribution.
[0031] Based on the context-enhanced query representation, a knowledge reasoning chain is constructed on the implicit knowledge network to obtain a multi-path reasoning graph. Joint probability inference is then performed on the multi-path reasoning graph to obtain query extended data.
[0032] Furthermore, the step of calculating the similarity between the query extended data and the implicit knowledge network to obtain a candidate dataset includes:
[0033] The extended query data is subjected to semantic tensor decomposition to obtain multi-dimensional query feature vectors, and the multi-dimensional query feature vectors are dynamically weighted to obtain an adaptive query representation.
[0034] Based on the adaptive query representation, multi-scale semantic matching is performed on the implicit knowledge network to obtain candidate knowledge subgraphs, and graph neural network embedding is performed on the candidate knowledge subgraphs to obtain a knowledge graph embedding matrix.
[0035] The adaptive query representation and the knowledge graph embedding matrix are subjected to high-order similarity calculation using tensor decomposition technology to obtain a multimodal similarity tensor. Then, nonlinear dimensionality reduction and cluster analysis are performed on the multimodal similarity tensor to obtain the similarity distribution spectrum.
[0036] Based on the similarity distribution spectrum, the implicit knowledge network is subjected to adaptive threshold segmentation to obtain candidate datasets.
[0037] Furthermore, the step of performing multi-factor comprehensive ranking and extraction on the candidate dataset to obtain the target extraction data includes:
[0038] Multi-dimensional feature extraction is performed on the candidate dataset to obtain a candidate data feature matrix, and nonlinear feature interaction modeling is performed on the candidate data feature matrix to obtain an interaction feature tensor.
[0039] Based on the interactive feature tensor, dynamic temporal correlation analysis is performed on the candidate dataset to obtain a temporal evolution sequence, and multi-scale wavelet transform is performed on the temporal evolution sequence to obtain a multi-resolution time-frequency feature map.
[0040] An adaptive attention mechanism is used to weight the importance of the multi-resolution time-frequency feature map to obtain a fused feature vector, and a nonlinear projection transformation is performed on the fused feature vector to obtain a high-dimensional sorting space.
[0041] Based on the high-dimensional sorting space, the candidate dataset is sorted using multi-objective Pareto optimization to obtain preliminary sorting results. Target features are then extracted from the preliminary sorting results to obtain target extracted data.
[0042] The present invention also provides an Internet big data extraction device, comprising:
[0043] The data acquisition module is used to perform distributed crawling to collect data from the Internet data source, obtain the original network dataset, and perform heterogeneous data structuring transformation on the original network dataset to obtain a structured data matrix.
[0044] The parsing module is used to perform multi-level semantic parsing on the structured data matrix to obtain a semantic feature map;
[0045] The classification module is used to perform topic segmentation and classification on the semantic feature graph to obtain a topic domain knowledge tree, and to perform association rule mining on the topic domain knowledge tree to obtain a hidden knowledge network.
[0046] An extension module is used to perform semantic decomposition and expansion of preset query conditions based on the implicit knowledge network to obtain extended query data;
[0047] The calculation module is used to perform similarity calculation between the query extended data and the implicit knowledge network to obtain a candidate dataset;
[0048] The sorting module is used to perform multi-factor comprehensive sorting and extraction on the candidate dataset to obtain the target extraction data.
[0049] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.
[0050] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0051] This invention provides a method for extracting big data from the internet, comprising the following steps: distributed crawling of the internet data source to obtain an original network dataset; heterogeneous data structuring transformation of the original network dataset to obtain a structured data matrix; multi-level semantic parsing of the structured data matrix to obtain a semantic feature map; topic segmentation and classification of the semantic feature map to obtain a topic domain knowledge tree; association rule mining of the topic domain knowledge tree to obtain a latent knowledge network; semantic decomposition and expansion of preset query conditions based on the latent knowledge network to obtain query expansion data; similarity calculation between the query expansion data and the latent knowledge network to obtain a candidate dataset; and multi-factor comprehensive sorting and extraction of the candidate dataset to obtain the target extraction data. This method solves the technical problem in the used car market where the wide range of data sources and semantic ambiguity lead to significant deviations in price prediction and maintenance cost analysis in existing systems. It achieves semantic decomposition and expansion of user queries based on latent knowledge networks, which not only improves query accuracy but also automatically recommends expansion conditions based on context and related knowledge, significantly enhancing the system's semantic understanding and interactive capabilities. Attached Figure Description
[0052] Figure 1 This is a schematic diagram illustrating the steps of an Internet big data extraction method in one embodiment of the present invention;
[0053] Figure 2 This is a structural block diagram of an Internet big data extraction device according to an embodiment of the present invention;
[0054] Figure 3 This is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0055] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0057] like Figure 1 As shown, Figure 1 This invention provides a method for extracting big data from the internet, comprising the following steps:
[0058] Step S1: Collect data from Internet data sources using distributed crawlers to obtain the original network dataset, and perform heterogeneous data structuring transformation on the original network dataset to obtain a structured data matrix.
[0059] Specifically, in this step, distributed web crawling is first used to collect massive amounts of multi-source network information, forming a raw network dataset. In particular, a distributed architecture is used to deploy multiple crawler nodes to crawl data from different websites, platforms or databases in parallel, including unstructured or semi-structured web page information from used car trading platforms, car repair forums, and parts sales websites, to ensure that the data coverage is broad and timely. Subsequently, the collected raw network datasets, which typically contain various formats such as text, tables, and image descriptions, exhibit significant heterogeneity. Therefore, it is necessary to perform heterogeneous data structuring transformation on them. This involves using techniques such as natural language processing, HTML parsing, and pattern recognition to uniformly map various unstructured data into operable field formats and organize them according to a pre-defined structured model (such as relational database table structures or tensor matrices) to generate a structured data matrix. For example, an unstructured description such as "a certain car model was registered in 2018, has driven 120,000 kilometers, and has had its transmission replaced" can be transformed into field values such as "brand = A, registration year = 2018, mileage = 120,000 km, maintenance record = transmission replacement". Ultimately, this enables subsequent analysis modules to conduct in-depth semantic analysis and knowledge mining based on a unified format.
[0060] Step S2: Perform multi-level semantic parsing on the structured data matrix to obtain a semantic feature map.
[0061] Specifically, in the process of performing multi-level semantic parsing on a structured data matrix to obtain a semantic feature graph, the core lies in extracting semantically related feature nodes and their interrelationships from the structured field information, and modeling and representing them in the form of a graph structure. The specific implementation involves first performing lexical, syntactic, and contextual analysis on the field content of the structured data matrix using natural language processing and semantic analysis techniques to identify key entities and attributes such as vehicle model year, mileage, and maintenance records. Semantic labels are then further refined using domain knowledge bases (such as automotive parts terminology and used car evaluation rule bases). Subsequently, by constructing a semantic dependency model, the potential logical connections, causal relationships, or attribute dependencies between various fields are mapped to edges in the graph, forming a semantic feature graph composed of nodes and edges. For example, in the application scenario of used car price evaluation, if a record contains structured fields such as "registered in 2018, driven 130,000 kilometers, and engine replaced", the system will identify "registration year", "mileage" and "engine replacement" as semantic nodes, and judge the weight of the three factors on the residual value of the vehicle based on industry experience, thereby establishing the semantic association path between them, and finally forming a knowledge graph that can reflect the multi-dimensional characteristics of the vehicle, providing semantic support for subsequent topic classification and implicit knowledge mining.
[0062] Step S3: Perform topic segmentation and classification on the semantic feature graph to obtain a topic domain knowledge tree, and perform association rule mining on the topic domain knowledge tree to obtain a hidden knowledge network.
[0063] Specifically, in the process of segmenting and classifying semantic feature graphs to obtain topic domain knowledge trees, the key is to identify subgraph structures with common semantic characteristics by analyzing the association strength, semantic similarity, and contextual dependencies between nodes in the semantic feature graph, and cluster them into several topic domains, thereby constructing a hierarchical knowledge tree structure. The specific implementation involves using graph clustering algorithms (such as the Louvain algorithm or spectral clustering) to perform community detection on nodes in the semantic feature graph. Based on semantic density and connectivity, the entire graph is divided into multiple semantically relatively independent topic regions. Each topic region corresponds to a specific business dimension, such as "vehicle year and mileage relationship," "parts replacement frequency," and "accident record impact." Topic domain knowledge trees are formed by setting topic labels and hierarchical relationships. Subsequently, association rule mining techniques (such as FP-Growth or Apriori algorithms) are applied to this topic domain knowledge tree to further reveal the potential co-occurrence patterns and causal relationships between different topics and between nodes within a topic, ultimately generating a hidden knowledge network. For example, in the application scenario of used car appraisal, if the topic domain knowledge tree contains multiple topic nodes such as "high mileage", "transmission replacement" and "water damage repair", the system can discover implicit knowledge relationships such as "high mileage vehicles are more prone to transmission failure" or "the residual value of a vehicle decreases significantly after water damage repair", thereby supporting subsequent intelligent query expansion and data extraction tasks.
[0064] Step S4: Based on the implicit knowledge network, perform semantic decomposition and expansion of the preset query conditions to obtain extended query data.
[0065] Specifically, based on the implicit knowledge network, the preset query conditions are semantically decomposed and expanded. This aims to automatically identify extended information related to the query intent but not explicitly mentioned by mining the core semantic elements in the user's input query and combining this with existing semantic association paths in the implicit knowledge network, thereby generating more comprehensive and semantically coverive extended query data. The specific implementation is as follows: First, the original query conditions input by the user (such as "find a used car in good condition, with low mileage and few repair records") are transformed into a structured semantic expression using natural language understanding technology, identifying key query nodes such as "good condition," "low mileage," and "few repair records." Then, using these nodes as starting points, other semantic nodes with strong correlations are searched in the implicit knowledge network, such as "no major accidents," "regular maintenance," and "no structural damage." Using graph traversal algorithms and semantic similarity calculation methods, a set of logically supported additional query conditions is expanded, ultimately forming extended query data that includes both the original and extended semantics. For example, in the application scenario of used car appraisal, if a user enters "query vehicles with replaced transmissions", the system can identify the potential relationship between the query and "high mileage", "long service life" and "decreased engine performance" through implicit knowledge network, and then incorporate these factors into the extended query conditions, making the subsequent candidate dataset selection more accurate and more relevant to business.
[0066] Step S5: Calculate the similarity between the query extended data and the implicit knowledge network to obtain a candidate dataset.
[0067] Specifically, the process of calculating the similarity between the extended query data and the implicit knowledge network to obtain a candidate dataset essentially involves measuring the degree of matching between the user's query intent and the existing semantic structures in the knowledge network, and selecting the most relevant set of data as the basis for subsequent sorting and extraction. The specific implementation is as follows: After generating the extended query data, the system maps each semantic node in the query (including the original query node and the extended semantic nodes) to its corresponding position in the implicit knowledge network, and calculates its association strength with each data node in the network based on path distance in the graph structure and semantic similarity algorithms (such as cosine similarity, Jaccard coefficient, or graph embedding vector matching). Subsequently, data entries that highly match the current query are selected according to a set similarity threshold to form a candidate dataset. For example, in the application scenario of used car appraisal, if the user's query extended data contains semantic nodes such as "transmission replacement", "high mileage", and "deteriorating engine performance", the system will identify vehicle records with similar feature combinations in the implicit knowledge network. For instance, although a vehicle does not explicitly indicate transmission repair, its information such as "mileage exceeding 120,000 kilometers", "low maintenance frequency", and "abnormal engine noise" has a strong semantic correlation with the query conditions, and therefore it is included in the candidate dataset, thereby improving the breadth and depth of data retrieval and enhancing the relevance and practicality of the results.
[0068] Step S6: Perform multi-factor comprehensive sorting and extraction on the candidate dataset to obtain the target extraction data.
[0069] Specifically, the candidate dataset undergoes multi-factor comprehensive ranking and extraction. This aims to select the target data that best meets user needs or business objectives by quantitatively analyzing and weighting multiple key dimensions affecting the evaluation results. The specific implementation is as follows: After obtaining the candidate dataset, the system calculates a comprehensive score for each data item based on a pre-set evaluation model and core influencing factors in the used car transaction scenario, such as vehicle year, mileage, maintenance records, parts replacement status, and price fluctuations of similar models in the market. This process typically employs a weighted scoring method, machine learning ranking models (such as RankNet), or multi-criteria decision-making methods (such as TOPSIS) to weight and fuse various factors according to their importance to the final evaluation result, and then normalizes them to form a unified ranking index. For example, in the application scenario of used car price assessment, if the candidate dataset contains several vehicles with similar characteristics such as "transmission replacement" and "high mileage", the system will further consider factors such as "whether it is regularly maintained", "whether there is structural damage" and "market price differences in the region", and score and rank each vehicle accordingly. Finally, the system will extract the information of the top-ranked and most valuable vehicles as the target extraction data, providing users with more accurate assessment basis and decision support.
[0070] In a specific embodiment, the step of performing heterogeneous data structuring transformation on the original network dataset to obtain a structured data matrix includes:
[0071] Heterogeneous data parsing and classification labeling are performed on the original network dataset to obtain type identification results. Based on the type identification results, the original network dataset is cleaned and format unified to obtain a standardized data stream.
[0072] The standardized data stream is subjected to feature extraction and dimension mapping using a deep semantic vectorization method to obtain a high-dimensional feature tensor. The high-dimensional feature tensor is then subjected to sparsity optimization and dimensionality reduction to obtain a compact feature representation.
[0073] Based on the compact feature representation, a heterogeneous data association network is constructed to obtain a data element relationship graph. The data element relationship graph is then subjected to a structured transformation to obtain a structured data matrix.
[0074] Specifically, in the above technical solution, the step of "converting the original network dataset into a structured data matrix through heterogeneous data transformation" is a key preprocessing step in the entire Internet big data extraction method. Its purpose is to transform the collected original network data from an unstructured or semi-structured multi-source heterogeneous state into a unified, analyzable structured format. Specifically, this process includes three core sub-steps: heterogeneous data parsing and classification labeling, deep semantic vectorization and feature optimization, and heterogeneous data association network construction and structure transformation. First, in the process of heterogeneous data parsing and classification labeling of the original network dataset to obtain type recognition results, the system uses appropriate parsers to extract content based on the different formats of the data source (such as HTML web pages, JSON interfaces, text descriptions, table information, etc.), and automatically classifies and semantically labels the data fields using natural language processing and pattern recognition technologies. For example, a vehicle information retrieved from a used car trading platform might include information such as "registration year," "mileage," "accident record," and "repair items." However, due to different platforms using different formats, it might appear in various ways, such as "registered in 2018," "registered in 2018," or "purchased in 2018." In this case, the system uses Named Entity Recognition (NER) and syntactic analysis techniques to identify the semantic categories of these fields and uniformly label them, thus forming a type identification result. Subsequently, based on this type identification result, the system cleans the raw data, removing duplicates, missing characters, illegal characters, and other abnormal data, and standardizes all fields to a uniform format, such as unifying "120,000 km" and "120,000 km" into "120,000 km," thereby generating a standardized data stream. Secondly, in the process of extracting features and mapping dimensions of the standardized data stream using deep semantic vectorization methods to obtain high-dimensional feature tensors, the system employs deep learning techniques such as word embedding models (e.g., Word2Vec, BERT), graph neural networks (GNN), or convolutional neural networks (CNN) to map different types of data, such as text, numerical, and categorical fields, into a unified numerical vector space representation, i.e., a high-dimensional feature tensor. For example, for the maintenance record field "engine replaced," the system not only encodes it as a binary variable (0 / 1), but also combines it with contextual semantics to jointly model it with fields such as "mileage," "years of use," and "maintenance frequency," forming a semantically related multi-dimensional feature vector. Based on this, to further improve computational efficiency and model generalization ability, the system performs sparsity optimization and dimensionality reduction on this high-dimensional feature tensor, typically using techniques such as principal component analysis (PCA), t-SNE, or autoencoders to remove redundant information and retain the most representative feature dimensions, ultimately forming a compact feature representation.Finally, in the process of constructing a heterogeneous data association network based on the compact feature representation to obtain a data element relationship graph, and further transforming it into a structured data matrix, the system models the semantic dependencies, causal relationships, and co-occurrence patterns between various fields as nodes and edges in a graph structure, constructing a data element relationship graph that reflects the inherent logical relationships of the data. For example, in a used car evaluation scenario, if a record contains fields such as "high mileage," "transmission replacement," and "no major accidents," the system will identify potential semantic connections between these fields and model the correlation paths between them through a graph structure. Subsequently, the system transforms the data element relationship graph according to a preset structured pattern (such as a two-dimensional table, triplet form, or tensor structure) to finally generate a structured data matrix. Each row of this matrix corresponds to a data record, and each column represents a field dimension after semantic enhancement and feature optimization, enabling subsequent multi-level semantic parsing, topic classification, and knowledge mining tasks to be carried out efficiently on a unified and high-quality data basis. In summary, this step, through deep analysis, semantic vectorization modeling, and structured reconstruction of the original network dataset, achieves an intelligent transformation from massive heterogeneous data to a structured data matrix. This not only improves data quality and usability but also provides a solid foundation for subsequent knowledge graph construction and intelligent query expansion. In the practical application scenario of used car appraisal, this structured transformation capability is particularly crucial. It can effectively integrate fragmented information from multiple platforms, accurately depict the historical condition and market value of each vehicle, thereby supporting more intelligent and refined decision analysis.
[0075] In a specific embodiment, the step of performing topic segmentation and classification on the semantic feature map to obtain a topic domain knowledge tree includes:
[0076] Local neighborhood topology construction is performed on the high-dimensional semantic nodes in the semantic feature map to obtain the adjacency similarity matrix, and the adjacency similarity matrix is decomposed into feature subspace based on spectral analysis to obtain a low-rank embedding vector field.
[0077] Based on the low-rank embedding vector field, the high-dimensional semantic nodes are dynamically density-awarely partitioned to obtain an initial clustering seed cluster. Then, an adaptive diffusion fusion operation is performed on the initial clustering seed cluster to obtain a multi-level semantic condensation structure.
[0078] The topic discrimination function is calculated for the semantic regions in the multi-level semantic cohesion structure to obtain the topic attribution probability distribution. Based on the topic attribution probability distribution, the semantic regions are propagated in a hierarchical manner to obtain a preliminary topic domain hierarchical mapping. The preliminary topic domain hierarchical mapping includes a top-level abstract topic, a middle-level classification topic, and a bottom-level fine-grained topic label.
[0079] Based on the preliminary topic domain hierarchical mapping, the semantic feature graph is optimized for topic structure to obtain a topic domain knowledge tree.
[0080] Specifically, in the step of "segmenting and classifying the semantic feature graph to obtain a topic domain knowledge tree," the core objective is to identify semantically cohesive topic regions by performing structured analysis and semantic clustering on high-dimensional semantic nodes in the semantic feature graph, and to construct a hierarchically clear and logically defined topic domain knowledge tree based on these regions. This process includes several key technical steps: local neighborhood topology construction, low-rank embedding vector field generation, dynamic density-aware partitioning, adaptive diffusion fusion, topic attribution probability distribution calculation, and hierarchical label propagation. First, in the process of constructing local neighborhood topology for high-dimensional semantic nodes in the semantic feature graph to obtain an adjacency similarity matrix, the system extracts the adjacency relationship of each semantic node in the graph structure (such as "mileage," "transmission replacement," and "number of repairs"), and constructs an adjacency similarity matrix by calculating the semantic similarity between nodes (e.g., using cosine similarity or graph path similarity methods). This matrix reflects the local connection strength and semantic relevance between semantic nodes, providing a basis for subsequent subspace decomposition. Next, the adjacency similarity matrix is decomposed into a feature subspace based on spectral analysis. The aim is to map the complex relationships in the high-dimensional semantic space to a low-dimensional and more interpretable feature space, thus forming a low-rank embedding vector field. Specifically, the system employs spectral clustering to decompose the adjacency matrix into eigenvalues, extracting the top few largest eigenvectors to form a low-dimensional embedding space. This ensures that each semantic node is represented by a low-dimensional vector in this space, preserving the main semantic structure information of the original graph while reducing noise interference and the curse of dimensionality. Subsequently, based on this low-rank embedding vector field, the high-dimensional semantic nodes are dynamically density-awarely partitioned. The system can automatically identify semantically dense regions in the graph as initial clustering seed clusters. This step typically uses algorithms such as Density Peaks Clustering (DPSC) or DBSCAN to select appropriate cluster centers based on the distribution density and distance characteristics of nodes in the low-dimensional space. For example, in the application scenario of used car appraisal, certain nodes (such as "accident records," "engine replacement," and "maintenance frequency") may cluster together in the semantic space, reflecting their common influence mechanism in vehicle residual value assessment, and thus are classified into the same initial cluster. Further, the system performs an adaptive diffusion fusion operation on these initial seed clusters to discover a broader semantic association structure, forming a multi-level semantic cohesive structure. The diffusion fusion process simulates the propagation behavior of semantic information in the graph, controlling the propagation range by setting a threshold, gradually merging adjacent clusters with high semantic similarity, and ultimately constructing semantic regions with different levels of abstraction.For example, an initial cluster might only contain "transmission replacement" and "high mileage," but after diffusion, it might incorporate nodes such as "decreased engine performance" and "lack of regular maintenance," thus forming a more complete "powertrain health status" topic region. Next, the system calculates topic discriminant functions for the semantic regions within the formed multi-level semantic agglomeration structure to determine the topic affiliation probability distribution for each region. The topic discriminant function is typically based on Bayesian inference or neural network classification models, combined with existing domain knowledge bases (such as automotive repair terminology and used car evaluation standards), assigning a set of possible topic labels and their corresponding probabilities to each semantic region. For example, a semantic region might be classified as having a probability of 0.7 for the topic "vehicle condition influencing factors" and 0.3 for the topic "parts depreciation analysis," thus achieving soft topic segmentation. Based on this, the system further propagates hierarchical topic labels across semantic regions, generating a preliminary topic domain hierarchy mapping, including top-level abstract topics (such as "vehicle value assessment"), mid-level categorical topics (such as "vehicle condition analysis" and "price fluctuation factors"), and bottom-level fine-grained topic labels (such as "whether an accident has occurred" and "whether there is structural damage"). This hierarchical structure not only helps in understanding the semantic organization behind the data but also provides a clear semantic navigation system for subsequent knowledge mining and intelligent querying. Finally, based on this initial topic domain hierarchical mapping, the system optimizes the topic structure of the semantic feature graph, integrates the topic division results from each level, removes redundant, conflicting, or ambiguous topic nodes, and organizes them into a clearly structured and semantically coherent topic domain knowledge tree through graph reconstruction technology. For example, in used car transaction evaluation, this knowledge tree can clearly express the causal chain of "vehicle year → mileage → maintenance records → market pricing," helping users more accurately understand the intrinsic relationship between vehicle condition and its market value. In summary, this step, through a series of graph-based semantic modeling and topic clustering techniques, achieves an efficient conversion from a semantic feature graph to a topic domain knowledge tree. Throughout the process, the rich semantic relationships in the original semantic graph are preserved, while a hierarchical, easy-to-understand, and applicable topic knowledge system is constructed, laying a solid foundation for subsequent implicit knowledge mining and intelligent query expansion. In practical applications of used car price assessment, the ability to construct such a thematic structure is particularly important. It can help the system accurately identify key factors affecting the residual value of a vehicle and improve the interpretability and practicality of data analysis through a structured thematic organization.
[0081] In a specific embodiment, the step of performing association rule mining on the topic domain knowledge tree to obtain a hidden knowledge network includes:
[0082] The branch topic nodes in the topic domain knowledge tree are extracted based on the information entropy increase of the context co-occurrence pattern to obtain the high-frequency semantic co-occurrence sequence, and the high-frequency semantic co-occurrence sequence is probabilistically dependent modeled based on the conditional random field to obtain the semantic dependency probability map.
[0083] Based on the semantic dependency probability graph, multi-granularity path sampling is performed on leaf-level specific topic instances in the topic domain knowledge tree to obtain a set of topic evolution paths. Knowledge flow propagation modeling is then performed on the set of topic evolution paths to obtain the causal transmission weight matrix between topics.
[0084] The causal transmission weight matrix is calculated using the maximum likelihood estimation method to obtain a set of potential association rules. Based on the set of potential association rules, a bidirectional reasoning chain is constructed for the root topic node in the topic domain knowledge tree to obtain a cross-domain knowledge reasoning path.
[0085] Based on the cross-domain knowledge reasoning path, the topic domain knowledge tree is extended by knowledge entity mapping to obtain the implicit knowledge network.
[0086] Specifically, in the step of "mining association rules from the topic domain knowledge tree to obtain an implicit knowledge network," the core objective is to uncover potential knowledge associations hidden beneath the explicit structure by deeply analyzing the semantic dependencies and contextual co-occurrence patterns among nodes within the topic domain knowledge tree, and to construct an implicit knowledge network with causal reasoning capabilities. This process includes several key implementation steps: high-frequency semantic co-occurrence sequence extraction, semantic dependency probability graph modeling, topic evolution path sampling and propagation modeling, causal transmission weight matrix calculation, potential association rule discovery, cross-domain reasoning chain construction, and knowledge entity mapping expansion. First, in the process of extracting high-frequency semantic co-occurrence sequences from the branch topic nodes in the topic domain knowledge tree based on information entropy increase, the system starts from each branch node of the knowledge tree, scans the semantic combinations that frequently co-occur in their context, and identifies co-occurrence patterns with statistical significance. For example, in the application scenario of used car evaluation, certain topic nodes such as "high mileage," "transmission replacement," and "engine abnormal noise" may frequently appear in the same vehicle record, forming a stable semantic co-occurrence sequence. This process typically employs a sliding window mechanism combined with incremental information entropy analysis to filter out co-occurring sequences that are semantically highly relevant and have a high frequency of distribution, providing data support for subsequent probabilistic modeling. Subsequently, based on these high-frequency semantic co-occurrence sequences, the system further utilizes sequence modeling techniques such as Conditional Random Fields (CRFs) to construct a semantic dependency probability graph. This graph structure not only characterizes the co-occurrence frequency between different topic nodes but also represents the strength of their dependencies through probability edges. For example, if a sequence shows that "high mileage" always appears before "gearbox replacement," the system will establish a directional semantic dependency edge between the two and assign a corresponding transition probability value, thus constructing a semantic dependency probability graph that reflects the semantic evolution trend. Based on this, the system will perform multi-granularity path sampling on leaf-level specific topic instances in the topic domain knowledge tree based on this semantic dependency probability graph to obtain a set of topic evolution paths. The so-called "leaf-level specific topic instances" refer to the lowest level and finest-grained topic expressions in the knowledge tree, such as "whether an accident has occurred" or "whether chassis components have been replaced." By simulating multiple possible paths from the root node to the leaf node and performing weighted sampling based on the transition probabilities set in the semantic dependency probability graph, the system can generate a large set of topic evolution paths with practical semantic meaning. These paths reflect the typical evolutionary order and logical relationship between different topics in real-world data. Next, the system performs knowledge flow propagation modeling on the collected set of topic evolution paths to quantify the causal transmission strength between topics and ultimately form a causal transmission weight matrix.This process typically employs Bayesian network propagation models or Markov chain methods to simulate how knowledge flows along paths between topics and calculates the indirect influence weights between each pair of topics. For example, in the used car market, "high mileage" may indirectly affect "increased maintenance costs" through "gearbox wear," and this transmission relationship is captured and reflected in the causal transmission weight matrix. Subsequently, the system calculates the rule confidence of this causal transmission weight matrix based on the maximum likelihood estimation method, identifying potential association rule sets with high confidence. These rules are not merely simple co-occurrence relationships but knowledge expressions with a certain causal explanatory power, such as "if a vehicle has structural damage, its residual value will decrease by more than 15%" or "if the maintenance cycle exceeds the recommended time by 30%, the risk of engine performance degradation increases by 40%." These potential association rules constitute the core reasoning units of the knowledge network. Furthermore, based on these potential association rules, the system constructs a bidirectional reasoning chain for the root topic node in the topic domain knowledge tree, that is, it performs forward reasoning from the root node to the leaf node, and simultaneously performs reverse verification from the leaf node back to the root node, thus forming a cross-domain knowledge reasoning path. For example, when assessing the overall value of a vehicle, the system can start from the top-level topic "vehicle residual value prediction," and reason along the path to deduce the impact of sub-topics such as "high mileage," "transmission replacement," and "lack of regular maintenance." It can then verify whether these sub-topics reasonably support the residual value prediction result, thus forming a closed-loop reasoning path. Finally, based on these cross-domain knowledge reasoning paths, the system expands the topic domain knowledge tree by mapping knowledge entities, mapping abstract semantic nodes that originally only existed in the topic hierarchy to concrete knowledge entities in the real world (such as part names, repair item numbers, price indices, etc.), thereby constructing a comprehensive and semantically rich implicit knowledge network. For example, a topic node "powertrain failure rate" can be mapped to a specific repair record database, corresponding to entity indicators such as "transmission replacement frequency" and "clutch replacement frequency," making the entire knowledge network not only possess logical reasoning capabilities but also real-world data support. In summary, this step, through a series of techniques based on probabilistic graphs, path modeling, and knowledge reasoning, achieves an intelligent transformation from a topic domain knowledge tree to an implicit knowledge network. Throughout the process, the hierarchical and interpretable nature of the topic structure was preserved, while causal reasoning and entity mapping mechanisms were introduced, resulting in a latent knowledge network with powerful semantic understanding capabilities and business application value. In practical applications of used car appraisal, this latent knowledge network can help the system more accurately identify key factors affecting vehicle residual value and reveal the complex relationship between parts replacement, repair history, and market prices through cross-topic reasoning paths, thereby providing more intelligent and accurate support for user decision-making.
[0087] In a specific embodiment, the semantic decomposition and expansion of the preset query conditions based on the implicit knowledge network to obtain extended query data includes:
[0088] Deep syntactic dependency analysis is performed on the preset query conditions to obtain a query semantic structure tree. Key entities are extracted and relationships are identified on the query semantic structure tree to obtain a query intent skeleton. The query intent skeleton includes core query entities, query constraints, inter-entity relationship links, and implicit query intent.
[0089] Based on the query intent skeleton, semantic neighborhood activation is performed on the implicit knowledge network to obtain a related knowledge subgraph. Then, semantic path traversal and aggregation are performed on the related knowledge subgraph to obtain an extended semantic set, which includes synonymous concept groups, hierarchical concept chains, cross-domain related concepts, and temporal evolution variants.
[0090] The extended semantic set is transformed into a vector representation by distributed semantic coding to obtain a semantic embedding matrix. The semantic embedding matrix is then fused with attention weighting to obtain a context-enhanced query representation, which includes an explicit query vector, an implicit intent vector, a relational constraint vector, and a timeliness weight distribution.
[0091] Based on the context-enhanced query representation, a knowledge reasoning chain is constructed on the implicit knowledge network to obtain a multi-path reasoning graph. Joint probability inference is then performed on the multi-path reasoning graph to obtain query extended data.
[0092] Specifically, in the step of "semantically decomposing and expanding preset query conditions based on implicit knowledge networks to obtain extended query data," the core objective is to achieve multi-dimensional semantic enhancement and expansion of the query intent by deeply analyzing the original query semantics input by the user and combining it with the structured semantic relationships in the constructed implicit knowledge network, thereby generating extended query data with higher coverage and business relevance. This process consists of four key stages: query semantic structure analysis, knowledge subgraph activation and semantic aggregation, context-enhanced representation learning, and multi-path reasoning and joint inference. First, in the process of performing deep syntactic dependency analysis on preset query conditions to obtain a query semantic structure tree, the system uses natural language processing technology, especially dependency parsing-based methods, to model the syntactic structure of the user's original query statement. For example, in a used car evaluation scenario, if a user inputs "I want to find a car with a replaced gearbox but still in good condition," the system will identify "find" as the core verb, "vehicle" as the search object, "replaced gearbox" as a constraint attribute, and "good condition" as an additional condition, and construct a query semantic structure tree reflecting the semantic master-slave relationship accordingly. Subsequently, the system further extracts key entities and identifies relationships within the tree structure, extracting core query entities such as "transmission replacement," "mileage," and "maintenance records," as well as the logical relationships between them (e.g., parallel, limiting, causal). This ultimately forms a query intent skeleton, which includes not only explicit query content but also implicit intents that the user may not have explicitly expressed, such as potential needs like "reasonable price" and "no structural damage." Next, the system performs semantic neighborhood activation on the implicit knowledge network based on the extracted query intent skeleton. The aim is to quickly locate knowledge subgraphs related to the current query within the vast knowledge graph. Specifically, the system uses the core entity in the query as a starting point and performs a breadth-first or depth-first search within the implicit knowledge network, activating semantic neighbor nodes within a certain range around it to form a localized related knowledge subgraph. For example, for the core entity "transmission replacement," the system might activate related nodes such as "high mileage," "missing regular maintenance," and "decreased engine performance," constructing a semantic neighborhood network around this entity. Based on this, the system performs semantic path traversal and aggregation operations on the relevant knowledge subgraph to identify concept sets that are directly or indirectly related to the query topic, ultimately forming an extended semantic set. This set typically includes synonymous concept groups (such as "transmission replacement"), hierarchical concept chains (such as "powertrain failure" and "component aging"), cross-domain related concepts (such as "increased maintenance costs"), and temporal evolution variants (such as "uptime after replacement" and "subsequent changes in failure rate"). Subsequently, the system performs vector representation transformation on the above extended semantic set through distributed semantic encoding, mapping it to a unified semantic embedding space to form a semantic embedding matrix.This process typically employs semantic encoding models such as BERT, TransE, and GraphSAGE, ensuring that each semantic concept has a corresponding low-dimensional dense representation in the vector space. To further enhance the context adaptability of the semantic representation, the system performs attention-weighted fusion on the semantic embedding matrix, dynamically adjusting the weights of each semantic component by incorporating factors such as the user's historical query behavior, regional preferences, and timeliness, ultimately generating a context-enhanced query representation. This representation not only includes explicit query vectors (e.g., "transmission replacement") but also integrates implicit intent vectors (e.g., "hope for good vehicle condition"), relational constraint vectors (e.g., "no major accidents"), and timeliness weight distributions (e.g., "recent repairs are more important"), thus forming a highly semantic composite query expression. Finally, based on this context-enhanced query representation, the system constructs a knowledge reasoning chain for the implicit knowledge network, aiming to discover multiple reasoning paths highly relevant to the current query intent, thereby forming a multi-path reasoning graph. For example, when evaluating a vehicle whose transmission has been replaced, the system might infer multiple causal chains such as "high mileage → transmission wear → replacement records → increased maintenance costs → market price reduction," and simulate the semantic transmission mechanism of these paths using graph neural networks or probabilistic graphical models. Based on this, the system performs joint probabilistic inference on the multi-path reasoning graph, comprehensively considering the confidence, support, and relevance of each path to calculate the contribution of each path to the query results, and generates the final extended query data accordingly. This extended data not only includes the content mentioned in the user's original query but also relevant information derived from knowledge network reasoning, such as "it is recommended to check maintenance records" and "pay attention to whether there is structural damage," thus significantly improving the intelligence level and completeness of the query results. In summary, this step, through a series of techniques based on semantic structure analysis, knowledge graph activation, vectorized representation learning, and multi-path reasoning, achieves the intelligent transformation from the user's original query statement to semantically rich and comprehensive extended query data. Throughout the process, the semantic integrity of the original query was preserved, while the structured semantic relationships of the implicit knowledge network were leveraged to uncover a wealth of potential business-related information, providing a solid foundation for subsequent data matching and sorting. This query expansion capability is particularly important in the practical application of used car price assessment. It helps the system accurately capture the user's true intent and recommend vehicle information that better meets actual needs through a knowledge-driven approach, thereby improving the overall intelligence level of the service and the quality of the user experience.
[0093] In a specific embodiment, the step of calculating the similarity between the query extended data and the implicit knowledge network to obtain a candidate dataset includes:
[0094] The extended query data is subjected to semantic tensor decomposition to obtain multi-dimensional query feature vectors, and the multi-dimensional query feature vectors are dynamically weighted to obtain an adaptive query representation.
[0095] Based on the adaptive query representation, multi-scale semantic matching is performed on the implicit knowledge network to obtain candidate knowledge subgraphs, and graph neural network embedding is performed on the candidate knowledge subgraphs to obtain a knowledge graph embedding matrix.
[0096] The adaptive query representation and the knowledge graph embedding matrix are subjected to high-order similarity calculation using tensor decomposition technology to obtain a multimodal similarity tensor. Then, nonlinear dimensionality reduction and cluster analysis are performed on the multimodal similarity tensor to obtain the similarity distribution spectrum.
[0097] Based on the similarity distribution spectrum, the implicit knowledge network is subjected to adaptive threshold segmentation to obtain candidate datasets.
[0098] Specifically, in the step of "calculating similarity between the query extended data and the implicit knowledge network to obtain a candidate dataset," the core lies in achieving efficient matching between user query intent and entities in the knowledge graph through multi-level, multi-dimensional semantic modeling and matching mechanisms, and thereby selecting the most relevant set of data as the basis for subsequent sorting and extraction. This process consists of four key steps: multi-dimensional query feature vector construction, multi-scale semantic matching and graph embedding, high-order similarity tensor calculation, and adaptive threshold segmentation to generate a candidate dataset. First, in the process of performing semantic tensor decomposition on the query extended data to obtain multi-dimensional query feature vectors, the system converts the context-enhanced query representation (such as the explicit query vector and implicit intent vector generated in the previous step) into a high-dimensional tensor structure, and then decomposes it into multiple low-dimensional feature vectors with clear semantic meanings through tensor decomposition techniques (such as CP decomposition or Tucker decomposition). For example, in the application scenario of used car appraisal, if a user's query includes multiple extended semantic nodes such as "transmission replacement," "low mileage," and "good vehicle condition," the system can map these into feature vectors corresponding to different dimensions such as vehicle mechanical condition, usage intensity, and exterior maintenance, thus forming a multi-dimensional query feature vector set. Subsequently, the system further dynamically assigns weights to these feature vectors, adjusting the importance of each dimension in the overall query based on factors such as the user's historical preferences, query timeliness, and regional differences, ultimately generating an adaptive query representation that more accurately reflects the core requirements of the current task. Next, based on this adaptive query representation, the system performs multi-scale semantic matching on the implicit knowledge network, aiming to identify local subgraph structures highly relevant to the current query within the knowledge graph. Specifically, the system employs a graph search strategy, selecting knowledge nodes from the knowledge network that have direct or indirect connections with the query entity to construct candidate knowledge subgraphs. For example, if the query involves "transmission replacement," the system will not only match vehicle entries that explicitly indicate this repair record, but also identify entries that, while not directly mentioned, contain related semantics such as "powertrain failure" or "engine performance degradation," thereby expanding the matching scope and improving the recall rate. Subsequently, the system uses a graph neural network (GNN) to embed the candidate knowledge subgraph, mapping all node and edge information into a shared low-dimensional vector space to form a knowledge graph embedding matrix. Each vector in this matrix represents a knowledge entity or relation, preserving its semantic location information in the original graph. Based on this, the system uses tensor decomposition technology to perform high-order similarity calculations between the adaptive query representation and the knowledge graph embedding matrix to capture the complex semantic interactions between them.Specifically, the system combines the query feature vector with the knowledge graph embedding vector into a three-dimensional tensor structure. Then, it calculates the multimodal similarity score between each query-knowledge entity pair using tensor multiplication, cosine similarity, or other neural network scoring functions. These scores are organized into a multimodal similarity tensor, where each layer corresponds to a semantic matching mode (e.g., word-level matching, relation-level matching, path-level matching). To further improve the interpretability and practicality of the results, the system performs nonlinear dimensionality reduction (e.g., t-SNE or UMAP) and cluster analysis on the multimodal similarity tensor to identify data clusters with high similarity, ultimately forming a similarity distribution spectrum. This spectrum reflects the semantic affinity between the query and various entities in the knowledge graph, aiding in subsequent precise filtering. Finally, based on this similarity distribution spectrum, the system performs adaptive threshold segmentation on the implicit knowledge network. Specifically, it dynamically sets a reasonable similarity threshold based on the current query's confidence level, the urgency of user needs, and historical behavioral data, including all knowledge entities above this threshold in the candidate dataset. For example, in the practical application of used car price evaluation, if a vehicle does not explicitly record "transmission replacement," but its indicators such as "low powertrain health score" and "high repair frequency" highly match the query conditions, it will still be included in the candidate dataset, thereby improving the comprehensiveness and intelligence of the recommendation results. This adaptive mechanism ensures that the candidate dataset covers the user's main concerns while avoiding the introduction of excessive noise interference, achieving a balance between accuracy and recall. In summary, this step, through semantic tensor modeling, graph neural network embedding, multimodal similarity calculation, and adaptive segmentation, achieves an efficient semantic matching process from user query to candidate dataset. Throughout the process, the multidimensional characteristics of query semantics are preserved, and the structured semantic information of implicit knowledge networks is fully utilized, resulting in a candidate dataset with high semantic consistency and business relevance. In the practical application scenario of used car evaluation, this mechanism can help the system accurately identify vehicle information that highly matches user needs, providing a high-quality data foundation for subsequent data sorting and target extraction, significantly improving the overall system's intelligent recommendation capabilities and user experience quality.
[0099] In a specific embodiment, the step of performing multi-factor comprehensive ranking and extraction on the candidate dataset to obtain the target extraction data includes:
[0100] Multi-dimensional feature extraction is performed on the candidate dataset to obtain a candidate data feature matrix, and nonlinear feature interaction modeling is performed on the candidate data feature matrix to obtain an interaction feature tensor.
[0101] Based on the interactive feature tensor, dynamic temporal correlation analysis is performed on the candidate dataset to obtain a temporal evolution sequence, and multi-scale wavelet transform is performed on the temporal evolution sequence to obtain a multi-resolution time-frequency feature map.
[0102] An adaptive attention mechanism is used to weight the importance of the multi-resolution time-frequency feature map to obtain a fused feature vector, and a nonlinear projection transformation is performed on the fused feature vector to obtain a high-dimensional sorting space.
[0103] Based on the high-dimensional sorting space, the candidate dataset is sorted using multi-objective Pareto optimization to obtain preliminary sorting results. Target features are then extracted from the preliminary sorting results to obtain target extracted data.
[0104] Specifically, in the step of "comprehensively ranking and extracting candidate datasets from multiple factors to obtain target extraction data," the core objective is to construct a complex ranking model that integrates semantics, temporal sequence, and multidimensional features to select the target extraction data that best meets user needs or business objectives from the candidate datasets. This process includes several key technical steps: multidimensional feature extraction and interactive modeling, dynamic temporal correlation analysis and time-frequency feature extraction, feature fusion and spatial mapping under an attention mechanism, and multi-objective optimization ranking based on a high-dimensional ranking space. First, in the process of extracting multidimensional features from the candidate datasets to obtain the candidate data feature matrix, the system extracts various types of features from the candidate vehicle records, including structured fields (such as mileage and registration year), semantic tags (such as whether an accident has occurred or whether the engine has been replaced), price fluctuation trends, and regional market average prices, and organizes these heterogeneous information into a structured feature matrix. Each row corresponds to a candidate data record, and each column represents a standardized feature dimension. Subsequently, the system further performs nonlinear feature interaction modeling on the candidate data feature matrix. Utilizing cross layers, product layers, or multi-head interaction modules in deep neural networks, it automatically identifies and models complex combinations of different features. For example, "high mileage + transmission replacement" may have a greater impact on vehicle residual value than a single feature, thus generating an interaction feature tensor to enhance the model's understanding of the inherent patterns in the data. Next, based on this interaction feature tensor, the system performs dynamic temporal correlation analysis on the candidate dataset, aiming to capture the evolution trend of candidate vehicles over historical time. Specifically, the system arranges each vehicle's historical repair records, price changes, maintenance cycles, etc., in chronological order, constructing a temporal evolution sequence reflecting its state evolution. For example, in a used car evaluation scenario, a candidate vehicle might have had its transmission replaced in 2021, experienced engine failure in 2023, and had a decreased maintenance frequency in 2024; these events constitute an evolution path with temporal logic. To more precisely characterize this evolutionary feature, the system further performs multi-scale wavelet transform on the temporal evolution sequence, extracting its feature representation at different frequency levels to form a multi-resolution time-frequency feature map. This operation helps identify the difference between short-term fluctuations and long-term trends, improving the ranking model's sensitivity to dynamic data changes. Subsequently, the system uses an adaptive attention mechanism to perform importance-weighted processing on the obtained multi-resolution time-frequency feature maps, aiming to identify the most influential feature segments in the current query context. Specifically, the system dynamically adjusts the attention weight distribution based on contextual preferences in the query's extended data (e.g., whether the user's focus is on "low mileage" or "no major accidents"), ensuring that certain key time points or specific frequency components occupy a higher proportion in the overall feature representation.For example, if a user is particularly concerned about the recent usage of a vehicle, the system will assign a higher attention weight to maintenance records from 2024, while relatively reducing the impact of earlier repair events. Ultimately, the system fuses the weighted time-frequency features with interaction features to form a fused feature vector, and maps it to a high-dimensional ranking space through a nonlinear projection transformation (such as a fully connected network or kernel mapping), giving each candidate data point a comparable ranking representation within this space. Based on this, the system performs multi-objective Pareto optimization ranking on the candidate datasets using this high-dimensional ranking space to consider multiple conflicting but equally important evaluation metrics. For example, in used car transactions, users may simultaneously focus on multiple aspects such as price reasonableness, vehicle condition stability, maintenance costs, and geographical convenience. These objectives may be contradictory (e.g., a vehicle in good condition but with a high price), making ranking impossible using a single scoring function. In this case, the system employs a multi-objective evolutionary algorithm (such as NSGA-II) or a gradient-based Pareto front search method to find a set of data points in the high-dimensional ranking space that satisfy the optimal trade-off between the objectives, forming a preliminary ranking result. These results not only excel in a single dimension but also achieve a balance among multiple objectives overall. Finally, the system extracts target features from the preliminary ranking results. Based on actual business needs or user preferences, it extracts several key feature fields from the top-ranked candidate data as the final output target extraction data. For example, in a user's query task, the system might return several vehicles as recommendations that are "in good condition, have less than 80,000 kilometers on the odometer, have no structural damage, and have not undergone major repairs in the past year." Each result includes detailed feature descriptions and scoring criteria, providing users with clear, reliable, and highly relevant decision support. In summary, this step, by integrating multi-dimensional feature modeling, temporal evolution analysis, attention mechanisms, and multi-objective optimization, achieves intelligent ranking and accurate extraction from candidate datasets to target extraction data. Throughout the process, the semantic integrity of the original data is preserved while introducing complex interactive modeling and dynamic reasoning mechanisms, resulting in highly personalized matching capabilities and business applicability of the final target extraction data. In practical applications of used car evaluation, this multi-factor comprehensive ranking mechanism is particularly important. It can help the system quickly locate the high-quality vehicle resources that best match the user's true intentions from massive candidate data, significantly improving the platform's intelligent service capabilities and user experience quality.
[0105] The above describes the method for extracting internet big data in the embodiments of the present invention. The following describes the device for extracting internet big data in the embodiments of the present invention. Please refer to [link / reference]. Figure 2 One embodiment of the Internet big data extraction device in this invention includes:
[0106] The data acquisition module 21 is used to perform distributed crawling to collect data from the Internet data source, obtain the original network dataset, and perform heterogeneous data structuring transformation on the original network dataset to obtain a structured data matrix.
[0107] Parsing module 22 is used to perform multi-level semantic parsing on the structured data matrix to obtain a semantic feature map;
[0108] The classification module 23 is used to perform topic segmentation and classification on the semantic feature map to obtain a topic domain knowledge tree, and to perform association rule mining on the topic domain knowledge tree to obtain a hidden knowledge network.
[0109] Extension module 24 is used to perform semantic decomposition and expansion of preset query conditions based on the implicit knowledge network to obtain query extended data;
[0110] The calculation module 25 is used to perform similarity calculation between the query extended data and the implicit knowledge network to obtain a candidate dataset;
[0111] The sorting module 26 is used to perform multi-factor comprehensive sorting and extraction on the candidate dataset to obtain the target extraction data.
[0112] In this embodiment, the specific implementation of each unit in the above device embodiment is described in the above method embodiment, and will not be repeated here.
[0113] Reference Figure 3 This invention also provides a computer device whose internal structure can be as follows: Figure 3 As shown, the computer device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the data corresponding to this embodiment. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.
[0114] Those skilled in the art will understand that Figure 3 The structures shown are merely block diagrams of some structures related to the present invention and do not constitute a limitation on the computer devices on which the present invention is applied.
[0115] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0116] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the present invention and embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0117] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0118] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. An Internet big data extraction method, characterized in that, The method comprises the following steps: Distributed crawler collection is performed on an Internet data source to obtain an original network data set, and heterogeneous data structural conversion is performed on the original network data set to obtain a structured data matrix; Multi-level semantic analysis is performed on the structured data matrix to obtain a semantic feature map; Theme segmentation and classification are performed on the semantic feature map to obtain a theme domain knowledge tree, and association rule mining is performed on the theme domain knowledge tree to obtain an implicit knowledge network; Based on the implicit knowledge network, semantic decomposition and expansion are performed on a preset query condition to obtain query expansion data; Similarity calculation is performed on the query expansion data and the implicit knowledge network to obtain a candidate data set; Multi-factor comprehensive sorting and extraction are performed on the candidate data set to obtain target extraction data; The theme segmentation and classification of the semantic feature map to obtain the theme domain knowledge tree comprises: Local neighborhood topology construction is performed on high-dimensional semantic nodes in the semantic feature map to obtain an adjacency similarity matrix, and feature subspace decomposition based on spectral analysis is performed on the adjacency similarity matrix to obtain a low-rank embedding vector field; Based on the low-rank embedding vector field, dynamic density perception division is performed on the high-dimensional semantic nodes to obtain initial clustering seed clusters, and adaptive diffusion fusion operations are performed on the initial clustering seed clusters to obtain a multi-level semantic condensed structure; Theme discriminant function calculation is performed on semantic regions in the multi-level semantic condensed structure to obtain a theme attribution probability distribution, and hierarchical theme label propagation is performed on the semantic regions based on the theme attribution probability distribution to obtain a preliminary theme domain hierarchical mapping; wherein the preliminary theme domain hierarchical mapping comprises top-level abstract themes, middle-level classification themes, and bottom-level fine-grained theme labels; Based on the preliminary theme domain hierarchical mapping, theme structure optimization is performed on the semantic feature map to obtain the theme domain knowledge tree; The association rule mining of the theme domain knowledge tree to obtain the implicit knowledge network comprises: Context co-occurrence pattern extraction based on information entropy increase is performed on branch theme nodes in the theme domain knowledge tree to obtain high-frequency semantic co-occurrence sequences, and probability dependence modeling is performed on the high-frequency semantic co-occurrence sequences based on a conditional random field to obtain a semantic dependence probability graph; Multi-granularity path sampling is performed on leaf-level specific theme instances in the theme domain knowledge tree based on the semantic dependence probability graph to obtain a theme evolution path set, and knowledge flow propagation modeling is performed on the theme evolution path set to obtain an inter-theme causal conduction weight matrix; Rule confidence calculation is performed on the causal conduction weight matrix based on a maximum likelihood estimation method to obtain a set of potential association rules, and bidirectional reasoning chain construction is performed on root theme nodes in the theme domain knowledge tree based on the set of potential association rules to obtain a cross-domain knowledge reasoning path; Based on the cross-domain knowledge reasoning path, knowledge entity mapping expansion is performed on the theme domain knowledge tree to obtain the implicit knowledge network. 2.The Internet big data extraction method of claim 1, wherein, The heterogeneous data structural conversion of the original network data set to obtain the structured data matrix comprises: The original network data set is subjected to heterogeneous data parsing and classification annotation to obtain a type recognition result, and the original network data set is subjected to data cleaning and format unification processing based on the type recognition result to obtain a standardized data stream; The standardized data stream is subjected to feature extraction and dimension mapping by a deep semantic vectorization method to obtain a high-dimensional feature tensor, and the high-dimensional feature tensor is subjected to sparsity optimization and dimension reduction processing to obtain a compact feature representation; A heterogeneous data association network is constructed based on the compact feature representation to obtain a data element relationship graph, and the data element relationship graph is subjected to structured conversion to obtain a structured data matrix. 3.The Internet big data extraction method of claim 1, wherein, The query expansion data is obtained by performing semantic decomposition and expansion on the preset query condition based on the implicit knowledge network, including: Deep syntactic dependency analysis is performed on the preset query condition to obtain a query semantic structure tree, and key entity extraction and relationship recognition are performed on the query semantic structure tree to obtain a query intent skeleton, wherein the query intent skeleton includes a core query entity, a query constraint condition, an entity relationship link, and an implicit query intent; The implicit knowledge network is activated based on the query intent skeleton to obtain a related knowledge subgraph, and semantic path traversal and aggregation are performed on the related knowledge subgraph to obtain an expanded semantic set, wherein the expanded semantic set includes a synonymous concept group, an upper and lower concept chain, a cross-domain associated concept, and a time sequence evolution variant; The expanded semantic set is converted into a vector representation by distributed semantic coding to obtain a semantic embedding matrix, and the semantic embedding matrix is subjected to attention weighted fusion to obtain a context enhanced query representation, wherein the context enhanced query representation includes an explicit query vector, an implicit intent vector, a relationship constraint vector, and a timeliness weight distribution; The knowledge reasoning chain of the implicit knowledge network is constructed based on the context enhanced query representation to obtain a multi-path reasoning graph, and joint probability inference is performed on the multi-path reasoning graph to obtain query expansion data. 4.The Internet big data extraction method of claim 1, wherein, The similarity between the query expansion data and the implicit knowledge network is calculated to obtain a candidate data set, including: The query expansion data is subjected to semantic tensor decomposition to obtain a multi-dimensional query feature vector, and the multi-dimensional query feature vector is subjected to dynamic weight distribution to obtain an adaptive query representation; The adaptive query representation is subjected to multi-scale semantic matching on the implicit knowledge network to obtain a candidate knowledge subgraph, and the candidate knowledge subgraph is subjected to graph neural network embedding to obtain a knowledge graph embedding matrix; The adaptive query representation and the knowledge graph embedding matrix are subjected to high-order similarity calculation by tensor decomposition technology to obtain a multi-modal similarity tensor, and the multi-modal similarity tensor is subjected to nonlinear dimension reduction and clustering analysis to obtain a similarity distribution spectrum; The candidate data set is obtained by adaptively thresholding the implicit knowledge network based on the similarity distribution spectrum. 5.The Internet big data extraction method of claim 1, wherein, The target extraction data is obtained by performing multi-factor comprehensive sorting and extraction on the candidate data set, including: Multi-dimensional feature extraction is performed on the candidate data set to obtain a candidate data feature matrix, and nonlinear feature interaction modeling is performed on the candidate data feature matrix to obtain an interaction feature tensor; Based on the interaction feature tensor, dynamic time-series correlation analysis is performed on the candidate data set to obtain a time-series evolution sequence, and multi-scale wavelet transform is performed on the time-series evolution sequence to obtain a multi-resolution time-frequency feature map; An importance weighting is performed on the multi-resolution time-frequency feature map through an adaptive attention mechanism to obtain a fusion feature vector, and a nonlinear projection transformation is performed on the fusion feature vector to obtain a high-dimensional ranking space; Based on the high-dimensional ranking space, multi-objective Pareto optimization ranking is performed on the candidate data set to obtain a preliminary ranking result, and target feature extraction is performed on the preliminary ranking result to obtain target extraction data.
6. An Internet big data extraction device, characterized by, It includes: The acquisition module is used for distributed crawler acquisition of the Internet data source to obtain an original network data set, and is used for heterogeneous data structured conversion of the original network data set to obtain a structured data matrix; The analysis module is used for multi-level semantic analysis of the structured data matrix to obtain a semantic feature map; The classification module is used for topic segmentation and classification of the semantic feature map to obtain a topic domain knowledge tree, and is used for association rule mining of the topic domain knowledge tree to obtain an implicit knowledge network; The expansion module is used for semantic decomposition and expansion of a preset query condition based on the implicit knowledge network to obtain query expansion data; The calculation module is used for similarity calculation of the query expansion data and the implicit knowledge network to obtain a candidate data set; The sorting module is used for multi-factor comprehensive sorting and extraction of the candidate data set to obtain target extraction data; Wherein, the topic segmentation and classification of the semantic feature map to obtain the topic domain knowledge tree includes: Local neighborhood topology construction is performed on the high-dimensional semantic nodes in the semantic feature map to obtain an adjacency similarity matrix, and feature subspace decomposition based on spectral analysis is performed on the adjacency similarity matrix to obtain a low-rank embedding vector field; Based on the low-rank embedding vector field, dynamic density perception division is performed on the high-dimensional semantic nodes to obtain initial clustering seed clusters, and adaptive diffusion fusion operation is performed on the initial clustering seed clusters to obtain a multi-level semantic condensation structure; Topic discriminant function calculation is performed on the semantic regions in the multi-level semantic condensation structure to obtain a topic attribution probability distribution, and hierarchical topic label propagation is performed on the semantic regions based on the topic attribution probability distribution to obtain a preliminary topic domain hierarchical mapping; wherein, the preliminary topic domain hierarchical mapping includes top-level abstract topics, middle-level classification topics and bottom-level fine-grained topic labels; Based on the preliminary topic domain hierarchical mapping, topic structure optimization is performed on the semantic feature map to obtain a topic domain knowledge tree; The association rule mining of the topic domain knowledge tree to obtain the implicit knowledge network includes: extracting a context co-occurrence pattern based on information entropy increase on a branch topic node in the subject domain knowledge tree to obtain a high-frequency semantic co-occurrence sequence, and performing probability dependent modeling on the high-frequency semantic co-occurrence sequence based on a conditional random field to obtain a semantic dependent probability graph; performing multi-granularity path sampling on a leaf-level specific topic instance in the subject domain knowledge tree based on the semantic dependent probability graph to obtain a topic evolution path set, and performing knowledge flow propagation modeling on the topic evolution path set to obtain a causal conduction weight matrix between topics; performing rule confidence calculation on the causal conduction weight matrix based on a maximum likelihood estimation method to obtain a set of potential association rules, and constructing a bidirectional reasoning chain for a root topic node in the subject domain knowledge tree based on the set of potential association rules to obtain a cross-domain knowledge reasoning path; performing knowledge entity mapping expansion on the subject domain knowledge tree based on the cross-domain knowledge reasoning path to obtain an implicit knowledge network. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 5.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Implementation method of science and technology project text mining
CN114265936A
System and method of computational social network development environment for human intelligence
US9317567B1