Competitive intelligence multi-source heterogeneous collection and quantification method and system for dynamic information targets
Patent Information
- Application Number
- CN202611307314.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-27
- Publication Date
- 2026-09-25
AI Technical Summary
传统的情报采集方法主要依赖于人工设定关键词进行定期检索与筛选,其不足之处在于:首先,该方法高度依赖专家经验,难以覆盖目标实体的隐性关联与动态演变关系,导致情报视野狭窄且易遗漏关键信息;其次,面对多源异构数据,传统方法缺乏统一的知识表示与融合机制,信息整合困难,形成“信息孤岛”;最后,采集过程是开环的,对采集结果的质量缺乏客观、量化的评估标准,无法验证情报的有效性和对新动态的预测能力
[0016]与现有技术相比,本发明通过构建显性子图谱与隐性子图谱,并采用基于注意力机制的双通道知识蒸馏网络进行深度融合,深入挖掘并融入了深层次的潜在关联;利用融合了时序权重的图嵌入算法生成语义向量,并基于语义向量的信息熵和边的时效权重智能构建语义关联导航图,自动规划最优探查路径并生成精准查询指令,从被动检索变为主动探查,提高了采集的针对性与效率;引入了预测性质量评估模型,将第二知识图谱输入时序图神经网络,预测未来时刻的知识图谱,并将其作为基准,与新采集的第二情报数据进行多维度量化比对,解决了情报质量评估缺乏前瞻性标准的问题,实现了对情报质量的前瞻性、可量化评估,提升了多源异构数据采集的准确性和全面性。
Smart Images

Figure CN122819261A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information retrieval and data processing technology, and more specifically, to a method and system for the multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets. Background Technology
[0002] The effectiveness of competitive intelligence directly depends on the timely and accurate acquisition, integration, and evaluation of relevant information from massive, multi-source, and heterogeneous public information sources such as government websites, news media, and corporate announcements. Traditional intelligence gathering methods mainly rely on manually setting keywords for periodic retrieval and screening. Their shortcomings are: First, this method heavily relies on expert experience, making it difficult to cover the implicit connections and dynamic evolution of target entities, resulting in a narrow intelligence perspective and the easy omission of key information; second, faced with multi-source heterogeneous data, traditional methods lack a unified knowledge representation and fusion mechanism, making information integration difficult and creating "information silos"; finally, the collection process is open-loop, lacking objective and quantitative evaluation standards for the quality of the collected results, making it impossible to verify the effectiveness of the intelligence and its predictive ability for new developments.
[0003] In recent years, mainstream methods have attempted to introduce knowledge graph technology to structurally represent intelligence entities and their relationships, and have improved the systematic nature and efficiency of information organization by automating data collection through web crawlers and API interfaces. However, existing methods still have the following bottlenecks to be addressed: First, the constructed knowledge graphs are mostly based on explicit textual relationships, making it difficult to mine and integrate potential and implicit connections between entities, resulting in limited depth and insight. Second, the construction and utilization of the graphs are often separate, lacking a feedback mechanism for intelligent, proactive, and focused collection based on the semantic and timeliness characteristics of the graph, making collection activities still somewhat blind and inefficient. Third, the quality assessment of intelligence data often remains at the level of superficial consistency and completeness checks, lacking a predictive and quantifiable deep assessment mechanism based on the evolutionary patterns of intelligence.
[0004] Therefore, designing an integrated method and system that can automatically mine and integrate explicit and implicit knowledge, intelligently guide subsequent collection directions based on the integrated knowledge, and conduct forward-looking quantitative evaluation of the collection results has become a key issue in improving the intelligence level and practical effectiveness of dynamic intelligence target competitive intelligence work. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this disclosure provides a method and system for the multi-source heterogeneous acquisition and quantification of competitive intelligence on dynamic intelligence targets.
[0006] In a first aspect, embodiments of this disclosure provide a method for quantifying competitive intelligence from multiple heterogeneous sources for dynamic intelligence targets, including: Based on the dynamic intelligence target, first intelligence data is collected, a first knowledge graph is constructed, the first knowledge graph includes explicit subgraphs and implicit subgraphs, and a second knowledge graph is generated through a preset knowledge graph fusion model; The second knowledge graph is encoded using a graph embedding algorithm to obtain semantic vectors; A semantic association graph is constructed based on the semantic vector, and the node series of the semantic association graph is converted into query instructions through a preset mapping rule. The query instructions are executed to collect data and obtain the second intelligence data. The second intelligence data is input into a preset quality assessment model to generate a quality assessment result.
[0007] In one optional implementation, the step of collecting first intelligence data based on dynamic intelligence targets and constructing a first knowledge graph, wherein the first knowledge graph includes explicit subgraphs and implicit subgraphs, includes: Based on the dynamic intelligence target, a keyword set is generated and collected in parallel from public information sources, including at least government department websites, news media platforms and corporate websites, to obtain the first intelligence data; Named entity recognition is performed on the first intelligence data to extract entities including at least organizations, people and products, and the relationships between the entities are extracted. Using the entities as nodes and the extracted relationships between entities as edges, an explicit subgraph is constructed.
[0008] In one optional implementation, the first knowledge graph includes a dominant subgraph and a recessive subgraph, comprising: The dominant subgraph is represented and learned using a graph convolutional network to obtain entity node vectors; Calculate the cosine similarity between any two entity node vectors, and treat the two entities as an entity pair, and count the co-occurrence frequency of the entity pair in the first intelligence data; The cosine similarity and the co-occurrence frequency are weighted by a preset weighting coefficient to obtain the association score; If the association score exceeds a preset association threshold, the entity pair is treated as an edge with potential association, and a latent subgraph is constructed.
[0009] In one optional implementation, generating the second knowledge graph through a preset knowledge graph fusion model includes: The vector representations of the entity nodes of the dominant subgraph and the hidden subgraph are obtained respectively, and used as the feature vectors of the dominant nodes and the feature vectors of the hidden nodes. The explicit node feature vector and the implicit node feature vector are input into the independent graph encoding channels of the attention-based dual-channel knowledge distillation network model, including the first graph encoding channel and the second graph encoding channel; The explicit node feature vector and the implicit node feature vector are respectively subjected to graph attention aggregation processing of corresponding sub-graph structures to output explicit node aggregation features and implicit node aggregation features; The explicit node aggregation features and the implicit node aggregation features are subjected to first-direction distillation and second-direction distillation, respectively, and then fused to generate a second knowledge graph.
[0010] In one optional implementation, the first directional distillation and the second directional distillation of the dominant node aggregation features and the recessive node aggregation features respectively include: The first directional distillation calculates the first correlation degree between the explicit node aggregation feature and the implicit node aggregation feature to obtain the first attention weight, and then performs a weighted sum with the implicit node aggregation feature to generate a context enhancement vector; The context enhancement vector is fused with the explicit node aggregation feature to obtain the first distillation node feature; The second directional distillation calculates the second correlation degree between the latent node aggregation feature and the explicit node aggregation feature to obtain the second attention weight, and then performs a weighted sum with the explicit node aggregation feature to generate a feature supplement vector; The feature supplement vector is fused with the latent node aggregate feature to obtain the second distillation node feature; The features of the first distillation node and the features of the second distillation node are fused to generate a fused feature vector of the node. The second knowledge graph is generated by combining the edge relationships in the explicit subgraph and the implicit subgraph.
[0011] In one optional implementation, encoding the second knowledge graph using a graph embedding algorithm includes: Based on the fusion feature vectors of nodes and the set of edge relationships in the second knowledge graph, the timeliness weights corresponding to each edge are calculated through a preset time-series evaluation function; The timeliness weight is negatively correlated with the collection time of the first intelligence data associated with the edge relationship; Using the topology of the second knowledge graph and the time-sensitivity weight as constraints, the nodes are modulated through a graph attention network to generate neighborhood vectors. The neighborhood vector is fused with the fusion feature vector of the corresponding node to output the semantic vector of each node.
[0012] In one optional implementation, constructing a semantic association graph based on the semantic vector includes: Calculate the information entropy value of the semantic vector of each node, mark the nodes with entropy values lower than the preset sparse threshold as the first node, and extract the nodes connected to the edges with timeliness weights higher than the preset dynamic threshold and mark them as the second node. Merge the first node and the second node to form the target node set. Based on the topological structure of the target node set and the second knowledge graph, a semantic association navigation graph is constructed; The associated navigation map is calculated using a preset path planning algorithm to generate one or more optimal exploration paths; Based on the optimal exploration path, a subset of collection instructions for publicly available information sources is generated.
[0013] In one optional implementation, generating a subset of collection instructions for publicly available information sources based on the optimal exploration path includes: Extract the node sequence on each of the optimal exploration paths, and obtain the edge relationships between the nodes from the second knowledge graph; Based on the edge relationship between the node sequence and the node, corresponding query elements are generated according to the preset mapping rules, and the query elements of all nodes on the same path are merged into a subset of the collection instructions corresponding to the path. Based on predefined interface protocol adaptation rules, the subset of collection instructions is instantiated into a collection task of a public information source interface to collect an intelligence data stream, which includes text, images and audio. The intelligence data stream is formatted, parsed, and its content extracted to obtain standardized second intelligence data.
[0014] In one optional implementation, the step of inputting the second intelligence data into a preset quality assessment model to generate a quality assessment result includes: The second knowledge graph is input into a preset time-series graph neural network prediction model to generate a third knowledge graph for a preset future time. The standardized second intelligence data is structured and transformed, and then compared with the third knowledge graph by calculating predefined multi-dimensional quantitative indicators to generate comparison results. The predefined multi-dimensional quantitative indicators include: prediction accuracy score and intelligence completion score. Based on the quantitative score of the comparison results, a quality assessment result of the second intelligence data is generated.
[0015] Secondly, this disclosure also provides a competitive intelligence multi-source heterogeneous acquisition and quantification system for dynamic intelligence targets, including: a target perception module, an encoding module, an execution module, and an evaluation module; The target perception module is used to collect first intelligence data based on dynamic intelligence targets, construct a first knowledge graph, the first knowledge graph includes explicit subgraphs and implicit subgraphs, and generate a second knowledge graph through a preset knowledge graph fusion model; The encoding module is used to encode the second knowledge graph using a graph embedding algorithm to obtain a semantic vector; The execution module is used to construct a semantic association graph based on the semantic vector, and convert the node series of the semantic association graph into query instructions through a preset mapping rule, execute the query instructions to collect data, and obtain the second intelligence data. The evaluation module is used to input the second intelligence data into a preset quality evaluation model and generate quality evaluation results.
[0016] Compared with existing technologies, this invention constructs explicit and implicit subgraphs and deeply integrates them using a dual-channel knowledge distillation network based on an attention mechanism, thereby deeply mining and incorporating deep-seated potential connections. It utilizes a graph embedding algorithm incorporating temporal weights to generate semantic vectors, and intelligently constructs a semantic association navigation graph based on the information entropy of the semantic vectors and the temporal weights of the edges. This automatically plans the optimal exploration path and generates precise query instructions, transforming passive retrieval into active exploration, improving the targeting and efficiency of data collection. Furthermore, it introduces a predictive quality assessment model, inputting the second knowledge graph into a temporal graph neural network to predict the knowledge graph at future moments. This prediction is then used as a benchmark for multi-dimensional quantitative comparison with newly collected second intelligence data, solving the problem of a lack of forward-looking standards for intelligence quality assessment. This achieves a forward-looking and quantifiable assessment of intelligence quality, improving the accuracy and comprehensiveness of multi-source heterogeneous data collection. Attached Figure Description
[0017] Figure 1 A flowchart of a method for quantifying multi-source heterogeneous acquisition of competitive intelligence for dynamic intelligence targets provided in this embodiment of the disclosure; Figure 2 A flowchart of bidirectional knowledge distillation for competitive intelligence of dynamic intelligence targets provided in this embodiment of the disclosure; Figure 3 A flowchart illustrating the generation of competitive intelligence data for dynamic intelligence targets provided in this embodiment of the disclosure; Figure 4 A schematic diagram of a multi-source heterogeneous acquisition and quantification system for competitive intelligence of dynamic intelligence targets provided in an embodiment of this disclosure. Detailed Implementation
[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0019] See Figure 1 The diagram shows a flowchart of a method for multi-source heterogeneous acquisition and quantification of competitive intelligence on dynamic intelligence targets provided in this embodiment of the present disclosure. The method includes steps S101 to S104, wherein: S101: Collect first intelligence data based on dynamic intelligence targets, construct a first knowledge graph, the first knowledge graph includes explicit subgraphs and implicit subgraphs, and generate a second knowledge graph through a preset knowledge graph fusion model; S102: Encode the second knowledge graph using a graph embedding algorithm to obtain a semantic vector; S103: Construct a semantic association graph based on the semantic vector, and convert the node series of the semantic association graph into query instructions through a preset mapping rule, execute the query instructions to collect data, and obtain the second intelligence data; S104: Input the second intelligence data into the preset quality assessment model to generate quality assessment results.
[0020] In practice, a set of keywords is generated based on dynamic intelligence targets, and collected in parallel from public information sources, including at least government department websites, news media platforms, and corporate websites, to obtain the first intelligence data.
[0021] For example, the dynamic intelligence target refers to the monitoring object whose status, attributes or related information evolves over time, including competitor entities such as specific companies, enterprise groups or strategic business units, technology fields, market or industrial chain links, key figures or teams, and hot events.
[0022] Core terms are obtained by directly parsing the dynamic intelligence target. For example, for the dynamic intelligence target "Company A's new energy vehicle business", the core terms include: "Company A", "A vehicle", and "A new energy". As an optional implementation, the dynamic intelligence target's domain entity knowledge base, business knowledge graph, or historical intelligence data can be used for correlation expansion. Examples of related terms include: products and models such as the names of models already launched by Company A, technical terms such as 800V high-voltage platform, key figures, partners and suppliers, competitors, and policies and standards.
[0023] In specific implementation, the core words and related words are semantically expanded, and the semantics include synonyms, near-synonyms, abbreviations and full names, to obtain semantically expanded words; Based on the core words, related words, and semantically extended words, a keyword set is generated and stored in key-value pair format.
[0024] As an optional implementation, multiple collection sub-processes are executed in parallel, wherein each collection sub-process initiates a query against one or more public information sources, and selects the corresponding collection method to execute the query according to the interface rules supported by the assigned target information source; For example, for publicly available information sources that provide API interfaces, data requests are made using API calling methods, and data is obtained by constructing request parameters that conform to the requirements of the interface documentation; For websites that do not have open APIs and publish information in the form of web pages, a web crawler method based on the HTTP / HTTPS protocol is used to crawl the content. This method simulates a browser request to download the web page content and uses HTML parsing technology to extract text and data from the page.
[0025] Collect unstructured text data, text data with specific structures, and semi-structured data as primary intelligence data; For example, free text obtained from news reports, industry analysis reports, official company press releases, and comments is referred to as unstructured text data; Text data containing fixed fields or patterns obtained from publicly available bidding documents, patent document abstracts and claims, academic paper titles, and other texts is text data with the specific structure described above. The information extracted from the product introduction page, management team introduction page, and financial statements of the company's official website constitutes the semi-structured data.
[0026] The first intelligence data collected is preprocessed to obtain preprocessed first intelligence data.
[0027] As an optional implementation, a deduplication method based on the SimHash algorithm is adopted to calculate the fingerprint value corresponding to each newly collected text data and store it in a predefined fingerprint database. The calculated new fingerprint value is compared with the value in the predefined fingerprint database; For example, if the new fingerprint value already exists in the fingerprint database, the corresponding text data is discarded; If the new fingerprint value does not exist in the fingerprint database, it is added to the fingerprint database, and the corresponding text data is retained.
[0028] As an optional implementation method, webpage cleaning is performed based on DOM tree parsing, and noise filtering is performed using regular expressions to remove garbled characters, irrelevant character sequences, and script code; By parsing the DOM structure of a webpage, non-content nodes such as advertisements and navigation can be identified and removed to obtain clean text data.
[0029] As an optional implementation, automatic encoding detection and conversion technology is used to automatically identify the original encoding of the text and convert it to UTF-8, and to identify date and time strings in various formats and parse them into a unified time object.
[0030] As an optional implementation method, a rule-driven automated metadata extraction and annotation method is adopted, which automatically records the source URL and collection timestamp during the collection process; Based on predefined URL pattern rules, such as labeling domains containing "news" as news and text keyword rules, such as labeling text containing "patent number" as patent information, the system automatically assigns type labels to each piece of data and outputs intelligence data records with standardized metadata.
[0031] Named entity recognition is performed on the preprocessed first intelligence data to extract entities including at least organizations, people and products, and the relationships between the entities are extracted.
[0032] Using the entities as nodes and the extracted relationships between entities as edges, an explicit subgraph is constructed.
[0033] As an optional implementation, the BERT model is used as the basic architecture, which includes a word embedding layer, a position encoding layer, and a 12-layer Transformer encoder.
[0034] In practice, the BERT base model is pre-trained using a general public text corpus. The batch size can be set according to the hardware configuration, and can be set to 32. The public text corpus is batch-processed, and the parameters of each layer of the model are updated through backpropagation to complete the learning of basic semantic features. For example, data from government announcements, industry news, and corporate annual reports are collected, and the data is cleaned and labeled for use in adaptive fine-tuning of the model.
[0035] In practice, a stochastic gradient descent optimizer is used with a fixed learning rate of 2e-5 and 4 iterations. After fine-tuning, a pre-trained named entity recognition model is obtained.
[0036] As an optional implementation, structured pruning is used to remove redundant attention heads in the Transformer encoder and to quantize the model parameters. The lightweight named entity recognition model is loaded onto the computing nodes corresponding to each thread, and each thread calls the model to perform entity recognition to obtain the first entity set.
[0037] As an optional implementation, entity relation extraction is performed by fusing dependency parsing with a BiLSTM network, wherein the BiLSTM network employs the dropout mechanism. Specifically, when constructing the BiLSTM model, the dataset with labeled entity relationships is divided into training set, validation set and test set in a ratio of 8:1:1; The model uses a bidirectional LSTM structure with a hidden layer dimension of 256, a dropout rate of 0.5, and the Adam optimizer. The initial learning rate is set to 0.001, and the training epochs are 50. The early stopping strategy stops training when the validation set loss does not decrease for 5 consecutive epochs.
[0038] In a specific implementation, the first entity set and the corresponding preprocessed first intelligence data are synchronously distributed to each thread to perform dependency parsing and extract semantic features through the BiLSTM network. Based on predefined entity relationship extraction rules, entity relationships are extracted, including: membership, competition, cooperation, technology application, and event impact.
[0039] The semantic features are input into the SVM classifier, which outputs the corresponding relationship types in the entity pairs to obtain the entity relationship set.
[0040] As an optional implementation, the entity set is matched with the domain entity knowledge base using an entity linking algorithm to determine the unique identifier of each entity, and the relationship types in the entity relationship set are named and classified to obtain an explicit subgraph.
[0041] Specifically, the name string of each entity to be linked is extracted, along with two words before and after it in the surrounding text as context, and the entity's type tag, such as organization, person, or product. A preliminary search is performed in the domain entity knowledge base based on the entity name and type, using fuzzy string matching, allowing a maximum edit distance of 2 for the name.
[0042] If the number of candidate entities retrieved exceeds 20, filtering is performed based on entity type, retaining only those with the same type. To control subsequent computational complexity while maintaining recall, 20 candidate entities is a common engineering practice value.
[0043] For each candidate entity, two types of similarity are calculated. First, string similarity, which uses the edit distance algorithm to measure the difference between the entity name and the candidate name. The edit distance is divided by the maximum of the lengths of the two names to obtain a normalized difference value, and then 1 is subtracted from this value to obtain the similarity score.
[0044] Secondly, semantic similarity involves concatenating the context words of the entity to be linked into a short text, converting the short text into a 384-dimensional vector using a pre-trained BERT model, and similarly converting the description text of the candidate entity into a 384-dimensional vector. The cosine similarity between the two vectors is then calculated as the semantic similarity score.
[0045] As an optional implementation, string similarity and semantic similarity can be weighted and summed with weights of 0.3 and 0.7 respectively to obtain a comprehensive similarity score. The weighting design can capture contextual information and perform entity disambiguation. If the comprehensive similarity score is greater than a preset similarity score threshold, the candidate entity with the highest score is selected as the link result, and a unique identifier is assigned to the current entity. If the highest comprehensive similarity score is lower than the preset similarity score threshold, the entity is marked as a new entity not included in the knowledge base, and a temporary unique identifier is generated, such as "NEW_" plus the first 8 digits of the entity name hash value.
[0046] The similarity score threshold can be set to 0.75 based on the experience of common similarity matching tasks.
[0047] After linking all entities, the extracted entity relationship types are uniformly named and categorized. For example, "competition," "rivalry," and "competitive products" are uniformly mapped to the standard relationship name "competitive relationship," and "cooperation," "collaboration," and "strategic partner" are mapped to "cooperative relationship." Based on the explicit subgraph, a corresponding adjacency matrix is generated. The rows and columns of the adjacency matrix correspond to all entity nodes, and the element values are 1 or 0, used to indicate whether there is a direct edge relationship between the corresponding two nodes. For example, if a direct edge relationship exists, the corresponding position is 1; otherwise, it is 0.
[0048] In practice, the explicit subgraph is represented and learned through a graph convolutional network to obtain entity node vectors.
[0049] Extract the first entity attribute from the explicit sub-graph. The first entity attribute includes: entity unique identifier, entity name, entity type such as organization, person, product, technical term, and domain. As an optional implementation, an entity matching algorithm is used to associate entities in the explicit sub-graph with entities in the preprocessed first intelligence data and the domain entity knowledge base, and to extract second entity attributes such as product category and technical parameter association information for product entities, industry affiliation and business scope association information for organization entities, and position and department association information for person entities.
[0050] Specifically, for each entity in the explicit sub-graph, three features are extracted. First, the entity name, represented by a segmented word set, is used to calculate Jaccard similarity. Second, the entity type label, such as organization, person, or product, is used to determine category consistency. Third, the contextual feature, which is a vector composed of surrounding words when the entity appears in the preprocessed first intelligence data, converted into a 384-dimensional vector using the BERT model.
[0051] For the entities to be matched, the same three features are extracted. The following are calculated: Name Jaccard similarity (the intersection of the two entity name sets divided by the union); Type consistency (a score of 1 if the types are completely identical, otherwise 0); Context cosine similarity (the cosine similarity score between the two context vectors).
[0052] The three similarity scores are weighted and summed with weights of 0.5, 0.2, and 0.3 to obtain the comprehensive matching score. The weights are set as follows: the entity name is the most distinctive feature in the matching, so it is given the highest weight of 0.5; the contextual features can reflect the semantic environment of the entity, so they are given the second highest weight of 0.3; the entity type is only used as a coarse-grained filtering condition, so it is given a lower weight of 0.2.
[0053] If the overall matching score exceeds the preset matching score threshold, the entities are determined to be the same entity. The preset matching score threshold can be based on the experience of common entity matching tasks and can be fine-tuned in practical applications according to the precision and recall curves on the validation set. It is usually set between 0.65 and 0.75.
[0054] For entities that are determined to be the same entity, perform an attribute merging operation.
[0055] Specifically, the merging rules adopt a priority-based comparison method: the first priority is to select attribute values from the domain entity knowledge base; if the knowledge base does not provide the attribute, the second priority is to select the attribute value that appears most frequently in the first intelligence data; if the frequencies are the same, the third priority is to select the attribute value with the highest authority score of the source website. The source authority score is preset according to the website domain level, for example, the score of government websites is 3, the score of well-known media is 2, and the score of ordinary enterprise websites is 1.
[0056] For example, suppose an organizational entity, "ABC Technology Company," needs to have its "Industry Affiliation" attribute merged. This attribute is not included in the domain entity knowledge base. The first intelligence data has three sources: Source 1 is a government announcement, labeled as "Software and Information Technology Services"; Source 2 is news media, also labeled as "Software and Information Technology Services"; Source 3 is the company's official website, labeled as "Science and Technology Promotion and Application Services." "Software and Information Technology Services" appears twice, while "Science and Technology Promotion and Application Services" appears once. Therefore, the most frequent "Software and Information Technology Services" is selected as the merged attribute value. If the frequencies are the same, the source authority scores are compared, with the higher score taking priority. The merged attribute set is used as the second entity attribute and associated with the corresponding entity node in the explicit sub-graph. Consistency checks are performed on the first entity attribute and the second entity attribute, eliminating duplicate and invalid attributes, standardizing attribute names and description formats, and forming a standardized attribute set for each entity node. Based on the standardized attributes, a corresponding transformation method is adopted, and node feature vectors are obtained through vector fusion. The standardized attribute set includes: categorical attributes, textual attributes, and relational attributes. For example, a category coding method is used to convert the category attributes such as entity type, domain, and product category into discrete coding vectors; The textual attributes, such as entity names and business scope descriptions, are transformed using a pre-trained word embedding method to obtain dense vectors; The associated attributes, such as technical parameter association information and business scope association information, are transformed into feature items using a feature decomposition method and then converted into vectors through one-hot encoding.
[0057] For example, the associated information "business scope associated information" can be decomposed into binary feature items "whether it involves artificial intelligence" and "whether it involves cloud computing". The existence of the binary feature item is determined by converting it into a two-dimensional vector representation through one-hot encoding, such as [1, 0] indicating yes and [0, 1] indicating no.
[0058] In practical implementation, splicing and fusion can be used to splice the vectors corresponding to various attributes of the same entity node in a preset order to obtain the entity node feature vector. For example, the splicing can be done in the order of categorical attribute vector, text attribute vector, and relational attribute vector.
[0059] In specific implementation, the feature vectors of the entity nodes and the adjacency matrix of the explicit subgraph are synchronously input into the graph convolutional network to aggregate the features of the neighboring nodes of each node. The node features are combined with the features of the neighboring nodes through feature fusion, and a nonlinear transformation is performed to generate updated node features.
[0060] The system propagates through graph convolutional layers, eventually outputting a vector of entity nodes.
[0061] Calculate the cosine similarity between any two entity node vectors, and treat the two entities as an entity pair, then count the co-occurrence frequency of the entity pair in the first intelligence data after preprocessing.
[0062] In specific implementation, co-occurrence frequency statistics are performed through predefined windows, including document-level and fixed-distance sliding windows; For example, each independent intelligence data record, such as a news headline, a product description, or a patent abstract text, is used as the document-level window. If two entities appear in the same intelligence data record, it is counted as one co-occurrence.
[0063] A sliding window based on the number of words can be used. A fixed window size can be set, such as a window containing 10 consecutive words. If two entities appear in the same sliding window at the same time, it is counted as one co-occurrence.
[0064] Based on the predefined window, the preprocessed first intelligence data is traversed, and a counter is maintained for each pair of entities; Whenever a pair of entities is simultaneously discovered in any window, its corresponding counter is incremented by one. After the traversal is complete, the value of each counter is the co-occurrence frequency of the entity pair.
[0065] As an alternative implementation, the co-occurrence frequency values are linearly scaled to the range of 0 to 1 using the max-min normalization method.
[0066] The cosine similarity and the co-occurrence frequency are weighted by a preset weighting coefficient to obtain the association score; The preset weighting coefficients are determined based on the characteristics of the preprocessed first intelligence data, resulting in cosine similarity weight and co-occurrence frequency weight. For example, the cosine similarity weight can be set to 0.6, and the co-occurrence frequency weight can be set to 0.4.
[0067] If the association score exceeds a preset association threshold, the entity pair is treated as an edge with potential association, and a latent subgraph is constructed.
[0068] As an optional implementation, the association threshold can be dynamically determined using a distribution-based quantile method; In practice, the association scores of all entity pairs are sorted, and the score corresponding to the preset percentage position is selected as the preset association threshold for this construction. The preset percentage is fine-tuned based on the distribution characteristics of the preprocessed first intelligence data. For example, the preset percentage can be increased when the data is sparse.
[0069] The preset percentage is fine-tuned based on the distribution characteristics of the preprocessed first intelligence data. For example, the preset percentage ranges from 60% to 85%.
[0070] Specifically, the standard deviation σ of the association scores of all entities is calculated. If σ < 0.1, it indicates that the score distribution is concentrated, so 80% is taken; if 0.1 ≤ σ ≤ 0.3, it indicates that the distribution is moderate, so 70% is taken; if σ > 0.3, it indicates that the distribution is dispersed, so 60% is taken.
[0071] See Figure 2 The diagram shows a flowchart of bidirectional knowledge distillation for competitive intelligence of dynamic intelligence targets provided in this embodiment of the disclosure, wherein: In practical implementation, a bidirectional cross-attention distillation model is constructed, including: an input layer, an encoding layer, a distillation layer, a fusion layer, and an output layer; The input layer is equipped with independent input port one and input port two, which respectively receive node feature vectors from the dominant subgraph and the hidden subgraph, and transmit them to the coding layer.
[0072] In specific implementation, the vector representations of the entity nodes of the explicit subgraph and the implicit subgraph are obtained respectively, and used as the feature vectors of explicit nodes and the feature vectors of implicit nodes; The input port one receives the feature vector of the dominant node and generates it through the dominant subgraph graph convolutional network encoder; As an optional implementation, the TransGCN model based on graph convolutional networks is used for vector representation to construct the adjacency matrix of the explicit subgraph, with entities as rows and columns, and matrix elements indicating whether there is an explicit edge relationship between entities; Each entity node is assigned an initial feature vector, which is obtained by word embedding encoding of the entity's attribute information, such as the industry category of the organization, the job information of the person, and the functional parameters of the product. As an optional implementation, the Word2Vec word embedding algorithm is used to encode and transform attribute information. The adjacency relation matrix and the initial feature vector are input into the TransGCN model, and the local neighborhood features of the nodes are aggregated through graph convolution kernels. After iterative calculations through three layers of graph convolutional layers, the low-dimensional dense vector of each entity node is output, which is the feature vector of the explicit node.
[0073] Similarly, the second input port receives the latent node feature vector, which is generated by the preceding latent subgraph shallow semantic encoder, which is calculated based on entity co-occurrence frequency and cosine similarity.
[0074] The explicit node feature vector and the implicit node feature vector are input into the independent graph encoding channels of the attention-based dual-channel knowledge distillation network model, including the first graph encoding channel and the second graph encoding channel; The first image encoding channel and the second image encoding channel are processed in parallel.
[0075] The explicit node feature vector and the implicit node feature vector are respectively subjected to graph attention aggregation processing of corresponding sub-graph structures to output explicit node aggregation features and implicit node aggregation features; The explicit node aggregation features and the implicit node aggregation features are subjected to first-direction distillation and second-direction distillation, respectively, and then fused to generate a second knowledge graph.
[0076] In a specific implementation, the first directional distillation calculates the first correlation degree between the explicit node aggregation feature and the implicit node aggregation feature to obtain the first attention weight, and then performs a weighted sum with the implicit node aggregation feature to generate a context enhancement vector. For example, the explicit node feature vector is converted into a query vector through linear projection, and the implicit node features are simultaneously converted into key vectors and value vectors; As an optional implementation, dot product operation is used to calculate the similarity between each query vector and all key vectors, which is used as the first relevance. The first correlation is then converted into a normalized probability distribution using the Softmax function to obtain a set of first attention weights.
[0077] Based on the first attention weight, all the value vectors are weighted and summed to generate a context enhancement vector; Perform a residual join operation, add the context enhancement vector to the explicit node feature vector element by element, and output the enhanced first distillation node feature through normalization.
[0078] Similarly, the second directional distillation projects the latent node aggregate features into a query vector and the explicit node aggregate features into a key vector and a value vector through linear projection. The second attention weight is obtained by calculating the second correlation degree between the latent node aggregation feature and the explicit node aggregation feature, and then weighted and summed with the explicit node aggregation feature to generate a feature supplement vector. The feature supplement vector is fused with the latent node aggregate feature to obtain the second distillation node feature.
[0079] As an optional implementation, a scaled dot product attention mechanism is used to calculate the matrix dot product between the query vector and the key vector to obtain the second similarity score matrix; Divide the second similarity score matrix by a preset scaling factor, and then input the scaled second similarity score matrix into the Softmax function for normalization to obtain the second attention weight.
[0080] For example, the scaling factor is set to the square root of the dimension of the query vector or key vector. Taking the graph attention network configuration as an example, if the dimension of the node feature vector is 256, and the dimension of the query vector and key vector is 64 after linear transformation, then the scaling factor is 8; if the dimension is 28, then the scaling factor is 11.3. In actual calculation, an approximate integer value of 11 or 12 can be used.
[0081] The value vector is weighted and summed using the second attention weight to obtain the explicit feature supplement vector; The second distillation node features are output by performing residual connections with the aggregated features of the latent nodes and performing layer normalization.
[0082] The features of the first distillation node and the features of the second distillation node are fused to generate a fused feature vector of the node. The second knowledge graph is generated by combining the edge relationships in the explicit subgraph and the implicit subgraph.
[0083] As an optional implementation, for each entity node, its corresponding first distillation node features and second distillation node features are concatenated and input into a fully connected layer for fusion and dimensionality reduction to generate a fused feature vector of the node.
[0084] All verified entity relationship edges, such as "belonging" and "competition," in the explicit subgraph are merged with the "potential association" edges selected by threshold filtering in the implicit subgraph. Duplicate edges are removed to form a unified edge set.
[0085] Using all entities carrying fused feature vectors as nodes and the merged unified edge set as the connection relationship, the final second knowledge graph is constructed.
[0086] As an optional implementation, a normalized text fingerprint generation method is used to generate a unique string identifier for the entity as the entity's ID.
[0087] As an optional implementation, a feature dimension splicing method is adopted, in which the feature vectors of the first distillation node and the feature vectors of the second distillation node belonging to the same entity are directly connected end to end along their feature dimensions. For example, if the first distillation node feature is a 256-dimensional vector and the second distillation node feature is a 256-dimensional vector, then concatenating them will generate a 512-dimensional combined feature vector.
[0088] In a specific implementation, the combined feature vector is input into a fully connected feedforward neural network for transformation. The fully connected feedforward neural network includes a fusion layer and a dimensionality reduction layer. For example, the combined feature vector is linearly transformed through the fusion layer and then fused using the ReLU activation function; The feature dimension of the combined feature vector is compressed by linear projection through the dimensionality reduction layer, and finally the fused feature vector of the node is output.
[0089] As an optional implementation, all edges labeled with relation types such as "membership", "competition", "cooperation" and "supply" are extracted from the explicit subgraph. Each edge is stored in the form of a triple (head entity ID, relation type, tail entity ID) to obtain the first edge set. Extract all potential associated edges that pass the association score threshold from the hidden subgraph, mark them as potential association types, and store them in triplet form to obtain the second edge set.
[0090] As an optional implementation, a deduplication method based on entity ID and relation type is adopted to perform a union operation on the first edge set and the second edge set; for example, by comparing whether any two edges have the same head entity ID, tail entity ID and relation type, if they are exactly the same, only one of them is retained, thereby eliminating duplicate edges and forming a third edge set.
[0091] In a specific implementation, the entity nodes with the fused feature vectors are taken as a node set, and the third edge set is taken as the connection relationship between the nodes to construct a directed attribute graph, which serves as the second knowledge graph.
[0092] Based on the fusion feature vectors of nodes and the set of edge relationships in the second knowledge graph, the timeliness weights corresponding to each edge are calculated through a preset time-series evaluation function; The timeliness weight is negatively correlated with the collection time of the first intelligence data associated with the edge relationship.
[0093] As an optional implementation, a mapping relationship between explicit and potential associated edges and the collection time is established respectively; For example, for explicit relationship edges, the publication time of news reports or official documents mentioning the relationship is extracted as the first timestamp; For potential associated edges, the collection time of the co-occurring text of the entity on which the association is based is used as the second timestamp.
[0094] As an optional implementation, the exponential time decay algorithm can be used as the time series evaluation function to calculate the time weight. Based on the current time, the time difference between each first timestamp, second timestamp and the current time is calculated. The time difference is then substituted into the exponential decay function to obtain the time weight value of each edge. The timeliness weight value is normalized to between 0 and 1.
[0095] Specifically, determine the decay half-life parameter. , The value is determined based on the update frequency of the dynamic intelligence target; for example, for news media data, We use 1 day as the update frequency, meaning it's updated daily. For industry reports or company announcements with a weekly update frequency, Use data from the past 7 days; for patent or policy document data updated monthly or more frequently, Use 30 days. In general competitive intelligence gathering scenarios, if there is no explicit data classification, the default is... Take 7 days.
[0096] For each edge, obtain the collection time of its associated intelligence data, and calculate the time difference between that collection time and the current time. , Take the integer value in days. For example, if the news report associated with a certain edge was collected on May 1, 2026, and the current time is May 11, 2026, then Δt equals 10 days.
[0097] time difference and Substituting into the exponential decay function, the original value of the time-related weight is obtained. ,Right now The value is approximately 0.37, and if it falls between 0 and 1, it can be directly used as the time-sensitivity weight value. Using the topology of the second knowledge graph and the time-sensitivity weight as constraints, a graph attention network is used to modulate the nodes and generate neighborhood vectors. As an optional implementation, a multi-head temporal graph attention network is used to encode nodes, wherein the attention system between a node and its neighboring nodes is determined by the semantic similarity between the node's feature vectors and the temporal weight of the connecting edges. The attention coefficients are obtained by calculating feature-based attention scores and weighting the attention scores using the time-sensitivity weights of the corresponding edges; then normalizing the modulated attention scores using the softmax function. The fusion feature vectors of neighboring nodes are weighted and summed based on the attention coefficients to generate the neighborhood vector of the node.
[0098] In specific implementation, the method of feature concatenation and linear transformation is used to concatenate the fused feature vector of each node with the neighborhood vector; The concatenated long vector is input into a single-layer fully connected neural network. Through linear transformation and nonlinear activation functions such as ReLU, the concatenated vector is fused and projected onto the target dimension, and the semantic vector of the node is output. The target dimension is set according to the specific engineering situation.
[0099] Calculate the information entropy value of the semantic vector of each node, mark the nodes with entropy values lower than the preset sparse threshold as the first node, and extract the nodes connected to the edges with timeliness weights higher than the preset dynamic threshold and mark them as the second node. Merge the first node and the second node to form the target node set. As an optional implementation method, the Shannon entropy calculation method is used to calculate the information entropy value of the semantic vector of each node by statistically analyzing the distribution probability of the features of each dimension of the semantic vector. Nodes whose information entropy value is lower than a preset sparse threshold are marked as first nodes, where a low entropy value indicates that the node's feature information is concentrated and it is a core intelligence node. Simultaneously, nodes connected to edges whose timeliness weight is higher than the preset dynamic threshold are extracted and marked as second nodes; The first node and the second node are integrated by a set merging operation, and duplicate nodes are removed to form the target node set.
[0100] The preset sparse threshold and the preset dynamic threshold are obtained by acquiring a set of historical information entropy values and calculating their percentiles. Based on actual engineering needs, the entropy value corresponding to a selected percentile is chosen.
[0101] Specifically, the preset sparsity threshold is set to the 30th percentile of the information entropy value set, and the preset dynamic threshold is set to the 70th percentile of the time-weight set. Depending on actual engineering needs, the percentiles can be adjusted accordingly; for example, the sparsity threshold can be selected within the range of 20% to 40%, and the dynamic threshold can be selected within the range of 60% to 80%. See also... Figure 3 The diagram shown is a flowchart illustrating the generation of competitive intelligence data for dynamic intelligence targets according to an embodiment of this disclosure, wherein: Based on the topological structure of the target node set and the second knowledge graph, a semantic association navigation graph is constructed; In specific implementation, the nodes in the target node set are used as nodes of the navigation graph. All edge relationships between target nodes in the second knowledge graph are retained, and edge relationships between target nodes and non-target nodes that have a timeliness weight higher than the dynamic threshold are supplemented to form an edge set of the semantic association navigation graph. A graph database is used to store the semantic association navigation graph, ensuring fast querying and traversal of node and edge relationships.
[0102] The semantic association navigation graph is calculated using a preset path planning algorithm to generate one or more optimal exploration paths.
[0103] As an optional implementation, the semantic association navigation graph is calculated using an improved Dijkstra algorithm. By simultaneously considering the cost of nodes and the cost of edges, the information entropy value of the semantic vector of a node and the timeliness weight of the edge are weighted and summed as the path weight. The lower the path weight, the higher the intelligence value corresponding to the path. The semantic association navigation graph is traversed with each target node as the starting node, the path weight from the starting node to all other target nodes is calculated, and the path with the lowest path weight is selected as the optimal exploration path. If there are multiple paths with the same lowest weight, all of them should be retained as the optimal exploration path.
[0104] Specifically, a total cost value is initialized for each node, with the starting node's total cost value set to 0. For the remaining nodes, the initial total cost value should be greater than the actual total cost value of any possible path. Since the number of edges in each path does not exceed the total number of nodes minus 1, the information entropy value of each node ranges from 0 to 1, and the time-sensitivity weight of each edge also ranges from 0 to 1, the upper limit of the total cost value of any path is the total number of nodes multiplied by 1.5.
[0105] For safety reasons, the initial total generation value of non-starting nodes is set to twice the total number of nodes. For example, if the total number of nodes is 100, the initial total generation value is set to 200.
[0106] Simultaneously, a priority queue is maintained to extract the node with the smallest current total generation value for expansion each time. For the extracted current node, all its neighboring nodes are examined in turn, and the total generation value from the starting node through the current node to that neighboring node is calculated. This total generation value is equal to the total generation value accumulated by the current node plus the time weight of the edge connecting the current node and the neighboring node, plus the product of the information entropy value of the neighboring node and the balance coefficient of 0.5.
[0107] If the newly calculated total generation value is less than the current total generation value recorded by the neighboring node, then update the total generation value of the neighboring node and record the current node as the predecessor node of that neighboring node. Repeat the above process until the priority queue is empty.
[0108] The optimal exploration path generation process is as follows: For each node marked as a target node, starting from that target node, backtrack sequentially according to the recorded predecessor node relationships until returning to the starting node, thus obtaining a complete node sequence. This sequence is the optimal exploration path from the starting node to the target node. The above process is performed on all starting nodes to obtain multiple candidate paths. From these, one or more paths with the lowest total agent value are selected: if multiple paths have the same and lowest total agent value, all are retained as the final optimal exploration path; otherwise, only the path with the unique lowest total agent value is retained as the optimal exploration path. Based on the optimal exploration path, a subset of collection instructions for public information sources is generated.
[0109] As an optional implementation, the node sequence on each of the optimal exploration paths is extracted, and the edge relationships between the nodes are obtained from the second knowledge graph; Based on the edge relationship between the node sequence and the node, corresponding query elements are generated according to the preset mapping rules, and the query elements of all nodes on the same path are merged into a subset of the collection instructions corresponding to the path. As an optional implementation, the node sequence on each of the optimal exploration paths is extracted, and the edge relationships between the nodes are obtained by querying a graph database; Based on the node sequence and the edge relationships between nodes, the corresponding query elements are generated using the entity-relationship-entity triplet mapping method. Specifically, adjacent nodes on the path and their edge relationships are formed into triplets, and each triplet is converted into a query element. The query element includes the core entity, the associated entity, and the relationship type. Combine all query elements on the same path in the order of the path, add information source filtering conditions, such as determining the priority public information sources to be collected based on the first intelligence data source associated with the path node, and merge them into a subset of collection instructions corresponding to that path.
[0110] As an optional implementation, thread pool technology is used to achieve multi-threaded parallel processing. Thread groups are divided according to the type of the public information source. Each thread group corresponds to a type of public information source, such as a government department website thread group, a news media platform thread group, and a corporate website thread group.
[0111] Based on predefined interface protocol adaptation rules, the subset of collection instructions is instantiated into a collection task of a public information source interface to collect an intelligence data stream, which includes text, images and audio. In practice, based on predefined interface protocol adaptation rules, each subset of collection instructions is instantiated into a collection task corresponding to the public information source interface; For example, government websites use the HTTP / HTTPS protocol for GET requests, news media platforms use the RESTful API protocol, and enterprise websites use web crawler protocols such as those that follow the robots.txt protocol. The acquisition task is assigned to the corresponding thread group, and the threads in the thread pool execute the acquisition task in parallel to acquire the intelligence data stream, which includes text, images and audio.
[0112] As an optional implementation, a multi-threaded parallel parsing strategy is used for parsing; For example, the text data is formatted and parsed using the Beautiful Soup parsing tool to extract information such as text content, publication time, source, and author; Image data was converted and metadata extracted using OpenCV. Images were uniformly converted to JPEG format, and metadata such as shooting time and image size were extracted. The audio data was parsed using the Librosa tool, converted to WAV format, and information such as audio duration and sampling rate was extracted.
[0113] In practice, the parsed data is standardized, and the data format and field naming are unified to obtain standardized second intelligence data.
[0114] The second knowledge graph is input into a preset time-series graph neural network prediction model to generate a third knowledge graph for a preset future time. As an optional implementation, an evolutionary graph convolutional network is used as a temporal graph neural network prediction model. The model consists of a graph convolutional network module and a recurrent neural network module. The graph convolutional network captures the structural information of the graph at each time step, and the recurrent neural network models the dynamic evolution of the graph convolutional network parameters over time.
[0115] In practice, a second knowledge graph snapshot sequence arranged in chronological order is collected as training data. Each snapshot contains the graph structure (nodes and edges) and node features at that moment. The graph sequence of consecutive historical moments is used as input to train the model to predict the graph state at the next future moment. The loss function uses contrastive loss, which encourages the predicted node embeddings generated by the model to be as close as possible to the node embeddings of the real future graph in the vector space.
[0116] The current second knowledge graph and its previous time step graph snapshots are input into the trained evolution graph convolutional network model. Through the forward propagation process, the model outputs a prediction of the graph at a future preset time, including the predicted node feature vectors and the probability that there are connections between nodes. Edges whose predicted relationship probability is higher than a preset probability threshold are retained to form the third knowledge graph obtained from the prediction.
[0117] The preset probability threshold is determined according to the accuracy requirements of the actual project. For example, it can be set to 0.7 to 0.9, with a default value of 0.8. If a high recall rate is required, it can be reduced to 0.6; if a high precision rate is required, it can be increased to 0.95.
[0118] As an optional implementation, the normalized second intelligence data is input into a BERT-based joint entity and relation extraction model to extract structured fact triples in the form of (head entity, relation, tail entity) from the text sentences, forming a second intelligence structured fact set.
[0119] The second set of structured facts is compared with the third knowledge graph. By traversing each triple in the second set of structured facts, the head and tail entity nodes with the same entity ID are located in the node set of the third knowledge graph by precise matching of entity IDs. Based on successfully locating the head and tail nodes, query the third knowledge graph to check if there is a predicted edge record that starts with the head entity ID and ends with the tail entity ID.
[0120] If the aforementioned predicted edge exists, the relation in the fact triple is compared with the relation type in the predicted edge record using a preset relation synonym mapping table. This table defines equivalence or similarity relations between different relation expressions. For example, "competition" and "market competition" are considered to be the same, and "cooperation" and "strategic cooperation" are considered to be the same. This is determined by querying this mapping table. For example, if the relation strings are completely identical or are defined as equivalent in the relation synonym mapping table, then the relation types are determined to be the same.
[0121] Specifically, we collected common relationship types and their synonyms within the field, such as "competition" being synonymous with "market competition," "competitive relationship," and "rivalry." Establish equivalence relation groups and map each group to a standard relation name.
[0122] Examples of mapping relationships are as follows: {"competition", "market competition", "competitive relationship"} -> "competition"; {"cooperation", "strategic cooperation", "collaboration"} -> "cooperation"; {"acquisition", "merger and acquisition", "holding"} -> "acquisition".
[0123] The mapping table is stored in key-value pair format. As an optional implementation, if the relation strings are inconsistent and not in the relation synonym mapping table, a BERT-based model is invoked to calculate the semantic similarity score of the two relation phrases. If the score exceeds a preset similarity threshold, the relation is determined to be semantically similar. The similarity threshold is determined according to the accuracy requirements of the actual project.
[0124] Specifically, the semantic similarity score output by the BERT model ranges from 0 to 1, with a higher score indicating greater similarity between the two relation phrases. To achieve high precision and avoid misclassification as similar, the similarity threshold can be set to 0.9; to achieve high recall and identify semantically similar relations as much as possible, the similarity threshold can be set to 0.7; by default, the similarity threshold can be set to 0.8.
[0125] As an optional implementation, accuracy is used as the scoring metric. The number of fact triples that are successfully predicted by the third knowledge graph is divided by the total number of facts in the second intelligence structured fact set, and the ratio is between 0 and 1.
[0126] In specific implementation, each predicted edge in the third knowledge graph is accompanied by an existence probability value generated by a temporal graph neural network model. Based on a preset probability threshold, all predicted edges with existence probabilities higher than the probability threshold are selected to form a set of high-confidence predicted edges. The probability threshold is determined according to the accuracy requirements of the actual project.
[0127] For each edge in the set of high-confidence predicted edges, a query is performed in the adjacency table of the current second knowledge graph; If no edge record with the same head entity ID, tail entity ID, and relation type can be found, the predicted edge is marked as an emerging relation edge, and all emerging relation edges constitute the emerging relation edge set.
[0128] Traverse each edge in the set of emerging relation edges, and in the second set of structured facts, search for whether there exists a fact triple whose head entity ID and tail entity ID are the same as the emerging edge, and whose relation is statistically consistent or semantically similar to the emerging edge. If the fact triple exists, the emerging relation edge is counted as a verified emerging edge.
[0129] In practice, the total number of edges in the set of emerging relation edges is counted, along with the number of emerging edges that have been verified. As an optional implementation, the recall rate calculation formula is used, which divides the number of verified emerging edges by the total number of edges in the set of emerging relational edges. The resulting ratio is the intelligence completion score, and the score value is between 0 and 1.
[0130] As an optional implementation method, a weighted linear combination method is used to sum the prediction accuracy score and the intelligence completion score according to preset weights to generate a comprehensive quality score between 0 and 1. The weighting can be adjusted according to business needs. For example, the weight of the prediction accuracy score can be set to 0.6, and the weight of the intelligence completion score can be set to 0.4; if the business focuses more on the discovery of emerging intelligence, the weight of the intelligence completion score can be increased to 0.7, and the weight of the prediction accuracy score can be decreased to 0.3.
[0131] Based on the same inventive concept, this disclosure also provides a multi-source heterogeneous acquisition and quantification system for competitive intelligence of dynamic intelligence targets, which is used in the multi-source heterogeneous acquisition and quantification method for competitive intelligence of dynamic intelligence targets. Since the principle of the control method in this disclosure is similar to the multi-source heterogeneous acquisition and quantification method for competitive intelligence of dynamic intelligence targets described above, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.
[0132] See Figure 4 The diagram shown is a schematic of a multi-source heterogeneous acquisition and quantification system for dynamic intelligence targets provided in this embodiment of the present disclosure. The system includes: a target perception module 10, an encoding module 20, an execution module 30, and an evaluation module 40, wherein: Target perception module 10: used to collect first intelligence data based on dynamic intelligence targets, construct a first knowledge graph, the first knowledge graph including explicit subgraphs and implicit subgraphs, and generate a second knowledge graph through a preset knowledge graph fusion model; Encoding module 20: Used to encode the second knowledge graph using a graph embedding algorithm to obtain a semantic vector; Execution module 30: is used to construct a semantic association graph based on the semantic vector, and convert the node series of the semantic association graph into query instructions through preset mapping rules, execute the query instructions to collect data, and obtain the second intelligence data; Evaluation module 40: Used to input the second intelligence data into a preset quality evaluation model and generate a quality evaluation result. Those skilled in the art will understand that, in the methods described above in the specific embodiments, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic. It should be understood that determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information.
[0133] In the description of this specification, the terms "exemplary," "for example," "specifically," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
Claims
1. A method for quantifying competitive intelligence from multiple sources and heterogeneous sources for dynamic intelligence targets, characterized in that: include: Based on the dynamic intelligence target, first intelligence data is collected, a first knowledge graph is constructed, the first knowledge graph includes explicit subgraphs and implicit subgraphs, and a second knowledge graph is generated through a preset knowledge graph fusion model; The second knowledge graph is encoded using a graph embedding algorithm to obtain semantic vectors; A semantic association graph is constructed based on the semantic vector, and the node series of the semantic association graph is converted into query instructions through a preset mapping rule. The query instructions are executed to collect data and obtain the second intelligence data. The second intelligence data is input into a preset quality assessment model to generate a quality assessment result.
2. The method for multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets according to claim 1, characterized in that, The process of collecting first intelligence data based on dynamic intelligence targets and constructing a first knowledge graph includes: Based on the dynamic intelligence target, a keyword set is generated and collected in parallel from public information sources, including at least government department websites, news media platforms and corporate websites, to obtain the first intelligence data; Named entity recognition is performed on the first intelligence data to extract entities including at least organizations, people and products, and the relationships between the entities are extracted. Using the entities as nodes and the extracted relationships between entities as edges, an explicit subgraph is constructed.
3. The method for multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets according to claim 1, characterized in that, The process of collecting first intelligence data based on dynamic intelligence targets and constructing a first knowledge graph also includes: The dominant subgraph is represented and learned using a graph convolutional network to obtain entity node vectors. Calculate the cosine similarity between any two entity node vectors, and treat the two entities as an entity pair, and count the co-occurrence frequency of the entity pair in the first intelligence data; The cosine similarity and the co-occurrence frequency are weighted by a preset weighting coefficient to obtain the association score; If the association score exceeds a preset association threshold, the entity pair is treated as an edge with potential association, and a latent subgraph is constructed.
4. The method for multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets according to claim 1, characterized in that, The step of generating the second knowledge graph through a preset knowledge graph fusion model includes: The vector representations of the entity nodes of the dominant subgraph and the hidden subgraph are obtained respectively, and used as the feature vectors of the dominant nodes and the feature vectors of the hidden nodes. The explicit node feature vector and the implicit node feature vector are input into the independent graph encoding channels of the attention-based dual-channel knowledge distillation network model, including the first graph encoding channel and the second graph encoding channel; The explicit node feature vector and the implicit node feature vector are respectively subjected to graph attention aggregation processing of corresponding sub-graph structures to output explicit node aggregation features and implicit node aggregation features; The explicit node aggregation features and the implicit node aggregation features are subjected to first-direction distillation and second-direction distillation, respectively, and then fused to generate a second knowledge graph.
5. The method for multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets according to claim 4, characterized in that, The first-direction distillation and the second-direction distillation of the dominant node aggregation features and the recessive node aggregation features respectively include: The first directional distillation calculates the first correlation degree between the explicit node aggregation feature and the implicit node aggregation feature to obtain the first attention weight, and then performs a weighted sum with the implicit node aggregation feature to generate a context enhancement vector; The context enhancement vector is fused with the explicit node aggregation feature to obtain the first distillation node feature; The second directional distillation calculates the second correlation degree between the latent node aggregation feature and the explicit node aggregation feature to obtain the second attention weight, and then performs a weighted sum with the explicit node aggregation feature to generate a feature supplement vector; The feature supplement vector is fused with the latent node aggregate feature to obtain the second distillation node feature; The features of the first distillation node and the features of the second distillation node are fused to generate a fused feature vector of the node. The second knowledge graph is generated by combining the edge relationships in the explicit subgraph and the implicit subgraph.
6. The method for multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets according to claim 1, characterized in that, The encoding of the second knowledge graph using a graph embedding algorithm includes: Based on the fusion feature vectors of nodes and the set of edge relationships in the second knowledge graph, the timeliness weights corresponding to each edge are calculated through a preset time-series evaluation function; The timeliness weight is negatively correlated with the collection time of the first intelligence data associated with the edge relationship; Using the topology of the second knowledge graph and the time-sensitivity weight as constraints, the nodes are modulated through a graph attention network to generate neighborhood vectors. The neighborhood vector is fused with the fusion feature vector of the corresponding node to output the semantic vector of each node.
7. The method for multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets according to claim 6, characterized in that, The construction of the semantic association graph based on the semantic vector includes: Calculate the information entropy value of the semantic vector of each node, mark the nodes with entropy values lower than the preset sparse threshold as the first node, and extract the nodes connected to the edges with timeliness weights higher than the preset dynamic threshold and mark them as the second node. Merge the first node and the second node to form the target node set. Based on the topological structure of the target node set and the second knowledge graph, a semantic association navigation graph is constructed; The associated navigation map is calculated using a preset path planning algorithm to generate one or more optimal exploration paths; Based on the optimal exploration path, a subset of collection instructions for publicly available information sources is generated.
8. The method for multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets according to claim 7, characterized in that, The step of generating a subset of collection instructions for publicly available information sources based on the optimal exploration path includes: Extract the node sequence on each optimal exploration path and obtain the edge relationships between nodes from the second knowledge graph; Based on the edge relationship between the node sequence and the node, corresponding query elements are generated according to the preset mapping rules, and the query elements of all nodes on the same path are merged into a subset of the collection instructions corresponding to the path. Based on predefined interface protocol adaptation rules, the subset of collection instructions is instantiated into a collection task of a public information source interface to collect intelligence data streams, which include text, images and audio. The intelligence data stream is formatted, parsed, and its content extracted to obtain standardized second intelligence data.
9. The method for multi-source heterogeneous collection and quantification of competitive intelligence on dynamic intelligence targets according to claim 8, characterized in that, The step of inputting the second intelligence data into a preset quality assessment model to generate a quality assessment result includes: The second knowledge graph is input into a preset time-series graph neural network prediction model to generate a third knowledge graph for a preset future time. The standardized second intelligence data is structured and transformed, and then compared with the third knowledge graph by calculating predefined multi-dimensional quantitative indicators to generate comparison results. The predefined multi-dimensional quantitative indicators include: prediction accuracy score and intelligence completion score. Based on the comparison results, a quality assessment result for the second intelligence data is generated.
10. A multi-source heterogeneous acquisition and quantification system for competitive intelligence of dynamic intelligence targets, used to implement the multi-source heterogeneous acquisition and quantification method for competitive intelligence of dynamic intelligence targets as described in any one of claims 1-9, characterized in that, The system includes: The target perception module is used to collect first intelligence data based on dynamic intelligence targets, construct a first knowledge graph, the first knowledge graph includes explicit subgraphs and implicit subgraphs, and generate a second knowledge graph through a preset knowledge graph fusion model; The encoding module is used to encode the second knowledge graph using a graph embedding algorithm to obtain a semantic vector; The execution module is used to construct a semantic association graph based on the semantic vector, and convert the node series of the semantic association graph into query instructions through a preset mapping rule, execute the query instructions to collect data, and obtain the second intelligence data. The evaluation module is used to input the second intelligence data into a preset quality evaluation model and generate quality evaluation results.