Open source intelligence intelligent acquisition method and system based on multi-source heterogeneous data fusion

By using multi-source heterogeneous data fusion technology, noise is dynamically modeled and stable evidence nodes are constructed, solving the problem of open-source intelligence evidence chains breaking in noisy environments and realizing the construction and updating of stable evidence chains under high noise conditions.

CN121524947AInactive Publication Date: 2026-02-13NANJING LEKBELL INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511713007.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing open-source methods for constructing intelligence evidence chains cannot dynamically model noise, cannot determine whether noise fluctuations cause the evidence chain to break erroneously, and cannot guarantee the stability and reliability of the evidence chain structure under unstable noise conditions.

Method used

A multi-source heterogeneous data fusion method is adopted. Through multi-level noise dynamic modeling and noise weight generation, stable evidence nodes are constructed and the evidence chain is dynamically updated. This includes data preprocessing, noise representation vector construction, noise weight generation and node fusion, and dynamic updating of the evidence chain.

Benefits of technology

It enables the construction of a stable chain of evidence that is updatable, traceable, and unaffected by noise fluctuations in a high-noise environment, thereby improving the structural stability and correlation reliability of the chain of evidence and enabling the link to maintain adaptive adjustment and continuous evolution under noise instability conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121524947A_ABST
    Figure CN121524947A_ABST
Patent Text Reader

Abstract

The invention discloses an open source intelligence intelligent collection method and system based on multi-source heterogeneous data fusion, and relates to the technical field of multi-source heterogeneous data fusion processing, and the method comprises the steps: carrying out the multi-source heterogeneous data collection and preprocessing of a sample, and carrying out the construction of a noise representation vector; performing multi-level noise dynamic modeling on the samples, and performing noise weight generation and node fusion; and establishing stable evidence nodes based on the nodes, and updating the dynamic structure of the evidence chain. According to the method, a stable evidence node construction and evidence chain dynamic structure updating mechanism of noise constraint is introduced, so that the generation and adjustment of a link depend on the comprehensive judgment of time consistency, semantic consistency and noise credibility; and local self-adaptive repair of the node association relationship is realized under the condition of unstable noise or continuous data increment, so that the evidence chain is ensured to still have structural stability, association reliability and continuous evolution ability in a high-noise environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-source heterogeneous data fusion processing, in particular to an open source intelligence intelligent collection method and system based on multi-source heterogeneous data fusion. BACKGROUND

[0002] With the continuous expansion of the application of open source intelligence (OSINT) in national security, public opinion management, anti-fake information analysis and cross-platform event tracking, multi-source heterogeneous data fusion technology has developed rapidly. The academia and industry have introduced cross-modal feature fusion, deep fake detection, graph structure relationship modeling and online data quality evaluation technologies to cope with the high heterogeneity of open source data in terms of source, time span, modal structure and content authenticity. In recent years, researchers have begun to try to use time series feature analysis, noise uncertainty modeling and event link relationship mining to improve the reliability of open source intelligence. However, due to the frequent update of Internet data and the significant dynamic change of noise over time, the noise processing mechanism of existing technologies still cannot adapt to the evidence chain construction needs in a high volatility environment.

[0003] Although the research on OSINT gradually pays attention to data noise processing, there are still obvious defects in the existing technology on the key issues of "dynamic noise modeling", "evidence chain stability under noise disturbance" and "cross-time period link reliability judgment". Firstly, existing noise processing generally stays in static or batch-level modeling, which cannot dynamically sequence the time evolution of source noise, modal noise and instance noise, especially cannot detect short-term mutations and sustained fluctuations of noise, resulting in the inability to construct a time series noise curve reflecting the credibility change of real data, and the inability to answer the core question of "how to dynamically model the time series of noise". Secondly, the existing evidence chain construction algorithm usually uses fixed threshold, homogeneous similarity or simple topological structure, and does not have noise sensitivity analysis capability, which cannot distinguish whether the link breakage is caused by real event change or noise disturbance, so it cannot judge "whether noise fluctuation leads to false breakage of evidence chain", resulting in the decline of link stability and accuracy. Thirdly, the existing OSINT evidence chain generation method lacks a real-time noise feedback mechanism, and the link structure is prone to chain breakage, chain error or false association when facing cross-platform information conflict, sudden appearance of fake text, timestamp drift or data inconsistency, and also lacks the ability to adaptively adjust the node relationship in an unstable noise environment, so it cannot realize "stable evidence chain under unstable noise". SUMMARY

[0004] In view of the above problems, the present application is proposed.

[0005] Therefore, the technical problem solved by the present application is that the existing open source intelligence evidence chain construction and noise processing method cannot dynamically time sequence model the noise, cannot determine whether noise fluctuation causes the evidence chain to be broken, cannot guarantee the stability of the evidence chain structure under unstable noise conditions, and cannot construct a stable evidence chain that can be updated, tracked and not affected by noise fluctuations in a complex multi-source and high-noise environment.

[0006] To solve the above technical problems, the present application provides the following technical solutions: an open source intelligence intelligent collection method based on multi-source heterogeneous data fusion, comprising performing multi-source heterogeneous data collection and preprocessing on samples, and constructing a noise representation vector; performing multi-level noise dynamic modeling on the samples, and generating noise weights and node fusion; constructing stable evidence nodes based on nodes, and updating the dynamic structure of the evidence chain.

[0007] As a preferred scheme of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion, wherein: the multi-source heterogeneous data collection and preprocessing comprises: performing encoding unification, basic denoising, word segmentation and field checking on text data, and triggering a re-parsing operation when abnormal characters appear; performing size normalization and brightness standardization on image data, and extracting brightness consistency and edge gradient statistics as basic features for noise analysis; performing key frame extraction on video data and extracting basic features consistent with images from the key frames; performing field type checking and time field formatting on structured data; after completion, all data is uniformly packaged as a parseable sample object, and a fallback strategy is triggered when there is a field defect or format anomaly.

[0008] As a preferred scheme of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion, wherein: the noise representation vector construction comprises: quantizing the source-level noise, the modal-level noise and the instance-level noise respectively; linearly combining the quantized source-level noise, the modal-level noise and the instance-level noise to generate the final noise value of the sample.

[0009] As a preferred scheme of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion, wherein: the multi-level noise dynamic modeling comprises: performing exponential weighted moving average on the source-level noise to obtain a smoothed trend value, performing batch mean trend calculation on the modal-level noise, performing time series fluctuation detection on the instance-level noise, and combining the three types of trend items to obtain a dynamic noise value.

[0010] As a preferred scheme of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion, the noise weight generation and node fusion comprises: noise weight is calculated according to a dynamic noise value, and node aggregation is performed according to the noise weight; in the node fusion process, weighted summation is performed on samples belonging to the same event category to form a node representation vector, and when the sample semantics deviate from the main sample set, whether the sample is retained to participate in fusion is adaptively determined according to the noise weight.

[0011] As a preferred scheme of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion, the stable evidence node construction comprises: correlation between nodes is judged based on time attributes and semantic attributes of the nodes; a timestamp difference value of two nodes is calculated, and if the timestamp difference value is lower than a preset time threshold, it is considered that the two nodes have correlation in time; similarity calculation is performed on semantic vectors of the two nodes, and the semantic consistency degree is judged by measuring the included angle or inner product size between the two semantic vectors; only when the two nodes simultaneously satisfy the time correlation and the semantic consistency, the two nodes can be considered to have effective correlation; after the effective correlation is established, the connection strength between the nodes is calculated according to the noise weight and the semantic similarity of the nodes, and if the connection strength exceeds a link establishment threshold, a link relationship is generated between the two nodes and becomes part of the evidence chain.

[0012] As a preferred scheme of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion, the evidence chain dynamic structure updating comprises: when a new node is received or the noise level of an existing node is updated, incremental updating is performed according to the local change of the link structure; when the new node is added, the time proximity and the semantic similarity between the new node and the existing node are analyzed, the connection strength is calculated according to the noise weights of the two nodes, and if the connection strength is higher than a link establishment threshold, the new node is added to the link and inserted into a corresponding position; when the noise weight of the existing node is significantly reduced due to noise change, the connection strength between the node and the adjacent nodes is recalculated, if the connection strength is insufficient to maintain the link relationship, the node connection is removed and the order of the surrounding nodes is reorganized; when the event order is disordered or the semantics conflict in the link, the link structure is re-adjusted according to the correlation between the nodes and the noise weight.

[0013] Another object of the present application is to provide an open source intelligence intelligent collection system based on multi-source heterogeneous data fusion.

[0014] As a preferred scheme of the open source intelligence intelligent collection system based on multi-source heterogeneous data fusion, wherein: including data processing module, noise processing module, evidence chain construction module; the data processing module is used for executing multi-source heterogeneous data collection and preprocessing to the sample, and noise characterization vector construction is carried out;The noise processing module is used for executing multi-level noise dynamic modeling to the sample, and noise weight generation and node fusion are carried out;The evidence chain construction module is used for constructing stable evidence node based on node, and the dynamic structure of evidence chain is updated.

[0015] A computer device comprising a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion.

[0016] A computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion.

[0017] The beneficial effects of the present application: the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion provided by the present application realizes the structure unification and noise feature explicitness of different modalities and different source data through the multi-source heterogeneous data collection and noise characterization vector construction steps, so that the originally indescribable noise factors can be quantified at the source level, the modal level and the instance level, thereby providing a calculable and traceable noise basic expression for subsequent dynamic time series modeling. Through multi-level noise dynamic modeling and node fusion based on noise weight, continuous monitoring and adaptive adjustment of noise changes over time are realized, so that the system can actively suppress the influence of high noise samples on node representation in the case of noise surge, modal drift or source anomaly, thereby avoiding node deviation and false link breakage caused by noise fluctuation. Through the introduction of noise-constrained stable evidence node construction and evidence chain dynamic structure updating mechanism, the generation and adjustment of the link rely on the comprehensive judgment of time consistency, semantic consistency and noise credibility, and the local adaptive repair of node association relationship is realized under the condition of noise instability or continuous data increment, so that the evidence chain still has structural stability, association reliability and continuous evolution ability in a high noise environment. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0019] Figure 1A whole flow chart of an open source intelligence intelligent collection method based on multi-source heterogeneous data fusion is provided for the embodiment 1 of the present application. DETAILED DESCRIPTION

[0020] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present application.

[0021] Embodiment 1, refer to Figure 1 For an embodiment of the present application, an open source intelligence intelligent collection method based on multi-source heterogeneous data fusion is provided, comprising:

[0022] S1: Multi-source heterogeneous data collection and preprocessing are performed on the sample, and a noise characterization vector is constructed.

[0023] Further, the multi-source heterogeneous data collection and preprocessing comprises: encoding unification, basic denoising, word segmentation and field checking are performed on the text data, and a re-parsing operation is triggered when an abnormal character appears;

[0024] Size normalization and brightness standardization are performed on the image data, and brightness consistency and edge gradient statistics are extracted as basic features for noise analysis;

[0025] Key frame extraction is performed on the video data, and the same basic features as the image are extracted from the key frame;

[0026] Field type checking and time field formatting are performed on the structured data;

[0027] After completion, all data are uniformly packaged into a parseable sample object, and a fallback strategy is triggered when the field is missing or the format is abnormal.

[0028] It should be noted that the noise characterization vector construction comprises: the source-level noise, the modal-level noise and the instance-level noise are quantified respectively;

[0029] The quantized source-level noise, modal-level noise and instance-level noise are linearly combined to generate the final noise value of the sample, and the calculation of the final noise value is represented as:

[0030]

[0031] wherein, the final noise value is, the noise dimension weight parameter is, the source noise quantization value is, a modal noise quantization value, an example noise quantization value, an index number of the sample.

[0032] It should also be noted that the source-level noise is based on the historical reliability statistics of the platform or account to which the sample belongs, and is formed by calculating long-term abnormality rate, publishing behavior instability and other factors to form a source noise indicator; the modal-level noise is based on the overall noise performance of the text, image, video or structured data in the current batch, and is formed by calculating the abnormality proportion within the modal to form a modal noise; the example-level noise is based on the content abnormality degree of a single sample, such as content loss amount, light inconsistency, frame jitter degree or field missing proportion, etc.

[0033] It should also be noted that by performing quantization on the source-level noise, the modal-level noise and the example-level noise respectively, and then combining the three types of noise through a weight ratio to generate a final noise value, the noise features of each sample can be abstracted into a unified noise representation vector. This not only realizes the fine splitting of the noise dimension, but also realizes the aggregated expression of the noise attribute.

[0034] S2: performing multi-level noise dynamic modeling on the sample, and generating noise weight and node fusion.

[0035] Further, the multi-level noise dynamic modeling includes performing exponential weighted moving average on the source-level noise to obtain a smoothed trend value, performing batch mean trend calculation on the modal-level noise, performing time series fluctuation detection on the example-level noise, and combining the three types of trend items to obtain a dynamic noise value.

[0036] It should also be noted that one preferred solution of the dynamic noise value specifically includes performing exponential weighted moving average processing on the source-level noise to obtain a source-level noise smoothed sequence, denoted as:

[0037]

[0038] wherein, is the smoothed source noise, is the smoothed source noise value of the previous sample, is a smoothing factor; performing modal internal mean trend modeling on the modal-level noise to obtain a modal noise trend item ; performing noise change detection on the example-level noise according to the position of the sample in the time series to form an example trend item . Then, the three types of trend items are combined to calculate the dynamic noise value , denoted as:

[0039]

[0040] wherein, For the noise dimension weight parameter, the dynamic noise value will be used for noise weight generation and node fusion to determine the contribution proportion of different samples when participating in node construction.

[0041] It should be noted that noise weight generation and node fusion includes calculating noise weight according to dynamic noise value, and performing node aggregation according to noise weight;

[0042] In the node fusion process, weighted summation is performed on samples belonging to the same event category to form a node representation vector. When the sample semantics deviates significantly from the main sample set, the noise weight is used to adaptively determine whether to retain the sample for fusion.

[0043] It should also be noted that one preferred solution of node aggregation includes mapping the contribution degree of the sample to a weight value that decreases with the increase of noise according to the dynamic noise value. Through an inverse proportional mapping mechanism, the larger the noise of a sample, the lower its weight, and the smaller the noise of a sample, the higher its weight. This mapping ensures that different samples will not deviate from the node representation due to abnormal noise during node fusion. The noise weight is calculated as follows:

[0044]

[0045] In the node fusion process, weighted summation is performed on samples belonging to the same event category to form a node representation vector , and the calculation process is as follows:

[0046]

[0047] wherein, is the sample semantic vector. When the semantic deviation of a certain sample is significantly higher than that of other samples, its influence will be automatically reduced according to its lower noise weight, and the sample will be removed if necessary, to maintain the consistency and structural stability of the node representation.

[0048] ​It should also be noted that by using multi-level noise dynamic modeling, noise weight generation, and node fusion, continuous tracking of noise temporal changes and adaptive control of the node fusion process are achieved, thereby improving the noise resistance stability of the node representation. This avoids node offset or link error breakage caused by noise fluctuations. By performing exponential weighted moving average on source-level noise, batch mean trend modeling on modal-level noise, and time series fluctuation detection on instance-level noise, dynamic expression of different noise sources in the time dimension is achieved. This approach breaks through the traditional analysis methods based solely on static noise values ​​or single-batch noise distribution, allowing noise to be input into subsequent decisions as a "change curve" rather than an "instantaneous value." Subsequently, by synthesizing the three noise trend terms into a dynamic noise value, this invention can output differentiated judgment results based on noise short-term surges, continuous fluctuations, or stable convergence patterns. This dynamic noise value can comprehensively reflect the noise behavior of samples during the actual acquisition process, which is something that traditional static noise models cannot provide. During the node fusion stage, by mapping dynamic noise to weights that automatically decay as noise increases, the semantic contribution of samples with higher noise levels automatically decreases, while the contribution of samples with lower noise levels automatically increases, thus achieving noise robustness adjustment in the node construction process. Simultaneously, when a sample's semantic direction deviates from the main sample set, this invention can automatically identify, weaken, or even isolate that sample using its noise weights, thereby preventing noise-induced node representation shifts. Fundamentally avoiding node-level mislinks, broken links, or false associations caused by noise fluctuations is a key aspect of solving the problem of "how to determine whether noise fluctuations lead to erroneous breaks in the chain of evidence."

[0049] S3: Perform multi-level noise dynamic modeling on the samples, and generate noise weights and fuse nodes.

[0050] Furthermore, the construction of stable evidence nodes includes determining the correlation between nodes based on their temporal and semantic attributes;

[0051] Calculate the timestamp difference between two nodes. If the timestamp difference is lower than a preset time threshold, the two nodes are considered to be related in time.

[0052] The similarity between the semantic vectors of two nodes is calculated, and the degree of semantic consistency is judged by measuring the angle or the size of the inner product between the two semantic vectors.

[0053] Two nodes are considered to be validly associated only if they simultaneously satisfy the conditions of temporal relevance and semantic consistency.

[0054] After establishing a valid association, the connection strength between nodes is calculated based on the noise weight and semantic similarity of the nodes. If the connection strength exceeds the link establishment threshold, a link relationship is generated between the two nodes and becomes part of the evidence chain.

[0055] It should be noted that a preferred solution for constructing stable evidence nodes specifically includes,

[0056] First, calculate the node time difference:

[0057]

[0058] wherein, is the time difference value of two evidence nodes, used to assess whether two intelligence events occur within a time proximity, is the timestamp of the th evidence node, used to mark the time of occurrence of the event corresponding to the node; is the timestamp of the th evidence node; and when is less than a preset time threshold , it is determined that the two nodes have a time correlation. Then calculate the semantic similarity:

[0059]

[0060] wherein, is the semantic similarity of node and node , used to measure the semantic consistency of the events corresponding to the two nodes; is the semantic vector of node ; is the semantic vector of node ; and when exceeds a set similarity threshold , it is determined that the two nodes are semantically consistent. When both the time correlation and the semantic consistency are met, calculate the connection strength:

[0061]

[0062] wherein, is the connection strength of node and node , used to describe the stability of the link between the two nodes; is the noise weight of node ; is the noise weight of node ; when is greater than a link establishment threshold , a formal connection is established to form an evidence chain structure.

[0063] It should be noted that the dynamic structure update of the evidence chain includes, when a new node is received or the noise level of an existing node is updated, incremental update is performed according to the local changes of the link structure; represented as:

[0064]

[0065] wherein, is the link connection strength of the new node with the existing node , used to measure whether the new node should join the existing evidence chain structure; is the noise weight of the new node , used to reflect its reliability in open source intelligence collection; is the semantic similarity between the new node and the existing node ;

[0066] It should also be noted that in the evidence chain construction process, the time threshold, the semantic similarity threshold and the link establishment threshold are all set adaptively and dynamically based on the noise trend, so as to solve the problem that static thresholds are difficult to adapt to noise fluctuations of multi-source heterogeneous data. Specifically, the dynamic noise value composed of the source-level noise trend, the modal-level noise trend and the instance-level noise trend is used to perform dynamic adjustment on the three types of thresholds respectively, so that the evidence chain can still maintain stability and continuity in a high noise environment.

[0067] Among them, the semantic similarity threshold adopts a noise-sensitive adjustment mechanism. When the dynamic noise value of the node shows an upward trend, the similarity threshold is automatically increased, so that only the nodes with significant semantic consistency can pass the similarity judgment, thereby avoiding the convergence of pseudo semantics caused by noise; when the dynamic noise value of the node decreases or remains stable, the similarity threshold is correspondingly reduced, so as to avoid that the real event link is incorrectly disconnected due to slight noise disturbance. For example, when the dynamic noise value of node i increases from 0.25 to 0.45, the semantic similarity threshold is automatically adjusted from 0.78 to 0.86; when the noise returns to 0.22, the similarity threshold is lowered to 0.75, so that the link judgment remains stable.

[0068] The time threshold adopts a dynamic setting method based on the time credibility of the source. For modalities with high time stamp source reliability (such as video frame time), the time threshold is automatically tightened; for modalities with unstable time stamp source (such as social media text), the time offset allowable range is automatically relaxed. This setting method can effectively avoid the link error breakage caused by the difference in synchronization accuracy of different sources. For example, when node j is derived from a video modality with high credibility, its time credibility score is 0.92, and the time threshold is automatically tightened to 30 seconds; while another node k is derived from a social media text with a credibility of only 0.61, the corresponding time threshold is extended to 180 seconds to accommodate potential time stamp drift or publication delay.

[0069] The link establishment threshold adopts a dynamic confidence interval setting mechanism to ensure that the link generation process is not affected by noise fluctuations. Specifically, according to the statistical distribution of the current inter-node connection strength, a confidence interval is determined as the threshold, and is adjusted upward or downward according to the overall noise level. When the overall noise level of the system is high, the connection strength distribution moves down as a whole, and the link establishment threshold is correspondingly increased to avoid the emergence of a large number of false links in a high noise scenario; when the noise level decreases, the threshold is automatically lowered to improve the link connectivity. For example, in a certain data batch, the 70% quantile of the connection strength is 0.52, and the overall noise mean is 0.41, the link establishment threshold is set to 0.52×(1+0.41)≈0.73; when the next batch noise mean decreases to 0.18, the threshold is automatically adjusted to 0.52×(1+0.18)≈0.61, so that the link establishment process has both noise sensitivity and self-balancing ability.

[0070] When a new node joins, the time proximity and semantic similarity between the new node and the existing nodes are analyzed, and the connection strength is calculated in combination with the noise weights of the two nodes. If the connection strength is higher than the link establishment threshold, the new node is added to the link and inserted into the corresponding position;

[0071] When an existing node's noise weight significantly decreases due to noise changes, the connection strength between the node and its neighboring nodes is recalculated. If the connection strength is insufficient to maintain the link relationship, the node connection is removed and the order of the surrounding nodes is reorganized;

[0072] When the event order within the link is disordered or there is a semantic conflict, the link structure is re-adjusted according to the association between nodes and noise weights.

[0073] It should also be pointed out that by constructing node association with node time attribute and semantic attribute as double constraints, the connection relationship between nodes is established on the common basis of event occurrence time and semantic consistency, avoiding errors caused by relying on single distance or similarity. Subsequently, by introducing node noise weight into connection strength calculation, nodes with high noise are automatically weakened in link connection, thereby reducing their destructive effect on evidence chain structure. This design upgrades the link establishment logic from the traditional "similarity is enough to establish" to a decision mechanism driven by three conditions of "similarity consistency + time proximity + noise credibility", which greatly reduces false associations caused by noise interference. Further, by performing incremental update on the link, when new nodes are added or existing nodes are deteriorated, the node connection strength is re-evaluated, and if necessary, link local disconnection, rearrangement or splitting operations are performed, so that the evidence chain can maintain structural self-consistency in the case of continuous data inflow, noise mutation or content conflict. This local update mechanism avoids the inefficient behavior of triggering full link reconstruction due to single-point noise anomaly, and realizes the high maintainability of evidence chain structure. It realizes the generation of a structure coherent, semantically consistent and sustainable updated evidence chain in a highly unstable noise environment, and is the core creative contribution of the present application to solve the problem of how to realize stable evidence chain in a highly unstable noise environment.

[0074] In embodiment 2, an open source intelligence intelligent collection system based on multi-source heterogeneous data fusion is provided, which includes a data processing module, a noise processing module, and an evidence chain construction module.

[0075] The data processing module is used for multi-source heterogeneous data acquisition and preprocessing on the sample, and noise representation vector construction; the noise processing module is used for multi-level noise dynamic modeling on the sample, and noise weight generation and node fusion; the evidence chain construction module is used for stable evidence node construction based on nodes, and dynamic structure update of the evidence chain.

Claims

1. An open source intelligence intelligent collection method based on multi-source heterogeneous data fusion, characterized in that, The method comprises the following steps: Performing multi-source heterogeneous data collection and preprocessing on the sample, and constructing a noise characterization vector; Performing multi-level noise dynamic modeling on the sample, and generating noise weight and node fusion; Based on the node, stable evidence node construction is performed, and the dynamic structure of the evidence chain is updated.

2. The open source intelligence intelligent collection method based on multi-source heterogeneous data fusion according to claim 1, characterized in that: The multi-source heterogeneous data collection and preprocessing comprises, Performing encoding unification, basic denoising, word segmentation and field checking on text data, and triggering re-parsing operation when abnormal characters appear; Performing size normalization and brightness standardization on image data, and extracting brightness consistency and edge gradient statistics as basic features for noise analysis; Performing key frame extraction on video data and extracting basic features consistent with image from key frames; Performing field type checking and time field formatting on structured data; After completion, all data is uniformly packaged into a parseable sample object, and a fallback strategy is triggered when the field is missing or the format is abnormal.

3. The open source intelligence intelligent collection method based on multi-source heterogeneous data fusion according to claim 2, characterized in that: The noise characterization vector construction comprises, Quantizing the source-level noise, the modal-level noise and the instance-level noise respectively; Linearly combining the quantized source-level noise, the modal-level noise and the instance-level noise to generate the final noise value of the sample.

4. The open source intelligence intelligent collection method based on multi-source heterogeneous data fusion according to claim 3, characterized in that: The multi-level noise dynamic modeling comprises, Performing exponential weighted moving average on the source-level noise to obtain a smoothed trend value, performing batch mean trend calculation on the modal-level noise, and performing time series fluctuation detection on the instance-level noise, and combining the three types of trend items to obtain a dynamic noise value.

5. The open source intelligence intelligent collection method based on multi-source heterogeneous data fusion according to claim 4, characterized in that: The noise weight generation and node fusion comprises, Calculating the noise weight according to the dynamic noise value, and performing node aggregation according to the noise weight; In the node fusion process, weighted summation is performed on samples belonging to the same event category to form a node representation vector, and when the sample semantics deviate from the main sample set, it is determined whether to retain the sample for fusion according to the noise weight.

6. The open source intelligence intelligent collection method based on multi-source heterogeneous data fusion according to claim 5, characterized in that: The stable evidence node construction comprises, Performing relevance judgment between nodes based on the time attribute and the semantic attribute of the node; Calculate the timestamp difference of two nodes, if the timestamp difference is less than the preset time threshold, it is considered that the two nodes have relevance in time; Calculate the similarity of the semantic vectors of the two nodes, and judge the semantic consistency degree by measuring the included angle or inner product size between the two semantic vectors; Only when the nodes meet the time relevance and semantic consistency at the same time, it is determined that the two nodes can establish effective association; After establishing effective association, the connection strength between nodes is calculated according to the noise weight and the semantic similarity of the nodes, and if the connection strength exceeds the link establishment threshold, a link relationship is generated between the two nodes and becomes part of the evidence chain.

7. The open source intelligence intelligent collection method based on multi-source heterogeneous data fusion according to claim 6, characterized in that: The dynamic structure updating of the evidence chain comprises, When a new node is received or the noise level of an existing node is updated, incremental updating is performed according to the local changes of the link structure; When a new node is added, analyze the time proximity and semantic similarity between the new node and the existing node, calculate the connection strength combining the noise weights of the two nodes, and if the connection strength is higher than the link establishment threshold, add the new node to the link and insert it to the corresponding position; When the existing node has a significant decrease in noise weight due to noise change, the connection strength between the node and the adjacent node is recalculated, if the connection strength is not enough to maintain the link relationship, the node connection is removed and the order of the surrounding nodes is reorganized; When the event sequence inside the link is disordered or there is semantic conflict, the link structure is adjusted according to the correlation between nodes and noise weight.

8. An open source intelligence intelligent collection system based on multi-source heterogeneous data fusion, adopting the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion according to any one of claims 1-7. It comprises a data processing module, a noise processing module, and an evidence chain construction module. The data processing module is used for performing multi-source heterogeneous data acquisition and preprocessing on samples, and constructing noise characterization vectors. The noise processing module is used for performing multi-level noise dynamic modeling on samples, and generating noise weight and node fusion. The evidence chain construction module is used for constructing stable evidence nodes based on nodes, and updating the dynamic structure of the evidence chain. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to realize the steps of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion in any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the open source intelligence intelligent collection method based on multi-source heterogeneous data fusion in any one of claims 1 to 7.