Intelligent aggregation method and engine for multi-source data

By performing text segmentation and vectorization mapping on multi-source data streams, and using a large language model to parse atomic fact tuples and perform cross-source entity alignment and logical reasoning, the problem of identifying subtle differences and detecting conflicts in multi-source data aggregation is solved, achieving logically consistent and reliable intelligent aggregation.

CN121808703AInactive Publication Date: 2026-04-07HANGZHOU YOUCAI INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-09
Publication Date
2026-04-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing multi-source data aggregation technologies struggle to accurately identify subtle differences in the same event when faced with semantically complex and highly unstructured text. They lack deep logical reasoning capabilities and conflict detection mechanisms, making it difficult to distinguish the authenticity of the generated aggregation results. Furthermore, they lack the ability to trace the source of original evidence, which limits their credibility and practical application in serious business and legal scenarios.

Method used

By performing text segmentation and vectorization mapping on multi-source data streams, parsing atomic fact tuples using a large language model, performing cross-source entity alignment and fact clustering, executing logical implication reasoning and conflict detection, and generating an intelligent aggregated report containing a traceable chain of evidence.

Benefits of technology

It achieves logically consistent, credible and traceable intelligent aggregation, effectively resolving conflicts and model illusions in the spatiotemporal dimensions of multi-source heterogeneous data, and generating highly confident aggregation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808703A_ABST
    Figure CN121808703A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent aggregation method and engine for multi-source data, and relates to the technical field of data aggregation, and the method comprises the steps: firstly carrying out the segmentation and vectorization mapping of a multi-source data stream containing resumes, news and documents, retaining an original index, and analyzing a fine-grained atomic fact through a fine-tuning large model; then, cross-source alignment and clustering are carried out on the fact tuples, deep logic implication reasoning and conflict detection are executed on the basis, and the support and exclusion states among facts of different sources are quantified; furthermore, truth value optimization is completed by using a global consistency scoring mechanism, and a high-confidence core fact is locked from complex contradictory information. And finally, generating an abstract based on a preferred fact, and establishing an evidence traceability link from the generated content to the source data by using the original index. Thus, logic verification can be fused into the generation process, conflicts and model illusion of multi-source heterogeneous data in the space-time dimension are effectively eliminated, and logic self-consistent, true, credible and traceable intelligent aggregation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data aggregation technology, and more specifically, to an intelligent aggregation method and engine for multi-source data. Background Technology

[0002] With the deepening of digital transformation, the internet and various professional databases have accumulated massive amounts of multi-source heterogeneous data, covering various textual forms such as job seekers' resumes, extensive news media reports, and legally binding court documents. This data is not only vast in quantity but also contains crucial intelligence regarding the behavioral patterns, social evaluations, and potential legal risks of individuals or businesses, serving as an important basis for background checks, talent profiling, and financial risk control. Faced with the current information explosion, relying solely on manual screening or traditional search tools is insufficient to meet the demand for rapid extraction and integration of high-value information. Building an intelligent system capable of deep semantic understanding, cross-source information integration, and outputting high-quality aggregated reports is of significant practical importance for breaking down data silos and supporting accurate decision-making.

[0003] However, existing multi-source data aggregation technologies still face numerous bottlenecks in practical applications. Traditional data processing methods mostly rely on keyword matching or rule-based shallow extraction. While they can accomplish basic information aggregation, they often struggle to accurately identify subtle differences in the same event across different sources when faced with semantically complex and highly unstructured text. More importantly, different data sources are often constrained by recording time, observation perspectives, and subjective stances, frequently resulting in inconsistent or even contradictory descriptions. For example, self-statements in a resume may differ from established facts in court documents in terms of timeline or nature of actions. Existing aggregation schemes typically lack deep logical reasoning capabilities and conflict detection mechanisms, making it difficult to effectively resolve contradictions in multi-source heterogeneous data within complex spatiotemporal contexts. They also cannot dynamically assess the credibility weights of different data sources, leading to aggregated results that are often merely simple information accumulations, making it difficult to distinguish truth from falsehood. Furthermore, current mainstream generative models are prone to "illusion" when integrating information and lack the ability to trace the source of original evidence. Users cannot quickly verify the authenticity of the generated content, which severely restricts the credibility and practical application of intelligent aggregation systems in serious business and legal scenarios.

[0004] Therefore, there is an urgent need for an optimized intelligent aggregation method for multi-source data. Summary of the Invention

[0005] This application is made in order to solve the above-mentioned technical problems.

[0006] According to one aspect of this application, an intelligent aggregation method for multi-source data is provided, comprising: S1: performing text segmentation and vectorization mapping on the acquired original multi-source data stream to obtain a vectorized context data block containing original index pointers, wherein the original multi-source data stream includes resume text, news webpage text, and court judgment text; S2: inputting the vectorized context data block into a large language model fine-tuned by instructions to obtain atomic fact tuples containing subject, predicate, and object information; S3: performing cross-source entity alignment and fact clustering on the atomic fact tuples to obtain clustered fact groups containing descriptions of the same event from different sources; S4: performing logical implication reasoning and conflict detection on the clustered fact groups to obtain a logical relation matrix describing the support and conflict states between facts; S5: performing global consistency scoring and truth value optimization on the logical relation matrix to select a set of verification golden facts with confidence levels meeting a preset threshold; S6: generating a natural language summary based on the set of verification golden facts, and establishing evidence tracing links using the original index pointers in the vectorized context data block to obtain an intelligent aggregation report containing a traceable evidence chain.

[0007] According to another aspect of this application, an intelligent aggregation engine for multi-source data is provided, comprising: a preprocessing and vectorization module for performing text segmentation and vectorization mapping on the acquired raw multi-source data stream to obtain vectorized context data blocks containing raw index pointers, wherein the raw multi-source data stream includes resume text, news webpage text, and court judgment text; an atomic fact extraction module for inputting the vectorized context data blocks into a large language model fine-tuned by instructions to obtain atomic fact tuples containing subject, predicate, and object information; and a cross-source alignment and clustering module for performing cross-source entity alignment and fact extraction on the atomic fact tuples. Clustering is used to obtain clustered fact groups containing descriptions of the same event from different sources; a logical reasoning detection module is used to perform logical implication reasoning and conflict detection on the clustered fact groups to obtain a logical relationship matrix describing the support and conflict states between facts; a truth evaluation and screening module is used to perform global consistency scoring and truth optimization on the logical relationship matrix to screen out the verification golden fact set with confidence levels meeting preset thresholds; an intelligent report generation module is used to generate natural language summaries based on the verification golden fact set and establish evidence tracing links using the original index pointers in the vectorized context data block to obtain an intelligent aggregated report containing a traceable evidence chain.

[0008] Compared with existing technologies, this application provides an intelligent aggregation method and engine for multi-source data. First, it segments and vectorizes multi-source data streams covering resumes, news, and documents, preserving the original index and using a fine-tuned large model to parse out fine-grained atomic facts. Then, it performs cross-source alignment and clustering of fact tuples, and on this basis, performs deep logical implication reasoning and conflict detection to quantify the support and rejection states between facts from different sources. Furthermore, it uses a global consistency scoring mechanism to optimize truth values, locking in high-confidence core facts from complex contradictory information. Finally, it generates a summary based on the optimized facts and uses the original index to establish an evidence-based traceability link from the generated content to the source data. This integrates logical verification into the generation process, effectively resolving conflicts and model illusions in the spatiotemporal dimensions of multi-source heterogeneous data, achieving logically consistent, credible, and traceable intelligent aggregation. Attached Figure Description

[0009] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0010] Figure 1 This is a flowchart of an intelligent aggregation method for multi-source data according to an embodiment of this application.

[0011] Figure 2 This is a data flow diagram of an intelligent aggregation method for multi-source data according to an embodiment of this application.

[0012] Figure 3 This is a flowchart of sub-step S1 of the intelligent aggregation method for multi-source data according to an embodiment of this application.

[0013] Figure 4 This is a flowchart of sub-step S2 of the intelligent aggregation method for multi-source data according to an embodiment of this application.

[0014] Figure 5 This is a flowchart of sub-step S3 of the intelligent aggregation method for multi-source data according to an embodiment of this application.

[0015] Figure 6 This is a flowchart of sub-step S4 of the intelligent aggregation method for multi-source data according to an embodiment of this application.

[0016] Figure 7 This is a flowchart of sub-step S5 of the intelligent aggregation method for multi-source data according to an embodiment of this application.

[0017] Figure 8 This is a flowchart of sub-step S6 of the intelligent aggregation method for multi-source data according to an embodiment of this application.

[0018] Figure 9 This is a block diagram of an intelligent aggregation engine for multi-source data according to an embodiment of this application. Detailed Implementation

[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0020] To address the problems mentioned above in the background technology, this application proposes an intelligent aggregation method for multi-source data. Figure 1 This is a flowchart of an intelligent aggregation method for multi-source data according to an embodiment of this application. Figure 2 This is a data flow diagram for an intelligent aggregation method for multi-source data according to an embodiment of this application. (See also...) Figure 1 and Figure 2 As shown, the intelligent aggregation method for multi-source data includes the following steps: S1, performing text segmentation and vectorization mapping on the acquired original multi-source data stream to obtain a vectorized context data block containing original index pointers. The original multi-source data stream includes resume text, news webpage text, and court judgment text; S2, inputting the vectorized context data block into a large language model that has been fine-tuned by instructions to obtain atomic fact tuples containing subject, predicate, and object information; S3, performing cross-source entity alignment and fact clustering on the atomic fact tuples to obtain clustered fact groups containing descriptions of the same event from different sources; S4, performing logical implication reasoning and conflict detection on the clustered fact groups to obtain a logical relation matrix describing the support and conflict states between facts; S5, performing global consistency scoring and truth value optimization on the logical relation matrix to select a set of verification golden facts with confidence levels meeting a preset threshold; S6, generating a natural language summary based on the set of verification golden facts, and establishing evidence tracing links using the original index pointers in the vectorized context data block to obtain an intelligent aggregation report containing a traceable evidence chain.

[0021] In the aforementioned intelligent aggregation method for multi-source data, step S1 involves performing text segmentation and vectorization mapping on the acquired original multi-source data stream to obtain a vectorized context data block containing the original index pointers. The original multi-source data stream includes resume text, news webpage text, and court judgment text. It should be understood that because the original multi-source data stream encompasses various heterogeneous formats such as resume text, news webpage text, and court judgment text, its unstructured nature and lack of traceability mechanisms can lead to context loss or factual illusions during subsequent model inference. Therefore, this application performs text segmentation and vectorization mapping on the original multi-source data stream while simultaneously retaining the original index pointers, thereby transforming the chaotic heterogeneous data into a standardized vector format that is machine-understandable and traceable. This provides high-quality input with semantic integrity and positional certainty for subsequent atomic fact extraction, ensuring that the intelligent aggregation system maintains accurate traceability of the original evidence chain while understanding semantics.

[0022] In particular, in one specific embodiment, Figure 3 This is a flowchart of sub-step S1 of the intelligent aggregation method for multi-source data according to an embodiment of this application. Figure 3 As shown, step S1 includes: S11, performing multimodal format parsing and cleaning on the original multi-source data stream to obtain cleaned text corpus; S12, performing sliding segmentation and byte offset calculation on the cleaned text corpus based on preset window size and step size parameters to obtain a list of indexed text blocks containing content and position indexes; S13, performing high-dimensional feature embedding mapping on the list of indexed text blocks to obtain vectorized context data blocks.

[0023] Specifically, step S11 involves performing multimodal format parsing and cleaning on the original multi-source data stream to obtain cleaned text corpus. It should be understood that the original multi-source data stream contains widespread non-semantic noise such as HTML tags, ad placeholders, garbled characters, and invisible characters, and the encoding formats and layouts of files from different sources differ significantly. Direct processing would severely interfere with the accuracy of subsequent semantic feature extraction. Therefore, this application performs multimodal format parsing and cleaning on the original multi-source data stream to remove redundant information unrelated to business logic and unify heterogeneous encodings into a standard format. This significantly improves the purity and standardization of the text corpus, eliminates the obstacle of format noise to model understanding, and lays a solid data foundation for subsequent refined segmentation and feature extraction.

[0024] Specifically, in one possible embodiment, the multimodal attributes of the input data stream are determined by reading the file header signature, and the data is then distributed to the corresponding parsing logic. For PDF resumes and documents, optical character recognition or document object modeling is used to extract plain text content; for HTML news web pages, scripts, style sheets, and tag elements are removed by traversing the document tree structure. Based on this, regular expressions are used to match and remove advertising links and special interfering symbols, and all extracted text content is converted to a unified character encoding format. The final output is a cleaned text corpus containing only valid semantic information, thus ensuring the semantic consistency of the input data.

[0025] Specifically, in step S12, based on preset window size and step size parameters, sliding segmentation and byte offset calculation are performed on the cleaned text corpus to obtain a list of indexed text blocks containing content and position indices. It should be understood that since the text length of court documents or long news reports often exceeds the context window limit of large language models, and simple hard truncation would disrupt the semantic coherence across paragraphs, leading to logical breaks in key information. Therefore, this application further performs sliding segmentation and byte offset calculation on the cleaned text corpus based on preset window size and step size parameters to generate continuous text fragments with overlapping areas and accurately record their physical locations. This ensures that long texts maintain semantic integrity and smooth transitions after segmentation, while assigning precise coordinate attributes to each text block, achieving byte-level positioning of the original evidence.

[0026] Specifically, in one possible embodiment, the window length and step length configuration are first loaded (e.g., setting the window length to 512 characters or tokens, and the step length to 128 characters or tokens), constructing a sliding window that moves across the cleaned text corpus. As the window moves forward according to the step length, the text sequence within the window's coverage area is truncated, ensuring that a certain proportion of overlapping text is retained between adjacent windows to maintain semantic association. Simultaneously, based on the mapping relationships or document structure parsing information maintained during the cleaning process, the precise physical location of the currently truncated segment in the original multi-source data stream (including file identifier, page number, and corresponding original byte offset) is reverse-anchored and determined. The truncated text content and its corresponding position index are packaged and encapsulated to generate a series of indexed text block lists that contain both local semantic content and global position information, ensuring that subsequent logical reasoning can be mapped back to the original text context.

[0027] Specifically, step S13 involves performing high-dimensional feature embedding mapping on the list of indexed text blocks to obtain vectorized context data blocks. It should be understood that since natural language text is an unstructured sequence of symbols, computers cannot directly perform efficient mathematical operations and logical reasoning on it; it must be converted into a machine-understandable numerical form for deep semantic mining. Therefore, this application further performs high-dimensional feature embedding mapping on the list of indexed text blocks to utilize a pre-trained model to convert discrete text symbols into dense vectors containing rich semantic information. This allows the semantic features of text blocks to be mapped into a unified high-dimensional vector space, enabling subsequent steps to accurately quantify the semantic similarity and logical connections between texts from different sources through the geometric distance between vectors.

[0028] Specifically, in one possible embodiment, a Transformer-based deep language model (such as BERT or RoBERTa) pre-trained on a large-scale corpus is loaded, and a list of indexed text blocks is fed into the model in batches as input. The model's internal multi-head self-attention mechanism captures the bidirectional contextual dependencies between characters in the text, performing deep feature extraction on each text block, and encoding the variable-length text sequence into fixed-dimensional real-number vectors containing rich contextual semantic information. Subsequently, the generated feature vectors, along with the original text content and previously calculated position index pointers, are encapsulated at the object level to construct indivisible vectorized contextual data blocks, which serve as standard input units for subsequent fact extraction and logical reasoning, supporting high-performance vector retrieval and computation.

[0029] In the aforementioned intelligent aggregation method for multi-source data, step S2 involves inputting the vectorized context data block into a finely tuned large language model to obtain atomic fact tuples containing subject, predicate, and object information. It should be understood that although the vectorized context data block retains the semantic features and physical location index of the original text, its essence remains unstructured natural language representation. Computers find it difficult to directly perform precise logical reasoning and fact comparison based on such fuzzy, continuous text features. Therefore, this application further inputs the vectorized context data block into a finely tuned large language model to leverage the powerful semantic understanding and information extraction capabilities of the large language model, transforming unstructured text into standardized structured data. This allows complex, long texts to be deconstructed into atomic fact tuples containing subject, predicate, and object information, providing finely granular and semantically clear computational units for subsequent cross-source entity alignment and conflict detection, thereby significantly improving the accuracy and automation level of information extraction in background investigations.

[0030] In particular, in one specific embodiment, Figure 4This is a flowchart of sub-step S2 of the intelligent aggregation method for multi-source data according to an embodiment of this application. Figure 4 As shown, step S2 includes: S21, combining the text content in the vectorized context data block with system instructions and their source pointers based on a predefined background investigation fact extraction pattern to obtain a reasoning prompt context sequence; S22, inputting the reasoning prompt context sequence into a large language model to obtain the original model output text that conforms to the structured conventions; S23, performing structured parsing and format verification on the original model output text to obtain atomic fact tuples.

[0031] Specifically, in step S21, based on a predefined background check fact extraction pattern, the text content in the vectorized context data block is combined with system instructions and their source pointers to obtain a reasoning hint context sequence. It should be understood that large language models are prone to uncontrollable formatting or content illusions in free-generation mode, and simple text generation often loses the physical location information of the original evidence, leading to an inability to trace the final conclusion. Therefore, this application further combines the text content in the vectorized context data block with system instructions and their source pointers based on a predefined background check fact extraction pattern to construct a reasoning hint context sequence with strict constraints. This forces the model to follow specific data structure specifications at the input end and strongly binds the original index pointer to the text to be processed, ensuring that each subsequently generated fact tuple has a clear source basis, thereby effectively suppressing model illusions and ensuring the integrity of the evidence chain.

[0032] Specifically, in one possible embodiment, a pre-defined background check fact extraction pattern is first read. This pattern explicitly specifies that the output data must include five core fields: subject, predicate, object, time range, and confidence level. Next, system-level instructions are concatenated with the text block content to be processed using prompt word template logic. The system instructions explicitly require extraction based solely on the given text, strictly prohibiting the addition of unmentioned information. Simultaneously, the source pointer of the text block in the original file is embedded as metadata into the prompt word structure. Finally, the composite object containing structured constraint instructions, the original natural language text, and source index information is encapsulated into a reasoning prompt context sequence, providing input data with complete contextual constraints for subsequent model inference.

[0033] Specifically, in step S22, the inference prompt context sequence is input into the large language model to obtain the original model output text that conforms to structured conventions. It should be understood that since the inference prompt context sequence is only an input instruction in natural language form, the computer program cannot directly store it in a database or perform logical operations on it; it needs to be inferred by a deep neural network to be converted into a machine-readable structured string. Therefore, this application further inputs the inference prompt context sequence into the large language model to perform an autoregressive generation task, thereby utilizing the model's generation capabilities to output the original model output text that conforms to structured conventions. This ensures that the content output by the model strictly follows the preset JSON or tuple format, explicitly expressing the semantic information implicit in natural language as a computer-parsable string, providing standard intermediate data for subsequent programmatic processing and logical verification.

[0034] Specifically, in one possible embodiment, the constructed inference prompt context sequence is input into a finely tuned large language model via an application programming interface (API) or local call. During inference, the model parameters are set to a low-randomness mode, such as reducing temperature parameters, to ensure the stability and determinism of the output results. Subsequently, an autoregressive decoding strategy is executed to predict and generate character sequences conforming to a preset syntax one by one, while a forced structured output mode is enabled, restricting the generation of only strings conforming to JSON syntax. After a complete inference cycle, the output is the original text string containing the extracted information, which strictly corresponds to the input context content and presents the extracted factual elements in a structured key-value pair format.

[0035] Specifically, step S23 involves performing structured parsing and format validation on the original model output text to obtain atomic fact tuples. It should be understood that since the original model output text is essentially still in string format, and the model may occasionally contain syntax errors, missing fields, or meaningless filler content during generation, direct use may lead to system crashes or data pollution. Therefore, this application further performs structured parsing and format validation on the original model output text to deserialize the string into a program object and filter out noisy data that does not conform to business rules. This ensures that the final generated atomic fact tuples are fully compliant in terms of data type and field integrity, while eliminating invalid entries that cannot be mapped back to the original text or lack key elements, ensuring that the data flowing into subsequent stages has a high degree of cleanliness and usability.

[0036] Specifically, in one possible embodiment, the string output by the model is first converted into a list of objects in memory using parsing logic. Next, each object is iterated over to check if it contains predefined necessary fields such as subject, predicate, object, and time, and the data type of each field is verified to conform to the specifications. For entries with missing fields or malformed formats, automatic discarding or correction operations are performed. Finally, the valid data objects that pass the verification are further standardized, and the original source pointer passed down in step S1 is appended to the object, ultimately encapsulating it into an atomic fact tuple with complete metadata support, which serves as the standard input for subsequent entity alignment and conflict detection.

[0037] In the aforementioned intelligent aggregation method for multi-source data, step S3 involves performing cross-source entity alignment and fact clustering on atomic fact tuples to obtain clustered fact groups containing descriptions of the same event from different sources. It should be understood that because atomic fact tuples originate from heterogeneous data sources such as resumes, news articles, and documents, and are independent of each other, the same real event often exhibits differences in expression or semantic fragmentation in different texts, making direct comparison and verification extremely difficult. Therefore, this application further implements cross-source entity alignment and fact clustering processing on atomic fact tuples to consolidate descriptions pointing to the same entity or event scattered across different data sources into a unified logical set. This effectively eliminates the interference of the diversity of multi-source expressions on fact identification, constructs semantically relevant analytical units for subsequent in-depth logical conflict detection and truth value determination, and ensures that the comparative analysis is conducted on the same event dimension.

[0038] In particular, in one specific embodiment, Figure 5 This is a flowchart of sub-step S3 of the intelligent aggregation method for multi-source data according to an embodiment of this application. Figure 5 As shown, step S3 includes: S31, concatenating the subject, predicate, and object fields of the atomic fact tuple into a semantic string and mapping it into a dense vector, and obtaining a pairwise similarity matrix describing the semantic distance between tuples by calculating the cosine similarity between the vectors; S32, based on the similarity matrix, performing density-based clustering on all atomic fact tuples to obtain clustered fact groups.

[0039] Specifically, in step S31, the subject, predicate, and object fields of the atomic fact tuples are concatenated into a semantic string and mapped to a dense vector. The cosine similarity between the vectors is then calculated to obtain a pairwise similarity matrix describing the semantic distance between tuples. It should be understood that traditional keyword matching techniques struggle to capture synonym substitutions or complex sentence transformations in natural language, and literal character matching alone cannot accurately measure the deep semantic connections between different tuples. Therefore, this application further concatenates the subject, predicate, and object fields of the atomic fact tuples into a semantic string and maps it to a dense vector. A pairwise similarity matrix is ​​constructed by calculating the cosine similarity between the vectors, thereby transforming discrete symbol comparison into geometric distance calculation in a continuous vector space. This allows for precise quantification of the deep semantic affinity between tuples from different sources, providing a mathematically rigorous and robust metric for subsequent clustering.

[0040] Specifically, in one possible embodiment, each atomic fact tuple is first read, and the subject name, predicate action, and object contained therein are concatenated into a complete declarative sentence according to natural language word order. Next, a pre-trained semantic encoding model based on a Siamese network architecture is invoked. This model has been pre-tuned in supervised manner on large-scale natural language inference datasets such as SNLI or MNLI using a triplet loss function or a contrastive loss function to narrow the distance between semantically similar texts in the vector space and widen the distance between semantically irrelevant texts, thereby optimizing the isotropy of the sentence embedding space. Then, the declarative sentence is input into the model, and the text is encoded in parallel using a dual-tower structure with shared weights. All token vectors from the output layer are aggregated using an average pooling strategy, transforming them into fixed-dimensional real-valued feature vectors that contain the contextual semantic features of the sentence. Subsequently, a dot product operation is performed on all generated vectors pairwise and normalized to generate a pairwise similarity matrix describing the semantic distance between tuples. The values ​​in the matrix intuitively reflect the closeness of any two fact tuples in the semantic space.

[0041] Specifically, in step S32, density-based clustering is performed on all atomic fact tuples based on the similarity matrix to obtain clustered fact groups. It should be understood that, due to the complexity of candidates' backgrounds and the unknown number of events, traditional clustering algorithms based on a preset number of categories are ill-suited to this unstructured and dynamically changing data distribution, easily leading to distorted clustering results. Therefore, this application further performs density-based clustering on all atomic fact tuples based on the similarity matrix to automatically discover high-density connected semantic regions in the data and remove discrete noise. In this way, without needing to pre-specify the number of clusters, multiple descriptions belonging to the same event can be adaptively merged into independent and tightly clustered fact groups, ensuring that all relevant evidence for the same event is completely aggregated for subsequent verification.

[0042] Specifically, in one possible embodiment, the similarity matrix is ​​first converted into a distance metric matrix, and neighborhood radius and minimum point threshold parameters are set (e.g., a normalized distance value between 0.15 and 0.25 is set for the neighborhood radius, and a minimum point count of 2 is set to accommodate the clustering requirements of small sample events). Next, a density-based spatial clustering algorithm is applied to scan the entire dataset, connecting the core points of tuples whose distance is less than the neighborhood radius and whose density meets the requirements, expanding them into clusters. Isolated tuples that do not belong to any high-density region are marked as noise or processed separately. Finally, all tuples marked with the same cluster identifier are encapsulated into a clustered fact group, outputting a set containing multiple such groups, thus achieving the ordered organization and classification of disordered facts.

[0043] In the aforementioned intelligent aggregation method for multi-source data, step S4 involves performing logical implication reasoning and conflict detection on the clustered fact groups to obtain a logical relationship matrix describing the support and conflict states between facts. It should be understood that although the clustered facts point to the same event, heterogeneous data from different sources often contain subtle semantic conflicts or timeline contradictions, and simple text clustering cannot automatically identify these deep logical oppositions. Therefore, this application further performs logical implication reasoning and conflict detection on the clustered fact groups to deeply analyze the semantic interaction relationships between the facts within the group and establish their logical compatibility. This transforms qualitative textual descriptions into quantitative logical support or exclusion weights, constructing a rigorous logical topology structure and providing a solid mathematical foundation and data support for subsequent graph-based global consistency scoring and truth value selection.

[0044] In particular, in one specific embodiment, Figure 6 This is a flowchart of sub-step S4 of the intelligent aggregation method for multi-source data according to an embodiment of this application. Figure 6 As shown, step S4 includes: S41, performing pairwise permutations and combinations of facts within the clustered fact group and assigning premise and hypothesis roles to obtain a batch sequence of inference pairs that conforms to the model input specifications; S42, calling the natural language inference model to perform natural language inference model inference on the batch sequence of inference pairs to obtain the original inference probability vector set; S43, performing logical matrix mapping and sparsification on the original inference probability vector set to obtain the logical relation matrix.

[0045] Specifically, step S41 involves performing pairwise permutations and combinations of facts within a clustered fact group and assigning premise and hypothesis roles to obtain a batch sequence of inference pairs that conforms to the model input specifications. It should be understood that since the fact set within a clustered fact group is unordered, and natural language inference models require explicit directed text pairs to determine asymmetric logical relationships, and exhausting all possible interaction paths is a prerequisite for discovering implicit conflicts, this application further performs pairwise permutations and combinations of facts within the clustered fact group and assigns premise and hypothesis roles to transform discrete fact items into standardized logical verification units and cover all possible inference directions. This ensures that the model can comprehensively evaluate the logical relationship between any two facts, eliminate directional bias, and provide the neural network with structured sequence data that meets the requirements of its input layer.

[0046] Specifically, in one possible implementation, each generated cluster of facts is traversed, and all atomic fact tuples contained therein are extracted. Then, a full permutation algorithm is used to pair these tuples together, generating a list of pairs containing all possible directions. In each pair, one fact is explicitly designated as the logical premise, and the other as the logical assumption. Finally, strictly adhering to the input format requirements of the pre-trained model, these paired fact texts are concatenated and truncated with specific delimiters to construct a batch sequence of inference pairs that conforms to sequence length constraints and vocabulary specifications, ready to be fed into the model for computation.

[0047] Specifically, in step S42, a natural language inference model is invoked to perform natural language inference on the batch sequence of inference pairs to obtain the original inference probability vector set. It should be understood that since the constructed text pairs are merely combinations at the character level, the computer cannot directly quantify their inherent semantic consistency or degree of contradiction, requiring the use of deep neural networks to capture complex contextual interaction features. Therefore, this application further invokes a natural language inference model to perform natural language inference on the batch sequence of inference pairs, thereby utilizing a cross-attention mechanism to deeply analyze the subtle connections between premises and assumptions in the semantic space. In this way, unstructured text pairs can be mapped to a high-dimensional probability space, accurately outputting the numerical confidence level of each fact pair in the three dimensions of implication, neutrality, and contradiction, transforming fuzzy linguistic logic into precise computational vectors.

[0048] Specifically, in one possible embodiment, the constructed batch sequences of inference pairs are input into a deep neural network model based on an interactive encoder architecture, fine-tuned for natural language inference tasks (e.g., using models such as DeBERTa or ALBERT as a base, and pre-trained and fine-tuned using large-scale natural language inference datasets such as MNLI or SNLI). The premise and hypothesis texts are concatenated into a single sentence sequence using a special delimiter. The concatenated sequence is jointly encoded using a multi-layer transformer structure within the model, achieving deep interaction between premises and hypotheses at the token level through a multi-head attention mechanism. Subsequently, the classification identifier vector at the beginning of the sequence is extracted as a global semantic representation. This high-dimensional feature is mapped to a three-dimensional logical space through the model's classification layer, and a normalized exponential function is applied to calculate the posterior probability values ​​belonging to the implication, contradiction, and neutral categories. Finally, the output is a set of original inference probability vectors containing the logical judgments of all input pairs, which digitally represents the logical form between facts.

[0049] Specifically, step S43 involves performing logical matrix mapping and sparsification on the original inferred probability vector set to obtain a logical relation matrix. It should be understood that since the original inferred probability vector set is a discrete list of values, it lacks a topological structure that intuitively reflects the global relationships between facts, and it contains a large amount of low-confidence noise data that can interfere with the final truth decision. Therefore, this application further performs logical matrix mapping and sparsification on the original inferred probability vector set to reconstruct the probability list into a two-dimensional adjacency relationship and eliminate ambiguous weak connections. This allows for the construction of a logical topological graph that clearly describes the interactions between facts. Sparsification significantly reduces the complexity of subsequent graph calculations and highlights strong logical relationships, ensuring that the final consistency assessment is based on reliable and significant evidence.

[0050] Specifically, in one possible embodiment, the implied and contradictory probabilities are first extracted from the inference vector and converted into logical weight values ​​with positive and negative signs, respectively. Next, based on the index positions of the fact pairs, these weight values ​​are filled into the corresponding coordinates of a two-dimensional matrix to construct a preliminary association matrix. Subsequently, a preset confidence threshold is applied to traverse all elements in this matrix, forcing weights with absolute values ​​below the threshold to zero, thereby severing connections with ambiguous logical relationships. Finally, a sparse logical relationship matrix retaining only significant logical features is generated as the mathematical basis for global truth inference.

[0051] In the aforementioned intelligent aggregation method for multi-source data, step S5 involves performing global consistency scoring and truth value optimization on the logical relationship matrix to filter out a set of verifiable golden facts with confidence levels meeting a preset threshold. It should be understood that since the preceding steps only identified local logical relationships between facts and did not combine evidence from the entire network to form a final judgment on the truth or falsehood of the facts, and isolated evidence from a single data source is often insufficient to support high-risk decisions, this application further performs global consistency scoring and truth value optimization on the logical relationship matrix to transform discrete binary logical relationships into quantitative confidence indicators that reflect the overall picture of the facts. In this way, the collective intelligence and overall topological structure of the logical network can be used to offset noise or bias from individual data sources, accurately filtering out a set of highly credible facts that can withstand scrutiny from a complex network of contradictory information, providing a solid and error-free data foundation for generating the final report.

[0052] In particular, in one specific embodiment, Figure 7 This is a flowchart of sub-step S5 of the intelligent aggregation method for multi-source data according to an embodiment of this application. Figure 7 As shown, step S5 includes: S51, converting the logical relation matrix into a directed weighted graph, and assigning initial trust values ​​to the fact nodes in the directed weighted graph according to the source authority table to obtain an initial confidence graph; S52, performing iterative trust propagation and convergence calculation on the initial confidence graph to obtain a converged fact scoring table; S53, performing truth threshold filtering and mutual exclusion conflict resolution on the converged fact scoring table to obtain a verification golden fact set.

[0053] Specifically, in step S51, the logical relationship matrix is ​​converted into a directed weighted graph, and initial trust values ​​are assigned to the fact nodes in the directed weighted graph according to the source authority table to obtain an initial confidence graph. It should be understood that while a simple numerical matrix records logical weights, it lacks a topological structure that can intuitively express the dependencies and propagation paths between nodes, and it does not reflect the differences in authority of different sources at the legal or business level. Therefore, this application further converts the logical relationship matrix into a directed weighted graph and assigns initial trust values ​​to the fact nodes in the graph according to the source authority table to construct an initial confidence graph with prior knowledge and structured connections. This allows abstract logical judgments to be mapped to a traversable graph structure, ensuring that facts from authoritative channels such as courts or governments receive higher initial weights, providing an accurate starting point and topological foundation for subsequent trust propagation algorithms.

[0054] Specifically, in one possible embodiment, the input logical relation matrix is ​​first parsed, mapping the row and column indices of the matrix to nodes in a graph structure. Non-zero values ​​in the matrix elements are converted into directed edges connecting the nodes, where positive values ​​are defined as supporting edges and negative values ​​as repulsive edges. Next, a pre-defined source authority table is read, which defines the basic credit scores for different channels such as court documents, news media, and resume texts. Subsequently, each fact node in the graph is traversed, its original source attribute is identified, and the corresponding score is matched from the authority table as the initial trust value for that node. Finally, an initial confidence graph containing both logical topology and prior trust weights is constructed, ready for the iterative calculation phase.

[0055] Specifically, step S52 involves iterative trust propagation and convergence calculation on the initial confidence graph to obtain a converged fact score table. It should be understood that since the initial trust value only represents the credibility of isolated sources and does not yet reflect the mutual corroboration or falsification effects generated by logical chains between facts, it cannot identify the deeper truth supported by multiple pieces of evidence. Therefore, this application further implements iterative trust propagation and convergence calculation on the initial confidence graph to simulate the flow of confidence in the logical network, allowing nodes supported by high-confidence neighbors to receive score bonuses, while nodes opposed by conflicting nodes are suppressed. In this way, the trust distribution of the entire system can be stabilized through multiple rounds of mathematical deduction, calculating the final score that integrates prior authority and posterior logical verification, thereby quantifying the probability of the truthfulness of each fact.

[0056] Specifically, in one possible embodiment, the tail-blocking coefficient and convergence threshold parameters for iterative calculation are first set (e.g., the tail-blocking coefficient is set to 0.85, and the convergence error threshold is set to 0.001), and the trust propagation algorithm is started. In each iteration, all nodes in the graph are traversed, and according to the preset propagation rules, each node collects trust votes from its incoming neighbor nodes, where supporting edges transmit positive trust and conflicting edges transmit negative penalties. Next, the collected trust values ​​are weighted and updated based on the node's initial authority to generate a new confidence level for the current round. Subsequently, the maximum change in scores for all nodes between two iterations is calculated and compared with the convergence threshold. If the change is greater than the threshold, the next iteration continues until the system reaches dynamic equilibrium, ultimately outputting a converged fact score table containing stable scores for all fact nodes.

[0057] Here, based on the logical relationship matrix describing the support and conflict states between facts, the resulting initial confidence graph provides a static prior score, which is solely based on the authority of the source, such as a court judgment having a higher score than a blog post. Therefore, if a linear weighted iteration, similar to page ranking, is used to iteratively calculate the dynamic confidence of fact nodes, it may have flaws in background investigation scenarios. Specifically, linear weighted iteration has two algorithmic blind spots. First, there is the temporal blind spot, which assumes that the weight of all evidence is determined solely by the source, ignoring the time dimension. For example, in a background check, a recent low-confidence source can overturn a previous high-confidence source (such as the latest departure status compared to an employment record from ten years ago). Second, there is the risk equivalence fallacy. That is, if "support" and "conflict" are treated linearly symmetrically, in the field of risk control, the destructive power of "conclusive opposing evidence" far outweighs the constructive power of "supporting evidence." Linear accumulation cannot simulate this non-linear characteristic of "one-vote veto" or "high-risk sensitivity."

[0058] Therefore, when iteratively calculating the dynamic confidence of fact nodes, it is preferable to introduce a time-decaying kernel function and a nonlinear logarithmic activation mechanism to construct the confidence update process as a Bayesian log-odds update process, enabling the algorithm to prioritize the acceptance of new data and be more sensitive to conflicts. That is, the dynamic confidence of fact nodes is iteratively calculated by accumulating evidence in the logarithmic space, specifically including the following dimensions: Log-odds evidence transformation: In the logarithmic space, the initial trust value of the fact node is transformed into an initial evidence value through a log-odds function.

[0059] Time-series dynamic weights: for each neighbor node The transmitted vote of confidence is multiplied by a time decay factor based on the time difference. Because background checks are more sensitive to recent circumstances, The smaller the value, the more weight is retained.

[0060] Nonlinear asymmetric aggregation: dividing neighboring nodes into "support sets". With "Conflict Set" And set a conflict penalty coefficient. Greater than the support reward coefficient This ensures that the points deducted from conflicting evidence for each unit are significantly greater than the points gained from supporting evidence, adhering to the prudent principle of "better to misjudge than to overlook risks," and imposing asymmetric penalty weights on conflicting evidence.

[0061] Sigmoid normalization mapping: to ensure the final score The probability is represented by the fact that it always falls within the range of 0 to 1. The Sigmoid activation function is used to perform a nonlinear superposition mapping of the initial evidence value, the sum of supporting evidence after time decay and reward coefficient weighting, and the sum of conflicting evidence after time decay and penalty coefficient weighting, so as to obtain the dynamic confidence in the converged fact score table.

[0062] Therefore, the first The fact node at the th Confidence in round iteration The sigmoid activation function (log odds of base trust + (sum of supporting evidence × time factor × reward coefficient) - (sum of conflicting evidence × time factor × penalty coefficient)) is specifically expressed as: ,here, It is the logistic sigmoid activation function, used to map evidence to probabilities. It is a logarithmic probability function, i.e. Used to set the initial prior probability Convert to initial evidence value. It is a fact The static source prior authority (e.g., court = 0.95). and They respectively point to facts Supportive neighbor sets (implied relations) and conflicting neighbor sets (contradictory relations) are derived from the logical relation matrix. and It is the logical edge weight (i.e., NLI prediction probability). Neighboring nodes The confidence level in the previous iteration, Neighboring nodes The confidence level in the previous iteration. It is the time sensitivity decay constant (which can take values ​​from 0.1 to 1.0). The larger the value, the faster the algorithm forgets old data and the more inclined it is to accept the latest data. Current time and evidence The time difference between generation times Current time and evidence The time difference between generation times. and These are the support reward coefficient and the conflict penalty coefficient, where the setting is... 1 (if possible values) =1.1, =2.2), in order to mathematically establish the risk control logic that negative evidence has a greater weight than positive evidence.

[0063] In this way, graph computation acquires time awareness and risk aversion, significantly improving the accuracy of conflict resolution, especially in time-sensitive single-fact determinations such as former departures and subsequent arrivals. The algorithm is no longer biased by a large amount of old data and accurately identifies the latest state. Simultaneously, through asymmetric penalties ( The system is able to more sensitively detect subtle signs of fraud, avoid concealing the truth, strengthen the system's compliance and security, and directly enhance the authority of background investigation reports.

[0064] Specifically, step S53 involves performing truth threshold filtering and mutual exclusion conflict resolution on the converged fact scoring table to obtain a verification golden fact set. It should be understood that the converged scoring table may still contain long-tailed noise data with low confidence, and in some extreme cases, there may be situations where two mutually exclusive facts have high scores but cannot coexist physically. Therefore, this application further performs truth threshold filtering and mutual exclusion conflict resolution operations on the converged fact scoring table to determine the final facts according to strict business standards and forcibly resolve residual logical paradoxes. This completely eliminates substandard and questionable information and ensures that the output results have absolute logical uniqueness and self-consistency, forming a high-quality verification golden fact set, providing accurate material for generating unambiguous intelligent aggregation reports.

[0065] Specifically, in one possible embodiment, a preset truth-based threshold is first loaded. Each fact entry in the scoring table is then iterated through, and entries with confidence scores below the threshold are deemed unreliable and discarded. Next, conflict group verification is performed on the candidate facts filtered by the threshold, identifying fact pairs belonging to the same cluster but with mutually exclusive content. For such residual conflicts, a score-maximizing resolution strategy is implemented, retaining the fact with the highest score as the final truth value while removing mutually exclusive entries with lower scores. Finally, all the filtered and resolved fact nodes are encapsulated, along with their final confidence scores and original source indexes, constructing a rigorously structured verification golden fact set.

[0066] In the aforementioned intelligent aggregation method for multi-source data, step S6 involves generating a natural language summary based on the verified golden fact set and establishing evidence tracing links using the original index pointers in the vectorized context data block to obtain an intelligent aggregation report containing a traceable evidence chain. It should be understood that while the verified golden fact set possesses high logical accuracy, its essence remains discrete structured data tuples, lacking the coherence and readability of natural language, and simply presenting facts cannot intuitively prove their original source to the user. Therefore, this application further generates a natural language summary based on this fact set and establishes evidence tracing links using the original index pointers to reconstruct fragmented facts into a coherent business report and endow it with evidence tracing capabilities. This ensures that the final deliverable is both readable and verifiable, allowing users to access the original evidence at any time when reading the conclusions, significantly enhancing the report's credibility.

[0067] In particular, in one specific embodiment, Figure 8 This is a flowchart of sub-step S6 of the intelligent aggregation method for multi-source data according to an embodiment of this application. Figure 8 As shown, step S6 includes: S61, grouping the verification golden fact set according to business dimensions, and processing it using generative large model reconstruction to obtain a draft portrait with placeholders containing reference placeholders; S62, dynamically binding evidence anchors to the draft portrait with placeholders and vectorized context data blocks to obtain fully traceable interactive content; S63, calculating the global risk index based on the negative fact weights in the verification golden fact set, and encapsulating it with the risk summary and fully traceable interactive content to obtain an intelligent aggregation report.

[0068] Specifically, in step S61, the set of key facts for verification is grouped according to business dimensions, and processed using a generative large model reconstruction to obtain a draft profile with placeholders for citations. It should be understood that business scenarios such as background investigations have strict requirements for report structure, and directly generating long text can easily lead to misplaced citations, making it impossible to accurately insert subsequent evidence links into the corresponding text descriptions. Therefore, this application further groups the set of key facts for verification according to business dimensions and processes it using a generative large model reconstruction to construct a hierarchical text framework that reserves precise citation interfaces. This ensures that the generated report logically conforms to business reading habits, and at the same time, the placeholder mechanism forces the model to retain evidence insertion points when generating content, avoiding the risk of misplacement during subsequent source tracing and binding, and ensuring the accuracy of citation positions.

[0069] Specifically, in one possible embodiment, the entries in the verification golden fact set are first divided into subsets based on multiple dimensions, such as work experience, educational background, legal proceedings, and social evaluation, according to preset business logic rules. Next, these grouped factual data and their corresponding unique index identifiers are input as context into a generative large language model, and specific prompts are configured, requiring the model to retain a standardized citation mark containing the corresponding unique identifier at the end of each factual statement when generating narrative paragraphs. Subsequently, the model performs a text reconstruction task, outputting a draft profile containing specific placeholder symbols carrying identifier information. This draft maintains the fluency of natural language while clearly marking the specific locations where evidence links need to be inserted later and their corresponding index relationships.

[0070] Specifically, in step S62, the placeholder-based portrait draft and the vectorized context data block are dynamically bound to evidence anchors to obtain fully traceable interactive content. It should be understood that because the placeholder-based portrait draft only has a text structure and is not yet linked to actual physical evidence, users cannot view the original file by clicking, resulting in a broken chain of evidence in the report. Therefore, this application further dynamically binds evidence anchors to the placeholder-based portrait draft and the vectorized context data block to replace abstract text reference markers with interactive entity evidence links. This transforms static text reports into rich media content with deep interactive capabilities, enabling users to jump to the original screenshot or text fragment with a click, completely eliminating the black-box effect in the information aggregation process and achieving true fully traceable interactive content.

[0071] Specifically, in one possible embodiment, the draft image with placeholders is first traversed, and all reference marker symbols are parsed using regular expressions. Next, based on the unique identifier in the marker, the corresponding original file metadata, including the file ID, start byte offset, and end byte offset, is searched in the vectorized context data block index. Subsequently, standardized jump protocol links are constructed using this physical coordinate information, and the text placeholders in the draft are replaced with these link objects. Simultaneously, different highlighting styles are configured for the generated links based on the confidence attribute of the facts, ultimately outputting fully tracing interactive content that includes fluent text and precise evidence jump functionality.

[0072] Specifically, in step S63, a global risk index is calculated based on the negative fact weights in the verification gold fact set, and this index is then encapsulated with a risk summary and fully traceable interactive content to obtain an intelligent aggregated report. It should be understood that while purely narrative reports are detailed, they lack quantitative decision-making basis. Users find it difficult to quickly assess the overall risk level when faced with a large amount of factual details, and the dispersed data objects are not convenient for inter-system transmission and archiving. Therefore, this application further calculates a global risk index based on the negative fact weights and encapsulates it as a whole to provide an intuitive quantitative risk rating and package all results into a standard format. This provides users with easily understandable decision support indicators, while integrating a risk overview, detailed report, and underlying evidence chain into a unified delivery object, greatly improving the usability and workflow efficiency of the intelligent aggregated report in actual business processes.

[0073] Specifically, in one possible embodiment, the golden fact set is first scanned and verified, and a pre-trained risk classification model is used to identify risk fact items marked with negative attributes, such as criminal records or falsified resumes. This risk classification model is built on a large-scale sensitive domain corpus (covering court documents, administrative penalty announcements, and negative news reports). During the training phase, a supervised fine-tuning paradigm is adopted, where manually labeled risk category tags and text samples are input into the basic pre-trained language model. By optimizing the multi-class cross-entropy loss function, it is equipped with the ability to capture semantic features and perform qualitative analysis of unseen risk events. Next, a score is assigned to each identified negative fact according to a pre-set risk weight table, and a weighted cumulative algorithm is used to calculate the global risk index of the target subject. Subsequently, a risk warning list containing summaries of high-risk items is generated, and the index, list, and full-source interactive content are structurally assembled. Finally, all the assembled data is serialized into a standard JSON or PDF format file and output as the final intelligent aggregation report to the client or downstream business system.

[0074] In summary, the intelligent aggregation method for multi-source data based on the embodiments of this application is explained. First, it segments and vectorizes multi-source data streams covering resumes, news, and documents, preserving the original index and using a fine-tuned large model to parse out fine-grained atomic facts. Then, it performs cross-source alignment and clustering of fact tuples, and on this basis, performs deep logical implication reasoning and conflict detection to quantify the support and rejection states between facts from different sources. Furthermore, it utilizes a global consistency scoring mechanism to achieve truth value optimization, locking in high-confidence core facts from complex contradictory information. Finally, it generates a summary based on the optimized facts and uses the original index to establish an evidence traceability link from the generated content to the source data. In this way, logical verification is integrated into the generation process, effectively resolving conflicts and model illusions in the spatiotemporal dimensions of multi-source heterogeneous data, achieving logically consistent, credible, and traceable intelligent aggregation.

[0075] Figure 9 This is a block diagram of an intelligent aggregation engine for multi-source data according to an embodiment of this application. Figure 9 As shown, the intelligent aggregation engine 100 for multi-source data according to an embodiment of this application includes: a preprocessing and vectorization module 110, used to perform text segmentation and vectorization mapping processing on the acquired original multi-source data stream to obtain a vectorized context data block containing original index pointers, wherein the original multi-source data stream includes resume text, news webpage text, and court judgment text; an atomic fact extraction module 120, used to input the vectorized context data block into a large language model fine-tuned by instructions to obtain atomic fact tuples containing subject, predicate, and object information; and a cross-source alignment and clustering module 130, used to perform cross-source entity alignment and fact clustering on the atomic fact tuples. The system obtains clustered fact groups containing descriptions of the same event from different sources; a logical reasoning detection module 140 is used to perform logical implication reasoning and conflict detection on the clustered fact groups to obtain a logical relationship matrix describing the support and conflict states between facts; a truth value evaluation and screening module 150 is used to perform global consistency scoring and truth value optimization on the logical relationship matrix to screen out the verification golden fact set with confidence levels meeting preset thresholds; and an intelligent report generation module 160 is used to generate natural language summaries based on the verification golden fact set and establish evidence tracing links using the original index pointers in the vectorized context data block to obtain an intelligent aggregation report containing a traceable evidence chain.

[0076] As described above, the intelligent aggregation engine 100 for multi-source data according to the embodiments of this application can be implemented in various wireless terminals, such as servers with intelligent aggregation algorithms for multi-source data. In one possible implementation, the intelligent aggregation engine 100 for multi-source data according to the embodiments of this application can be integrated into the wireless terminal as a software module and / or a hardware module. For example, the intelligent aggregation engine 100 for multi-source data can be a software module in the operating system of the wireless terminal, or it can be an application developed for the wireless terminal; of course, the intelligent aggregation engine 100 for multi-source data can also be one of many hardware modules of the wireless terminal.

[0077] Alternatively, in another example, the intelligent aggregation engine 100 for multi-source data and the wireless terminal can also be separate devices, and the intelligent aggregation engine 100 for multi-source data can be connected to the wireless terminal via wired and / or wireless networks, and transmit interactive information in accordance with an agreed data format.

[0078] Here, those skilled in the art will understand that the specific operations of each step in the intelligent aggregation engine for multi-source data described above have been referenced above. Figures 1 to 8The method for intelligent aggregation of multi-source data is described in detail in the description of the method, and therefore, its repeated description will be omitted.

Claims

1. A method for intelligent aggregation of multi-source data, characterized in that, include: S1: Perform text segmentation and vectorization mapping on the acquired raw multi-source data stream to obtain a vectorized context data block containing the original index pointer. The raw multi-source data stream includes resume text, news webpage text, and court judgment text. S2: Input the vectorized context data block into the instruction-fine-tuned large language model to obtain atomic fact tuples containing subject, predicate, and object information; S3: Perform cross-source entity alignment and fact clustering on the atomic fact tuples to obtain clustered fact groups containing descriptions of the same event from different sources; S4: Perform logical implication reasoning and conflict detection on clustered fact groups to obtain a logical relationship matrix describing the supporting and conflicting states between facts; S5: Perform global consistency scoring and truth value optimization on the logical relationship matrix to select the set of golden facts for verification that meet the preset confidence threshold; S6: Generate a natural language summary based on the verified golden fact set, and establish evidence tracing links using the original index pointers in the vectorized context data block to obtain an intelligent aggregated report containing a traceable evidence chain.

2. The intelligent aggregation method for multi-source data according to claim 1, characterized in that, Step S1 includes: performing multimodal format parsing and cleaning on the original multi-source data stream to obtain cleaned text corpus; performing sliding segmentation and byte offset calculation on the cleaned text corpus based on preset window size and step size parameters to obtain a list of indexed text blocks containing content and position indexes; and performing high-dimensional feature embedding mapping on the list of indexed text blocks to obtain vectorized context data blocks.

3. The intelligent aggregation method for multi-source data according to claim 2, characterized in that, Step S2 includes: combining the text content in the vectorized context data block with system instructions and their source pointers based on a predefined background investigation fact extraction pattern to obtain an inference hint context sequence; inputting the inference hint context sequence into a large language model to obtain the original model output text that conforms to the structured conventions; and performing structured parsing and format verification on the original model output text to obtain atomic fact tuples.

4. The intelligent aggregation method for multi-source data according to claim 3, characterized in that, Step S3 includes: concatenating the subject, predicate, and object fields of the atomic fact tuples into a semantic string and mapping it into a dense vector; calculating the cosine similarity between the vectors to obtain a pairwise similarity matrix describing the semantic distance between tuples; and based on the similarity matrix, performing density-based clustering on all atomic fact tuples to obtain clustered fact groups.

5. The intelligent aggregation method for multi-source data according to claim 1, characterized in that, Step S4 includes: performing pairwise permutations and combinations of facts within the clustered fact groups and assigning premise and hypothesis roles to obtain a batch sequence of inference pairs that conforms to the model input specifications; calling the natural language inference model to perform natural language inference on the batch sequence of inference pairs to obtain the original inference probability vector set; and performing logical matrix mapping and sparsification on the original inference probability vector set to obtain the logical relation matrix.

6. The intelligent aggregation method for multi-source data according to claim 5, characterized in that, Step S5 includes: converting the logical relation matrix into a directed weighted graph, and assigning initial trust values ​​to the fact nodes in the directed weighted graph according to the source authority table to obtain an initial confidence graph; performing iterative trust propagation and convergence calculation on the initial confidence graph to obtain a converged fact scoring table; and performing truth threshold filtering and mutual exclusion conflict resolution on the converged fact scoring table to obtain a verification golden fact set.

7. The intelligent aggregation method for multi-source data according to claim 6, characterized in that, An iterative trust propagation and convergence calculation is performed on the initial confidence graph to obtain the converged fact score table. This includes: introducing a time-decaying kernel function and a nonlinear log activation mechanism to construct the confidence update process as a Bayesian log-odds update process; in the log space, converting the initial trust value of the fact node into the initial evidence value through the log-odds function; calculating the time-series dynamic weights of the trust votes passed by neighboring nodes; performing nonlinear asymmetric aggregation to divide neighboring nodes into support and conflict sets, and setting the conflict penalty coefficient to be greater than the support reward coefficient to apply asymmetric penalty weights to conflict evidence; and using the Sigmoid activation function to perform a nonlinear superposition mapping of the initial evidence value, the sum of support evidence weighted by time-decaying and reward coefficients, and the sum of conflict evidence weighted by time-decaying and penalty coefficients to obtain the dynamic confidence in the converged fact score table.

8. The intelligent aggregation method for multi-source data according to claim 1, characterized in that, Step S6 includes: grouping the verification golden fact set according to business dimensions, and processing it using generative large model reconstruction to obtain a draft portrait with placeholders containing reference placeholders; dynamically binding evidence anchors to the draft portrait with placeholders and vectorized context data blocks to obtain fully traceable interactive content; calculating the global risk index based on the negative fact weights in the verification golden fact set, and encapsulating it with the risk summary and fully traceable interactive content to obtain an intelligent aggregation report.

9. An intelligent aggregation engine for multi-source data, characterized in that, include: The preprocessing and vectorization module is used to perform text segmentation and vectorization mapping on the acquired raw multi-source data stream to obtain vectorized context data blocks containing the original index pointers. The raw multi-source data stream includes resume text, news webpage text, and court judgment text. The atomic fact extraction module is used to input vectorized context data blocks into a large language model that has been fine-tuned by instructions to obtain atomic fact tuples containing subject, predicate, and object information. The cross-source alignment and clustering module is used to perform cross-source entity alignment and fact clustering on atomic fact tuples to obtain clustered fact groups containing descriptions of the same event from different sources. The logical reasoning detection module is used to perform logical implication reasoning and conflict detection on clustered fact groups to obtain a logical relationship matrix describing the support and conflict states between facts; The truth evaluation and filtering module is used to perform global consistency scoring and truth optimization on the logical relationship matrix in order to filter out the set of golden facts for verification that meet the preset confidence threshold. The intelligent report generation module is used to generate natural language summaries based on the verified golden fact set and to establish evidence tracing links using the original index pointers in the vectorized context data block to obtain an intelligent aggregated report containing a traceable chain of evidence.

Citation Information

Patent Citations

  • Multi-source data fusion type intelligent data management system

    CN121051167A

  • Data knowledge-based method based on semantic fusion

    CN121257543A

  • System and method for conflict resolution and truth extraction from multi-source document ingestion

    IN202521068259A

Cited By

  • Rental state monitoring method and equipment for operational assets, and medium

    CN122089446A