Open source intelligence analysis system and method driven by large language model

By constructing a deep coupling mechanism between the domain knowledge graph and the LLM generation process, the credibility of intelligence assertions is evaluated in real time. A self-attention aggregation algorithm is used for verification, which solves the illusion problem of large language models in open source intelligence analysis and achieves efficient credibility analysis and multi-path correction.

CN120430416BActive Publication Date: 2025-11-28BEIJING ANRUISHENG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510576421.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-11-28
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

Large Language Models (LLMs) suffer from the illusion problem in open-source intelligence analysis, generating false content that does not conform to the facts, making the output difficult to trust and apply. Existing solutions are inefficient and have limited coverage.

Method used

By constructing a deep coupling mechanism between the domain knowledge graph and the LLM generation process, the credibility of intelligence assertions is evaluated in real time. A self-attention aggregation algorithm is used for semantic query verification, forming a closed-loop process of generation-verification-iterative optimization, thereby reducing the risk of illusion penetration.

Benefits of technology

It significantly improves the credibility analysis capability of intelligence assertions, reduces the risk of illusion infiltration, realizes multi-path correction of low-confidence assertions, and forms an efficient closed-loop optimization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430416B_ABST
    Figure CN120430416B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of open source intelligence analysis, and particularly discloses a large language model driven open source intelligence analysis system and method. Public information data is collected, multi-source heterogeneous original data is gathered, semantic ambiguity and noise are eliminated through data cleaning and preprocessing, and a structured domain knowledge graph is constructed as a fact benchmark by relying on authoritative databases, patent documents and other reliable sources. The preprocessed data is input into an information encoder based on a large language model to generate preliminary intelligence assertions. Then, the preliminary intelligence assertions and the associated knowledge graph are extracted, and the consistency and factuality thereof are dynamically checked based on a semantic query formula self-attention aggregation mechanism to determine the confidence of the preliminary intelligence assertions. In this way, the confidence of the intelligence assertions can be effectively analyzed and judged, which helps to trigger a multi-path correction strategy for low-confidence assertions, forms a closed-loop process of "generation-verification-iterative optimization", and significantly reduces the illusion penetration risk.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of open source intelligence analysis, and more specifically, to a large language model driven open source intelligence analysis system and method. BACKGROUND

[0002] As an important part of the intelligence system, open source intelligence analysis (OSINT) is facing multidimensional challenges in the era of information explosion. With the breakthrough progress of large language models (LLMs) in text generation, semantic understanding, etc., they are widely used in intelligence summary generation, entity relationship mining, multi-source information correlation, etc. However, the "hallucination" problem of LLMs, i.e. generating false content that seems reasonable but does not conform to the facts, has become a core bottleneck restricting its landing in the field of intelligence analysis. If hallucination cannot be effectively alleviated, the intelligence output by LLMs cannot be trusted and applied, which is directly related to the feasibility and value of the entire method.

[0003] This problem is caused by two technical characteristics: training data pollution: LLMs rely on internet public data for training, while there are a lot of noise, contradictory information and malicious false content (such as social media rumors, fake documents, etc.) in open networks. The "knowledge" learned by the model through probability statistics is essentially a fitting of the distribution of training data, rather than an accurate modeling of the real world. Defects in the generation mechanism: LLMs generate text word by word based on autoregressive probability prediction, and the optimization goal is language fluency rather than factual correctness. When the input context information is insufficient, the model tends to fill in the information gap through "reasonable guesses" that prioritize semantic coherence, leading to factual errors.

[0004] Existing solutions focus on model fine-tuning (such as reinforcement training based on trusted data) or post-hoc manual verification, but have problems such as low efficiency, limited coverage, etc. For example, traditional rule engines can only detect explicit contradictions and cannot deal with potential errors at the complex semantic level; manual review is difficult to match the real-time generation speed of LLMs.

[0005] Therefore, an optimized large language model driven open source intelligence analysis scheme is expected. SUMMARY

[0006] To solve the above technical problems, the present application is proposed. The embodiments of the present application provide a large language model driven open source intelligence analysis system and method, which can effectively analyze and judge the confidence of intelligence assertions, and then help trigger a multi-path correction strategy for low-confidence assertions, forming a closed-loop process of "generation-verification-iterative optimization", significantly reducing the risk of hallucination penetration.

[0007] According to an aspect of the present application, a large language model driven open source intelligence analysis method is provided, comprising: collecting public information data; performing data cleaning and preprocessing on the public information data to obtain preprocessed public information data; extracting domain knowledge from a highly trusted source to construct a domain knowledge graph, and storing the domain knowledge graph in a trusted domain knowledge base; inputting the preprocessed public information data into an information encoder based on a large language model to obtain a set of preliminary intelligence assertions; extracting a first preliminary intelligence assertion from the set of preliminary intelligence assertions, and extracting a plurality of domain knowledge graphs from the trusted domain knowledge base; verifying the first preliminary intelligence assertion based on semantic query self-attention aggregation analysis between the first preliminary intelligence assertion and the plurality of domain knowledge graphs to determine the confidence of the first preliminary intelligence assertion.

[0008] According to another aspect of the present application, a large language model driven open source intelligence analysis system is provided, comprising: a public information collection module for collecting public information data; a public information processing module for performing data cleaning and preprocessing on the public information data to obtain preprocessed public information data; a domain knowledge graph construction module for extracting domain knowledge from a highly trusted source to construct a domain knowledge graph, and storing the domain knowledge graph in a trusted domain knowledge base; a preliminary intelligence assertion generation module for inputting the preprocessed public information data into an information encoder based on a large language model to obtain a set of preliminary intelligence assertions; a data extraction module for extracting a first preliminary intelligence assertion from the set of preliminary intelligence assertions, and extracting a plurality of domain knowledge graphs from the trusted domain knowledge base; a preliminary intelligence assertion verification module for verifying the first preliminary intelligence assertion based on semantic query self-attention aggregation analysis between the first preliminary intelligence assertion and the plurality of domain knowledge graphs to determine the confidence of the first preliminary intelligence assertion.

[0009] The large language model driven open source intelligence analysis system and method provided by the present application collects public information data, aggregates multi-source heterogeneous raw data, eliminates semantic ambiguity and noise through data cleaning and preprocessing. Relying on authoritative databases, patent documents and other trusted sources, a structured domain knowledge graph is constructed as a fact benchmark. The preprocessed data is input into an information encoder based on a large language model to generate preliminary intelligence assertions. Then, the preliminary intelligence assertions and associated knowledge graphs are extracted, and their consistency and factuality are dynamically verified based on a semantic query self-attention aggregation mechanism to determine the confidence of the preliminary intelligence assertions. In this way, the confidence of the intelligence assertions can be effectively analyzed and judged, which helps to trigger a multi-path correction strategy for low-confidence assertions, forming a closed-loop process of "generation-verification-iterative optimization", and significantly reducing the risk of illusion penetration. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described in the following description only relate to some embodiments of the present application, and are not a limitation on the present application.

[0011] Figure 1 A schematic flow chart of the open source intelligence analysis method driven by the large language model of the embodiments of the present application.

[0012] Figure 2 A schematic flow chart of step S6 in the open source intelligence analysis method driven by the large language model of the embodiments of the present application.

[0013] Figure 3 A schematic flow chart of step S63 in the open source intelligence analysis method driven by the large language model of the embodiments of the present application.

[0014] Figure 4 A schematic flow chart of step S631 in the open source intelligence analysis method driven by the large language model of the embodiments of the present application.

[0015] Figure 5 A schematic flow chart of step S632 in the open source intelligence analysis method driven by the large language model of the embodiments of the present application.

[0016] Figure 6 A schematic block diagram of the open source intelligence analysis system driven by the large language model of the embodiments of the present application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor also belong to the scope of protection of the present application.

[0018] In view of the above technical problems, in the technical scheme of the present application, a large language model driven open source intelligence analysis method is proposed, which realizes real-time reliability evaluation of intelligence assertions through the deep coupling mechanism of constructing field knowledge graph and LLM generation process. Specifically, the open source intelligence analysis method extracts entity relationships from high reliability sources (such as authoritative databases, patent documents, verified military / technology reports), constructs a multi-level field knowledge graph, and provides a structured fact benchmark for LLM output. And using self-attention aggregation algorithm, the preliminary assertions generated by LLM are dynamically semantically mapped and queried matched with the knowledge graph, and the reliability of the assertions is quantified through logical consistency analysis. In this way, the confidence of the intelligence assertions can be effectively analyzed and judged, and then it is helpful to trigger multi-path correction strategies (such as knowledge graph backtracking retrieval, multi-model cross-validation) for low confidence assertions, forming a closed-loop process of "generation-verification-iterative optimization", which significantly reduces the illusion penetration risk.

[0019] Specifically, Figure 1 The large language model driven open source intelligence analysis method of the embodiments of the present application is shown in the schematic flowchart. As shown in Figure 1 The large language model driven open source intelligence analysis method comprises: S1, collecting public information data; S2, performing data cleaning and preprocessing on the public information data to obtain preprocessed public information data; S3, extracting field knowledge from a highly trusted source to construct a field knowledge graph, and storing the field knowledge graph in a trusted field knowledge base; S4, inputting the preprocessed public information data into an information encoder based on a large language model to obtain a set of preliminary intelligence assertions; S5, extracting a first preliminary intelligence assertion from the set of preliminary intelligence assertions, and extracting a plurality of field knowledge graphs from the trusted field knowledge base; S6, verifying the first preliminary intelligence assertion based on semantic query self-attention aggregation analysis between the first preliminary intelligence assertion and the plurality of field knowledge graphs to determine the confidence of the first preliminary intelligence assertion.

[0020] Specifically, in step S1, public information data is collected. It should be understood that since the large language model aims to efficiently extract and reason about the complex, multi-source external world, first of all, the collection coverage is maximized. The public information data is extensive, multilingual, and multi-structured. Such data includes but is not limited to social media posts, news reports, government public documents, industry reports, patent databases, academic papers, technical white papers, and various legally accessible network content. The fundamental purpose of collecting public information data is to obtain factual materials, background clues, real-time dynamics, social opinions, etc., to provide sufficient original corpus for subsequent construction of knowledge graph and support for large model semantic learning and assertion generation.

[0021] It should be noted that the disclosed information data involved in the present application is authorized by the user or fully authorized by all parties, and the collection, use and processing of related data shall comply with relevant laws, regulations and standards of the country and region.

[0022] Specifically, in step S2, the disclosed information data is data cleaned and pre-processed to obtain pre-processed disclosed information data. It should be understood that in the open source intelligence analysis scene, the core value of data cleaning and preprocessing is to solve the problems of heterogeneity, redundancy and noise interference of the original disclosed information data. Since the disclosed information data sources cover social media text, news web pages, PDF reports, table data and other non-structured or semi-structured forms (such as Twitter tweets containing informal abbreviations, news web pages containing advertising interference text, and PDF documents containing layout noise), directly inputting a large language model will make it difficult for the information encoder to effectively extract semantic features, thereby affecting the generation quality of intelligence assertions. Therefore, the disclosed information data is further data cleaned and pre-processed to obtain pre-processed disclosed information data. Through the format standardization (such as converting HTML / PDF to pure text), entity normalization (such as unifying different expressions of "OpenAI" to a standard entity) and noise filtering (such as deleting advertising code and repeated paragraphs) implemented in the data cleaning stage, fragmented raw data can be converted into machine-parsable format with unified semantic representation. This process not only provides a structured data foundation for subsequent knowledge graph construction (such as extracting "company-technology-patent" triples through named entity recognition), but more importantly, it reduces the probability of false associations caused by input noise in the LLM generation process by eliminating data ambiguity.

[0023] Specifically, in step S3, domain knowledge is extracted from highly trusted sources to construct a domain knowledge graph, and the domain knowledge graph is stored in a trusted domain knowledge base. It should be understood that since the "hallucination" problem of LLMs is essentially caused by noise pollution of open network data and probabilistic generation mechanism, and traditional manual verification schemes cannot meet the real-time requirements, a verifiable knowledge benchmark independent of LLM training data system must be established. Therefore, in view of the inherent defects of large language model training data and the stringent requirements of intelligence analysis scene on factual accuracy, in the technical solution of the present application, domain knowledge is further extracted from highly trusted sources to construct a domain knowledge graph, and the domain knowledge graph is stored in a trusted domain knowledge base.

[0024] By selecting authoritative databases (such as the Global Patent Index), peer-reviewed scientific literature, and technical white papers, the authenticity of the entity relationships in the knowledge graph can be ensured. These sources have been reviewed by professional institutions and have the advantages of high information density, less noise interference, and traceable factual associations compared to open data such as social media. For example, in patent intelligence analysis, the "inventor-technology keyword-patent claim" triplets extracted from official patent data can form a semantic network covering the technology evolution path. The construction process of this knowledge graph involves structured processing such as entity disambiguation (e.g., distinguishing between inventors with the same name) and relationship verification (e.g., confirming the legal validity of patent citation relationships), forming machine-readable standardized knowledge units that support subsequent judgments on whether the assertions generated by large language models conflict with the technical roadmap.

[0025] In a specific embodiment, the construction of the domain knowledge graph first requires batch collection and high-precision extraction of highly reliable data sources. Taking the patent database as an example, the latest patent announcement text can be automatically collected, and structured elements such as technical invention points, patent claims, applicants, inventors, patent classifications, literature citations, technical effects, and application fields are identified. In the context of scientific papers, the title, author, abstract, key technical terms, method framework, experimental indicators, and cited literature metadata need to be accurately extracted. For such authoritative literature and databases, AI techniques such as semantic analysis, natural language processing, named entity recognition (NER), and relationship extraction (RE) are often used to automatically extract multi-dimensional entities (such as technical terms, company / organization names, domain experts, technical roadmaps, materials, and experimental methods) and semantic relationships between entities (such as "inventor-creation-technical solution," "organization-application-patent," "key term-association-patent claim," and "research method-validation-experimental results"). The entire modeling process integrates entity disambiguation (i.e., unified reference for homonyms and term variants), relationship accuracy verification, and temporal logic consistency discrimination.

[0026] The constructed domain knowledge graph is expressed in the form of RDF (Resource Description Framework) triplets "entity-relation-entity," which can be further expanded into a multi-level, multi-attribute, and multi-weight relationship network. For example, patent entities and their technical elements, inventors, affiliated organizations, citation history, technical roadmaps, and core keywords form multi-dimensional structured nodes; the association between nodes is refined to specific attributes such as time, space, ownership, process, citation, and negative limitations. To enhance retrieval efficiency and subsequent large model semantic decoding capabilities, the knowledge graph is also equipped with full-dimensional encoding and ontology hierarchy in the downstream, preparing for subsequent complex embedding reasoning, semantic comparison, and assertion structured mapping.

[0027] In a more specific embodiment, the goal is to build a knowledge graph in the field of high-end chip technology. First, automatically batch-grab the patent announcements of global mainstream chip manufacturers in the past three years. After the text is processed by the OCR and structure standardization module, the entities are identified according to the "inventor-key technology point-process route-patent number" dimension. Extract entities such as "TSMC-three-dimensional stacking process-patent number XXX", "Samsung-5nm EUV lithography technology-patent number YYY", "Intel-high-density packaging method-patent number ZZZ" and relationship instances. Then supplement the WIPO / EPO and CNKI scientific paper databases, and extract the "CMOS integration-performance improvement index", "production process-failure rate", "process node-scaling experiment" content in the papers in parallel, and couple the paper results with the patent technology route. At this time, the system uses the entity disambiguation module to automatically map "5nm EUV lithography", "5nm EUV lithography", and "high-resolution extreme ultraviolet lithography", and to attribute the key entity attributes, such as technology maturity, recent breakthrough time, and associated experts. The global knowledge built is finally automatically converted into RDF or graph databases such as Neo4j, OrientDB, etc. machine-readable structure, and is archived with hierarchical labels and numbers.

[0028] Specifically, in step S4, the pre-processed public information data is input into an information encoder based on a large language model to obtain a set of preliminary intelligence assertions. It should be understood that although the cleaned and structured public data has eliminated format noise, it still contains a large amount of implicit associated information (such as cross-document technical term co-occurrence, temporal implication of enterprise strategic trends), and it is difficult for traditional rule engines or statistical methods to effectively extract deep semantic features. Therefore, the pre-processed public information data is further input into an information encoder based on a large language model to obtain a set of preliminary intelligence assertions. Through the vector space mapping capability of the information encoder, the dispersed technical entities (such as "5G millimeter wave antenna array design") in the text, behavior descriptions (such as "a company submits a PCT patent application"), and context can be converted into high-dimensional semantic representations, and then the generation-based reasoning of LLMs is used to generate preliminary intelligence assertions (for example, inferring that "the enterprise is laying out the radio frequency front-end technology of the sixth generation communication base station"). This step essentially realizes the jump in information density through the probabilistic completion mechanism of the model, thereby fully utilizing the unique advantages of LLMs in unstructured text understanding and potential pattern discovery, and abstracting the implicit but not explicitly stated technical trends, competitive relationships, etc. in the original data into verifiable proposition assertions.

[0029] Specifically, in step S5, a first preliminary intelligence assertion is extracted from the set of preliminary intelligence assertions, and a plurality of domain knowledge graphs are extracted from the trusted domain knowledge base. It should be understood that when the LLM generates a set of preliminary assertions based on preprocessed public data, it is essentially a generalization of the implicit patterns in a large amount of text through a probabilistic model (such as inferring that "a certain enterprise is developing graphene battery technology"). These inferences, although semantically reasonable, lack factual anchors, and therefore there is a fundamental contradiction between the generation mechanism of large language models and the demand for intelligence credibility. Therefore, in the technical solution of the present application, a first preliminary intelligence assertion is further extracted from the set of preliminary intelligence assertions, and a plurality of domain knowledge graphs are extracted from the trusted domain knowledge base.

[0030] Specifically, in step S6, the first preliminary intelligence assertion is checked based on semantic query formula self-attention aggregation analysis between the first preliminary intelligence assertion and the plurality of domain knowledge graphs to determine the confidence of the first preliminary intelligence assertion. It should be understood that by extracting multi-dimensional knowledge graphs (such as the material science application records in the enterprise's historical patents, the research and development team's professional background graph, and the supply chain cooperation partner's technology roadmap) from the trusted domain knowledge base, a structured fact verification network can be constructed. The self-attention aggregation algorithm used in the execution process is not simply keyword matching, but by deep semantic alignment of assertion encoding vectors and knowledge graph vectors (such as identifying the logical association between the specific application scenario of "graphene" in the patent claim and the technical maturity expression in the assertion), the monomer semantic matching degree in high-dimensional space is calculated. The core value of this dynamic verification mechanism is to break through the limitations of traditional rule engines - not only can it detect explicit contradictions (such as the enterprise's official denial of the technology direction), but also can capture potential semantic biases (such as the knowledge graph showing that all related patents of the enterprise are focused on lithium-ion battery improvements, unrelated to the graphene technology path), providing a basis for confidence assessment and assertion correction.

[0031] In one embodiment, as Figure 2As shown, the checking the first preliminary intelligence assertion based on the semantic query self-attention aggregation analysis between the first preliminary intelligence assertion and the plurality of domain knowledge graphs to determine the confidence of the first preliminary intelligence assertion comprises: S61, structurally coding the first preliminary intelligence assertion to obtain a first preliminary intelligence assertion structured coding vector; S62, structurally coding the plurality of domain knowledge graphs to obtain a plurality of domain knowledge graph structured coding vectors; S63, performing semantic query self-attention feature retrieval processing on the first preliminary intelligence assertion structured coding vector and the plurality of domain knowledge graph structured coding vectors to obtain a preliminary intelligence assertion semantic feature query response semantic coding vector; S64, determining the confidence of the first preliminary intelligence assertion based on the preliminary intelligence assertion semantic feature query response semantic coding vector.

[0032] Specifically, in steps S61 and S62, the first preliminary intelligence assertion is structurally coded to obtain a first preliminary intelligence assertion structured coding vector, and the plurality of domain knowledge graphs are structurally coded to obtain a plurality of domain knowledge graph structured coding vectors. It can be understood that when the LLM generates an assertion such as "a company is developing a neuromorphic computing-based image sensor", the text form contains rich combinations of technical terms ("neuromorphic computing" + "image sensor"), but direct fact verification faces the problem of mismatched semantic granularity - structured entity relationships may be stored in the knowledge graph. Therefore, in order to solve the gap between the diversity of natural language expression and machine computable semantics, in the technical solution of the present application, the first preliminary intelligence assertion is further structurally coded to obtain a first preliminary intelligence assertion structured coding vector, and the plurality of domain knowledge graphs are structurally coded to obtain a plurality of domain knowledge graph structured coding vectors. By structurally coding, the assertion text is converted into a vector containing multi-dimensional features such as technical entities, function descriptions, and application scenarios, and the triples in the knowledge graph are coded into vectors containing dimensions such as entity types, relationship weights, and timing information, which essentially builds a bridge across the surface expression of text and the deep association of domain knowledge. This coding process not only unifies the representation form of heterogeneous data (projects free text and graph triples into the same vector space), but more importantly extracts semantic features that are critical to fact verification through feature engineering.

[0033] Those of ordinary skill in the art should know that structured coding belongs to a routine process, but in order to facilitate understanding, specific embodiments are provided as illustrations. First, a deep analysis is made on the first preliminary intelligence assertion automatically generated by a large language model. The first preliminary intelligence assertion is a natural language text, which contains complex information related to specific technical topics, behavioral subjects, functional scenarios, and time elements. For example, the assertion "a company is developing an image sensor based on neuromorphic computing" implicitly integrates enterprise entities, core technical terms, and application intentions, as well as optional time sequence descriptions within the text. To achieve structured coding, key entities (such as "neuromorphic computing", "image sensor", "a company") in the assertion, attributes (such as "development", "based on"), time sequence relationships, and context environment must be fully analyzed through entity recognition, syntax dependency analysis, semantic role labeling, and context disambiguation. Subsequently, these elements are mapped into a multi-dimensional attribute vector of "entity-attribute-relation-scenario-time" according to the set knowledge representation ontology. The vector generally uses distributed representation (embedding), which projects text units into a high-dimensional continuous space through an embedding layer. Deep semantic embedding models such as BERT and ERNIE can be used to bring semantically highly related words and structures closer in the vector space. At the same time, position encoding is used to map time sequence relationships into feature dimensions, thereby forming a high-dimensional structure vector that comprehensively considers technical semantics, entity relationships, time attributes, and application scenarios, i.e., the first preliminary intelligence assertion structured coding vector.

[0034] Correspondingly, the structured coding process of multiple domain knowledge graphs is based on standard triples (usually in the form of "entity A-relation-entity B", such as "company A-research-technology B") stored in authoritative knowledge graphs. Through ontology mapping, semantic normalization, and importance weight assignment, each triple and its attached attributes (such as technology maturity, legal status, citation count, historical timeline, and domain classification) are converted into structured vector units.

[0035] Specifically, in step S63, the first preliminary intelligence assertion structured coding vector and the multiple domain knowledge graph structured coding vectors are subjected to semantic query-based self-attention feature retrieval processing to obtain a preliminary intelligence assertion semantic feature query response semantic coding vector. It can be understood that when the assertion generated by the LLM (such as "a new quantum encryption chip uses topological insulator materials") needs to be verified against the patent data (such as the "quantum bit structure based on Majorana fermions" technical solution disclosed by a company) stored in the knowledge graph, simple cosine similarity calculation or keyword matching cannot capture the deep association between technical terms - although "topological insulator" and "Majorana fermion" do not have direct overlap on the surface text, there is a potential technical association in the field of condensed matter physics for quantum state manipulation.

[0036] Therefore, in the technical solution of the present application, the first preliminary intelligence assertion structured coding vector and the plurality of field knowledge graph structured coding vectors are further subjected to semantic query-based self-attention feature retrieval processing to obtain a preliminary intelligence assertion semantic feature query response semantic coding vector. Through the semantic query-based self-attention mechanism, the system can project the first preliminary intelligence assertion structured coding vector (containing material properties, technical route, etc. dimensional features) and the field knowledge graph structured coding vector (containing patent technology points, experimental data, etc. structured information) into a unified high-dimensional semantic space, and dynamically identify key feature dimensions (such as the functional relationship between material band structure parameters and qubit stability) using self-attention weights. This process is realized through fine-grained control by the relationship gate agent module - for example, when there are multiple related patents in the knowledge graph, the gating mechanism will strengthen the patent features of the same technical generation (such as third-generation quantum computing schemes) as the target assertion, and suppress the interference of outdated technical routes, in order to identify assertions that are superficially reasonable but have hidden technical contradictions.

[0037] In one embodiment, as shown in Figure 3 the first preliminary intelligence assertion structured coding vector and the plurality of field knowledge graph structured coding vectors are subjected to semantic query-based self-attention feature retrieval processing to obtain a preliminary intelligence assertion semantic feature query response semantic coding vector, comprising: S631, after feature enhancement of the first preliminary intelligence assertion structured coding vector, respectively perform single semantic query coding on each of the plurality of field knowledge graph structured coding vectors to obtain a set of first preliminary intelligence assertion single semantic query score coding vectors; S632, respectively calculate the single semantic weight of each first preliminary intelligence assertion single semantic query score coding vector in the set of first preliminary intelligence assertion single semantic query score coding vectors to obtain a set of first preliminary intelligence assertion single query semantic self-attention weights; S633, based on the set of first preliminary intelligence assertion single query semantic self-attention weights, aggregate the set of first preliminary intelligence assertion single semantic query score coding vectors to obtain the preliminary intelligence assertion semantic feature query response semantic coding vector.

[0038] In one embodiment, as shown in Figure 4As shown, after feature enhancement of the first preliminary information assertion structured encoding vector, each domain knowledge graph structured encoding vector in the plurality of domain knowledge graph structured encoding vectors is respectively subjected to single semantic query encoding to obtain a set of first preliminary information assertion single semantic query score encoding vectors, including: S6311, the first preliminary information assertion structured encoding vector is subjected to feature enhancement based on deconvolution coding to obtain a strengthened first preliminary information assertion structured encoding vector, and the strengthened first preliminary information assertion structured encoding vector has the same feature dimension as each domain knowledge graph structured encoding vector in the plurality of domain knowledge graph structured encoding vectors. Specifically, the process can be represented by the formula:

[0039]

[0040] Wherein, V1 is the first preliminary information assertion structured encoding vector, f deconv is deconvolution coding, W deconv is a deconvolution weight matrix, and ‖·‖ is a vector norm. ′ V1′ is the strengthened first preliminary information assertion structured encoding vector.

[0041] S6312, the strengthened first preliminary information assertion structured encoding vector and each domain knowledge graph structured encoding vector in the plurality of domain knowledge graph structured encoding vectors are respectively subjected to single semantic query encoding to obtain a set of first preliminary information assertion single semantic query score encoding vectors. Specifically, the process can be represented by the formula:

[0042] V2={V 21 ,V 22 ,...,V 2i ,...,V 2n}

[0043] R i =tanh{W i [V1′;V 2i ]+b i}

[0044] Wherein, V2 is a plurality of domain knowledge graph structured encoding vectors, V 21 , V 22 , V 2i , and V 2n are the 1st, 2nd, i-th, and n-th domain knowledge graph structured encoding vectors in the plurality of domain knowledge graph structured encoding vectors, respectively, W i and b i are a trainable weight matrix and a trainable bias vector, respectively, [·;·] is a vector concatenation operation, tanh(·) is a tanh function, and Ri encode a first preliminary intelligence assertion monomer semantic query score encoding vector.

[0045] Specifically, through the way of deconvolution coding, the preliminary intelligence assertion structured encoding vector is complemented in feature dimension and expanded in scale in the high-dimensional semantic space, not only increasing its ability to express complex semantic elements such as internal technical relationship, context dependence and time-space scene, but also fully considering the differences in information granularity and representation density of knowledge from different structured sources. In this way, the strengthened preliminary intelligence assertion structured encoding vector can realize complete feature dimension alignment with the structured encoding vectors of multiple domain knowledge graphs in the vector space. The consistency in dimension lays a foundation for the subsequent semantic retrieval process, so that the information interaction between the assertions generated from natural language and the structured knowledge can be conducted fairly and justly, and the incomplete expression or misjudgment phenomenon caused by feature scale mismatch is eliminated.

[0046] After the feature scale alignment, the strengthened first preliminary intelligence assertion structured encoding vector and the structured encoding vectors of multiple domain knowledge graphs enter the monomer semantic query encoding link respectively. At this stage, the system performs deep semantic matching modeling on the strengthened first preliminary intelligence assertion structured encoding vector and each domain knowledge graph structured encoding vector, and deep correspondence between the semantic information such as technical entities, domain relationships and time sequence dependence contained in the assertion and the specific technical route, patent theme and fact node in the knowledge graph.

[0047] In one embodiment, as shown in Figure 5 S6321, based on the feature set distribution characteristics of the set of first preliminary intelligence assertion monomer semantic query score encoding vectors, determine the monomer semantic matching degree of each first preliminary intelligence assertion monomer semantic query score encoding vector in the set of first preliminary intelligence assertion monomer semantic query score encoding vectors to obtain a set of first preliminary intelligence assertion monomer semantic matching degrees; S6322, input the set of first preliminary intelligence assertion monomer semantic matching degrees into the relationship gate agent module to obtain the set of first preliminary intelligence assertion monomer query semantic self-attention weights.

[0048] It should be understood that after the strengthened first preliminary intelligence assertion structured encoding vector and the multiple domain knowledge graph structured encoding vectors, the first preliminary intelligence assertion monomer semantic query score encoding vector is obtained. This set not only contains the fine-grained matching relationship between the assertion and each knowledge graph in the semantic space, but also is the most critical information carrier to drive the subsequent global aggregation mechanism.

[0049] In analyzing the feature set of the set of monomer semantic query score encoding vectors based on the feature set of the set of monomer semantic query score encoding vectors of the first preliminary intelligence assertion, through global statistics and relative position judgment, the relationship between a single score vector and a knowledge node is no longer considered in isolation, but the entire set of monomer semantic query score encoding vectors is taken as the context to analyze the prominence of each vector in the whole, the relative dispersion, and the distance from the set center trend. Such self-distribution analysis can reveal which score encoding vectors belong to high relevance and key results in all knowledge graph matching, and which may be at the average or edge (noise) level. The self-distribution feature acts on the matching degree calculation stage, so that the system does not treat all matches equally, but dynamically allocates attention to semantic alignment items with high scores or significant heterogeneity indicators. Finally, after inherent alignment calibration, each query score encoding vector obtains a quantitative monomer semantic matching degree, which not only reflects the semantic coupling strength with the corresponding knowledge graph vector, but also integrates dynamic weight adjustment in the overall feature distribution background, making the analysis more discriminative and contextually intelligent.

[0050] Here, in the direct feature splicing of the reinforced first preliminary intelligence assertion structured encoding vector and the domain knowledge graph structured encoding vector to perform monomer semantic query encoding, it is expected that the accuracy of monomer semantic query encoding can be improved through optimization of dynamic feature representation from the feature space to the semantic query encoding space and inherent alignment calibration between the feature space and the semantic query encoding space.

[0051] In one embodiment, based on the feature set of the set of monomer semantic query score encoding vectors of the first preliminary intelligence assertion, the monomer semantic matching degree of each monomer semantic query score encoding vector of the set of monomer semantic query score encoding vectors of the first preliminary intelligence assertion is determined to obtain a set of monomer semantic matching degrees of the first preliminary intelligence assertion, including: performing inherent alignment calibration on each monomer semantic query score encoding vector in the set of monomer semantic query score encoding vectors of the first preliminary intelligence assertion to obtain a set of calibrated monomer semantic query score encoding vectors of the first preliminary intelligence assertion.

[0052] Specifically, if the reinforced first preliminary intelligence assertion structured encoding vector V1 ′ and the corresponding domain knowledge graph structured encoding vector V 2i is denoted as V 3i , that is, V 3i = [V1 ′ ; V 2iFirst, the first preliminary intelligence assertion query potential vector under the mapping representation is calculated, which can be expressed by the formula:

[0053] V 4i =W i V 3i -V 3i

[0054] Among them, V 4i To establish the initial intelligence assertion query potential vector, W 4i This is the projection weight matrix.

[0055] That is, the first preliminary intelligence assertion query interaction encoding vector ontology is used as the dynamic eigenvector representation to obtain the projection representation of the eigenstate under the spatial mapping process.

[0056] Then, the inherent alignment between the feature space and the semantic query encoding space is calibrated:

[0057] R′ i =R i +gS i V 4i

[0058] Among them, S i To make the first preliminary intelligence assertion, query the covariance matrix, that is, let V 4i =S i R i To obtain, R i ′ To encode the first preliminary intelligence assertion single-unit semantic query score vector after calibration, the coupling constant g is calculated, for example, in the same way as during deconvolution enhancement, to maintain symmetry, i.e.:

[0059]

[0060] Therefore, in the first preliminary intelligence assertion query potential vector V 4i In the case of inherent alignment generators, the accuracy of the representation of monolithic semantic query encoding is improved by using covariant paths under canonical symmetry conditions as a fixed guarantee of the mapping canonicity.

[0061] Specifically, based on the distribution characteristics of the set of calibrated first preliminary intelligence assertion single semantic query score encoding vectors, the single semantic matching degree of each calibrated first preliminary intelligence assertion single semantic query score encoding vector is calculated to obtain the set of first preliminary intelligence assertion single semantic matching degrees. Specifically, this process can be expressed by the formula:

[0062]

[0063] Where n is R′ i The number of R′k exp is the natural exponential function with base e, softmax(·) is the softmax function, a i is the first preliminary intelligence assertion monomer semantic matching degree.

[0064] When the set of first preliminary intelligence assertion monomer semantic matching degrees is generated, the input relationship gate agent module becomes a key step to further improve the intelligence of the aggregation process. This module models the internal relationship and external interference of semantic association, dynamically controls the influence of each monomer matching degree in synthesizing the final aggregation weight through the gate mechanism, and realizes the fine management of information flow intensity and signal-to-noise ratio. For example, when the structured encoding vectors of the domain knowledge graph belong to technical homology, time sequence evolution or upstream and downstream logic, the gate mechanism will retain the relevant matching degrees; on the contrary, for the structured encoding vectors of the domain knowledge graph that are historical, fact conflict or semantic decoupling, the gate module directly shields the influence. This process not only ensures that the final self-attention weight can be concentrated on the most valuable feature matching, but also effectively avoids the penetration of edge noise, pseudo correlation and even interference information into the first preliminary intelligence assertion monomer semantic query score encoding vector.

[0065] Specifically, the calculation process of inputting the set of first preliminary intelligence assertion monomer semantic matching degrees into the relationship gate agent module to obtain the set of first preliminary intelligence assertion monomer query semantic self-attention weights can be represented by the formula as follows:

[0066]

[0067] wherein, mask is a mask function, τ is a preset threshold, w i is the first preliminary intelligence assertion monomer query semantic self-attention weight.

[0068] Finally, based on the set of first preliminary intelligence assertion monomer query semantic self-attention weights, the set of first preliminary intelligence assertion monomer semantic query score encoding vectors is aggregated to obtain the preliminary intelligence assertion semantic feature query response semantic encoding vector. The preliminary intelligence assertion semantic feature query response semantic encoding vector not only carries the key information of each domain knowledge graph related to the assertion, but also presents the deep synthesis expression between heterogeneous knowledge in the form of information embedding in high-dimensional space. Its embedded weight distribution mechanism significantly enhances the integration efficiency of multi-source heterogeneous information, maximizes the preservation of semantic features with fact support value, and suppresses misleading pseudo correlation factors.

[0069] Specifically, the calculation process of aggregating the set of the first preliminary intelligence assertion monomer semantic query score encoding vectors based on the set of the first preliminary intelligence assertion monomer query semantic self-attention weight to obtain the preliminary intelligence assertion semantic feature query response semantic encoding vector can be represented by a formula as follows:

[0070] v p =∑ i w i R i

[0071] wherein, v p is the preliminary intelligence assertion semantic feature query response semantic encoding vector.

[0072] Specifically, in step S64, the confidence of the first preliminary intelligence assertion is determined based on the preliminary intelligence assertion semantic feature query response semantic encoding vector. It should be understood that when the preliminary intelligence assertion semantic feature query response semantic encoding vector generated by the semantic query formula self-attention aggregation carries multi-dimensional semantic association information, it is essentially a complex pattern expression in a high-dimensional space, which cannot be directly understood and utilized by the downstream assertion revision process. Therefore, the confidence of the first preliminary intelligence assertion is determined again based on the preliminary intelligence assertion semantic feature query response semantic encoding vector. In particular, in one embodiment, the confidence of the first preliminary intelligence assertion is determined based on the preliminary intelligence assertion semantic feature query response semantic encoding vector, including: passing the preliminary intelligence assertion semantic feature query response semantic encoding vector through a classifier-based confidence analyzer to obtain an analysis result, wherein the analysis result is used to represent the confidence level of the first preliminary intelligence assertion.

[0073] In one specific embodiment, the preliminary intelligence assertion semantic feature query response semantic encoding vector is input into a trained multi-class classifier (such as a softmax classifier), and the mapping relationship between the input vector and the set of predefined confidence labels is modeled through the classifier. Taking the softmax classifier as an example, the core working mechanism is to activate the output layer nodes of the neural network for each input vector, normalize it into a multi-class probability distribution through the softmax function, and then output the corresponding confidence level prediction result.

[0074] In summary, the large language model driven open source intelligence analysis method provided in the application collects public information data, aggregates multi-source heterogeneous raw data, eliminates semantic ambiguity and noise through data cleaning and preprocessing, relies on authoritative databases, patent literature and other reliable sources to build a structured domain knowledge graph as a fact benchmark. The preprocessed data is input into an information encoder based on a large language model to generate preliminary intelligence assertions. Then, the preliminary intelligence assertions and the associated knowledge graph are extracted, and the consistency and factuality of the preliminary intelligence assertions are dynamically checked based on a semantic query self-attention aggregation mechanism to determine the confidence of the preliminary intelligence assertions. In this way, the confidence of the intelligence assertions can be effectively analyzed and judged, which helps to trigger a multi-path correction strategy for low-confidence assertions, forming a closed-loop process of "generation-verification-iterative optimization" to significantly reduce the illusion penetration risk.

[0075] The application also provides a large language model driven open source intelligence analysis system, as shown in Figure 6 The large language model driven open source intelligence analysis system 600 includes a public information collection module 610 for collecting public information data, a public information processing module 620 for cleaning and preprocessing the public information data to obtain preprocessed public information data, a domain knowledge graph construction module 630 for extracting domain knowledge from highly reliable sources to construct a domain knowledge graph and storing the domain knowledge graph in a reliable domain knowledge base, a preliminary intelligence assertion generation module 640 for inputting the preprocessed public information data into an information encoder based on a large language model to obtain a set of preliminary intelligence assertions, a data extraction module 650 for extracting a first preliminary intelligence assertion from the set of preliminary intelligence assertions and extracting a plurality of domain knowledge graphs from the reliable domain knowledge base, and a preliminary intelligence assertion checking module 660 for checking the first preliminary intelligence assertion based on semantic query self-attention aggregation analysis between the first preliminary intelligence assertion and the plurality of domain knowledge graphs to determine the confidence of the first preliminary intelligence assertion.

[0076] In an embodiment, the preliminary intelligence assertion verification module comprises: a preliminary intelligence assertion structured coding unit configured to perform structured coding on the first preliminary intelligence assertion to obtain a first preliminary intelligence assertion structured coding vector; a domain knowledge graph structured coding unit configured to perform structured coding on the plurality of domain knowledge graphs to obtain a plurality of domain knowledge graph structured coding vectors; a feature retrieval processing unit configured to perform semantic query formula self-attention feature retrieval processing on the first preliminary intelligence assertion structured coding vector and the plurality of domain knowledge graph structured coding vectors to obtain a preliminary intelligence assertion semantic feature query response semantic coding vector; and a confidence determination unit configured to determine the confidence of the first preliminary intelligence assertion based on the preliminary intelligence assertion semantic feature query response semantic coding vector.

[0077] In an embodiment, the feature retrieval processing unit is configured to: perform feature enhancement on the first preliminary intelligence assertion structured coding vector, and then perform single-body semantic query coding on the first preliminary intelligence assertion structured coding vector and each of the plurality of domain knowledge graph structured coding vectors to obtain a set of first preliminary intelligence assertion single-body semantic query score coding vectors; calculate a single-body semantic weight of each of the set of first preliminary intelligence assertion single-body semantic query score coding vectors to obtain a set of first preliminary intelligence assertion single-body query semantic self-attention weights; and aggregate the set of first preliminary intelligence assertion single-body semantic query score coding vectors based on the set of first preliminary intelligence assertion single-body query semantic self-attention weights to obtain the preliminary intelligence assertion semantic feature query response semantic coding vector.

[0078] Here, those skilled in the art can understand that the specific operations of each module and unit in the above-described large language model driven open source intelligence analysis system have been described in detail above with reference to the description of the large language model driven open source intelligence analysis method Figures 1 to 4 , and therefore, the repeated description thereof will be omitted.

[0079] The basic principles of the present application are described above in combination with specific embodiments, but it should be noted that the advantages, advantages, effects, etc. mentioned in the present application are only examples and are not limiting, and these advantages, advantages, effects, etc. cannot be considered as the must-have of each embodiment of the present application. In addition, the specific details of the above-described embodiments are only for the purpose of example and for the purpose of understanding, and the above-described details do not limit the present application to the must-use of the above-described specific details to realize.

[0080] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in a certain embodiment can be referred to the relevant description of other embodiments. In the several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. For example, the embodiments of the device described above are merely schematic, and the division of the modules is merely a logical function division, and there can be another division manner in actual implementation. The modules illustrated as separated components can or can not be physically separated, and the components illustrated as modules can or can not be physical units, i.e., they can be located in one place or distributed on a plurality of network units. In some embodiments, some or all of the modules can be selected according to actual needs to achieve the purposes of the embodiments.

[0081] In addition, each functional module in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional modules.

[0082] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0083] Furthermore, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The plurality of units stated in the device claim can also be implemented by one unit by means of software or hardware.

[0084] Finally, it should be noted that the above description is given for the purpose of illustration and description. In addition, the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. An open-source intelligence analysis method driven by a large language model, characterized in that, include: Collect publicly available information data, including social media text, news web pages, PDF reports, and tabular data; The publicly available information data is cleaned and preprocessed to obtain preprocessed publicly available information data; Domain knowledge is extracted from highly trusted sources to construct a domain knowledge graph, and the domain knowledge graph is stored in a trusted domain knowledge base; The preprocessed public information data is input into an information encoder based on a large language model to obtain a set of preliminary intelligence assertions; Extract a first preliminary intelligence assertion from the set of preliminary intelligence assertions, and extract multiple domain knowledge graphs from the trusted domain knowledge base; The confidence level of the first preliminary intelligence assertion is determined by verifying the first preliminary intelligence assertion based on semantic query-based self-attention aggregation analysis between the first preliminary intelligence assertion and the multiple domain knowledge graphs, including: The first preliminary intelligence assertion is structured and encoded to obtain the first preliminary intelligence assertion structured encoding vector; The knowledge graphs of the multiple domains are structured and encoded to obtain structured encoding vectors for the knowledge graphs of the multiple domains. The first preliminary intelligence assertion structured encoding vector and the multiple domain knowledge graph structured encoding vectors are subjected to semantic query-based self-attention feature retrieval processing to obtain the preliminary intelligence assertion semantic feature query response semantic encoding vector. Based on the semantic feature query response semantic encoding vector of the preliminary intelligence assertion, the confidence level of the first preliminary intelligence assertion is determined; The first preliminary intelligence assertion structured encoding vector and the multiple domain knowledge graph structured encoding vectors are subjected to semantic query-based self-attention feature retrieval processing to obtain the preliminary intelligence assertion semantic feature query response semantic encoding vector, including: After feature enhancement, the first preliminary intelligence assertion structured encoding vector is combined with each of the domain knowledge graph structured encoding vectors in the multiple domain knowledge graph structured encoding vectors to perform individual semantic query encoding to obtain a set of first preliminary intelligence assertion individual semantic query score encoding vectors. Calculate the individual semantic weight of each first preliminary intelligence assertion individual semantic query score encoding vector in the set of first preliminary intelligence assertion individual semantic query score encoding vectors to obtain the set of first preliminary intelligence assertion individual query semantic self-attention weights. The set of semantic self-attention weights for the first preliminary intelligence assertion single query is aggregated to obtain the semantic encoding vector of the first preliminary intelligence assertion semantic feature query response.

2. The open-source intelligence analysis method driven by a large language model according to claim 1, characterized in that, After feature enhancement of the first preliminary intelligence assertion structured encoding vector, it is then used to perform individual semantic query encoding with each of the multiple domain knowledge graph structured encoding vectors to obtain a set of first preliminary intelligence assertion individual semantic query score encoding vectors, including: The first preliminary intelligence assertion structured encoding vector is enhanced by deconvolutional coding to obtain an enhanced first preliminary intelligence assertion structured encoding vector. The enhanced first preliminary intelligence assertion structured encoding vector has the same feature scale as the structured encoding vectors of each of the multiple domain knowledge graphs. The enhanced first preliminary intelligence assertion structured encoding vector and each of the domain knowledge graph structured encoding vectors are respectively subjected to individual semantic query encoding to obtain a set of first preliminary intelligence assertion individual semantic query score encoding vectors.

3. The open-source intelligence analysis method driven by a large language model according to claim 2, characterized in that, Calculate the individual semantic weights of each first preliminary intelligence assertion individual semantic query score encoding vector in the set of first preliminary intelligence assertion individual semantic query score encoding vectors to obtain the set of first preliminary intelligence assertion individual query semantic self-attention weights, including: Based on the feature set self-distribution characteristics of the set of first preliminary intelligence assertion single semantic query score encoding vectors, the single semantic matching degree of each first preliminary intelligence assertion single semantic query score encoding vector in the set of first preliminary intelligence assertion single semantic query score encoding vectors is determined to obtain the set of first preliminary intelligence assertion single semantic matching degrees. The set of semantic matching degrees of the first preliminary intelligence assertion entity is input into the relation gating proxy module to obtain the set of self-attention weights of the query semantics of the first preliminary intelligence assertion entity.

4. The open-source intelligence analysis method driven by a large language model according to claim 3, characterized in that, Based on the self-distribution characteristics of the feature set of the set of first preliminary intelligence assertion single semantic query score encoding vectors, the single semantic matching degree of each first preliminary intelligence assertion single semantic query score encoding vector in the set of first preliminary intelligence assertion single semantic query score encoding vectors is determined to obtain the set of first preliminary intelligence assertion single semantic matching degrees, including: The inherent alignment of each first preliminary intelligence assertion single semantic query score encoding vector in the set of first preliminary intelligence assertion single semantic query score encoding vectors is calibrated to obtain the set of calibrated first preliminary intelligence assertion single semantic query score encoding vectors. Based on the distribution characteristics of the set of calibrated first preliminary intelligence assertion single semantic query score encoding vectors, the single semantic matching degree of each calibrated first preliminary intelligence assertion single semantic query score encoding vector is calculated to obtain the set of first preliminary intelligence assertion single semantic matching degrees.

5. The open-source intelligence analysis method driven by a large language model according to claim 4, characterized in that, Determining the confidence level of the first preliminary intelligence assertion based on the semantic encoding vector of the query response to the semantic features of the preliminary intelligence assertion includes: passing the semantic encoding vector of the query response to the semantic features of the preliminary intelligence assertion through a classifier-based confidence analyzer to obtain an analysis result, wherein the analysis result is used to represent the confidence level of the first preliminary intelligence assertion.

6. An open-source intelligence analysis system driven by a large language model, characterized in that, include: The public information collection module is used to collect publicly available information data, including social media text, news web pages, PDF reports, and tabular data. The public information processing module is used to perform data cleaning and preprocessing on the public information data to obtain preprocessed public information data. A domain knowledge graph construction module is used to extract domain knowledge from highly trusted sources to construct a domain knowledge graph, and store the domain knowledge graph in a trusted domain knowledge base; The preliminary intelligence assertion generation module is used to input the preprocessed public information data into the information encoder based on a large language model to obtain a set of preliminary intelligence assertions; The data extraction module is used to extract a first preliminary intelligence assertion from the set of preliminary intelligence assertions and to extract multiple domain knowledge graphs from the trusted domain knowledge base. The preliminary intelligence assertion verification module is used to verify the first preliminary intelligence assertion based on semantic query-based self-attention aggregation analysis between the first preliminary intelligence assertion and the multiple domain knowledge graphs to determine the confidence level of the first preliminary intelligence assertion, including: The first preliminary intelligence assertion is structured and encoded to obtain the first preliminary intelligence assertion structured encoding vector; The knowledge graphs of the multiple domains are structured and encoded to obtain structured encoding vectors for the knowledge graphs of the multiple domains. The first preliminary intelligence assertion structured encoding vector and the multiple domain knowledge graph structured encoding vectors are subjected to semantic query-based self-attention feature retrieval processing to obtain the preliminary intelligence assertion semantic feature query response semantic encoding vector. Based on the semantic feature query response semantic encoding vector of the preliminary intelligence assertion, the confidence level of the first preliminary intelligence assertion is determined; The first preliminary intelligence assertion structured encoding vector and the multiple domain knowledge graph structured encoding vectors are subjected to semantic query-based self-attention feature retrieval processing to obtain the preliminary intelligence assertion semantic feature query response semantic encoding vector, including: After feature enhancement, the first preliminary intelligence assertion structured encoding vector is combined with each of the domain knowledge graph structured encoding vectors in the multiple domain knowledge graph structured encoding vectors to perform individual semantic query encoding to obtain a set of first preliminary intelligence assertion individual semantic query score encoding vectors. Calculate the individual semantic weight of each first preliminary intelligence assertion individual semantic query score encoding vector in the set of first preliminary intelligence assertion individual semantic query score encoding vectors to obtain the set of first preliminary intelligence assertion individual query semantic self-attention weights. The set of semantic self-attention weights for the first preliminary intelligence assertion single query is aggregated to obtain the semantic encoding vector of the first preliminary intelligence assertion semantic feature query response.

7. The open-source intelligence analysis system driven by a large language model according to claim 6, characterized in that, The preliminary intelligence assertion verification module includes: A preliminary intelligence assertion structured coding unit is used to perform structured coding on the first preliminary intelligence assertion to obtain a first preliminary intelligence assertion structured coding vector. A domain knowledge graph structured coding unit is used to perform structured coding on the multiple domain knowledge graphs to obtain multiple domain knowledge graph structured coding vectors; The feature retrieval processing unit is used to perform semantic query-based self-attention feature retrieval processing on the first preliminary intelligence assertion structured encoding vector and the multiple domain knowledge graph structured encoding vectors to obtain the preliminary intelligence assertion semantic feature query response semantic encoding vector. The confidence determination unit is used to query the response semantic encoding vector based on the semantic features of the preliminary intelligence assertion to determine the confidence of the first preliminary intelligence assertion.

8. The open-source intelligence analysis system driven by a large language model according to claim 7, characterized in that, The feature retrieval processing unit is used for: After feature enhancement, the first preliminary intelligence assertion structured encoding vector is combined with each of the domain knowledge graph structured encoding vectors in the multiple domain knowledge graph structured encoding vectors to perform individual semantic query encoding to obtain a set of first preliminary intelligence assertion individual semantic query score encoding vectors. Calculate the individual semantic weight of each first preliminary intelligence assertion individual semantic query score encoding vector in the set of first preliminary intelligence assertion individual semantic query score encoding vectors to obtain the set of first preliminary intelligence assertion individual query semantic self-attention weights. The set of semantic self-attention weights for the first preliminary intelligence assertion single query is aggregated to obtain the semantic encoding vector of the first preliminary intelligence assertion semantic feature query response.

Citation Information

Patent Citations

  • Reliability evaluation method and system for open source threat intelligence

    CN117081810A

  • Semantic analysis driven model-based open source intelligence relevance identification method and system

    CN119149737A