Open source intelligence public opinion hotspot real-time monitoring and report generation method and system
By constructing an enhanced search model and vector database, combined with neural searchers and generators, the limitations of large-scale pre-trained models in accessing and understanding knowledge capabilities and the dependence correlation of traditional search methods are solved, and the function of automatically generating public opinion reports is realized, reducing costs and improving efficiency.
Patent Information
- Application Number
- CN202510039974.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively deal with and utilize the limitations of large-scale pre-trained models in accessing and understanding knowledge capabilities, as well as the problems of information overload and insufficient knowledge quality caused by dependency of traditional search methods.
Build an enhanced search model, combining pre-trained neural searchers and generators, and form an end-to-end probabilistic model through end-to-end fine-tuning. The model uses the information output by the neural searcher as additional text information and is fused into the generator in a marginal way to generate the final sequence. At the same time, data processing and metadata extraction are carried out through vector databases and knowledge graphs to quickly generate public opinion reports.
It realizes automatic generation of public opinion reports based on the report outline entered by users, which reduces labor and time costs, improves information processing efficiency and knowledge quality, helps enterprises and governments better understand social situations and improves decision-making efficiency.
Smart Images

Figure CN120067282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing, and particularly relates to a method and system for real-time monitoring and report generation of open-source intelligence public opinion hotspots. Background Art
[0002] With the rapid development of the Internet and social media, a large amount of content is generated on the Internet every day. These contents provide a large amount of resources for open-source public opinion intelligence, but at the same time, they also bring the problem of serious information overload.
[0003] Currently, large-scale pre-trained models represented by Bert and gpt store a lot of factual knowledge in their parameters, and their ability to access and accurately understand and apply knowledge is still limited. Although traditional retrieval can supplement knowledge to a certain extent, it is highly dependent on the relevance between the retrieved data and the content to be retrieved. First, it is similar to FAQ, asking and answering questions one by one, and one needs to combine all the answer contents by oneself; second, for a large number of pre-trained knowledge bases in the early stage, professional engineers are required to generate individual knowledge entries, including answers and multiple similar questions. At the current content generation speed, the workload is extremely large. If the retrieved database has noise or insufficient viewpoints, the retrieved results will not conform to the context or deviate from the facts, and insufficient knowledge injection quality will make it difficult to complete the task. In addition, existing real-time monitoring systems for public opinion hotspots only collect and analyze event information, and are less applied in further public opinion analysis and report generation. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and system for real-time monitoring and report generation of open-source intelligence public opinion hotspots, which can quickly generate public opinion reports by performing enhanced retrieval preprocessing on public opinion data.
[0005] In a first aspect, the present invention provides a method for real-time monitoring and report generation of open-source intelligence public opinion hotspots, including:
[0006] Model construction process: constructing an enhanced retrieval model, which is formed by combining a pre-trained neural retriever p η (z|x) with parameters η and a pre-trained generator p θ (y i |x,z,y 1:i-1 ) and performing end-to-end fine-tuning; the neural retriever p η (z|x) includes a query encoder and a document indexing module, and is used to return the top-k distribution of paragraph z given query x; the generator p θ (y i |x,z,y 1:i-1) It is a seq2seq model, which is used to return the current token generated by θ according to the context Y of the previous i-1 tokens, the original input, and the retrieved passage z given the query x; finally, the information output by the neural retriever is regarded as additional text information and integrated into the generator in a marginalized manner to generate the final sequence; 1:i-1 and generate the current token; finally, the information output by the neural retriever is regarded as additional text information and integrated into the generator in a marginalized manner to generate the final sequence;
[0007] Vector database construction process: collect open-source intelligence, then perform data cleaning, data processing, and metadata extraction, and then encode it through the query encoder in the neural retriever to form a vector database;
[0008] Report generation process: obtain the report outline, automatically determine whether there is enough data in the vector database to support report generation according to the keywords and parameters in the outline. If the data is sufficient, the keywords and parameters are used as the query x in the order of the sub-points above and below, and input into the enhanced retrieval model. Finally, the report is generated according to the outline and all output information.
[0009] Further, in the report generation process, if it is determined that the data in the vector database is not enough to support the generation of the current outline report, open-source intelligence is automatically collected according to the keywords and parameters, then data cleaning, data processing, and metadata extraction are performed, and then encoded through the query encoder in the neural retriever, so as to be able to quickly respond to queries in subsequent retrieval processes.
[0010] Further, in the report generation process, obtaining the report outline is directly input by the user or guiding the user to generate it;
[0011] Guiding the user to generate specifically includes: receiving the text directly input by the user, collecting the theme, sub-field, analysis dimension, and attention indicators, and generating and displaying input suggestions based on historical data and user behavior patterns; for widely or ambiguously input text by the user, pop up auxiliary text to prompt the user how to refine the input content, or provide options to refine the collection scope; after the user completes the input, the information provided by the user will be analyzed using a large language model to generate a draft outline, receive the user's editing on the draft outline, and generate a report outline according to the confirmation operation.
[0012] Further, in the process of integrating the information output by the neural retriever into the generator, a sequence model and a token model are adopted. The sequence model uses the same document to generate all sequences, and the token model uses all retrieved documents to generate sequences. The calculation method is as follows:
[0013]
[0014] Among them, N represents the total length of the target sequence y, i represents the i-th token currently generated in the target sequence y, and z ∈ top-k(p(·|x)) represents the k most relevant passages z selected from the retriever, according to the relevance score p(·|x) for the given query x;
[0015] Neural retriever p η (z|x), which encodes the input query statement x to obtain the encoded vector q(x). Additionally, the maximum inner product of the vector d(z) in the vector database and q(x) is used to search for the top-k relevant documents, and the output is used as p η (z|x);
[0016] d(z) = LLM d (z), q(x) = LLM d (x)
[0017] Among them, LLM d represents the large language model.
[0018] Furthermore, during the construction of the vector database, metadata extraction includes file name, time, chapter title, and picture alt information. By turning entities into nodes and relationships into relations, knowledge graphs are used for retrieval.
[0019] In a second aspect, the present invention provides an open-source intelligence public opinion hot spot real-time monitoring and report generation system, including:
[0020] A model construction module for constructing an enhanced retrieval model, which is formed by combining a pre-trained neural retriever p η (z|x) with parameters η and a pre-trained generator p θ (y i |x,z,y 1:i-1 ) with parameters θ, and performing end-to-end fine-tuning; the neural retriever p η (z|x) includes a query encoder and a file index module for returning the top-k distribution of the passage z given the query x; the generator p θ (y i |x,z,y 1:i-1 ) is a seq2seq model for returning the current token generated by θ according to the context y 1:i-1 of the previous i - 1 tokens, the original input, and the retrieved passage z; finally, the information output by the neural retriever is regarded as additional text information and fused into the generator in a marginalized manner to generate the final sequence;
[0021] The vector database construction module is used to collect open-source intelligence, then perform data cleaning, data processing, and metadata extraction, and then encode it through the query encoder in the neural retriever to form a vector database;
[0022] The report generation module is used to obtain a report outline. According to the keywords and parameters in the outline, it automatically determines whether there is enough data in the vector database to support report generation. If the data is sufficient, according to the order of the upper and lower sub-points of the keywords and parameters, it is used as a query x to input the enhanced retrieval model in turn. Finally, a report is generated according to the outline and all output information.
[0023] Further, in the report generation module, if it is determined that the data in the vector database is not enough to support the generation of the current outline report, it automatically collects open-source intelligence according to the keywords and parameters, then performs data cleaning, data processing, and metadata extraction, and then encodes it through the query encoder in the neural retriever, so as to be able to quickly respond to queries in subsequent retrieval processes.
[0024] Further, in the report generation module, obtaining the report outline is directly input by the user or guiding the user to generate it;
[0025] Guiding the user to generate specifically includes: receiving the text directly input by the user, collecting the theme, sub-field, analysis dimension, and attention indicators, and generating and displaying input suggestions based on historical data and user behavior patterns; for widely or ambiguously input text by the user, pop up auxiliary text to prompt the user how to refine the input content, or provide options to refine the collection scope; after the user completes the input, the information provided by the user will be analyzed using a large language model to generate a draft outline, receive the user's editing on the draft outline, and generate a report outline according to the confirmation operation.
[0026] Further, in the process of integrating the information output by the neural retriever into the generator, a sequence model and a tagging model are adopted. The sequence model uses the same document to generate all sequences, and the tagging model uses all retrieved documents to generate sequences. The calculation method is as follows:
[0027]
[0028] Among them, N represents the total length of the target sequence y, i represents the i-th token currently generated in the target sequence y, z ∈ top-k(p(·|x)) represents the k most relevant paragraphs z selected from the retriever, based on the relevance score p(·|x) of the given query x;
[0029] Neural retriever p η(z|x), the encoded vector q(x) is obtained by encoding the input query statement x. Additionally, the vector d(z) in the vector database is subjected to a maximum inner product search with q(x) to retrieve the top-k relevant documents, which are output as p η (z|x);
[0030] d(z) = LLM d (z), q(x) = LLM d (x)
[0031] Among them, LLM d represents a large language model.
[0032] Furthermore, during the construction of the vector database, metadata extraction includes file name, time, chapter title, and picture alt information. By turning entities into nodes and relationships into relations, knowledge graphs are used for retrieval.
[0033] The technical solutions provided in the embodiments of the present invention have at least the following technical effects:
[0034] By constructing an enhanced retrieval model, the information output by the neural retriever is regarded as additional text information and integrated into the generator in a marginalized manner to generate the final sequence. Thus, keywords and parameters can be automatically obtained according to the report outline input by the user, and then the enhanced retrieval model is input to automatically generate an opinion report, reducing labor costs and time costs, providing data support for enterprises and governments, helping them understand the social situation more objectively and comprehensively, and improving decision-making efficiency.
[0035] The above description is only an overview of the technical solutions of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented in accordance with the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically illustrates the embodiments of the present invention. Description of the Drawings
[0036] The present invention will be further described below with reference to the accompanying drawings in conjunction with the embodiments.
[0037] Figure 1 It is a schematic diagram of the enhanced retrieval model according to the embodiment of the present invention;
[0038] Figure 2 It is a flowchart of the method in the first embodiment of the present invention;
[0039] Figure 3 It is a schematic structural diagram of the system in the second embodiment of the present invention. Detailed Embodiments
[0040] An embodiment of the present invention provides a method and system for real-time monitoring and report generation of open-source intelligence public opinion hotspots. By performing enhanced retrieval preprocessing on public opinion data, an public opinion report can be quickly generated.
[0041] The overall idea of the technical solution in the embodiment of this application is as follows:
[0042] The pre-trained large model injects some domain-specific knowledge at the large model parameter level, which requires a large amount of computing power and time. The enhanced retrieval model is equivalent to being split. One side is a semantic vector database composed of various data vectorizations, and the other side is a large model. The two cooperate, so that while avoiding spending computing power and time on pre-training, there is also a great improvement in retrieval accuracy. The pre-trained neural retriever is much lighter and is suitable for calculating the similarity between query statements (outline information) and documents in a large amount of data to find relevant documents.
[0043] Through a general fine-tuning method, a non-parametric memory is given to the pre-trained parameter memory generation model, denoted as an enhanced retrieval model. In the established model, the parametric memory is a pre-trained seq2seq converter, and the non-parametric memory is a dense vector index of the public opinion database, which is accessed through a pre-trained neural retriever. These components are combined into an end-to-end trained probability model.
[0044] Combine the pre-trained retriever (query encoder + document index) with the pre-trained seq2seq model (generator) and perform end-to-end fine-tuning. For query x, we use maximum inner product search (MIPS) to find the top-K documents. For the final prediction y, we regard z as a latent variable and marginalize the seq2seq predictions of different documents.
[0045] As Figure 1 shown, it is a schematic diagram of the enhanced retrieval model constructed in the embodiment of the present invention. The query encoder q is used to construct a vector database, and the neural retriever and generator are used for output. Among them, the neural retriever is used to search for relevant paragraphs, and the generator is used to summarize and generate text. The model consists of two parts: (1) a retriever p η (z|x) that returns the top-k truncated distribution of text paragraphs given query x; (2) a generator p θ (y i |x,z,y 1:i-1 ) that returns, given query x, the text generated by θ based on the context y of the previous i-1 tokens 1:i-1, the original input and the retrieved paragraphs generate the current token. Finally, the information retrieved from its output is regarded as additional text information and integrated into the generator in a marginalized manner to generate the final sequence. During the integration process, a sequence model and a token model are used. The sequence model uses the same document to generate all sequences, and the token model uses all the retrieved documents to generate sequences. The calculation method is as follows:
[0046]
[0047] Regarding the retriever p η (z|x), the encoded vector q(x) is obtained by encoding the input query statement x. In addition, the documents in the knowledge base are pre-encoded to obtain the document encoding vector d(z) in advance. Then, the maximum inner product of q(x) and d(z) is used to search for the top-K relevant documents, and the output is used as p η (z|x).
[0048] d(z) = LLM d (z), q(x) = LLM d (x)
[0049] Regarding the generator p θ (y i |x,z,y 1:i-1 ), in this embodiment, ChatGLM3-base is used as the training model, and then the query statement x and the retrieved z are input into it to obtain the generated sequence text.
[0050] Example 1
[0051] This embodiment provides a method for real-time monitoring and report generation of open-source intelligence public opinion hotspots, as Figure 2 shown, including:
[0052] Model construction process: Construct an enhanced retrieval model, which is formed by combining a pre-trained neural retriever p η (z|x) with parameters η pre-trained and a pre-trained generator p θ (y i |x,z,y 1:i-1 ) with parameters θ, and perform end-to-end fine-tuning; the neural retriever p η (z|x) includes a query encoder and a file indexing module, which are used to return the top-k distribution of paragraph z given the query x; the generator p θ (y i |x,z,y 1:i-1 ) is a seq2seq model, which is used to return, given the query x, the context y of the previous i-1 tokens according to θ 1:i-1, the original input and the retrieved paragraphs generate the current token; finally, the information output by the neural retriever is regarded as additional text information and integrated into the generator in a marginalized manner to generate the final sequence;
[0053] In the process of integrating the information output by the neural retriever into the generator, a sequence model and a token model are adopted. The sequence model uses the same document to generate all sequences, and the token model uses all the retrieved documents to generate sequences. The calculation method is as follows:
[0054]
[0055] Among them, N represents the total length of the target sequence y, i represents the i-th token currently generated in the target sequence y, z ∈ top-k(p(·|x)) represents the k most relevant paragraphs z selected from the retriever, according to the relevance score p(·|x) of the given query x;
[0056] Neural retriever p η (z|x), is encoded from the input query statement x to obtain the encoded vector q(x). In addition, the vector d(z) in the vector database is used to perform a maximum inner product search with q(x) to find the top-k relevant documents, and the output is used as p η (z|x);
[0057] d(z) = LLM d (z), q(x) = LLM d (x)
[0058] Among them, LLM d represents the large language model.
[0059] The process of constructing the vector database: collect open source intelligence, then perform data cleaning, data processing, and metadata extraction, and then encode through the query encoder in the neural retriever to form a vector database.
[0060] Use a web crawler program to monitor specified websites, forums, Weibo, WeChat official accounts and other platforms in real time, and capture relevant news, articles, and discussion content. In addition, valuable information can be screened from open source intelligence by setting keywords and semantic analysis. Through web crawler technology and network monitoring technology, open source intelligence on the network is collected in real time, improving the efficiency and accuracy of intelligence collection.
[0061] Use a data cleaning tool (such as Loader) to process the collected data, including removing extra spaces, punctuation marks, special characters, etc. For files in formats such as PDF, Word, Markdown, use the corresponding tools for conversion to meet the requirements of subsequent processing. Use knowledge graph technology to encode entities and relationships for subsequent retrieval.
[0062] During the construction of the vector database, metadata extraction includes file name, time, chapter title, picture alt information. By turning entities into nodes and relationships into relations, use the knowledge graph for retrieval. Retrieval techniques include similarity retrieval, keyword retrieval, and SQL retrieval. Reorder the retrieval results, using planB reordering or combining factors such as relevance and matching degree. Query rotation uses methods such as subqueries and HyDE. Store entities and relationships in the graph database, and achieve efficient association queries and aggregation analysis through the graph database, thereby improving the retrieval efficiency. At the same time, combine technologies such as similarity retrieval, keyword retrieval, and SQL retrieval to provide accurate retrieval results for users.
[0063] Report generation process: Obtain the report outline. According to the keywords and parameters in the outline, automatically determine whether there is enough data in the vector database to support report generation. If the data is sufficient, use the keywords and parameters in the order of the upper and lower points as query x and input them into the enhanced retrieval model in turn. Finally, generate a report based on the outline and all output information.
[0064] During the above-mentioned report generation process, if it is determined that the data in the vector database is not enough to support the generation of the current outline report, automatically collect open-source intelligence according to the keywords and parameters, then perform data cleaning, data processing, and metadata extraction, and then encode through the query encoder in the neural retriever, so as to be able to quickly respond to queries during subsequent retrieval processes. The technologies used in the report generation process mainly include intent-based chunking, vectorization, and generation technologies. Intent-based chunking divides the text into chunks of a fixed size through sentence segmentation, recursive segmentation, and special segmentation. Vectorization converts text, images, audio, and video, etc. into vector matrices through an embedding model. Generation technologies are generated through large language models such as ChatGPT, ChatGLM, and Claude.ai.
[0065] During the report generation process, the report outline is obtained by direct input from the user or by guiding the user to generate it. Check whether the user has provided an outline. If the user has an outline, directly use the outline and the enhanced retrieval results to generate the public opinion report. If there is no outline, guide the user to generate an outline: Receive the text directly input by the user, collect the theme, sub - field, analysis dimension, and attention indicators, and generate and display input suggestions based on historical data and user behavior patterns; for broad or unclear user - input text, pop up auxiliary text to prompt the user on how to refine the input content, or provide options to narrow down the collection scope; after the user finishes inputting, use a large - language model to analyze the information provided by the user, generate a draft outline, receive the user's edits on the draft outline, and generate the report outline based on the confirmation operation.
[0066] For example, the user can operate according to the following steps: a. Input the theme: Ask the user to briefly describe the theme of the public opinion analysis, such as: "Analysis of the real - estate market trend in a certain city in China". b. Input the sub - field: Ask the user to select the sub - field for analysis, such as: "House price trend", "Policy impact", etc. c. Input the analysis dimension: Ask the user to select the analysis dimension, such as: "Time", "Region", "Policy factor", etc. d. Input the attention indicators: Ask the user to list the indicators that need to be concerned during the analysis process, such as: "Average house price", "Sales volume", etc. Then the system generates an outline based on the information input by the user and displays it for the user to confirm. After confirmation, proceed to the next step.
[0067] Based on the outline and the enhanced retrieval results, the system automatically generates a public opinion report. The report mainly includes the following content: a. Public opinion theme: Display the theme input by the user. b. Analysis dimension: List the analysis dimensions selected by the user. c. Attention indicators: Display the indicators concerned by the user. d. RAG level: Rate each analysis dimension and attention indicator according to the risk assessment results. e. Analysis results: Give detailed analysis results based on the ratings of each dimension and indicator. f. Conclusion: Give the final conclusion and suggestions based on the analysis results. During the report generation process, mark the output information of the generator and use it as additional text information to input into the enhanced retrieval model for the next query x, and loop through the steps until the last keyword and parameter are input to obtain the output.
[0068] Present the generated public opinion report to the user for reference and download. If the user has any questions or needs to modify the report, they can return to the previous step for adjustment.
[0069] Embodiments of the present invention can automatically generate an opinion report based on information such as the theme, sub - field, analysis dimension, and attention indicators input by the user, including the opinion theme, analysis dimension, attention indicators, RAG level, analysis results, and conclusions, etc., which reduces the labor cost and time cost and helps to improve the ability of enterprises and governments to grasp the opinion situation. By real - time monitoring and analyzing network opinions, it provides data support for enterprises and governments, helps them understand the social situation more objectively and comprehensively, and improves the decision - making efficiency.
[0070] Based on the same inventive concept, this application also provides a system corresponding to the method in Embodiment 1. For details, see Embodiment 2.
[0071] Embodiment 2
[0072] In this embodiment, an open - source intelligence opinion hot - spot real - time monitoring and report - generating system is provided. As Figure 3 shown, it includes:
[0073] A model construction module for constructing an enhanced retrieval model. The enhanced retrieval model is formed by combining a pre - trained neural retriever p η (z|x) with parameters η and a pre - trained generator p θ (y i |x,z,y 1:i-1 ) and performing end - to - end fine - tuning. The neural retriever p η (z|x) includes a query encoder and a document indexing module for returning the top - k distribution of paragraph z given the query x. The generator p θ (y i |x,z,y 1:i-1 ) is a seq2seq model for returning the current token generated by θ based on the context y 1:i-1 of the previous i - 1 tokens, the original input, and the retrieved paragraph z. Finally, the information output by the neural retriever is regarded as additional text information and fused into the generator in a marginalized manner to generate the final sequence;
[0074] A vector database construction module for collecting open - source intelligence, then performing data cleaning, data processing, and metadata extraction, and then encoding through the query encoder in the neural retriever to form a vector database;
[0075] A report generation module for obtaining a report outline, automatically determining whether there is sufficient data in the vector database to support report generation according to the keywords and parameters in the outline. If the data is sufficient, the keywords and parameters are used as the query x in the enhanced retrieval model in the order of the upper and lower sub - points of the outline, and finally a report is generated according to the outline and all output information.
[0076] In a possible implementation, in the report generation module, if it is determined that the data in the vector database is not sufficient to support the generation of the current outline report, open-source intelligence is automatically collected according to keywords and parameters, and then data cleaning, data processing, and metadata extraction are performed, and then encoded through the query encoder in the neural retriever, so as to be able to quickly respond to queries in subsequent retrieval processes.
[0077] In a possible implementation, in the report generation module, the report outline is obtained by direct input from the user or by guiding the user to generate it;
[0078] Guiding the user to generate specifically includes: receiving the text directly input by the user, collecting the theme, sub-field, analysis dimension, and attention indicators, and generating and displaying input suggestions based on historical data and user behavior patterns; for broad or unclear user input text, pop up auxiliary text to prompt the user how to refine the input content, or provide options to refine the collection scope; after the user completes the input, the information provided by the user will be analyzed using a large language model to generate a draft outline, receive the user's editing on the draft outline, and generate a report outline according to the confirmation operation.
[0079] In a possible implementation, in the process of fusing the information output by the neural retriever into the generator, a sequence model and a token model are adopted. The sequence model uses the same document to generate all sequences, and the token model uses all the retrieved documents to generate sequences. The calculation method is as follows:
[0080]
[0081] Among them, N represents the total length of the target sequence y, i represents the i-th token currently generated in the target sequence y, z ∈ top-k(p(·|x)) represents the k most relevant paragraphs z selected from the retriever, according to the relevance score p(·|x) of the given query x;
[0082] Neural retriever p η (z|x), is encoded by the input query statement x to obtain the encoded vector q(x). In addition, the maximum inner product of the vector d(z) in the vector database and q(x) is used to search for the top-k relevant documents, and the output is used as p η (z|x);
[0083] d(z) = LLM d (z), q(x) = LLM d (x)
[0084] Among them, LLM d represents a large language model.
[0085] In a possible implementation, during the construction of the vector database, metadata extraction includes file name, time, chapter title, and picture alt information. By turning entities into nodes and relationships into relations, a knowledge graph is used for retrieval.
[0086] Since the system introduced in Embodiment 2 of the present invention is the system adopted for implementing the method in Embodiment 1 of the present invention, based on the method introduced in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the system, so it will not be elaborated here. Any system adopted for the method in Embodiment 1 of the present invention falls within the scope of protection of the present invention.
[0087] The present invention constructs an enhanced retrieval model, treats the information output by the neural retriever as additional text information, and integrates it into the generator in a marginalized manner to generate the final sequence. Thus, it can automatically obtain keywords and parameters according to the report outline input by the user, and then input them into the enhanced retrieval model to automatically generate an opinion report, reducing the labor cost and time cost, providing data support for enterprises and governments, helping them understand the social situation more objectively and comprehensively, and improving the decision-making efficiency.
[0088] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0089] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a system for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0090] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction system that implements the functions specified in one or more of the procedures Figure 1 one or more of the procedures and / or blocks Figure 1 the functions specified in one or more of the blocks.
[0091] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the procedures Figure 1 one or more of the procedures and / or blocks Figure 1 the functions specified in one or more of the blocks.
[0092] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative only and not used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A method for real-time monitoring and report generation of open source intelligence public opinion hot spots, characterized in that: include: Model building process: Construct an enhanced retrieval model, which consists of a pre-trained neural retrieval engine p with parameter η η (z|x) and a pre-trained generator p with parameters θ θ (y i |x,z,y 1:i-1 ) are combined and fine-tuned end-to-end; the retrieval machine p η (z|x) includes a query encoder and a document indexing module to return the top-k distribution of paragraphs z given a query x; the generator p θ (y i |x,z,y 1:i-1 ) is a seq2seq model that returns the context y of the first i-1 tokens based on θ given a query x. 1:i-1 , the original input and the retrieved paragraph z generate the current tag; finally, the information output by the retriever is regarded as additional text information and fused into the generator through marginalization to generate the final sequence; Vector database construction process: collect open source intelligence, then perform data cleaning, data processing and metadata extraction, and then encode it through the query encoder in the neural searcher to form a vector database; Report generation process: obtain the report outline, and automatically determine whether there is sufficient data in the vector database to support report generation based on the keywords and parameters in the outline. If the data is sufficient, the keywords and parameters are input as query x in the order of upper and lower points in the enhanced retrieval model, and finally generate a report based on the outline and all output information.
2. The method according to claim 1, characterized in that: During the report generation process, if it is judged that the data in the vector database is not sufficient to support the generation of the current outline report, open source intelligence is automatically collected based on keywords and parameters, and then data cleaning, data processing and metadata extraction are performed, and then encoded through the query encoder in the neural retriever, so that the query can be responded to quickly in the subsequent retrieval process.
3. The method according to claim 1 or 2, characterized in that: During the report generation process, the report outline is obtained by being directly input by the user or guided by the user to generate; Guiding user generation specifically includes: receiving text directly input by users, collecting topics, sub-fields, analysis dimensions and focus indicators, generating and displaying input suggestions based on historical data and user behavior patterns; for broad or unclear user input text, popping up auxiliary text to prompt users how to refine the input content, or providing options to refine the collection scope; When the user completes the input, a large-scale language model will be used to analyze the information provided by the user, generate a draft outline, receive the user's edits on the draft outline, and generate a report outline based on the confirmation operation.
4. The method according to claim 1, characterized in that: In the process of integrating the information output by the neural retriever into the generator, a sequence model and a tag model are used. The sequence model uses the same document to generate all sequences, and the tag model uses all retrieved documents to generate sequences. The calculation method is as follows: Where N represents the total length of the target sequence y, i represents the i-th token currently generated in the target sequence y, and z∈top-k(p(·|x)) represents the k most relevant paragraphs z selected from the retriever, according to the relevance score p(·|x) for the given query x; Neural Retriever η (z|x), the encoded vector q(x) is obtained by encoding the input query statement x, and the vector d(z) and q(x) in the vector database are searched for the top-k related documents by performing the maximum inner product, and the output is used as p η (z|x); p η (z|x)∝exp(d(z)Tq(x)) d(z)=LLM d (z),q(x)=LLM d (x) Among them, LLM d Represents a large language model.
5. The method according to claim 1, characterized in that: During the construction of the vector database, metadata is extracted including file name, time, chapter title, and image alt information. By converting entities into nodes and relationships into relation, knowledge graphs are used for retrieval.
6. A system for real-time monitoring and report generation of open source intelligence public opinion hot spots, characterized in that: include: A model building module is used to build an enhanced retrieval model, which is composed of a pre-trained neural retrieval engine p with parameter η. η (z|x) and a pre-trained generator p with parameters θ θ (y i |x,z,y 1:i-1 ) are combined and fine-tuned end-to-end; the neural retrieval unit p η (z|x) includes a query encoder and a document indexing module to return the top-k distribution of paragraphs z given a query x; the generator p θ (y i |x,z,y 1:i-1 ) is a seq2seq model that returns the context y of the first i-1 tokens based on θ given a query x. 1:i-1 , the original input and the retrieved paragraph z generate the current tag; finally, the information output by the neural retriever is regarded as additional text information and fused into the generator through marginalization to generate the final sequence; The vector database building module is used to collect open source intelligence, then perform data cleaning, data processing and metadata extraction, and then encode it through the query encoder in the neural retriever to form a vector database; The report generation module is used to obtain the report outline, and automatically determine whether there is sufficient data in the vector database to support report generation based on the keywords and parameters in the outline. If the data is sufficient, the enhanced retrieval model is input as query x in the order of the upper and lower points of the keywords and parameters, and finally a report is generated based on the outline and all output information.
7. The system according to claim 6, characterized in that: In the report generation module, if it is judged that the data in the vector database is not sufficient to support the generation of the current outline report, open source intelligence is automatically collected according to keywords and parameters, and then data cleaning, data processing and metadata extraction are performed, and then encoded through the query encoder in the neural retriever, so that the query can be responded to quickly in the subsequent retrieval process.
8. The system according to claim 6 or 7, characterized in that: In the report generation module, the report outline is obtained by direct input by the user or guided by the user to generate; Guiding user generation specifically includes: receiving text directly input by users, collecting topics, sub-fields, analysis dimensions and focus indicators, generating and displaying input suggestions based on historical data and user behavior patterns; for broad or unclear user input text, popping up auxiliary text to prompt users how to refine the input content, or providing options to refine the collection scope; When the user completes the input, a large-scale language model will be used to analyze the information provided by the user, generate a draft outline, receive the user's edits on the draft outline, and generate a report outline based on the confirmation operation.
9. The system according to claim 6, characterized in that: In the process of integrating the information output by the neural retriever into the generator, a sequence model and a tag model are used. The sequence model uses the same document to generate all sequences, and the tag model uses all retrieved documents to generate sequences. The calculation method is as follows: Where N represents the total length of the target sequence y, i represents the i-th token currently generated in the target sequence y, and z∈top-k(p(·|x)) represents the k most relevant paragraphs z selected from the retriever, according to the relevance score p(·|x) for the given query x; Neural Retriever η (z|x), the encoded vector q(x) is obtained by encoding the input query statement x, and the vector d(z) and q(x) in the vector database are searched for the top-k related documents by performing the maximum inner product, and the output is used as p η (z|x); Among them, LLM d Represents a large language model.
10. The system according to claim 6, characterized in that: During the construction of the vector database, metadata is extracted including file name, time, chapter title, and image alt information. By converting entities into nodes and relationships into relation, knowledge graphs are used for retrieval.
Citation Information
Cited By
Image-text report generation method fusing multi-mode large language model and RAG mechanism
CN120995994A
Method for generating picture-text report by fusing multi-modal large language model and RAG mechanism
CN120995994B
Report generation method and device and electronic equipment
CN121093925A