An Open Source Intelligence Source Extension Method and System Oriented Towards Business Themes

CN122570799APending Publication Date: 2026-08-14NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-06
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0010]针对现有开源情报采集过程中存在的目标识别难、主题关联弱、人工成本高、动态更新不及时等问题,本申请提出一种面向业务主题的开源情报信源扩展方法及系统,用以融合规则匹配与图像识别能力,提出支持多源目标自动识别、扩展及动态更新的开源情报信源扩展方法

Benefits of technology

[0013]本申请实施例通过规则匹配和图像识别的技术组合,利用互联网探测与自动截屏技术,对网站、论坛、社交群组等目标进行普查与种子扩展采集,并基于大模型的图像识别与聚类算法,实现对业务主题的自动识别、分类与信源体系构建。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122570799A_ABST
    Figure CN122570799A_ABST
Patent Text Reader

Abstract

To address the problems of target identification difficulties, weak topic correlations, high manual costs, and untimely dynamic updates in existing open-source intelligence gathering processes, this application discloses an open-source intelligence source expansion method and system oriented towards business themes. Involving internet and data processing technologies, the method includes: collecting data on multiple types of targets containing business themes; determining and classifying the themes of the collected data using a combination of rule matching and image recognition; and generating a source target library based on the results of theme determination and classification according to different business themes. This application integrates rule matching and image recognition capabilities, proposing an open-source intelligence source expansion method that supports automatic identification, expansion, and dynamic updates of multi-source targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of Internet and data processing technology, and in particular to an open-source intelligence source extension method and system oriented towards business themes. Background Technology

[0002] With the popularization of the internet and the exponential growth of information, a large amount of publicly available information resources are constantly being generated in cyberspace, including websites, forums, social media platforms, news media, blogs, and online databases. After effective screening and analysis, this publicly available data can provide important intelligence references for fields such as intelligence analysis, national security, economic research, and public opinion monitoring, and is collectively referred to as Open Source Intelligence (OSINT).

[0003] However, in the process of monitoring and collecting open-source intelligence on the internet, due to the sheer number and frequent updates of internet targets, researchers and operational personnel often struggle to determine which targets should be monitored and collected, thus failing to efficiently obtain targeted and accurate intelligence information. This problem of "collecting inaccurate and incomplete information" has become a prominent challenge in current open-source intelligence monitoring work. In practical operations, due to a lack of understanding of the overall distribution and characteristics of internet target sites, staff often rely solely on experience to monitor individual known websites or databases, lacking systematic methods for expanding information sources. Expanding and sorting target information sources requires a significant investment of manpower and time. Meanwhile, the dynamic changes of network targets are frequent; some forums, social groups, or databases periodically migrate, are redesigned, or become invalid, causing interruptions to existing monitoring links, making it difficult to maintain effective collection and continuous monitoring of information source data.

[0004] Current open-source intelligence gathering efforts generally suffer from the following problems: First, targets are scattered and structurally complex. Internet targets are diverse, involving different forms such as web pages, forums, and social groups, making unified collection and processing difficult. Second, topic identification is difficult. Traditional methods rely heavily on structured analysis of web page text, but in reality, some website pages are dynamically generated or have closed access, making it difficult to extract effective text information. Third, manual labor costs are high and updates are not timely. Intelligence personnel need to manually screen and verify website categories and topics, resulting in a lag in information source updates. Fourth, websites change frequently. Targets such as forums and social groups may migrate, go offline, or be redesigned over time, causing the information source monitoring link to fail.

[0005] How to automatically identify and expand the effective sources of information relevant to business topics in large-scale Internet targets has become a key technical challenge in the field of open source intelligence.

[0006] Existing methods include rule-based topic identification methods and AI-based content analysis methods, among which: (1) Rule-based topic identification method Existing open-source intelligence gathering mainly adopts rule-based text retrieval methods [3]. By setting business-related keywords, regular expressions, or Boolean logic rules, web page text, metadata, titles, and link information are matched and filtered to identify target sites. This method is simple to implement, easy to deploy, and can quickly obtain a preliminary target set under a specific topic.

[0007] However, rule-based data collection methods have obvious limitations: First, the keyword coverage is limited, making it difficult to handle multilingual, multi-expression, and unstructured content; second, rules need to be set and maintained manually, and when the Internet environment changes dynamically, the rule base is not updated in time, which can easily lead to recognition errors or omissions; third, relying solely on text matching cannot analyze the themes of web pages containing images, videos, or dynamically loaded content, resulting in low recognition accuracy.

[0008] (2) Content analysis methods based on artificial intelligence In recent years, some studies have attempted to introduce machine learning algorithms into the field of open-source intelligence gathering. These studies utilize natural language processing and topic modeling techniques (such as LDA and BERT) to automatically extract latent topic features from web page text and determine the relevance of target websites to business topics based on similarity or semantic relevance. This allows for automatic classification of web page text, aiding in the identification of topic-related target websites. Other studies employ social network analysis and graph algorithms (such as node similarity calculation and PageRank) to automatically extract latent topic features from web page text and determine the relevance of target websites to business topics based on similarity or semantic relevance.

[0009] While AI-based content analysis methods outperform traditional rule-based methods in terms of semantic understanding and automation, they still have certain limitations. First, these models typically rely on large-scale, high-quality labeled data for training, but intelligence corpora are characterized by high confidentiality, scarce samples, and slow updates, making it difficult to continuously support model iteration. Second, existing models largely focus on text content analysis, neglecting the visual structure and multimodal features of web pages, resulting in poor performance when recognizing web pages containing images, videos, or dynamically loaded content. Third, deep learning models are costly to update and retrain, and their response to new website types or topic changes is slow, making it difficult to achieve real-time, continuous source expansion and dynamic monitoring. Summary of the Invention

[0010] To address the problems of target identification difficulty, weak topic correlation, high manual costs, and untimely dynamic updates in existing open-source intelligence gathering processes, this application proposes an open-source intelligence source extension method and system oriented towards business themes. This method integrates rule matching and image recognition capabilities, and proposes an open-source intelligence source extension method that supports automatic identification, extension, and dynamic updates of multi-source targets.

[0011] This application provides a method for extending open-source intelligence sources based on business themes, including: Data collection is performed on multiple types of targets that include business themes; A method combining rule matching and image recognition is used to determine and classify the themes of the collected data; Based on the results of topic determination and classification, a source target database is generated according to different business topics.

[0012] This application provides an open-source intelligence source extension system oriented towards business themes, including a processor and a memory. The memory stores a computer program, which, when executed by the processor, implements the steps of the aforementioned open-source intelligence source extension method oriented towards business themes.

[0013] This application embodiment uses a combination of rule matching and image recognition technologies, along with internet detection and automatic screenshot technology, to conduct a general survey and seed expansion collection of targets such as websites, forums, and social groups. Based on large-scale model image recognition and clustering algorithms, it achieves automatic identification, classification, and information source system construction of business topics.

[0014] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0015] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a schematic diagram of the overall process of the open-source intelligence source expansion method for business-oriented embodiments of this application; Figure 2 This application provides a target data collection and monitoring process based on a census and seed list expansion. Figure 3 This application provides a topic identification and classification process based on rule matching and image recognition in its embodiments. Figure 4 This is a schematic diagram of the process for constructing a source target library based on a business theme, as described in an embodiment of this application. Detailed Implementation

[0016] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0017] This application provides a method for extending open-source intelligence sources based on business themes, such as... Figure 1 As shown, it includes the following steps: In step S101, data is collected from multiple types of targets that include business themes. Specifically, this can be achieved through internet surveys and extended monitoring based on seed lists to quickly obtain entry information for target websites, forums, and social groups, and automatically complete screenshots of webpage homepages.

[0018] In step S102, a method combining rule matching and image recognition is used to determine and classify the themes of the collected data. Specifically, a rule + AI theme recognition mechanism can be used, combining keyword matching and image semantic recognition, to perform theme classification and cluster analysis on the screenshot images.

[0019] In step S103, based on the results of topic determination and classification, a target information source database is generated according to different business topics. That is, a multi-topic open-source intelligence information source database is constructed based on the identification results to achieve continuous monitoring and iterative updates of target information sources.

[0020] In some embodiments, data collection for multiple types of targets containing business themes includes two complementary methods: internet census data collection and extended monitoring based on seed lists. This achieves automated identification and preliminary screening of multiple types of targets such as websites, forums, and social groups. Figure 2 As shown, where: Internet census data collection includes using web crawlers and domain name detection technologies to conduct domain name surveys and site scans in the public space of the Internet based on set geographical, language, industry, or keyword ranges; and using methods such as HTTP response detection, port detection, and page title recognition to filter out non-business-related or invalid links and extract valid website and forum homepage information.

[0021] Expanded monitoring based on seed lists involves using manually provided or historically accumulated target site lists as initial input, and then expanding at multiple levels using hyperlink relationships, similar domain names, and social association analysis to continuously discover new potential sources of information.

[0022] It also includes: taking screenshots of the collected web pages after they have finished loading, converting the sites into static image files, and providing basic data for subsequent image recognition.

[0023] In some embodiments, such as Figure 3 As shown, the method of combining rule matching and image recognition is used to determine and classify the themes of the collected data, including: Based on a predefined business theme keyword library, such as categories including politics, economy, science and technology, and culture, text-level feature extraction and matching analysis are performed on the collected web page information. The business theme keyword library includes core words, extended words and synonyms related to the theme. The matching analysis process performs regularization and keyword retrieval on the webpage's related information, which includes the title, URL path, meta tags, description information, and a portion of the main text summary.

[0024] In some embodiments, the matching analysis process for regularizing the association information of web pages and retrieving keywords includes: for each web page With the topic Define a function that scores the rule matching score as the degree of keyword matching: in, Theme Keyword set, Keywords Frequency of appearance on a webpage Keyword weight (core keywords have higher weight). This represents the length of the webpage text, used for normalization.

[0025] If the target page meets the matching threshold for a certain topic keyword, If the match is unclear or the information is missing, it will be initially labeled as the corresponding topic category. If the match is unclear or the information is missing, it will proceed to the subsequent image recognition stage for further judgment.

[0026] In some embodiments, image recognition determination includes: An image semantic recognition method based on the CLIP model is adopted. The CLIP model consists of a visual encoder (ViT) and a text encoder (Transformer). A screenshot of a webpage is input into the visual encoder of the CLIP model, and a pre-defined business theme description is input. Input a text encoder to calculate the semantic similarity between webpage screenshots and topic tags, that is: in, It determines the topic category of a page, thereby compensating for the lack of analysis of web pages with missing text or complex structures.

[0027] In some embodiments, the matching analysis process further includes generating a final topic classification result by weighting the rule matching result and the image recognition result. In a specific example, low-confidence or conflicting samples can be automatically marked as "to be verified" for subsequent iterative optimization or manual review.

[0028] Automatically generate source target databases based on different business themes, such as Figure 4 As shown, in some embodiments, generating a source target database based on different business themes, according to the results of topic determination and classification, includes: The results of topic identification and classification are structured and organized into source records, and a website link is generated for each source record. Screenshot samples , topic tags Language type and update time The unified entry for metadata, each source record is represented as: Clustering and matching algorithms are used to merge related sites under the same topic and remove duplicate or invalid links, thereby achieving dynamic optimization of information sources.

[0029] In some embodiments, clustering and matching algorithms are used to group related sites under the same topic, including: For any two source records , Calculate multimodal similarity: If the result is higher than the threshold, it is determined that the two source records are duplicates or highly similar, and the record with the more recent timestamp is retained. Based on similarity matrix K-means clustering is performed on information sources under the same topic; Repeated data collection, judgment, and classification—that is, the data collection and identification process from the first stage to the second stage—is continuously iterated in the background, performing detection, calculation, and correction. When new data is collected, its similarity increment with existing source records is calculated. like If a record is found to be a new source, it is identified and added to the database. Conversely, if it is not, the metadata of the original record is updated, thereby continuously improving the accuracy of target site and group analysis.

[0030] Ultimately, the source target library can directly output a set of sources on the corresponding topics according to business needs, providing a real-time and accurate source foundation for subsequent open-source intelligence collection and analysis systems.

[0031] The method in this application adopts a dual mechanism combining "Internet census collection" and "expanded monitoring based on seed lists" to achieve wide-area discovery and accurate tracking of various Internet targets such as websites, forums, and social groups, significantly improving data coverage and target discovery efficiency. This application integrates rule matching and artificial intelligence image semantic recognition technology to construct a joint judgment model of "text features + visual features", which can effectively deal with multilingual, unstructured, and dynamically loaded web page scenarios, significantly improving the accuracy and robustness of topic recognition and classification.

[0032] The method described in this application can run continuously in the background, automatically executing the collection, detection, classification, and correction processes. It can adaptively update newly added, changed, and invalid information sources, ensuring the real-time nature and completeness of the information source target library, and realizing the automated and continuous maintenance of open-source intelligence information sources. The information source target library constructed in this application can be output by topic classification, supports seamless integration with various intelligence analysis platforms or security monitoring systems, and can flexibly support the open-source intelligence collection, analysis, and application needs in multiple business scenarios.

[0033] This application also proposes an open-source intelligence source extension system oriented towards business themes, including a processor and a memory. The memory stores a computer program, which, when executed by the processor, implements the steps of the aforementioned open-source intelligence source extension method oriented towards business themes.

[0034] It should be noted that, in the embodiments of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0035] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0036] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0037] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims. All of these forms are within the protection scope of this application.

Claims

1. A method for extending open-source intelligence sources oriented towards business themes, characterized in that, include: Data collection is performed on multiple types of targets that include business themes; A method combining rule matching and image recognition is used to determine and classify the themes of the collected data; Based on the results of topic determination and classification, a source target database is generated according to different business topics.

2. The open-source intelligence source extension method for business-themed applications as described in claim 1, characterized in that, Data collection for multiple types of targets encompassing business themes includes two methods: internet census data collection and extended monitoring based on seed lists. Internet census data collection includes using web crawlers and domain name detection technologies to conduct domain name surveys and site scans in the public space of the Internet based on set geographical, language, industry, or keyword ranges; and using HTTP response detection, port detection, and page title recognition methods to filter out non-business-related or invalid links and extract valid website and forum homepage information. Expansion monitoring based on seed lists includes using manually provided or historically accumulated target site lists as initial input, and conducting multi-level expansion using hyperlink relationships, similar domain names, and topic association analysis.

3. The open-source intelligence source extension method for business-themed applications as described in claim 2, characterized in that, Data collection for multiple types of targets that include business themes also includes: After the collected web pages have finished loading, a screenshot operation is performed to convert the site into a static image file.

4. The open-source intelligence source extension method for business-themed applications as described in claim 3, characterized in that, The method of combining rule matching and image recognition is used to determine and classify the themes of the collected data, including: Based on a predefined business theme keyword library, text-level feature extraction and matching analysis are performed on the collected web page information. The business theme keyword library includes core words, extended words and synonyms related to the theme. The matching analysis process performs regularization and keyword retrieval on the webpage's related information, which includes the title, URL path, meta tags, description information, and a portion of the main text summary.

5. The open-source intelligence source extension method for business-themed applications as described in claim 4, characterized in that, The matching analysis process includes regularization of webpage association information and keyword retrieval, including: For each webpage With the topic Define a function that scores the rule matching score as the degree of keyword matching: in, Theme Keyword set, Keywords Frequency of appearance on a webpage For keyword weight, The length of the webpage text; If the target page meets the matching threshold for a certain topic keyword, If the match is clear, it is marked as the corresponding topic category. If the matching result is unclear or the information is missing, image recognition is performed.

6. The open-source intelligence source extension method for business-themed applications as described in claim 5, characterized in that, Image recognition and determination include: An image semantic recognition method based on the CLIP model is adopted. The webpage screenshot is input into the visual encoder of the CLIP model, and the preset business topic description is input into the text encoder. The semantic similarity between the webpage screenshot and the topic tag is calculated to complete the matching analysis.

7. The open-source intelligence source extension method for business-themed applications as described in claim 6, characterized in that, The matching analysis process also includes generating the final topic classification result by weighting the rule matching results and the image recognition results.

8. The open-source intelligence source extension method for business-themed applications as described in claim 1, characterized in that, Based on the results of topic determination and classification, a source target database is generated according to different business topics, including: The results of topic identification and classification are structured and organized into source records, and a website link is generated for each source record. Screenshot samples , topic tags Language type and update time The unified entry for metadata, each source record is represented as: Clustering and matching algorithms are used to merge related sites under the same topic and remove duplicate or invalid links, thereby achieving dynamic optimization of information sources.

9. The open-source intelligence source extension method for business themes as described in claim 8, characterized in that, Clustering and matching algorithms are used to group related sites under the same topic together: For any two source records , Calculate multimodal similarity: If the result is higher than the threshold, it is determined that the two source records are duplicates, and the record with the more recent timestamp is retained. Based on similarity matrix K-means clustering is performed on information sources under the same topic; Repeatedly collect, judge, and classify data; when new data comes in, calculate its similarity increment with existing source records: like If it is identified as a new information source, it will be added to the database.

10. An open-source intelligence source extension system oriented towards business themes, characterized in that: It includes a processor and a memory, wherein a computer program is stored in the memory, and when executed by the processor, the computer program implements the steps of the business-oriented open-source intelligence source extension method as described in any one of claims 1 to 9.