A SOFTWARE ARCHITECTURE FOR COMPUTER-BASED OPEN-SOURCE INTELLIGENCE ANALYSIS AND A SUITABLE METHODOLOGY FOR THIS ARCHITECTURE.
Patent Information
- Authority / Receiving Office
- TR · TR
- Patent Type
- Patents
- Current Assignee / Owner
- HAVELSAN HAVA ELEKTRONIK SANAYI VE TICARET ANONIM SIRKETI
- Filing Date
- 2024-12-05
- Publication Date
- 2026-06-22
Smart Images

Figure 00000015_0000 
Figure 00000016_0000
Abstract
Description
1 TARIFF COMPUTER-BASED OPEN SOURCE INTELLIGENCE ANALYSIS A software architecture that will work and a suitable system for that architecture. METHOD Technical Area 5 The invention was created from publicly available internet resources by organizations for the purpose of acquiring information. Planning, data collection, data processing, data exploration on a defined topic, analysis, information acquisition and dissemination are asynchronous processes in the intelligence cycle. software that involves some form of automation and will run on a computer. It is related to architecture and a method that works in accordance with this architecture. 10 More specifically, the invention utilizes large languages on data obtained from open sources. sentiment analysis, text summarizing, text classification, and machine translation using various models. By performing analyses such as entity recognition, raw data will be transformed into information. It is related to a software architecture. Previous Technique 15 Open source intelligence involves data collection, data processing, and the analysis of collected data. It basically consists of four main stages, including reporting. Data is personal. any access restrictions such as web addresses, social media, open databases It is collected from non-internet sources. In the data collection process... libraries such as Selenium [1], BeautifulSoup [2] and Scrapy [3] and 20 APIs are used. Existing open-source intelligence systems can speed up the data collection process. typically computers with powerful hardware and advanced networking technologies They have preferred to use different types of databases. The collected data is stored in different types of databases. It is being kept hidden. 25 2 Different models and algorithms for data processing and analysis. Text-based data, natural language processing techniques, and machine learning are used. Meaningful information is obtained using algorithms and big data analytics methods. The data is being transformed. This involves cleaning, categorizing, and analyzing the data. Advanced software tools and artificial intelligence-powered systems are used for this. 5 When the findings are communicated to the end user, depending on the expected format... Different models are used for format conversions. In patent document number CN217386353, which is included in the prior art, a The text refers to an open-source intelligence information gathering system. Patent document number US2011258187, which is included in the prior art, states 10 accessing an information request, identifying a primary data source associated with the request. and a primary dataset collected from a primary data source through a data collection process This refers to an open-source intelligence (OSINT) method that involves accessing information. In the patent document numbered CN116186369, which is included in the prior art, a The text refers to open-source intelligence gathering methods. 15 Open source intelligence involves gathering information from numerous sources during the data collection phase. This stage takes time. Current systems handle the data collection process. To speed things up, run operations with multiple threads and use high-performance equipment. It uses computers and advanced network systems [4]. With this method, data The performance increase in the collection process will only occur due to the bottleneck at the network layer. It can progress that far. In other words, it has a limited reach despite the high cost. performance is achieved. Open source intelligence (OSINT) tools Another significant problem they face is language barriers. Worldwide Thousands of languages are spoken, and each language has its own unique structure and rules. 25 3 Because many open-source data collection tools are optimized for a specific language, Their performance may decrease when used in different languages. Different models are used in processing and analyzing the collected data, and The use of algorithms reduces system modularity. This situation leads to code... Increasing complexity while the incompatibility of different models compromises system-wide consistency. 5 This can lead to problems and risks of errors. Reduced system modularity, This negatively affects the system's scalability and maintenance costs. The interdependence of processes in open source data collection stages is varied. This leads to negative consequences. This dependency reduces the system's flexibility. This limits its ability to adapt quickly to changing conditions. A 10 Slowing down one stage can cause the entire process to slow down, and at one stage A potential error could halt the entire operation. Between stages. Dependencies negatively affect the system's fault tolerance and performance. together, monitoring and controlling the stages, identifying and correcting problems It makes things more difficult. 15 When collecting data from open sources with open source data collection systems, the goal is... not to be detected by the open source information system and not to leave a footprint Various methods are being tried in this regard. One of these methods is a proxy server. The goal is to use a proxy server. Using a proxy server is an additional step for an open-source data collection system. Hardware and maintainability costs can be incurred. 20 Furthermore, the dynamic conceptual structure of knowledge and the relationships between these structures There is currently no solution that represents their relationship. To eliminate all the disadvantages mentioned above, we need a new software architecture and a new computer-based method that works according to this architecture It was deemed necessary to carry it out. 25 4 Purposes of the Invention The purpose of this invention is to enable organizations to access information from the public internet. Planning, data collection, data processing, data on a topic determined from the sources. the intelligence cycle including discovery, analysis, information acquisition and information dissemination a software architecture that includes asynchronous automation and 5 It is the implementation of a method that works appropriately. Another purpose of this invention is to perform large-scale language tests on data obtained from open sources. sentiment analysis, text summarizing, text classification, and machine translation using various models. It enables raw data to be transformed into information by performing analyses such as entity recognition. It is the implementation of architecture and a method that works in accordance with this architecture. 10 In invention, the method involves data acquisition, parallel operation, and master-slave (master-slave) private virtual servers (VPS - Virtual Private Server) that conform to the (slave) architectural pattern This is provided by web scrapers running on the server, and depending on the volume of data. The number of virtual servers varies dynamically. Therefore, which The ability to dynamically determine which hardware resources will be used, 15 to be collected The data should have both a controllable load value and serve the research purpose. It enhances its quality in terms of proximity. For this reason, the invention is automatic and powerful. It has a design that is scalable and complies with Amdahl's law [6]. The target is data scraping agents working on open source websites, using unique IP addresses. open source data collection operation through virtual servers owned by 20 IP blocking situations by the target information system because it is provided. It can be overcome. Furthermore, it can be achieved at a higher level without encountering a bottleneck problem. Improved performance and lower network costs can be achieved. Another important advantage offered by distributed architecture is open-source knowledge gathering. It eliminates the dependency between its components. Communication between components 25 This is achieved through an event-based microservices architecture. Between components... Communications are transmitted asynchronously via message brokers. This is happening. This method involves any component during the analysis phase. In the event of a possible error, all remaining system components will continue to function. This will allow it to operate without being affected in a timely manner. Thus, "One Failure The "point" condition will be prevented, and the system's tolerance to errors will be increased. 5 Thanks to the OpenNMT [7] technology used in the invention, traditional statistical and Higher accuracy and fluency compared to rule-based translation methods. This allows for more natural and meaningful translations between language pairs. In this way, language barriers are overcome, and a wider data collection pool is created. Consistent, comprehensive, and in-depth analyses are obtained. 10 Models used in processing and analyzing collected data Combining them in a large language model increases the modularity of the system and the models By eliminating incompatibility problems between them, it reduces the risk of errors and more It creates a consistent system. With increased system modularity, lower costs are achieved. A more manageable system has been achieved with reduced maintenance costs. 15 Information engineering methods and the relationship between conceptual structures and information systems Dynamic derivation of its ontology as a relation, contribution to the literature as a solution. It will provide. The fact that the invention is domestic and national provides various benefits to the nation. It is of critical importance. Access to the information provided by unauthorized individuals can be prevented. The country's 20 reducing external dependence by decreasing the need to use external sources for intelligence. This can reduce the impact. Different security protocols can be created for different users. Detailed Description of the Invention The computer-based method used to achieve the purpose of this invention is attached. This is shown in the figures. 25 6 These shapes; Figure 1: Schematic representation of the system in which the invented method will be used. Figure 2: Flowchart of the method described in the invention. The parts shown in the figures are individually numbered, and these numbers... The corresponding answers are given below. 5 1. User 2. Open source data collection service 3. Orchestrator 4. Database 5. Message agent 10 6. Artificial intelligence services 7. Dedicated virtual server 8. URL tail 9. Content queue 10. Language recognition model 15 11. Language translation model 12. Major language model 13. Internet The system that is the subject of the invention, - An orchestrator (3) connected to the database (4), 20 - Connected to the orchestra (3) and in a parallel structure message intermediaries (5), - Each message medium (5) is connected to each other in a parallel structure and private virtual servers (7), - An AI service (2) connected to each message agent (5), 25 7 - Language recognition model (10) included in the artificial intelligence service (2), language Translation model (11) and big language model (12) It includes. The invention concerns the method, 5 - The user searches (2) from (1) open source data collection service and requesting analysis, - Orchestrator (3) found in the open source data collection service (2) via the message intermediaries (5) related to 10 conveying the request, - Request to private virtual servers (7) via message broker (5) transmission, - The relevant request is made via private virtual servers (7) over the internet (13) search is done via and the result is sent to the message agent (5) 15 transmission, - The relevant content is sent to AI services via the message agent (5) (6) transmission, - Data from artificial intelligence services (6) are reused coming to the orchestrator (3) and the result being transmitted to the user (1) 20 It includes the steps. In the system that is the subject of the invention, the orchestrator (3) master-slave software architecture pattern As the master component, interaction among other components, structural and workload deployment across behavioral subsystems, system performance, system security, Monitoring system outputs and storing data in short-term memory and disk 25 It ensures that it is held. 8 Message broker (5), data scraping via the internet (13) and artificial intelligence service (6) task queues (URL queues) for sub-processes that will perform analysis tasks (8) and content queue (9)) is an event-based managing component. Tasks “URL”, These include “Content”, “Content Language”, “Translated Content” and “AI Analytics Content”. It is generated for the queues hosted in the topic and data is retrieved from these queues. It is consumed. In order to prevent the “Single Point of Failure” problem in the message medium (5). Message forwarders are kept in multiple and redundant configurations, and when the lead server fails, the others take over. The message agent (5) is activated. This Event-Driven Architecture With the architecture style, data flow is distributed and asynchronous, highly scalable. is provided. 10 The dedicated virtual server (7) performs the data scraping task. Multiple parallel dedicated Packages that can be run on a distributed VPS designed using virtual servers (7) This is accomplished through agents deployed as messengers. Each agent acts as a message intermediary. (5) 15 to perform URL scraping tasks in the queue allocated to it “HTML Retrieval”, “HTML Parsing”, “URL Extraction”, “URL Filtering” and It performs “URL Restriction” operations [9]. Language recognition model (10), an advanced open source data collection analysis, to be acquired 20 without being dependent on the data language of the desired information found in open sources This can make it possible. In this context, obtained from open sources The content language for 157 different languages is Fasttext
[10] in the Language Recognition Model (10) component. It is detected using an artificial intelligence model. The detected content language and content are translated. To be done, send a message to the topic queue "Content Language" on the message agent (5). It is written as 25. Although the language of the data can vary, the system component to be analyzed Performing analyses in every language can be very costly and may yield inaccurate results. Therefore, the AI will perform the analysis while the data is obtained in its existing language at the source. The systems must be translated into the language in which they operate with optimum performance. 30 9 It is necessary. In this context, the OpenNMT model for the Language Translation Model (11) is required. using previously identified language content into English or Turkish. Translation is ensured. The Language Translation Model (11) component is "Content Language" for this task. After receiving the task from the queue in the topic and performing the translation process It produces the work outputs for the "Translated Content" topic. 5 Within the scope of OSINT, Sentiment Analysis is performed on data to transform data into information. Analysis such as Text Classification, Text Summarization, and Named Entity Recognition. It is widely used. With the Big Language Model (12), this DDL (Natural Language) Processing analyses are performed with deep learning models that have billions of parameters. This is done. Furthermore, unlike classic DDI models, OSINT analysis reports are provided. dynamically generated text generation capability of the large language model (12) This is provided. In this context, content consumed from the queue belonging to the "Translated Content" topic. DDI analyses are performed on the content, and the analysis results are posted in the "AI Analysis" topic. They are written as messages. 15 Large-scale data, in terms of data type and volume, can be obtained from open sources. For this function to be efficient, it must be performed simultaneously in parallel. This means it needs to happen. Many of the known systems It implements parallelism in multithreaded operation. However, the powerful 20 Processes executed as multiple threads from a server with a specific configuration, Although it can perform highly in computational data processing, at some point They are getting stuck in a network bottleneck. This is because these powerful servers are connected to a single Ethernet cable. They access the internet with the limited bandwidth of their adapter. If you have a single, powerful adapter... Instead of a server, processes are performed by multiple 25-bit systems with low-configuration specifications. If it is performed as a multi-process on the server, it will also be accessible at the internet exit point. Parallelism is thus ensured. The data retrieval function is performed using distributed - multiple processes, data retrieval The goal is to ensure that the target information system does not perceive the HTTP request as a threat and 30 It also ensures that it doesn't block access. This allows for the addition of a proxy server to the system. Integration costs can also be eliminated. Furthermore, distributed VPSs are suitable for enterprise use. If the data scraping process is located in a public cloud environment instead of on-site infrastructure, then... The resulting environmental footprint can be eliminated. The distributed design, which enables data acquisition across all these VPSs, is a key feature of the invention. This ensures parallel operation and strong scalability with the processes. The invention is a whole system that performs unique functions through its architectural design. It enables asynchronous communication between its components. Many OSINT 10 In contrast to the fact that the vehicle provides ease of use with its simple architectural design, the sustainability, testability, maintainability, and It ignores its modularity. OSINT tools are used every 15 years to transform raw data into meaningful information. It develops or uses different AI analysis models for an analysis process. However, the analyses should be carried out by every institution or organization that will use OSINT systems. The fact that these AI systems vary from person to person and are numerous makes development easier. This increases the cost. Therefore, Mixtral, Llama-3 or higher quality is preferable. These 20 were used with pre-trained versions of BDMs that had test results. Analyses can be provided using a single language model. Trainings are based on publicly available data. This is done, however, within the scope of OSINT, it is specific to a particular area or a specific data type. When category-specific analyses are required, BDM provides the desired accuracy. They may not produce results. Also, since training these models is very expensive, not every... It is not possible to retrain with the new dataset. At this point, previously 25 In the literature, retrieval is used to enable trained models to perform analysis with new data. Augmented Generation (RAG) method has been proposed. With the architecture we propose, open We vectorize the data we obtain from various sources using Embedding Models. the process of collecting, storing, and analyzing these vectors as parameters in BDM. It is ensured that it is used as such. 30 11 However, the BDMs found in the literature are predominantly in English content. They are being trained. This allows them to perform more successful analyses on English-language content. This means that, therefore, with the architecture we propose, we will first utilize open sources. The language of the content is determined using the Fasttext AI language detection model on the obtained text. is being processed, then the original text is rendered on this language using OpenNMT models 5 It is being translated into English. The translated text is Embedding. Thanks to their models, they become usable for analysis after being vectorized. is being brought. With this invention, social engineering, misleading news and disinformation efforts, 10 about terrorist attacks, cyberattacks and other organized crime groups Classified and structured information can be obtained. Networks of criminal organizations... Structures can be identified. Special cases such as natural disasters and global disease outbreaks. Data-driven solutions can be obtained in crisis situations. Defense industry and military. organizations should not openly disclose information about competing / enemy products, platforms, or entire countries. Information can be obtained by analyzing data collected from various sources. References: [1] The Selenium Browser Automation Project, 20 https: / / www.selenium.dev / documentation / (accessed Jun. 6, 2024). [2] Beautiful Soup Documentation - Beautiful Soup 4.12.0 documentation, https: / / www.crummy.com / software / BeautifulSoup / bs4 / doc / (accessed Jun. 6, 2024). [3] “Scrapy 2.11 documentation,” Scrapy 2.11 documentation - Scrapy 2.11.2 25 documentation, https: / / docs.scrapy.org / en / latest / (accessed Jun. 6, 2024). [4] R. Martínez-Castaáo, D. E. Losada and J. C. Pichel, "Real-Time Focused Extraction of Social Media Users," in IEEE Access, vol. 10, pp. 42607-42622, 2022, doi: 10.1109 / ACCESS.2022.3168977. keywords: {Real-time systems;Social networking (online);Task analysis;Crawlers;Computer architecture;Feature 30 12 extraction;Data mining;Big data;distributed systems;focused user extraction;supervised learning;information retrieval;real-time processing;social media}. [5] The Intelligence Cycle, https: / / irp.fas.org / cia / product / facttell / intcycle.htm [6] Amdahl GM (1967) Validity of the single-processor approach to achieve large 5 scale computing capabilities. AFIPS Joint Spring Conference Proceedings 30 (Atlantic City, NJ, Apr. 18–20), AFIPS Press, Reston VA, pp 483–485, At http: / / www-inst.eecs.berkeley.edu / ~n252 / paper / Amdahl.pdf [7] G. Klein, Y. Kim, Y. Deng, J. Senellart, and A.M. Rush. 2017. Opennmt: Open- source toolkit for neural machine translation. CoRR, abs / 1701.02810. 10 [8] S. Research, “Open-source intelligence market size, analysis, & share to [2022- 2030],” Straits Research, https: / / straitsresearch.com / report / open-source- intelligence- market#:~:text=Market%20Overview,period%20(2022%E2%80%932030 (accessed Jun. 6, 2024). 15 [9] Heydon, A., Najork, M. Mercator: A scalable, extensible Web crawler. World Wide Web 2, 219–229 (1999). https: / / doi.org / 10.1023 / A:1019213109274
[10] Word vectors for 157 languages, https: / / fasttext.cc / docs / en / crawl-vectors.html
[11] Open source models, https: / / mistral.ai / technology / #models
[12] Build the future of AI with Meta Llama 3, https: / / llma.mesa.com. / llama3 / 20
Claims
13 REQUESTS 1. The invention was determined from publicly available internet sources for the purpose of acquiring information. planning, data collection, data processing, data exploration, and analysis on the subject. asynchronous aspects of the intelligence cycle, such as information acquisition and dissemination some form of automation involved and 5 - An orchestrator (3) connected to the database (4), - Connected to the orchestra (3) and in a parallel structure message intermediaries (5), - Each message medium (5) is connected to each other in a parallel structure and private virtual servers (7), 10 - An AI service (2) connected to each message agent (5), - Language recognition model (10) included in the artificial intelligence service (2), language Translation model (11) and big language model (12) It relates to a method that will work on a system that includes 15 - The user searches (2) from (1) open source data collection service and requesting analysis, - Orchestrator (3) found in the open source data collection service (2) 20 related to the message intermediaries (5) that are parallel to each other. conveying the request, - Request to private virtual servers (7) via message broker (5) transmission, - The relevant request is made via private virtual servers (7) over the internet (13) search is done via and the result is sent to the message agent (5) 25 transmission, - The relevant content is sent to AI services via the message agent (5) (6) transmission, 14 - Data from artificial intelligence services (6) are reused coming to the orchestrator (3) and the result being transmitted to the user (1) It is characterized by including its steps.