AI-powered social media data collection and analysis system.

TR202614916A2Pending Publication Date: 2026-09-21INONU UNIVERSITESI REKTORLUGU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TR202614916
Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-09-01
Publication Date
2026-09-21

Smart Images

  • Figure 00000012_0000
    Figure 00000012_0000
Patent Text Reader

Abstract

The invention relates to an AI-powered social media data collection and analysis system (specifically in Turkish) designed for academics, master's and doctoral students, and independent researchers.
Need to check novelty before this filing date? Find Prior Art

Description

1 TARIFF AI-powered social media data collection and analysis system. Technical Area The invention benefits academics, master's and doctoral students, and independent researchers. It relates to a data collection and analysis system designed for this purpose. 5 The invention is particularly useful for AI-powered social media in the social sciences. It is related to data (especially Turkish) collection and analysis systems. State of the Art Today, academics, master's and doctoral students, and independent professionals... There are various research platforms available for researchers. 10 There are three types of vehicles currently used in the application: In the other current practice, • Commercial social listening platforms (Brandwatch, Talkwalker, Sprout Social, etc.): There is an annual licensing fee of $5,000-$50,000. This tool is for marketing. It is focused and not for academic research. 15 • Qualitative research software (NVivo, MAXQDA): Social media integration It is weak and lacks modules specific to the Turkish language. • Selenium-based mining scripts: They are open source and free. However, social... Media (X) is unstable to interface changes. It frequently malfunctions and is legally unstable. It is also debatable. 20 The disadvantages of current practices can generally be listed as follows: • They are not suitable for Turkish: The existing software is designed for English. For example, when you search for 'Alevism', you will see results like 'Alevism', 'to Alevism', 'Alevis', etc. Variants specific to Turkish cannot be captured. Turkish is an agglutinative language. Therefore, simple word matching leads to sociologically critical data loss. 25 It opens. 2 • The detection of profanity and insults is very crude: The current tools are "bad if these words are present." It operates on the assumption that "crazy man" and a serious insult are the same thing. It is not. The first is everyday jargon, the second is heavy profanity. Sociological anger. It is necessary to see this difference in the analysis. • Selenium dependency is unstable: Open source community solutions for Selenium 5 It uses, meaning it automatically controls the browser. The X social media interface every These scripts become corrupted during updates, resulting in significant data loss and process interruptions. This is causing problems. Furthermore, it is legally problematic / forbidden, etc. • Filter bubble effect: Data from a user's X social media search box When it collects them, X's algorithm applies a filter based on its profile. So, the collected 10 The data is not academically unbiased. This reflects sociological reality. It does not reflect this. • No location data: More than 90% of users have GPS location turned off. and it becomes impossible to keep track of and map social movements (rallies, festivals). It is coming. 15 • Software freezes: The interface locks up when pulling large amounts of data with traditional tools. Users can't do anything while the data is downloading. This reduces productivity. and it offers a poor user experience. • It is not possible to distinguish between relevant and irrelevant content: For example, searching for "Cemevi". When you do that, it's a genuine sociological discussion tweet with a 20 near a cemevi (Alevi place of worship). The restaurant's advertisement is yielding similar results. This is within the context of the available tools. He doesn't understand. • Algorithmic bias: Existing tools (API limitations or platform) (due to algorithms) data filtered according to the user's personal profile. and this ensures the academic and sociological impartiality of the collected data. 25 It has an effect. As a result of the research conducted on this matter, application number KR20240154740A This has been encountered. Here, a big data-driven K-content evaluation service is provided. The system is offered. This provides a big data-based K-content evaluation service. The system includes: a request for evaluation of K-content, which is Korean-language content, and 30 A user terminal that generates visualization data for an index for evaluation; and an aggregate that collects K-content data, including quantitative and qualitative data. The unit performs data preprocessing on quantitative data and natural language processing on qualitative data. a preprocessing unit that performs; an analytics unit that analyzes preprocessed K-content data 3 a visualization unit that includes a visualization unit that visualizes the analyzed K-content data. A server that provides an evaluation service. Another application encountered as a result of research on this subject is: This is application number US2025390896A1. The application describes individuals, organizations, and events. 5 including policies, products, places, ideas or other topics that can be evaluated a computer-based system for evaluating public opinion on a topic The system is explained. The system uses third-party platforms such as social media, news sites, etc. It collects and verifies data from various sources. The system uses natural language processing, rule-based logic, The issue can be evaluated using methods such as sentiment analysis and data modeling. It determines their actions and the corresponding public reactions. Public reactions, 10 For example, based on reliability, emotion, time-decreasing functions, or other factors. These factors are weighted and a public opinion response score is generated. A literature review on the subject revealed that software from X / Twitter social media platforms... English-language studies on data collection using this method are quite advanced, However, studies in the Turkish field remain both few in number and limited in scope. 15 It is observed that sample selections in Turkish studies are generally inadequate, and This reduces the reliability of the results. Looking at the literature, for example, Yang and Srinivasan published a paper in the field of Psychology in 2016. Although her research used a large sample size, it only utilized data from two years. (Yang, C., & Srinivasan, P. (2016). Life Satisfaction and the Pursuit of Happiness on 20 Twitter. PLoS ONE, 11(3), e0150881). This period reflects periodic social changes. It is insufficient to capture. Vicente's 2023 systematic review. However, this study is consistent with the vast majority of research conducted on Twitter. This reveals that there is no sampling frame (Vicente, 2023, p. 5468) (Vicente, P. (2023). Sampling Twitter users for social science research: evidence from a 25 systematic review of the literature. Qual Quant, 57, 5449–5489). In conclusion, due to the negative aspects described above and the current solutions being the subject of discussion... Due to its shortcomings, an improvement is needed in the relevant technical field. It has been made. 4 Purpose of the Invention The invention was created by drawing inspiration from existing situations and overcoming the aforementioned drawbacks. It aims to solve the problem. The main purpose of the invention is to enable researchers to collect Turkish social media data and The goal is to enable them to analyze. 5 Another goal of the invention is to search for all variants of words using natural language search. The goal is to ensure and prevent data loss. Another purpose of the invention is to label data by performing a three-level insult analysis. to provide. Another aim of the invention is to give the researcher the flexibility to choose a time interval, a wide and flexible 10 by offering the advantage of data collection opportunities and the selection of the right sampling strategy, It enables both the capture of periodic changes and interdisciplinary comparative analysis. The aim is to make it easier for them to conduct their studies. The structural and characteristic features and all the advantages of the invention are given in the figures below. Thanks to the detailed explanation written with references to these figures, it becomes clearer. 15 This will be understood as such, and therefore the evaluation will also take these forms and detailed explanations into account. This should be done taking that into consideration. Figures that will help understand the invention. Figure 1 shows the AI-powered social media data collection and analysis process that is the subject of the invention. This is a block diagram view of the system. 20 Explanation of Part References A. Official data collection layer B. Data cleaning and language processing layer C. Data storage and visualization layer D. Analysis and forecasting layer 25 E. Input interface F. Academic output G. It's over 1. Word expander 2. API connection 3. AI relevance filter 4. Three-level insult analysis module 5. Natural Language Search Assistant 5 6. Bot / Troll detection module 7. Trend forecasting module 8. Geographic map module 9. SQLite database 10. Asynchronous interface 10 11. Data assistant 12. Visualization module a. Output: Raw tweet list b. Yes c. No 15 d. Input e. Relevance score ≥ 5? f. AI manipulation report g. AI trend projection h. Interactive HTML map 20 i. AI sociological report Detailed Description of the Invention This detailed explanation describes the AI-powered social media data that is the subject of the invention. The preferred structures of the data collection and analysis system not only provide a better understanding of the subject. It is explained in order to facilitate understanding. 25 Official data collection layer (A): From the user login interface (E), they select a topic or submit a research request in natural language. It writes. The natural language search assistant (5) requests the social media (X / Twitter) search query. 6 It translates and expands the Turkish word (1) all Turkish variants of the keyword. It produces. The social media API connection (2) sends this query to the official social media. It sends the tweets to the server and retrieves them. All these tasks are done thanks to the asynchronous interface (10). It's done in the background and the user's screen doesn't freeze. Data cleaning and language processing layer (B): 5 The extracted raw tweets are first sent to an AI relevance filter (3) and then to an AI artificial intelligence relevance filter (3) each tweet according to its relevance to the research topic between 1-10 It assigns a score. Tweets with a score below 5 (advertisements, spam) are not written to the database. The remaining clean tweets are filtered through the three-level insult analysis module (4) They are filtered and each is given a toxicity label (clean / mild / moderate / severe) 10 is being appointed. Data storage and visualization layer (C): SQLite database (9) cleans all data by category, text, media, likes, retweets, date It saves the information to the local disk. The visualization module (12) creates a 3D scatter from this data. It produces charts, word clouds, and time series trend charts. 15 If the user wishes, they can chat with the data assistant (11) about the data (in Turkish language, etc.). is able to. Analysis and forecasting layer (D): This layer performs three main analyses: • Bot / Troll detection module (6) counts identical repeating messages and organic 20 It generates the score: • Trend forecasting module (7) with polarization and toxicity formulas for 48-hour viral It provides a prediction. • Geographic map module (8) extracts city names from tweet texts and Türkiye Place it on the map. 25 All three analyses are interpreted with an AI relevance filter (3) and reported in academic language. It is presented as such. 7 The invention, shown in the block diagram in Figure 1, is an AI-powered social system. media data collection and analysis system; • search for the request made by that user from the login interface (E) natural language search assistant (5) which translates the query into: (For example, 'last month 5 The author writes in colloquial Turkish, using phrases like "Alevism debates," and artificial intelligence does this. It converts professional social media search queries, automatically processing dates. (accounts, fills out the form.) a word expander that expands the searched word according to grammatical rules (1) (For example, the word “Alevism” can be translated as “Alevism’s, to Alevism, Alevis”, etc. It captures data in its current state and prevents data loss. that query is sent to the social media server and by the user themselves API link that retrieves all data officially from social media account (2), 15 and enables the user to perform analysis simultaneously. asynchronous interface (10) official data collection layer (A), which includes • How much of each tweet retrieved via that API connection (2) is related to the research topic? To assess relevance, please use a scale of 1 to 10 and a score of 5. Data includes ads, spam, and irrelevant content that are rated below a certain score. AI relevance filter that filters without entering the base (3), The slang and offensive content in those tweets were categorized as mild, moderate, and severe. Three-level insult analysis module (4) (Sociological anger 25 (In its measurement, it distinguishes everyday jargon from hate speech.) data cleaning and language processing layer (B), • It stores all the tweets it retrieves on the local disk and in CSV (Excel) format. SQLite database capable of exporting (9) (Academician's article 30 (can be submitted as an attachment.) 3D scatter plots, word clouds, and time series trend charts Drawing visualization module (12) (Ready for academic presentation and article) It produces images. 8 that user can chat about the data and ask questions data assistant (11) (User, with artificial intelligence relevance filter (3)) Ask Turkish questions about data by chatting through the data assistant (11) One might ask: "Which theme stands out the most in these tweets?" data storage and visualization layer (C), 5 • It captures accounts that repeat the exact same text and adds it to the total data. Bot / Troll detection module (6) which calculates the rate (For example, “What is this issue This answers the question, "How much of it is organic response, and how much is botnet?" by calculating the degree of polarization and toxicity level of that issue, 10 Trends that mathematically measure the probability of going viral in the next 48 hours. prediction module (7), that tweet used artificial intelligence to identify the city names in Türkiye A geographic map module that places colored pins on the map according to the type of event. (8) 15 including analysis and forecasting layer (D) It includes. The working principle of the system that is the subject of the invention: • From the user login interface (E), enter field information (Sociology, Psychology, etc.) and search They choose the concept they want or write in natural language, "I want to research this." 20 • AI-powered natural language search assistant (5) request for social media search Translates and expands Turkish word variants of the searched word (1) It produces. • API connection (2) sends this query to the social media's official server and It extracts tweets with information including text, media, likes, retweets, and date. 25 • Each tweet is sent to an AI relevance filter (3) and a relevance rating of 1 to 10 is given. A score is given. Submissions below 5 points are eliminated as spam / advertisement. • The remaining clean tweets are analyzed by the three-level insult analysis module (4) They are filtered and each is given a toxicity label (Clean / Mild / Moderate / Severe) It is labeled. 30 • All clean data is written to the SQLite database (9). • If the user wishes, the visualization module (12) can be used to create graphs and word clouds. It produces. 9 • Bot / Troll detection module (6) counts identical repeating messages and calculates their rate. and generates AI interpretation. • Trend prediction module (7), 48-hour viral with polarization + toxicity formula It provides a prediction. • The geographic map module (8) detects the cities in the tweet text using AI and Türkiye 5 It places it on the map. • User data collected via chat data assistant (11) with AI He speaks Turkish and asks questions. • Data in the SQLite database (9) is exported in CSV format and academic It can be included as an appendix to the article / thesis. 10 The invention is a computer-based scanner and a system that runs on a mobile device, Those who want to conduct research on Turkish social media data (e.g., X / Twitter) It can be used by everyone. Practical usage examples: • A doctoral student wrote a thesis on 'The Alevi-Bektaşi tradition in Turkish social life'. It aims to collect data on 'representation in the media'. The system in question is the subject of the invention. It sends him thousands of tweets, filters out the irrelevant ones, analyzes the remaining data, and prepares a report. It produces. • A communications academic wants to measure polarization during an election period. Meeting The system in question predicts which topics will go viral with a 48-hour trend forecast. He predicts. • A social psychologist measures hate speech in 'gender role' debates. He wants to analyze the hierarchical insults within the system, using everyday jargon. By separating it from hate speech, it provides measurable data on it. • An urban sociologist wants to study the geographical spread of protest movements. 25 The system in question extracts city names from Tweet texts in Türkiye. It is pinned to the map. • A research center is identifying bot operations on social media. He wants. The clone analysis module in the system that is the subject of the invention to be used for organic interaction. It separates the synthetic agenda. 30

Claims

REQUESTS 1. Academics, master's and doctoral students, and independent researchers. It is an AI-powered social media data collection and analysis system, and its features include: • Search 5 for the request that the user made from the login interface (E) natural language search assistant (5) which translates the query, word expander (1) that produces all variants of that query, that query is sent to the social media server and by the user themselves API that pulls all tweets officially from social media accounts connection (2) 10 official data collection layer (A), which includes • scoring each tweet retrieved via that API connection (2) between 1 and 10 and 5 Advertisements, spam, and irrelevant content that are rated below a certain score AI relevance filter (3), 15 The slang and offensive content in those tweets were categorized as mild, moderate, and severe. Three-level insult analysis module (4) which divides into three levels data cleaning and language processing layer (B), • SQLite database (9) which stores all the tweets that were taken on the local disk, 20 3D scatter plots, word clouds, and time series trend charts drawing visualization module (12) data storage and visualization layer (C), • It catches accounts that repeat the exact same text and adds 25 to the total data. Bot / Troll detection module (6) which calculates the ratio, by calculating the degree of polarization and toxicity level of that issue, Trends that mathematically measure the probability of going viral in the next 48 hours. prediction module (7), that tweet used artificial intelligence to identify the city names, and Türkiye 30 A geographic map module that places colored pins on the map according to the type of event. (8) including analysis and forecasting layer (D) It includes. 11 2. The system compliant with Request 1 is characterized by its ability to run in the background while data is being retrieved. It prevents interfaces from freezing and analyzes the user's data simultaneously. Official data collection layer (A) which includes an asynchronous interface (10) that enables it to do so It includes. 5 3. The system compliant with Request 1 is characterized by its ability to allow the user to chat about data and... Data storage and visualization including a data assistant (11) that enables asking questions It contains layer (C).