Textual patterns analyzer

The TPA system addresses computational inefficiencies in large-scale textual analysis by integrating hardware and software for efficient pattern recognition, enabling precise authorship identification through comprehensive analysis of complex linguistic patterns.

WO2026022801A1PCT designated stage Publication Date: 2026-01-29COPYLEAKS TECH LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2025/050398
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-21
Filing Date
2025-05-12
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing textual analysis methods face challenges in computational efficiency and scalability when processing large volumes of text, and they fail to systematically analyze complex interplays and co-occurrences of linguistic patterns, which are crucial for authorship identification.

Method used

A hardware-integrated solution with a scalable architecture that efficiently isolates unique linguistic patterns and performs comprehensive statistical validation of complex pattern combinations, using a Textual Patterns Analyzer (TPA) system for large-scale textual corpora analysis.

Benefits of technology

The TPA system provides precise and scalable authorship identification by analyzing elemental textual patterns across large datasets, ensuring computational efficiency and accuracy in identifying unique authorial fingerprints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050398_29012026_PF_FP_ABST
    Figure IL2025050398_29012026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for analyzing, identifying, and quantifying unique authorial elements in textual content, comprising: a) applying a Textual Patterns Analyzer (TPA) to at least two textual corpora to detect unique writing fingerprints within a text, b) identifying and analyzing the unique textual fingerprints and their different combinations in each corpus, c) applying the textual patterns analyzer to a new text, and d) performing data processing and outputting a detailed quantified statistical analysis of the writing elements that were found in the text based on the specific task, and determining their likelihood of appearing together in the same text, wherein the detailed statistical analysis indicates how likely these patterns and their different combinations are to appear, or not appear, in the specific author's writing, compared to other writers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TEXTUAL PATTERNS ANALYZER

[0002] Field of the Invention

[0003] The present invention relates to the field of textual analysis. More particularly, the invention relates to a method and system utilizing optimized computational hardware for analyzing, identifying, and quantifying unique characteristics in textual content, which may include but are not limited to authorial elements, stylistic features, or linguistic patterns.

[0004] Background for the Invention

[0005] Different types of textual data present unique linguistic and stylistic patterns, which may include, but are not limited to, syntactic structures, lexical choices, frequency of specific markers, or other linguistic indicators. These writing styles are characteristic of each human or non-human author, such as Artificial Intelligence (Al). In some cases, it is necessary to identify whether a particular piece of writing was written by a specific author, to determine from a group of authors which author is the most likely to have written a specific text, or whether a certain text was written by a human, or by Al (for example, by ChatGPT - a chatbot and virtual assistant developed based on large language models, that enables users to refine and steer a conversation towards a desired length, format, style, level of detail, and language). This capability may be important, for example, when there is a need to decide when textual content is computer-generated (for example, generated fake news), as well as in other tasks of authorship identification, such as identifying an anonymous author or a specific author, or detecting plagiarism.

[0006] Textual analysis for authorship attribution, authorship verification, and style characterization, traditionally employs different stylometric methods and techniques. These often rely on software-based analysis to quantify linguistic features such as lexical diversity, function word frequencies, sentence length distributions, and more. Machine learning models, such as Support Vector Machines (SVMs), Naive Bayes classifiers, Random Forests, and more recently, deep learning models like LSTM and Transformers, are also utilized to classify texts based on vectors derived from these features.

[0007] Current approaches have two main kinds of limitations:

[0008] 1. Computational efficiency and scalability, particularly required when processing extremely large volumes of text. Existing stylometric engines cannot exhaustively explore multilength patterns at the corpus scale within practical time limits.

[0009] 2. Existing engines focus primarily on individual, isolated linguistic markers, rather than systematically quantifying the complex interplay, combinations, and statistically significant co-occurrences of multiple, diverse patterns that can constitute a truly unique authorial "fingerprint". Specifically, many methods fail to systematically analyze or computationally handle the vast space of fine-grained, sub-structural textual patterns (such as elemental character sequences of varying lengths) and their differential frequencies across large corpora, which can be highly indicative of authorship or text source.

[0010] The present invention overcomes these challenges by introducing a hardware-integrated solution that efficiently isolates unique linguistic patterns, automates complex style differentiation, and achieves scalable, language-agnostic processing. Specifically, the novelty of the present invention lies in the synergistic integration of a specifically designed, scalable hardware-software architecture with computational techniques focused on the systematic identification and statistical validation of complex pattern combinations and co-occurrences, overcoming the limitations of conventional approaches in both computational efficiency and analytical depth. Addressing this scale is crucial, as the systematic analysis of all potential elemental patterns across multiple lengths and within large-scale textual corpora, as performed by the Textual Patterns Analyzer (TPA) system, leads to a combinatorial explosion of features, presenting a significant computational challenge that conventional methods struggle to address efficiently. It is therefore an object of the present invention to provide a method and a system for analyzing, identifying, and quantifying unique authorial elements in textual content, to estimate the probability that a specific analyzed text belongs or does not belong to a particular writer or group of writers. The TPA system is built on a layered hardware -software architecture where customized processing components handle segmentation and computational tasks. This setup enables concurrent processing across multiple nodes, minimizing bottlenecks and maximizing the system's analytical precision. In the context of this description, the term "writer" should be understood to refer to any entity capable of generating text, be it human or machine, such as Al.

[0011] It is another object of the present invention to provide a method and a system for analyzing, identifying, and quantifying unique authorial elements in textual content, which is able to efficiently process extremely large amounts of text and ensure scalability.

[0012] It is a further object of the present invention to provide a method and a system for analyzing, identifying, and quantifying unique authorial elements in textual content, which is language agnostic, i.e., can be applied to all languages, including programming languages.

[0013] Other objects and advantages of the invention will become apparent as the description proceeds.

[0014] Summary of the Invention

[0015] In one aspect, the invention relates to a system for analyzing, identifying, and quantifying unique authorial elements in textual content, comprising: a) Applying a Textual Patterns Analyzer (TPA), a novel system for the systematic examination of linguistic structures and frequencies across textual corpora, on at least two textual corpora to find unique writing fingerprints in a text, wherein the TPA systematically identifies elemental textual patterns (e.g., character sequences or other fundamental units) of varying lengths, analyzes their comparative frequency distributions between the corpora, and identifies patterns and combinations of patterns exhibiting statistically significant frequency divergence as unique writing fingerprints; b) Tokenizing the text into analyzable units, such as characters, words, phrases, or any structural elements relevant to the analytical task, which enables granular analysis; c) Categorizing each component by its frequency and distinctiveness within the author's corpus versus a control corpus, helps to isolate stylistic patterns unique to the author; d) The system applies multi-stage statistical analysis, beginning with relative frequency calculations for each unit of text, followed by significance testing to determine stylistic consistency within the author's corpus, applying these tests to both individual patterns and their statistically significant combinations and co-occurrences. By normalizing these results across varied text lengths, the system ensures accurate stylistic comparison, adjusting each statistical measure to account for structural variances between corpora; e) Applying the TPA system to a new text; and f) Compiling analytical results into a structured report by consolidating probability metrics, stylistic fingerprints, and statistical comparisons. This report, presented in a customizable format, outlines each fingerprint's relevance, likelihood ratios, and pattern combinations, providing a quantifiable measure of text distinctiveness. Wherein said detailed statistical analysis indicates how likely these patterns and their different combinations are to appear, or not appear, in the specific author's writing, compared to other writers. In this regard, "fingerprints" refer to measurable stylistic markers.

[0016] The system applies the TPA by first tokenizing the text, using high-throughput data pipelines within a purpose-built, hardware-optimized framework designed for intensive pattern discovery and comparison (which is necessary to handle the computationally intensive task of analyzing the differential frequencies of potentially billions of fine-grained textual patterns and their combinations across large datasets), into components such as words, phrases, and syntactical structures, then categorizing each component based on frequency and distinctiveness within the corpus. Specifically, this categorization involves analyzing the differential prevalence of finegrained sub-sequences across the author's corpus and a control corpus, considering multiple levels of granularity. A statistical analysis then calculates the normalized frequency of each component, adjusted to account for text length differences across corpora. This quantified output provides statistical indicators, such as likelihood ratios, which estimate the probability that certain patterns belong to a particular author compared to a control group. The system then compiles these results into a final report that quantifies the distinctiveness of an author's style.

[0017] In some embodiments, the system categorizes "fingerprints" by isolating distinctive patterns using a comparative frequency model, where each pattern's presence and unique sequence are logged against both the author's and control corpora. These fingerprints are ranked based on frequency and pattern consistency, allowing the system to create a probability profile that quantifies the distinctiveness of each stylistic element.;

[0018] In one embodiment, the TPA system includes an accessible interface, which may be an Application Programming Interface (API), a web-based interface, or any other protocol that enables integration with existing data processing and analytical platforms. This flexible interface supports seamless access to TPA's functionalities, allowing for versatile deployment and interaction within various software or hardware environments.

[0019] In some embodiments of the invention, the TPA system is language agnostic and can be applied to all languages, including programming languages.

[0020] In addition to that, the TPA system has an intuitive and interactive user interface that allows users to upload texts and view detailed analytics, including a graphical representation of the data.

[0021] According to some embodiments, users can use customizable filters in the TPA system that enable them to focus on specific types of patterns, styles, and writing signatures relevant to their specific needs. In some embodiments, the writing fingerprints are unique patterns or characteristics in writing, including one or more of the following: sequenced or unsequenced characters; recurring or non-recurring ideas; stylistic elements being unique to a specific writer or a group of writers; combinations of stylistic elements;

[0022] Within the TPA framework, the identification of 'recurring or non-recurring ideas' refers specifically to the detection of statistically significant, high-order combinations and cooccurrences of foundational textual patterns (such as specific sequences, lexical choices, and syntactic structures) that serve as measurable proxies for thematic consistency or shifts within the analyzed corpora.

[0023] In one aspect, the present invention relates to a computer-implemented method for analyzing, identifying, and quantifying unique authorial elements in textual content, comprising: applying a Textual Patterns Analyzer (TPA) to at least two textual corpuses to detect unique writing fingerprints within a text; identifying and analyzing the unique textual fingerprints and their different combinations in each corpus; applying the textual patterns analyzer to a new text; and performing data processing and outputting a detailed quantified statistical analysis of the writing elements that were found in the text based on the specific task, and determining their likelihood of appearing together in the same text, wherein said detailed statistical analysis indicates how likely these patterns and their different combinations are to appear, or not appear, in the specific author's writing, compared to other writers.

[0024] In yet another aspect, the present invention relates to a system for analyzing, identifying, and quantifying unique authorial elements in textual content, comprising: a processing module configured to execute a Textual Patterns Analyzer (TPA) on at least two textual corpora to detect unique writing fingerprints; a pattern recognition module configured to identify and analyze the unique textual fingerprints and their various combinations in the corpora; and a data processing module configured to apply the textual patterns analyzer to a new text and to output a detailed, quantified statistical analysis of the writing elements found in the text based on a specific task, wherein said statistical analysis indicates the likelihood that these patterns and their combinations will appear, or not appear, in the specific author's writing compared to other writers.

[0025] Brief Description of the Drawings

[0026] The above and other characteristics and advantages of the invention will be better understood through the following illustrative and non-limitative detailed description of preferred embodiments thereof, with reference to the appended drawings, wherein:

[0027] Figs. 1A and IB show an example of applying the TPA system on a corpus of Al writings and another control corpus of human writing;

[0028] Fig. 2 shows a flowchart of the text processing performed by the TPA system, according to an embodiment of the invention;

[0029] Fig. 3 illustrates the data preprocessing module, which is responsible for preparing raw textual data for subsequent analysis within the TPA system, according to an embodiment of the present invention;

[0030] Fig. 4 presents the pattern recognition module, which identifies and quantifies unique textual "fingerprints" or patterns that are characteristic of an author's style, according to an embodiment of the present invention;

[0031] Fig. 5 outlines the distributed processing module, which enhances the TPA system's efficiency through a cloud-based, hardware-optimized infrastructure, according to an embodiment of the present invention; and Fig. 6 shows the TPA system architecture, according to an embodiment of the present invention. It illustrates the integration of hardware and software components in the TPA, designed for efficient and large-scale text analysis.

[0032] Detailed Description of the Invention

[0033] The present invention provides a Textual Patterns Analyzer (TPA), incorporating a method and a system designed to analyze, identify, and quantify unique authorial elements in textual content. The system comprises processing elements associated with memory and data storage apparatus adapted to perform the method of the invention. The TPA system utilizes processing components carefully selected and potentially customized (which may include multi-core processors, GPUs, and potentially reconfigurable hardware) integrated within a co-designed hardware-software architecture optimized for data-intensive textual pattern and combination analysis. This optimization is necessary because of the significant computational load imposed by the TPA's approach of exhaustively analyzing the differential frequencies of elemental textual patterns across varying lengths within extensive corpora. The system architecture is adaptable to various processing environments, including cloud-based or decentralized hardware setups. Each component is equipped with adequate memory and bandwidth to manage high data throughput, enabling efficient analysis of extensive text corpora. The system dynamically manages processing resources based on computational demands and can use various high-speed data transfer technologies to ensure efficient communication between nodes and storage units. This adaptable hardware configuration allows the TPA to operate effectively in a range of high-demand data processing environments.

[0034] Cloud-based clusters: The TPA operates on a cloud-based distributed infrastructure designed to handle extensive data processing needs. This infrastructure includes hardware-optimized components to balance data loads dynamically across multiple nodes, each equipped with specialized processing units to support parallel processing and efficient data segmentation. The system employs high-speed, redundant storage with low-latency access pathways, ensuring swift retrieval of prioritized data and maintaining optimized performance and responsiveness across all nodes. Processing nodes utilize specialized, distributed data structures optimized for concurrent updates within their memory hierarchies to efficiently manage the potentially vast amount of frequency data generated during the analysis of elemental patterns.

[0035] The TPA's dynamic scaling protocol activates additional nodes or reallocates resources as data demands fluctuate. By redistributing processing power and storage access in real-time, the system maintains workload equilibrium, ensuring low latency even during peak processing periods. As demand increases, additional virtual machines or containers are activated to distribute tasks effectively across the infrastructure. Each virtual machine is assigned a specific data subset to analyze, ensuring balanced processing loads. This dynamic scaling approach maintains low latency and consistent performance, enabling the TPA to operate reliably under varying workloads.

[0036] Data storage: The TPA's distributed storage system incorporates high-speed storage designed to support rapid data access and maintain high availability. Storage components connect to processing nodes through low-latency data transfer pathways, allowing for minimized retrieval times and enhanced data throughput. This rapid access is crucial for retrieving corpus data and for reading / writing the large volumes of intermediate pattern frequency data generated by the PRM. This configuration employs techniques to ensure redundancy and resilience, optimizing data availability and consistency for uninterrupted processing. Such a hardware-oriented setup enables the TPA to access extensive text corpora efficiently, meeting the demands of frequent data queries and large-scale pattern recognition tasks.

[0037] Local, secured query interface: The TPA can be queried and accessed from a user interface connected to the underlying infrastructure via a secure network connection, which may be local, cloud-based, or a hybrid environment, providing flexibility in deployment. This connection enables the users to submit queries and receive real-time results without the need for any unique local hardware. The TPA system operates from data centers that are specifically optimized for high-performance computational workloads. The data is encrypted at rest and in transit to ensure its privacy and integrity, and adhere to relevant compliance standards and regulations.

[0038] Operation method: The TPA system initiates by isolating a writer's corpus, followed by baseline comparison against a control corpus. Elemental textual patterns, such as character subsequences of varying lengths, are identified based on their statistically significant differential frequency between the corpora. These identified patterns are then segmented into component units and normalized for consistent comparison. Final probability metrics are calculated based on these normalized units, forming the basis for the distinctiveness score. The "distinctiveness score" refers to a probabilistic metric for authorship likelihood. This initial comparison establishes unique 'fingerprints' in the author's writing, identifying distinctive patterns such as sequenced or unsequenced characters, recurring or non-recurring ideas, and stylistic elements. The TPA quantifies these fingerprints by segmenting each text into analyzable units, which may include characters, words, phrases, or other relevant structural elements, calculating the relative frequency of each element to account for corpus variations, and applying analytical measures— such as statistical tests or computational models— to assess the distinctiveness of identified patterns.

[0039] When applied to a new text, the TPA compares its features against the established patterns in the database, evaluating the frequency and structure of elements relative to the author's corpus and the control corpus. The system calculates likelihood ratios and other statistical values to determine whether specific patterns are distinctively associated with the author. Additionally, the TPA examines k-combinations of elements, identifying recurrent combinations that suggest unique stylistic tendencies. This process counts the absolute occurrences of each pattern and performs additional analysis to confirm their statistical relevance, providing a quantitative and statistically valid measure of the text's similarity to the author's style.

[0040] The final output is an aggregate score or metric that represents the distinctiveness of the text by analyzing patterns and recurring combinations. This metric can be customized to suit specific application requirements, providing a measure of text similarity or uniqueness. This score, along with the stored results of each analysis, serves as a comprehensive measure of the author's stylistic fingerprint, aiding in precise authorship determination.

[0041] The analysis by the TPA system can give an estimation regarding the probability that a specific analyzed text belongs to a particular writer or group of writers or regarding the probability that a particular text was written by a human or by artificial intelligence (for example, by ChatGPT) with or without presenting the intermediate calculation results according to which the decision was made, i.e., the parameters and factors themselves.

[0042] The TPA system can be accessed through an Application Programming Interface (API) for seamless integration with existing text analysis platforms or deployed as a standalone system for in-depth internal analysis.

[0043] The TPA system can process extremely large amounts of text efficiently, ensuring scalability.

[0044] It is also language agnostic: it can be applied to all languages, including programming languages. This is achieved by primarily operating on fundamental units like character sequences, token types, and / or generalized syntactic patterns, thus minimizing reliance on language-specific dictionaries, parsers, or semantic models.

[0045] Examples and illustrations of the TPA system

[0046] Figs. 1A and IB show an example of applying the TPA system on a corpus of artificial intelligence writings and another control corpus of human writing. The example shows that the TPA system could accurately detect artificial intelligence unique "fingerprints" (sequences of characters or ideas that repeat or do not repeat) and quantify them statistically. In Fig. 1A, block 101 shows an example sentence that contains text patterns occurring 150 times more frequently in Al-generated text. In Fig. IB, block 102 presents an example document featuring combinations of text patterns that are 330 times more prevalent in Al-generated text. Fig. 2 shows a flowchart of the text processing performed by the TPA system, according to an embodiment of the invention. In the first step, 201 (which is a learning phase), the TPA is applied to at least two textual corpora. In the next step, 202, the TPA system identifies and analyzes the unique textual fingerprints in each corpus. In the next step, 203, the TPA is applied to a new text. At step 204, the TPA system performs data processing and outputs quantified analysis results of the writing elements found in the new text, based on a predefined specific task (a specific author, for example), and provides a quantified analysis of their joint appearances in the same text.

[0047] Fig. 3 shows a flowchart of the Data Preprocessing Module (DPPM) 300, according to an embodiment of the invention. DPPM 300 prepares raw text data for analysis by the TPA system. The module is responsible for cleaning, segmenting, and standardizing the data, ensuring it is in an optimized format for further analysis. DPPM 300 may involve the following procedures:

[0048] Data Segmentation (step 301): At this step, the TPA separates structured data (e.g., metadata such as author and date) from unstructured data (e.g., raw text content). This enables the system to handle metadata and content differently, as metadata is often used for categorization or filtering, while unstructured content requires deeper linguistic analysis;

[0049] Noise Filtering and Cleansing (step 302): At this step, the TPA applies noise filters to remove irrelevant content, such as formatting characters and other non-textual elements. This reduces computational load and avoids skewing the analysis by removing elements that do not contribute meaningfully to pattern recognition;

[0050] Language Detection and Routing (step 303): At this step, the TPA detects the language of the text to apply specific language-based preprocessing steps. The TPA is languageagnostic, so detecting and handling multilingual content appropriately ensures consistency in analysis across different languages;

[0051] Tokenization (step 304): Tokenizing the text into smaller units (for example: characters, words, phrases, sentences) to allow the system to analyze text at a granular level, making it easier to identify unique patterns, syntactic structures, and writing signatures; Final Data Quality Check and Storage (steps 305, 306): The TPA performs a quality check (305) to ensure data consistency, completeness, and integrity after preprocessing. Then it stores the cleaned and structured data in a temporary storage location, ready for input into the Pattern Recognition Module.

[0052] Fig. 4 illustrates the Pattern Recognition Module (PRM) 400, according to an embodiment of the invention. PRM 400 identifies unique authorial elements or "fingerprints" in the text. This module performs in-depth syntactic, lexical, and frequency analysis to recognize patterns unique to the author's style. PRM 400 may involve the following procedures:

[0053] Identification of Unique Sequences (step 401): The TPA detects and records elemental textual patterns (e.g., character-level sub-sequences) across a predefined range of lengths or complexities. It analyzes the frequency of these patterns within the target and control corpora to detect those whose occurrence rates show statistically significant divergence, signifying patterns characteristic of the target source's style. This captures stylistic patterns, such as frequent use of certain phrases or expressions, as well as sentence structures that differentiate authors' writing habits;

[0054] Syntactic Structure Analysis (step 402): The TPA analyzes syntactic structures (e.g., sentence composition, clause complexity) to identify an author's characteristic sentence patterns. This step focuses on structural tendencies, such as preferred sentence lengths, types (e.g., simple vs. complex), and punctuation styles, helping to build a syntactic "signature" for the author;

[0055] Lexical Choice and Vocabulary Profiling (step 403): At this step, the TPA profiles the vocabulary, focusing on word choice, frequency, and word type (e.g., nouns, adjectives, specific jargon) for each author or author group;

[0056] Frequency Analysis and Pattern Counting (step 404): Based on the information so far, the TPA counts the occurrences of identified patterns across the author's corpus and the control corpus, and quantifies them for comparative assessments that highlight the distinctiveness of an author's style versus the control group. It then normalizes the frequencies of identified patterns based on the corpus size, in order to create a balanced score that accounts for varying lengths of text and corpora sizes between authors;

[0057] Statistical Significance Testing (step 405): Now the TPA conducts statistical tests (e.g., chi-square tests, t-tests, z-scores) on the frequency of patterns to determine which elements are statistically significant in distinguishing one author from another. This provides quantitative evidence of the uniqueness of detected patterns. These tests may incorporate methods robust to large numbers of simultaneous comparisons, such as adjustments based on False Discovery Rate (FDR) control, to ensure high confidence in identifying differentiating patterns from the multitude analyzed.;

[0058] Combination and Co-occurrence Analysis (step 406): Distinctively, TPA moves beyond analyzing features in isolation. It systematically identifies patterns (including sequences, syntactic structures, and lexical choices identified in previous steps) that frequently cooccur within the author's texts compared to the control corpus. This involves methods such as applying multivariate statistical tests, analyzing high-order combinations, and using a specific graph-based approach, to assess the likelihood and statistical significance of these complex stylistic signatures appearing together. This allows the system to detect subtle and complex writing habits often missed by conventional methods focused on individual feature frequencies. This may involve constructing graph-based representations where nodes represent statistically significant elemental patterns and edges signify validated co-occurrence within text segments, allowing the system to identify and score complex, higher-order stylistic signatures characteristic of the author or source.

[0059] Fig. 5 displays the Distributed Processing Module (DPM) 500, according to an embodiment of the invention. DPM 500 segments data across nodes through a multi-stage approach: data is first parsed at the primary node, then relayed to secondary nodes for parallel processing. This segmentation hierarchy enables efficient analysis and ensures that data is processed concurrently, improving system response times and accuracy. DPM 500 is designed to enhance the TPA system's efficiency by leveraging distributed, hardware-optimized cloud resources. This enables large-scale, parallel processing of text data, reducing computational time and ensuring scalability. This architecture is specifically tailored to handle the unique computational demands arising from the combinatorial complexity of TPA's exhaustive, differential frequency analysis of fine-grained textual patterns across multiple lengths and large datasets, enabling parallel processing and ensuring scalability beyond that typically achieved by general-purpose distributed systems applied to text analysis. It enables the TPA to distribute tasks across multiple hardware nodes, ensuring balanced load management and minimizing latency. DPM 500 may involve the following procedures:

[0060] Task Segmentation and Allocation (step 501): This step breaks down the data processing tasks (e.g., pattern recognition, statistical analysis) into smaller, manageable units and allocates them to individual nodes across the cloud infrastructure. This allows efficient parallel processing of the data;

[0061] Dynamic Resource Allocation with GPU Optimization (step 502): The TPA dynamically allocates resources (e.g., CPU, GPU) based on the computational demands of each processing task, optimizing GPU usage for compute-intensive operations;

[0062] Data Sharding and Distributed Storage Access (step 503): The TPA splits the text corpus across distributed storage nodes to ensure high data accessibility and minimize data retrieval time. The sharding of the data ensures that each processing node can access the specific data it requires without waiting for a central server. This reduces latency and enhances the system's performance for large data sets;

[0063] Load Balancing Across Cloud Nodes (step 504): At this step, the TPA implements a loadbalancing mechanism that monitors the computational load of each node in real time and redistributes tasks to prevent node overloads. This ensures that no single node becomes a bottleneck, while maximizing the throughput of the distributed analytics engine;

[0064] Parallel Processing with Memory Management (step 505): The TPA utilizes shared memory across distributed nodes for frequently accessed data, reducing redundant data transfer and caching essential information within local memory on each node; High-Speed Networking for Inter-Node Communication (step 506): The TPA uses highspeed networking protocols (e.g., InfiniBand or Ethernet with low-latency configurations) to optimize data transfer between nodes. This is critical for the distributed environment where nodes must communicate frequently;

[0065] Automated Scaling and Resource Reallocation (step 507): The TPA integrates an automated scaling feature that monitors real-time demand and dynamically allocates additional hardware resources, such as additional GPU nodes or CPU cores, when the system experiences high demand. This enables the system to respond to workload spikes without manual intervention;

[0066] Distributed Output Compilation and User Access (step 508): The TPA compiles results from distributed nodes into a coherent output format that is stored and accessible through a secure API interface. This allows the users to query the final analysis.

[0067] Fig. 6 illustrates the TPA System Architecture 600, according to an embodiment of the invention, which integrates preprocessing and real-time analysis within a scalable, hardware- optimized infrastructure. This architecture employs high-speed interconnects for efficient data flow and dynamically allocates processing resources to balance computational loads. Advanced storage systems with redundancy and caching mechanisms ensure rapid data access.

[0068] Fig. 6 schematically illustrates the system architecture for processing, analyzing, and reporting data in an optimized and resource-efficient manner. The system begins with a User Interface (Ul) 601, designed to provide an intuitive dashboard with real-time feedback capabilities and customizable filters, allowing users to interact with the system seamlessly. Ul 601 communicates with the system through an API Gateway 602, which ensures secure data transmission and handles user queries efficiently. The API provides specific endpoints for tasks such as corpus submission (accepting formats like plain text, archives), configuration of analysis parameters (e.g., defining target vs. control corpora, setting sensitivity thresholds for pattern differentiation), job status monitoring, and retrieval of results, typically formatted in a structured manner like JSON. Secure authentication mechanisms are employed for API access. According to an embodiment of the invention, the customizable filters are implemented as configurable software modules that enable users to fine-tune analysis parameters. For instance, these filters can be adjusted to set specific thresholds for the frequency, weight, or statistical significance of textual features, and they can be configured to include or exclude particular types of data. Such data may include syntactic patterns (e.g., may refer to the structural aspects of sentences, such as common parts-of-speech sequences or grammatical constructions, semantic markers (e.g., indicators of meaning within the text, such as the presence of specific keywords, sentiment indicators, or thematic elements) or author-specific stylistic elements (e.g., these may consist of distinctive features that characterize a writer's style, including unique vocabulary choices, punctuation patterns, idiomatic expressions, etc.).

[0069] According to some embodiments of the invention, these customizable filters can be integrated into one or more processing components of the TPA system, where they dynamically adjust the parameters used by the TPA algorithm during processing. Alternatively, they may reside within the user interface module, allowing users to directly set and modify filter criteria, with the configured settings subsequently applied during text analysis. In some embodiments, the filters are distributed between the user side and the system side; preliminary filter settings are processed on the user side (e.g., via a web interface or API), while further refinement and application occur within the system. This flexible placement may enhance responsiveness, optimize processing load, and facilitate ease of configuration.

[0070] Data from the API Gateway 602 is directed towards two parallel Processing Nodes— Processing Node 1 and Processing Node 2— connected via a High-Speed Interconnect to enable rapid data exchange and synchronization. Processing Node 1 is equipped with a CPU and GPU to handle tasks related to pattern recognition and data tokenization, essential for preprocessing and structuring raw input data. In parallel, Processing Node 2, also utilizing a CPU and GPU, performs statistical analysis and combination analysis, deriving meaningful insights from the preprocessed data. The system includes a Dynamic Load Balancer 603, which is responsible for resource allocation and task distribution between the processing nodes. This component optimizes system performance by dynamically adjusting workloads based on resource availability and processing demands. Data processed through the nodes is temporarily stored in Distributed SSD Storage 604, configured with RAID configurations for redundancy and speed, coupled with caching controllers and NVMe technology to ensure fast data retrieval and write operations.

[0071] Finally, the system culminates in the Final Report Compilation module 605, where data from the distributed storage is aggregated and analyzed to produce comprehensive reports. This module generates unified score likelihood ratios, which provide users with probabilistic insights and decision-support metrics. The integrated workflow depicted in Fig. 6 demonstrates a highly efficient system architecture designed for robust data processing, real-time feedback, and accurate reporting.

[0072] Fig. 7 illustrates a possible hardware configuration 700 of the TPA system in a simplified block diagram, according to an embodiment of the invention. The diagram shows key hardware components and their interconnections that facilitate efficient, automated textual analysis through computational methods. A user's terminal device 701 serves as the interface for submitting text analysis requests, which can be transmitted to the TPA system via any suitable data communication channel.

[0073] The TPA system may incorporate a dedicated processing module 702 that may integrate multiple specialized subsystems, as follows:

[0074] The Data Preprocessing Module (DPPM) 300 within the unit is responsible for executing a series of algorithms to clean, segment, and normalize raw text data. This may include performing tokenization, stop word removal, character encoding conversion, and other preprocessing operations that standardize diverse textual inputs into a consistent format suitable for further analysis; Following preprocessing, the Pattern Recognition Module (PRM 400) may apply advanced natural language processing techniques, including statistical modeling, syntactic parsing, and feature extraction, to identify and quantify unique textual fingerprints and stylistic patterns inherent in the data. PRM 400 may utilize both rulebased and machine learning algorithms to perform detailed lexical and syntactic analysis, thereby generating a robust digital representation of the text; and

[0075] The Distributed Processing Module (DPM 500) may leverage a distributed computing architecture to manage large-scale text analysis. DPM 500 may employ parallel processing techniques, such as multi-threading, asynchronous processing, and load balancing across multiple processors, to efficiently handle extensive data sets and optimize computational resources. Data exchanges between these subsystems can be facilitated via high-speed internal data buses and managed by a central control unit that oversees synchronization, error detection, and correction.

[0076] The TPA features an intuitive, interactive user interface that allows users to upload and analyze texts with ease. The interface provides detailed, graphical representations of the data, showcasing patterns, frequencies, and stylistic elements detected in the analysis. Users can leverage customizable filters to narrow down their focus to specific patterns, stylistic elements, or writing signatures relevant to their analysis goals. Additionally, the TPA's dashboard offers real-time feedback, allowing users to refine parameters, adjust statistical thresholds, and toggle between different views— such as lexical profiles, syntactic structures, and co-occurrence analysis— to gain deeper insights into authorial style. This flexible interface is designed for a user-friendly experience while ensuring access to robust, detailed analytics tailored to individual needs. As various embodiments and examples have been described and illustrated, it should be understood that variations will be apparent to one skilled in the art without departing from the principles herein. Accordingly, the invention is not limited to the specific embodiments described and illustrated in the drawings.

Claims

CLAIMS1. A computer-implemented method for analyzing, identifying, and quantifying unique authorial elements in textual content, comprising: a) applying a Textual Patterns Analyzer (TPA) to at least two textual corpora to detect unique writing fingerprints within a text; b) identifying and analyzing the unique textual fingerprints and their different combinations in each corpus; c) applying the textual patterns analyzer to a new text; and d) performing data processing and outputting a detailed quantified statistical analysis of the writing elements that were found in the text based on the specific task, and determining their likelihood of appearing together in the same text, wherein said detailed statistical analysis indicates how likely these patterns and their different combinations are to appear, or not appear, in the specific author's writing, compared to other writers.

2. A computer-implemented method according to claim 1, wherein the analysis is performed using a system with an accessible interface, which may include an Application Programming Interface (API), web interface, or other protocols to enable integration with text analysis systems.

3. A computer-implemented method according to claim 1, wherein the TPA is language agnostic and can be applied to all languages, including programming languages.

4. A computer-implemented method according to claim 1, wherein the textual patterns analyzer has an intuitive and interactive user interface that allows users to upload texts and view detailed analytics, including a graphical representation of the data.

5. A computer-implemented method according to claim 1, further comprising enabling a user to apply one or more customizable filters within the TPA, wherein the customizablefilters are configured to focus the analysis on specific types of patterns, styles, or writing signatures relevant to the user's specific needs.

6. A computer-implemented method according to claim 1, wherein the writing fingerprints are unique patterns or characteristics in writing, including one or more of the following:- sequenced or unsequenced characters;- recurring or non-recurring ideas;- stylistic elements being unique to a specific writer or a group of writers.

7. A system for analyzing, identifying, and quantifying unique authorial elements in textual content, comprising: a) a processing module configured to execute a Textual Patterns Analyzer (TPA) on at least two textual corpora to detect unique writing fingerprints; b) a pattern recognition module configured to identify and analyze the unique textual fingerprints and their various combinations in the corpora; and c) a data processing module configured to apply the textual patterns analyzer to a new text and to output a detailed, quantified statistical analysis of the writing elements found in the text based on a specific task, wherein said statistical analysis indicates the likelihood that these patterns and their combinations will appear, or not appear, in the specific author's writing compared to other writers.

8. The system of claim 7, further comprising an accessible interface configured to receive text analysis requests, wherein the interface comprises an Application Programming Interface (API), a web interface, or other communication protocols enabling integration with text analysis systems.

9. The system of claim 7, further comprising an intuitive and interactive user interface configured to allow users to upload texts and to view detailed analytics, including graphical representations of the data.

Citation Information

Patent Citations

  • Variables and method for authorship attribution

    US20180101518A1