System and method for generating structured information through text based analysis
The system uses a generative large language model to analyze user-generated text data, addressing the challenge of unreliable external data sources and privacy restrictions, enabling accurate user insight generation and personalized services.
Patent Information
- Application Number
- PCT/US2025/031049
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-05-27
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional software platforms struggle to effectively analyze and utilize free text data from users to generate structured information for improving user experiences and platform functionality due to unreliable external data sources and data privacy restrictions, making it difficult to tailor services and maintain user engagement.
A system and method utilizing a generative large language model (LLM) to analyze, summarize, and classify user-generated text data, combined with cross-domain integration techniques, to generate structured information from first-party data, including user keystrokes and metadata, providing actionable insights for platform administrators.
Enables accurate prediction of user needs and behavior, enhances user satisfaction, and maintains scalability across large user populations by extracting reliable and quantifiable information from user-generated text, facilitating personalized services and improved platform performance.
Smart Images

Figure US2025031049_05022026_PF_FP_ABST
Abstract
Description
Attorney Docket No. TAPP-OOIWOSYSTEM AND METHOD FOR GENERATING STRUCTURED INFORMATION THROUGH TEXT BASED ANALYSISCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 651,676 entitled “SYSTEM AND METHOD FOR GENERATING STRUCTURED INFORMATION THROUGH TEXT BASED ANALYSIS,” filed on May 24, 2024, the content of which is incorporated by reference herein in its entiretyTECHNICAL FIELD
[0002] The present disclosure generally relates to systems and methods for generating structured information based on text analysis. More particularly the disclosure relates to systems and methods for multi-modal data integration, user behavior determination, user behavior prediction, customer intelligence processing, and / or adaptive computational resource management.BACKGROUND
[0003] Software platforms can benefit from user data to tailor and customize the software platform experience for particular user segments and applications. Existing sources of user data can include receiving data from external sources or networks, which may be unreliable or provide inaccurate information. Furthermore, desktop and mobile ecosystems are changing to make this data unavailable, or much more difficult to source, due to restrictions on data privacy and data sharing. Thus, it can be challenging to source information from a particular user base that can be used to improve the overall experience and output of the software platform as a whole.
[0004] The foregoing discussion, including the description of motivations for some embodiments of the invention, is intended to assist the reader in understanding the present disclosure, is not admitted to be prior art, and does not in any way limit the scope of any of the claims.1IPTS / 1289821 57.1Attorney Docket No. TAPP-001WOSUMMARY
[0005] A system and method for generating structured information including user properties is disclosed. The method can include a computer-implemented method.
[0006] The computer-implemented method can include receiving, via a server, text data collected from users. The method can include generating, using the server, identifiers associated with the text data and associated with the users. The method can include receiving a prompt having instructions for a large language model (LLM) to determine structured information including user properties based on the text data. The method can include transmitting the prompt, the text data, and identifiers to the LLM. The method can include determining, using the LLM, the structured information based on the prompt, the text data and identifiers. The method can include transmitting, using an aggregation component, the structured information and the identifiers to an administrator.
[0007] Various embodiments of the method can include one or more of the following steps.
[0008] In some embodiments, the method can include prior to receiving the text data, automatically recording, using a data gathering application, user keystrokes to generate the text data associated with the users. In some examples, the method can include transmitting, using the server, the text data and identifiers to a text store for storage. The prompt can include instructions for the LLM to analyze, summarize, classify and assess the text data to determine the structured information. The prompt can include a pre-defined prompt. The method can include extracting the prompt from an LLM prompt store. The method can include generating the prompt instructions for a large language model (LLM) to determine structured information based on the text data.
[0009] The system for generating structured information can include a processor. The system can include a memory storing instructions that, when executed by the processor, configure the system to: receive, via a server, text data collected from users, generate, using the server, identifiers associated with the text data and associated with the users, receive a prompt having instructions for a large language model (LLM) to determine structured information having user properties based on the text data, transmit the prompt, the text data, and identifiers to the LLM, determine, using the LLM, structured information based on the prompt, the text data and identifiers, and transmit, using an aggregation component, the structured information and the identifiers to an administrator.
[0010] Various embodiments of the system can include one or more of the following features.2IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0011] In some embodiments, the system can include a data gathering application that automatically records user keystrokes to generate the text data associated with the users. In some examples, the system can include a text store that stores the text data and identifiers. The prompt can include instructions for the LLM to analyze, summarize, classify and assess the text data to determine the structured information. The prompt can include a pre-defined prompt. The system can include an LLM prompt that stores that stores the prompt.
[0012] The computer-implemented method, can include receiving, via a data collection component, text data and user actions collected from users. The method can include transmitting the text data and user actions to a text assessment component and a behavioral assessment component, where the text assessment component and the behavioral assessment component each include respective large language models (LLMs). The method cam include generating, using the text assessment component, a prompt based on text characteristics of the text data, the prompt including instructions for a cross-domain assessment component. The method can include determining, using the behavioral assessment component, user behavioral patterns based on the user actions. The method can include determining, using the cross-domain assessment component, structured information including properties of the users based on the prompt, text data, and user interaction sequences.
[0013] Various embodiments of the method can include one or more of the following steps.
[0014] In some embodiments, the method can include storing, via a storage and preprocessing component, the text data. In some examples, generating, using the text assessment component, the prompt can include generating a parameterized prompt. Generating, using the text assessment component, the parameterized prompt can include initializing a prompt template to generate the parameterized prompt. Determining, using the behavioral assessment component, user behavioral patterns can include encoding user action sequences. Determining, using the behavioral assessment component, user behavioral patterns can include encoding temporal patterns between the user action sequences. Determining, using the cross-domain assessment component, structured information including properties of the users based on the prompt, text data, and user interaction sequences can include aligning user sequences using dynamic time warping.3IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying figures, which are included as part of the present specification, illustrate the presently preferred embodiments and together with the general description given above and the detailed description of the preferred embodiments given below serve to explain and teach the principles described herein. Furthermore, like reference numbers refer to similar or the same components within the figures.
[0016] FIG. 1 illustrates a system for generating structured information through text based analysis, according to some embodiments.
[0017] FIG. 2A illustrates a first exemplary prompt, according to some embodiments.
[0018] FIG. 2B illustrates a second exemplary prompt, according to some embodiments.
[0019] FIG. 3 illustrates exemplary user interface for the system for generating structured information through text based analysis, according to some embodiments.
[0020] FIG. 4 illustrates a method for generating structured information based on text analysis, according to some embodiments.
[0021] FIG. 5 illustrates a system for generating structured information, according to some embodiments.
[0022] FIG. 6 illustrates a system for generating structured information through text based analysis, according to some embodiments.
[0023] FIG. 7A illustrates a flowchart of a method for generating a parameterized prompt, according to some embodiments.
[0024] FIG. 7B illustrates exemplary pseudocode for the method of FIG. 7 A, according to some embodiments.
[0025] FIG. 8A illustrates a flowchart of a method for extracting attention signals, according to some embodiments.
[0026] FIG. 8B illustrates exemplary pseudocode for the method of FIG. 8 A, according to some embodiments.
[0027] FIG. 9A illustrates a flowchart of a method for processing long-form text that maintains context across window boundaries, according to some embodiments.
[0028] FIG. 9B illustrates exemplary pseudocode for the method of FIG. 9 A, according to some embodiments.
[0029] FIG. 10A illustrates a flowchart of a method for encoding user interaction sequences, according to some embodiments.4IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0030] FIG. 10B illustrates exemplary pseudocode for the method of FIG. 10 A, according to some embodiments.
[0031] FIG. 11 A illustrates a flowchart 1100 of a method for self-supervised pre-training, according to some embodiments.
[0032] FIG. 1 IB illustrates exemplary pseudocode for the method of FIG. 11 A, according to some embodiments.
[0033] FIG. 12A illustrates a flowchart of a method for identifying periodic patterns, according to some embodiments.
[0034] FIG. 12B illustrates exemplary pseudocode for the method of FIG. 12A, according to some embodiments.
[0035] FIG. 13 A illustrates a flowchart of a method for aligning user sequences using dynamic time warping, according to some embodiments.
[0036] FIG. 13B illustrates exemplary pseudocode for the method of FIG. 13A, according to some embodiments.
[0037] FIG. 14A illustrates a flowchart of a method for creating cross domain mapping, according to some embodiments.
[0038] FIG. 14B illustrates exemplary pseudocode for the method of FIG. 14A, according to some embodiments.
[0039] FIG. 15A illustrates a flowchart of a method for combining predictions using a Bayesian ensemble, according to some embodiments.
[0040] FIG. 15B illustrates exemplary pseudocode for the method of FIG. 15 A, according to some embodiments.
[0041] FIG. 16A illustrates a flowchart of a method for optimizing multiple system objectives, according to some embodiments.
[0042] FIG. 16B illustrates exemplary pseudocode for the method of FIG. 16 A, according to some embodiments.
[0043] FIG. 17A illustrates a flowchart 1700 of a method for optimizing batch parameters, according to some embodiments.
[0044] FIG. 17B illustrates exemplary pseudocode for the method of FIG. 17 A, according to some embodiments.
[0045] FIG. 18A illustrates a flowchart of a method for assessing feedback signals, according to some embodiments.5IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0046] FIG. 18B illustrates exemplary pseudocode for the method of FIG. 18 A, according to some embodiments.
[0047] FIG. 19A illustrates a flowchart of a method for applying a timed data transformation, according to some embodiments.
[0048] FIG. 19B illustrates exemplary pseudocode for the method of FIG. 19 A, according to some embodiments.
[0049] FIG. 20A illustrates a flowchart of a method for detecting user behavior, according to some embodiments.
[0050] FIG. 20B illustrates exemplary pseudocode for the method of FIG. 20A, according to some embodiments.
[0051] FIG. 21A illustrates a flowchart of a method for clustering and prioritizing user issues, according to some embodiments.
[0052] FIG. 21B illustrates exemplary pseudocode for the method of FIG. 21A, according to some embodiments.
[0053] FIG. 22A illustrates a flowchart 2200 of a method for optimizing support resource allocation, according to some embodiments.
[0054] FIG. 22B illustrates exemplary pseudocode for the method of FIG. 22A, according to some embodiments.
[0055] FIG. 23 A illustrates a flowchart of a method for building a churn prediction model, according to some embodiments.
[0056] FIG. 23B illustrates exemplary pseudocode for the method of FIG. 23 A, according to some embodiments.
[0057] FIG. 24A illustrates a flowchart of a method for identifying causal factors, according to some embodiments.
[0058] FIG. 24B illustrates exemplary pseudocode for the method of FIG. 24A, according to some embodiments.
[0059] FIG. 25A illustrates a flowchart of a method for optimizing churn interventions, according to some embodiments.
[0060] FIG. 25B illustrates exemplary pseudocode for the method of FIG. 25 A, according to some embodiments.
[0061] FIG. 26 A illustrates a flowchart of a method for detecting a purchase intent of a user, according to some embodiments.6IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0062] FIG. 26B illustrates exemplary pseudocode for the method of FIG. 26A, according to some embodiments.
[0063] FIG. 27A illustrates a flowchart of a method for identifying user needs, according to some embodiments.
[0064] FIG. 27B illustrates exemplary pseudocode for the method of FIG. 27A, according to some embodiments.
[0065] FIG. 28A illustrates a flowchart of a method for optimizing user engagement timing, according to some embodiments.
[0066] FIG. 28B illustrates exemplary pseudocode for the method of FIG. 28 A, according to some embodiments.
[0067] FIG. 29A illustrates a flowchart of a method for extracting high value patterns, according to some embodiments.
[0068] FIG. 29B illustrates exemplary pseudocode for the method of FIG. 29A, according to some embodiments.
[0069] FIG. 30A illustrates a flowchart of a method for building similarity models, according to some embodiments.
[0070] FIG. 30B illustrates exemplary pseudocode for the method of FIG. 30A, according to some embodiments.
[0071] FIG. 31 illustrates an exemplary detection diagram for the systems and methods described herein, according to some embodiments.
[0072] FIG. 32 illustrates an exemplary detection diagram for the systems and methods described herein, according to some embodiments.
[0073] FIG. 33 illustrates an exemplary detection diagram for the systems and methods described herein, according to some embodiments.
[0074] FIG. 34 illustrates an exemplary detection diagram for the systems and methods described herein, according to some embodiments.
[0075] FIG. 35 illustrates an exemplary detection diagram for the systems and methods described herein, according to some embodiments.
[0076] FIG. 36 illustrates an exemplary detection diagram for the systems and methods described herein, according to some embodiments.
[0077] FIG. 37 illustrates a diagram of an exemplary hardware and software systems implementing the systems and methods described herein, according to some embodiments.7IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0078] While the present disclosure is subject to various modifications and alternative forms, specific embodiments thereof have been shown by way of example in the drawings and will herein be described in detail. The present disclosure should be understood to not be limited to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure.DETAILED DESCRIPTION
[0079] Conventional software platforms have not been able to automatedly analyze, and / or classify free text data (e.g., text data) received from users of their respective platforms. For example, the free text data can contain a wealth of information that can be used to build a user profile for each user. In some examples, platform managers can use structured information about users to improve on usability and expound on the user base of their respective platforms. The structured information can include useful information about the users. Structured information can be generated from text data received from the users. Text data can include data that is structured, e.g., organized, or unstructured. In some examples, raw text data may include text data that does not have sufficient information to provide for useful information about one or more users. The raw text data may not have enough information to provide for a full description and / or a complete profile of a user. The structured information can include reliable, quantifiable, and / or actionable information about the user. In one example, platform managers can use structured information provide services to the users of the platform, to support the needs of the users so that the users are retained, and to improve on the overall user experience of the platform itself. In one example, platform managers can benefit from structured information about the users of their platform to efficiently run their respective platform, including to accurately target and / or provide custom advertising and marketing to certain users, to maintain and retain users by providing a tailored platform experiences, and to improve on the functionality of the platform as a whole. Sources of raw data describing the users of their respective platforms can include first party data input by users themselves, along with third-party data that can be sourced through devices such as third-party cookies (e.g., browser cookies), and information received from external data networks linked by other shared platforms (e.g., advertising identifiers and device identifiers). First party data can include data received and / or gathered directly from the users. In some examples, first party data can include text data describing a user's device settings, device usage, and / or device meta-data, among other data. First party data can include text data typed by the user, e.g., referred to herein as free text data. In some8IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO examples, the first party data can be in the form of messages, social media posts, diary entries, among other formats. Sources of the third party data can be unavailable, and be much more difficult to source due to restrictions on data privacy and data sharing. Third party data can include data about the users that was received and / or gathered from an external source. In some examples, third party data can include data about the users that was received and / or gathered from external data networks, external software, other software platforms, data received from other shared platforms, among other third-party data sources. There is a benefit for gathering first party data directly from the users themselves. In some examples, first party data received directly from the users can be used by the platform managers to improve on usability of their respective platforms, and expound on the platform's user base.
[0080] The present disclosure provides a technical solution to extracting, normalizing, and meaningfully combining heterogeneous signal types (textual, behavioral, temporal) to create unified user insights that exceed the predictive capabilities of an individual signal source. The system includes cross-domain integration techniques and computational efficiency improvements that allow for previously improved accuracy in predicting a user's needs, behavior, and satisfaction levels while maintaining scalability across large user populations.
[0081] In some embodiments, the present disclosure generally relates to systems and methods for generating structured information from text data received from users. In some embodiments, the systems presented herein can include software and / or hardware platforms having one or more users. In some examples, platform can be used to describe hardware and / or software for analyzing text to generate structured information. The structured information can be analyzed, e.g., using the systems and methods presented herein, to determine useful structured information about the user. The platform can include, but is not limited to, computing systems for gaming, programming, advertising, marketing, sales, among other applications. The platform can also be referred to as a computing platform, a digital platform, a software platform, an analysis platform, among other terms. Users can include customers, clients, subscribers, among other users of the platform. The systems and methods presented herein can provide information about the users through analysis of text data that are input by the users. In some embodiments, the systems and methods presented herein generate structured information from the text data, and provide the structured information to administrators of the system for analysis. For example, user entered text can be available via desktop computer, and / or mobile device software, which other conventional systems may find difficult to extract structured information from. In some examples, the systems and methods described herein can extract reliable, quantifiable, and / or actionable9IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO structured information from first-party information, such as user entered text input to a respective platform by the user. The systems and methods provide valuable insights about the users based on the input text data. In some embodiments, the systems and methods presented herein can include a custom artificial neural network to perform text analysis. The artificial neural network can include a large language model (LLM), among other machine learning algorithms. The systems and methods presented herein can use natural language processing performed by the generative large language model. As described herein the large language model can be referred to as a larger language model, a generative large language model, among other terms.
[0082] For example, the present disclosure describes a platform providing comprehensive signals describing software users through multi-modal analysis of user-generated data. The platform can use text data as a primary input source but incorporating multiple data streams through an integration architecture. The system can employ specialized artificial neural networks, including generative large language models acting as encoders, alongside other machine learning models in a tightly integrated processing pipeline with bidirectional feedback mechanisms and adaptive computational resource allocation.A FIRST EXEMPLARY SYSTEM FOR GENERATING STRUCTURED INFORMATION BASED ON TEXT ANALYSIS
[0083] Referring to FIG. 1, a system 100 for generating structured information through text based analysis is shown, according to some embodiments. In some embodiments, the system 100 can include an end user device 102, server 106, text store 112, prompt store 114, metadata store 116, scheduled processing system 120, determined user properties datastore 124, aggregation system 126, aggregation system 126, aggregation datastore 128, audience datastore 130, administrator device 132, among other components. The scheduled processing system 120 can include a generative large language model 122 (LLM). In some embodiments, the system 100 includes a software analytics platform. The system 100 can include one or more computer systems and / or computer components. The system 100 can include one or more servers, and / or include local computers. In some examples, each component of the system 100 can include computer software and / or hardware components.
[0084] In some embodiments, the end user device 102 can automatically record user input to generate the text data 104 and metadata 108 associated with one or more users 140 of the system 100. In some examples, the end user device 102 can automatically record keystrokes of the user 140 and locally store the keystroke data as text data 104. In some examples, a data gathering application can be used to gather the text data 104 from the user 140. The10IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO data gathering application can be loaded onto the end user device 102. In some embodiments, the text data 104 can be encrypted by the end user device 102. In some examples, the text data 104 can include encrypted text data 104. In some examples the metadata 108 can include device and user data. The device and user data can be unique to each individual user 140 of the system 100. The end user device 102 can store the metadata 108 onto a metadata store 116. As described herein, the system 100 can be referred to as a first exemplary system for generating structured information.
[0085] In some embodiments, the end user device 102 can securely transmit the text data 104 to the server 106. In some examples, the end user device 102 can transmit the text data 104 to the server 106 via a wired connection, and / or a wireless communication. The end user device 102 can transmit the text data 104 over a local network and / or over the internet. The text data 104 can include, but is not limited to, keystroke data received from a keyboard of the end user device 102, text data 104 input by the user 140 into form fields from a webpage, and / or text data 104 received from software applications loaded on the end user device 102. The server 106 can assess and / or analyze the text data 104, and create an identifier associated with each of the users 140 and / or source of the text data 104. The end user device 102 can include mobile devices, smart phones, tablets, keyboards, laptops, desktops computers, among other devices. Upon receipt of the text data 104, the server 106 can store 110 the text data 104 along with a corresponding user identifier (id), and / or an archive of the text data 104, onto the text store 112. The text store 112 can include an encrypted text store 112. The text store 112 can transmit the text data 104 to the scheduled processing system 120.
[0086] In some embodiments, the end user device 102 can securely transmit the metadata 108 to the aggregation system 126. In some examples, the end user device 102 can transmit the metadata 108 to the aggregation system 126 via a wired connection, and / or a wireless communication. The end user device 102 can transmit the metadata 108 over a local network and / or over the internet. The metadata 108 can include, but is not limited to, an identifier (id), device information, user information, among other descriptive data associated with the user 140 and / or the end user device 102. The metadata 108 can be stored onto a metadata store 116. The metadata store 116 can transmit the metadata 108 to the aggregation system 126.
[0087] In some embodiments, the scheduled processing system 120 can receive the text data 104 and prompts 118 for assessment and / or analysis. In some examples, an administrator 142 of the system 100 can send prompts 118 to the scheduled processing11IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO system 120 via an administrator device 132 and through a prompt store 114. The administrator device 132 can include an analytics user interface 134 used for transmitting prompts 118 from the administrator 142 to the prompt store 114. The prompt store 114 can transmit prompts 118 to the scheduled processing system 120. In some embodiments, the administrator 142 can include a platform manager, a client manager, a trusted user, among other users authorized to access the scheduled processing system 120.
[0088] In some embodiments, the prompts 118 can include an instruction for the generative large language model 122 to analyze, summarize, classify and / or evaluate the text data 104. In some examples, the prompts 118 can include an instruction for a generative large language model 122 of the scheduled processing system 120. The prompts 118 can be entered by the administrator 142 via the analytics user interface 134. In some embodiments, the administrator 142 can define custom prompts 118. In some examples, the prompts 118 received from the administrator device 132 can be stored onto the prompt store 114. The administrator 142 can enter and / or edit 138 the prompts 118 stored on the prompt store 114. The prompts 118 can include pre-defined prompts, e.g., prompts 118 that are preset for a particular platform and / or application. In some examples, the prompt store 114 can store both prompts 118 received from the administrator device 132 and the pre-defined prompts. In some embodiments, the prompt store 114 can include a collection of large language model prompts 118. In some examples, the prompts 118 can be stored as strings. The collection of prompts 118 can be entered by the administrator 142 based on predefined templates, and / or include static, e.g., fixed and / or predefined prompts 118. As shown, the prompt store 114 can be part of the scheduled processing system 120. The prompt store 114 can include a software and / or hardware component of the scheduled processing system 120. Each of the prompts 118 can include a set of configurations, which can define a quantity of text input, the frequency of evaluation, a type of large language model to use for evaluation, targeting parameters and / or the structure of the expected output for the system 100. In some examples a structure of the output of the system 100 can include a CSV file format, JSON file format, among other formats.
[0089] In some embodiments, the scheduled processing system 120 can generate structured information 144 based on the prompts 118 and the text data 104. In some examples, the structured information 144 can include determined user properties of each user 140 of the system 100. The structured information 144 can be stored on the determined user properties datastore 124. The system 100 and / or the administrator 142 themselves can define what structured information 144 to extract from the text data 104. In some examples, the generative large language model 122 can receive the prompts 118 as input, and generate12IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO and / or extract structured information 144 from the text data 104 based on the prompts 118. In this way, extraction of the structured data from the text data 104 by the generative large language model 122 can be guided by the prompts 118. The scheduled processing system 120 can store 110 the structured information 144 into determined user properties datastore 124. The aggregation system 126 can receive the metadata 108 and structured information 144, and combine and associate the data together. In some examples, the aggregation system 126 can combine the metadata 108 and the text data 104 together. The aggregation system 126 can store all the aggregated data 146 to an aggregation datastore 128. The aggregation datastore 128 transmit the aggregated data 146 to the administrator device 132. The administrator device 132 may store the resulting output aggregated data 146 to an audience datastore 130. The aggregated data 146 can include the structured information 144, and / or metadata 108.
[0090] In some embodiments, once the structured information 144 is determined, the scheduled processing system 120 can make the structured information 144 available for analysis by the administrator 142. The aggregated data 146 transmitted to the administrator device 132 can include the structured information 144. In some examples, the structured information 144 can be made available to the administrator 142 via an application, an analytics user interface 134. The analytics user interface 134 can also be referred to as a graphical user interface, an application programming interface, among other terms. In some examples, structured information 144 can be associated with a particular user id of a user 140. The scheduled processing system 120 can organize the structured information 144 based on the user id. The structured information 144 can be sent to the administrator 142 to allow the administrator 142 to analyze the structured information 144. In some examples, the system 100 can provide structured information 144 to the administrator device 132 via an application processing interface, and / or a dashboard. The application processing interface, and / or a dashboard can be located on a remote server accessible by the administrator 142 via the second end user device 102. To efficiently analyze and / or search through the structured information 144, the scheduled processing system 120 can analyze and / or search through the structured information 144 based on a given and / or predetermined criteria. The criteria can include, but is not limited to, user gender, demographic information, age, affluence, among others. The given criteria can be input by the administrator 142 via the administrator device 132. The predetermined criteria can be stored on the scheduled processing system 120 and / or stored with the prompts 118 on the prompt store 114. The scheduled processing system 120 can transmit the structured information 14413IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO for storage onto the text store 112. An exemplary analytics user interface 134 is shown at FIG. 3.
[0091] The scheduled processing system 120 can evaluate the prompts 118 based on predetermined and / or target parameters. The target parameters can include given, and / or predetermined boundary conditions for the evaluation of the prompts 118. The scheduled processing system 120 can perform the evaluation based on the text data 104 collected from users 140. In some embodiments, the results of the evaluation can be stored onto the aggregation datastore 128. The results of the evaluation can be stored along with the user id. The system 100 can schedule the evaluation of the prompts 118 based on provided scheduling parameters. In some examples, the administrator 142 can provide a schedule for when to execute the evaluation of the prompts 118.
[0092] In some embodiments, an exemplary function the system 100 allows is for administrators 142 to upload an audience file 148, e.g., a data file including user metadata 108 and / or including a list of user ids, and use the audience file to filter the resulting structured information 144. In some examples, consider a software platform and / or a company having a mobile application. The software platform is aware that 5% of their users 140 are of particularly high value, e.g., are always actively using the system, are high purchasers via the platform, make a substantial impact on contributing to the platform, among other reasons. The software platform would like to be able to invite more users 140 like the determined 5%. The system 100 can be used to and / or configured to for an administrator 142 to upload user ids (e.g., metadata 108) for corresponding to the determined 5% of users 140, and then based on the structured information 144, prompts 118 provided and user ids, the system 100 can determine any similarities or differences between those the determined 5% and the other 95% of the user base. In some examples, the determination can include determining the similarities and / or differences in terms of personality, language, interests, among others, and use that analysis to improve inviting more users similar to the determined 5% to the software platform. The audience file 148 can be stored and / or received (e.g., read and / or write) from the audience datastore 130.
[0093] In some embodiments, text 104 and metadata 108 can be processed and / or assessed by scheduled processing system 120, and the scheduled processing system 120 is not limited to an LLM and also not limited to scheduled processing, e.g., some of the processing can be scheduled, and some of the processing can be triggered. In some examples, some data can be processed in as close to real time using streaming, and a variety of heuristics, LLMs, among other machine learning models. In an example, the aggregation system 126 can read14IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO from both raw and / or processed outcomes. Although input into the system 100 is shown to come from the end user device 102, input can come from a third-party platform. In a particular non-limiting example, input for the system 100 can come from software managing user support tickets. In some examples, data can be supplied to the system via an application programming interface (API), email, software development kit (SDK), webhook, among other sources. In some embodiments, in addition to the scheduled processing system 120, one or more of the components of the system 100 can include LLMs and / or other machine learning models.
[0094] Referring to FIG. 2A and FIG. 2B, a first exemplary prompt 218 and second exemplary prompt 220 are shown respectively, according to some embodiments. The first exemplary prompt 218 can include “Answer very briefly in a json format using the tag Personality Trait: What personality trait is indicated by the text that is mentioned here:{input_text}." When passed the first exemplary prompt 218, the generative large language model (e.g., described in FIG. 2) can respond with a continuation of the first exemplary prompt 218, e.g., providing a response for the “{input text}”. In some examples, the generative large language model can provide a succinct answer to the first exemplary prompt 218, in a format that has been requested. In the case of the first exemplary prompt 218, the format requested is JSON.
[0095] Referring to FIG. 3, an exemplary user interface 302 for a system for generating structured information through text based analysis is shown, according to some embodiments. A processing module (e.g., the processing module 212 described in FIG. 2) can aggregate structured information 304 based on the text data from the users. The structured information 304 can be shown via either a programming or user interface 302. Administrators can examine and associate the structured information 304 with particular users based on individual user identifiers. In some examples, the structured information 304 can include determined characteristics of the users. In one example, the administrators can examine the determined characteristics of the users across a user base, a user group, or per user, by providing a list of user identifiers in a query to the user interface 302. Th user interface 302 can include bar graphs, pie charts, among other graphical representations of data. The user interface can include a toolbar 306, e.g., as shown in FIG. 3.METHOD FOR GENERATING STRUCTURED INFORMATION BASED ON TEXT ANALYSIS
[0096] Referring to FIG. 4, a method for generating structured information based on text analysis 400 is shown, according to some embodiments. At step 402, the method can15IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO include aggregating text provided by users 402. In some embodiments, step 402 can include automatically recording user input to generate the text data associated with one or more users. In some examples, step 402 can include automatically recording keystroke data of the users. At step 404, the method can include storing the text data and user identification to a data storage 404. In some embodiments, step 404 can include transmitting the text data over a local network and / or the internet to the data storage. In some examples, the transmitting can be performed via a wired and / or via wireless communication. At step 406, the method can include providing prompts to a generative large language model of a processing module 406. In some embodiments, step 406 can include automatically instructing the generative large language model using prompts. In some examples, the prompts can be provided by an administrator and / or the prompts can be pre-defined. At step 408, the method can include determining structured information from the text data based on the prompts using the generative large language model 408. In some embodiments, step 408 can include using the generative large language model to generate the structured information from the text data. In some examples, step 408 can include using the generative large language model to determine structured data from the text data. At step 410, the method can include providing the structured information for analysis to an administrator 410. In some embodiments, step 410 can include presenting the structured information to the administrator using a user interface, a graphical user interface, an application programming interface, among other interfaces.A SECOND EXEMPLARY SYSTEM FOR GENERATING STRUCTURED INFORMATION BASED ON TEXT ANALYSIS
[0097] Referring to FIG. 5, a system 500 for generating structured information is shown, according to some embodiments. In some embodiments, the system 500 can include data collection component 502, a storage and pre-preprocessing component 504, a text assessment component 506, a behavioral assessment component 508, a cross-domain assessment component 510, and an insight output component 512. Each of the components 502, 504, 506, 508, 510, 512 can be referred to as layers, among other terms. For example, as described herein, the data collection component 502 can be referred to as a data collection layer 502, among other terms. In some examples, each of the components 502, 504, 506, 508, 510, 512 can operate within a bidirectional information flow framework where insights from each component inform and enhance the processing of the other components. In one example, each of the components 502, 504, 506, 508, 510, 512 can send and / or receive data between each other component. As described herein, the system 500 can16IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO be referred to as a second exemplary system for generating structured information. In some examples, the first exemplary system, e.g., the system 100 of FIG. 1, and second exemplary system, e.g., the system 500 of FIG. 5, can refer to example systems for generating structured information.
[0098] In some embodiments, the data collection component 502 and the storage and prepreprocessing component 504 can collect data and store data received from users of the system 500. The data collection component 502 and storage and pre-preprocessing component 504 can securely collect and store multiple data types including: text data received from various sources, e.g., messages, support tickets, etc., interaction event streams recording user behavior, temporal metadata describing user activities, and contextual information describing the users' environment. The data collection component 502 and storage and pre-preprocessing component 504 can receive and transmit encrypted data using ephemeral key rotation techniques, validate and normalize incoming data in realtime, automatically classify data for privacy based on the content of the data, use a timebased data degradation pipeline having progressive anonymization stages. The storage and pre-preprocessing component 504 can store data in a specialized hybrid storage architecture that can combine: time-series optimized data stores for sequential behavior data, vector databases for embedding representations of text content, graph structures for relationship modelling between entities, and use traditional relational storage for structured metadata. The storage and pre-preprocessing component 504 can use the hybrid storage approach that can provide optimal performance for different query patterns while maintaining referential integrity across storage types.
[0099] In some embodiments, the data collection component 502 can gather and / or return user data from one or more sources. In some examples, the data collection component 502 can include aggregation endpoints for collecting user data. For example, the sources of data and / or aggregation endpoints can include an application programming interface (API), a webhook, a software development kit (SDK), 3rd party integration (e.g., via batch integration or email integration), among others. The data collection component 502 can include integrations with user platforms where the users post data using webhooks and / or an APIs. The data collection component 502 can include integrations with software platforms where users poll the platforms for data. In some examples, the data collection component 502 can include a software based keyboard, e.g., a mobile SDK keyboard that sends data to servers when users type. The data collection component 502 can include email integration that allows emails to be forwarded or copied to the system 500. The data collection component 502 can include a file based integration option where partners (e.g., other17IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO software platforms) can upload files into a storage system to be processed. In an example, uploading files to the storage system can be useful for uploading large amounts of data. The data collection component 502 can include a standalone REST HTTP API to support any other type of data collection integration. In some embodiments, data passed to the system 500 can be encrypted during transit and at rest, e.g., using rotating keys. In some examples, no data is stored longer than required. Incoming data for the system 500 can include text data with metadata. The incoming data can include a message from person A to persons B and C with text XYZ sent at certain time in response to another message X. The incoming data can include raw signal data e.g. events that have occurred. Events that have occurred can include a user logging in at this time, a user taking a particular action with a particular ticket at a particular time, a user taking another action at another later time, among other events. An input stream can contain aggregated information from a source system e.g., daily totals of activities, total keystrokes typed, applications installed on device, applications used that day, among others.
[0100] The storage and pre-preprocessing component 504 can store and prepare data. In some examples, the storage and pre-preprocessing component 504 can deduplicate, redact, and / or format data. In some examples, the storage and pre-preprocessing component 504 can deduplicate, redact, and / or format data received from the data collection component 502 for text and behavior analysis. In one non-limiting example, the storage and prepreprocessing component 504 can be used to detect boilerplate like text e.g., such as email footers, signatures, standard message introductions, etc., among other commonly repeated or similar sections. The storage and pre-preprocessing component 504 can mark up the detected boiler-plate text so that different weightings or policies can be applied to each category.
[0101] In some embodiments, the text assessment component 506 can include LLMs, neural networks, machine learning algorithms, natural language processing (NLP), heuristics. In some examples, the text assessment component 506 can use the LLMs, neural networks, machine learning algorithms, NLPs, heuristics in a hierarchical fashion to determine characteristics of text received from the storage and pre-preprocessing component 504. In an example, the storage and pre-preprocessing component 504 can use the LLMs, other neural network and machine learning algorithms, NLPs and heuristics to determine characteristics of the text. The text assessment component 506 can include single standalone processes or hierarchical processes, i.e., there can be processes that take as input combinations of raw input and outputs from previous text assessments. For example, hierarchical processing can occur at any depth. The outcome of processing routines used by18IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO the text assessment component 506 can affect the processing routines themselves. For example, the system 500 can categorize input text, but at the beginning of the input the system 500 may not have sufficient data to determine what all of the categories should be. Over time though the system 500 may be able to refine a list of categories - and the list can continue to change and / or evolve as the system 500 runs.
[0102] In some embodiments, the behavioral assessment component 508 can provide quantitative signals from the non-text input data using neural networks and / or heuristic techniques. In some examples, the behavioral assessment component 508 can provide quantitative signals from the non-text input data. The behavioral assessment component 508 can use both neural network and heuristic techniques and can additionally look at sequences of user actions by creating a vector space with those actions encoded within it.
[0103] In some embodiments, the cross-domain assessment component 510 can combine text and behavioral insights. In some examples the cross-domain assessment component 510 can combine behavioral insights to detect trends, risks, and / or areas for improvement. For example, the cross-domain assessment component 510 can combine output received from the text assessment component 506 and behavioral assessment component 508 to detect trends, risks, and / or areas for improvement.
[0104] In some embodiments, the insight output component 512 can deliver structured information. In some examples, the insight output component 512 can deliver insights, such as alerts and performance trends. The insight output component 512 can deliver insights via dashboards and / or administrative tools.
[0105] In some embodiments, the system 500 of FIG. 5 can encompass and / or include the system 100 of FIG. 1. In some examples, the system 100 of FIG. 1 can be a species of the genus of system 500 of FIG. 5. For example, the data collection component 502 of FIG. 5 can include the end user device 102 and the server 106 of FIG. 1. The storage and prepreprocessing component 504 of FIG. 5 can include the text store 112, prompt store 114, and the metadata store 116 of FIG. 1. The text assessment component 506 can include the scheduled processing system 120 and generative large language model 122 of FIG. 1. The behavioral assessment component 508 can include the metadata store 116, and aspects of user interaction patterns stored and tracked via the server 106 or during aggregation in aggregation system 126 of FIG. 1. The cross-domain assessment component 510 can include the aggregation system 126 and the aggregation datastore 128 of FIG. 1. The insight output component 512 of FIG. 5 can include the administrator device 132, analytics user interface19IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO 134, and the structured information 144 of FIG. 1. In an embodiment, the system 100 of FIG. 1 can include an exemplary implementation of the system 500 of FIG. 5.
[0106] Referring to FIG. 6, a system 600 for generating structured information through text based analysis is shown, according to some embodiments. The system 600 includes similar and / or the same components to the system 100 of FIG. 1. Therefore, the description for the same or similar labeled features in FIG. 1 apply to the same or similar labeled components in FIG. 6. For example, the end user device 602 of FIG. 6 can be the same or a similar end user device 102 from FIG. 1, and thus the description for the end user device 102 can apply to the end user device 602. In some embodiments, the system 500 of FIG. 5 can encompass and / or include the system 600 of FIG. 6, similar to that of FIG. 1.
[0107] In some embodiments, the system 600 can include an end user device 602, a server 606, an encrypted text store 612, an LLM prompt store 614, a signal and metadata store 616, a processing system 620, an LLM 622, a determined user properties store 624, an aggregation system 626, an aggregation data store 628, an audience datastore 630, among other components. In some embodiments, the end user device 602 can securely transmit text data 604, signal and metadata 608 to the server 606. The server 606 can assess and / or analyze the text data 104, the signal and metadata 608. The server 606 can store 610 the text data 604, signal and metadata 608 onto the encrypted text store 612 and signal and metadata store 616. The encrypted text store 612 and the signal and metadata store 616 can transmit the text data 604 to the processing system 620 and the aggregation system 626. Prompts from the LLM prompt store 614 can be read and / or written 618 to the processing system 620. The processing system 620 can include an LLM 622. In some embodiments, the prompts from the prompts store 614 can include an instruction for the LLM 622 to analyze, summarize, classify, assess, and / or evaluate the text data 604, signal and metadata 608. In some examples, the prompts can include an instruction for the LLM 622 of the processing system 620. In some embodiments, the scheduled processing system 620 can generate structured information based on the prompts, text data 604, signal and metadata 608. In some examples, the structured information can include determined user properties of each user 640 of the system 600. The structured information can be stored on the determined user properties store 624. The aggregation system 626 can receive structured information from the determined user properties store 624, and combine and associate the structured information into aggregated data. The aggregation system 626 can store the aggregated data to an aggregation data store 628. The aggregation data store 628 can transmit aggregated data to the administrator device 632. The administrator device 632 may store the resulting output aggregated data to an audience datastore 630. The administrator device 632 can20IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO include an analytics interface 634 for inspecting the aggregated data. The analytics interface 634 can include a user interface (UI). The aggregated data can include the structured information, signals, and / or metadata 608.
[0108] In some examples, text data 604, signal, and metadata 608 can be processed and / or assessed by the processing system 620. Processing, analysis, and / or assessment of the text data 604, signal, and metadata 608 received by the processing system 620 can be scheduled, and some of the processing can be triggered. In some examples, some data can be processed using streaming, heuristics, LLMs, among other machine learning models. In an example, the aggregation system 626 can read from both raw and / or processed outcomes. Although input into the system 600 is shown to come from the end user device 602, input can come from a third-party platform. In a particular non-limiting example, input for the system 600 can come from software managing user support tickets. In some examples, data can be supplied to the system via an application programming interface (API), email, software development kit (SDK), webhook, among other sources. In some embodiments, in addition to the processing system 620, one or more of the components of the system 600 can include LLMs and / or other machine learning models.
[0109] In some embodiments, the text assessment component 506 of FIG. 5 can include an LLM hierarchical prompt architecture. For example, the text assessment component 506 can include a directed acyclic graph (DAG) of parameterized LLM prompts having inheritance relationships that can allow the component 506 to assess complex workflows while maintaining computational efficiency. The text assessment component 506 can generate a parameterized prompts, e.g., based on received text data.
[0110] Referring to FIG. 7A, a flowchart 700 of a method for generating a parameterized prompt is shown, according to some embodiments. At a step 702, the method can include initializing a prompt template. At a step 704, the method can include applying contextspecific parametric adjustments to the prompt template to generate the prompt. At a step 706, the method can include setting confidence thresholds for the prompt based on an application context. At a step 708, the method can include adding specific instructions to the prompt based on text characteristics of received text data. At a step 710, the method can include assembling a final executable including all determined parameters from the previous steps. FIG. 7B shows exemplary pseudocode for the method of FIG. 7A.[OHl] In some embodiments, the text assessment component 506 of FIG. 5 can employ LLM based signal extraction mechanisms. For example, the text assessment component 506 can include an attention to weight assessment component, a confidence estimate assessment21IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO component, and a multi-scale sliding window assessment component. The attention to weight assessment component can extract attention signals from LLM model output. The attention to weight assessment component can capture and analyze attention weights across transformer layers to identify textual elements receiving highest model focus. The information determined by the attention to weight assessment component can be useful for the assessment of received text data because (a) text segments with high attention may contain key indicators of user sentiment, intent, or needs, (b) sections that consistently draw high attention across multiple prompts may be particularly informative for user analysis, and (c) attention patterns can help identify which parts of user text are most predictive for specific outcomes (churn, sales readiness, etc.). The attention weight assessment component can extract attention signals from output received the storage and pre-preprocessing component 504. The attention signals can include parts of the output received from the storage and pre-preprocessing component 504 that are most relevant based on the input text data.
[0112] Referring to FIG. 8 A, a flowchart 800 of a method for extracting attention signals is shown, according to some embodiments. At a step 802, the method can include processing attention weights from specified layers, e.g., received from the data collection component 502 and / or the storage and pre-preprocessing component 504. At a step 804, the method can include aggregating across attention heads, e.g., based on the attention weights. At a step 806, the method can include identifying high-attention tokens, e.g., specific relevant words. At a step 808, the method can include mapping back to original text spans. At a step 810, the method can include normalizing and ranking the attention signals by an attention score. FIG. 8B shows exemplary pseudocode for the method of FIG. 8 A.
[0113] In some embodiments, the confidence estimate assessment component can evaluate the reliability of LLM outputs based on input characteristics, model behavior, and historical performance. For example, certain types of inputs can be inherently more challenging for LLMs to process accurately. Exemplary input characteristics assessed by the confidence estimate assessment component can include: (a) text length (e.g., text that is too short may lack context, too long may dilute focus), (b) linguistic complexity (e.g., highly technical or specialized vocabulary), (c) ambiguous text (e.g., text with multiple possible interpretations), (d) contextual consistency (e.g., text with contradictory statements), (e) domain relevance (e.g., how closely the text aligns with the model's training domains, and (f) language patterns (e.g., unusual syntax, slang, or regional dialects).22IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0114] In some embodiments, the multi-scale sliding window assessment component can process long-form text that maintains context across window boundaries. For example, the multi-scale sliding window assessment component processes long-form text in text segments while maintaining context.
[0115] Referring to FIG. 9A, a flowchart 900 of a method for processing long-form text that maintains context across window boundaries is shown, according to some embodiments. At a step 902, the method can include generating overlapping windows, e.g., text segments, from the long-form text. At a step 904, the method can include stopping if the window reaches an end of the text. At a step 906, the method can include processing each window. At a step 908, the method can include merging overlapping results with special handling for boundary entities. FIG. 9B shows exemplary pseudocode for the method of FIG. 9 A.
[0116] In some embodiments, the behavioral assessment component 508 of FIG. 5 can extract meaningful patterns from user interactions, e.g., based on output received from the storage and pre-preprocessing component 504 and / or text assessment component 506. The behavioral assessment component 508 can be used for vector representation of interaction sequences. In an example, the behavioral assessment component 508 can encode user interaction sequences, e.g., based on output received from the storage and pre-preprocessing component 504 and / or text assessment component 506. In some examples, the behavioral assessment component 508 can transform sequential user actions into fixed-dimensional vector representations using encoding. The behavioral assessment component 508 converts users' sequences of actions or interactions into vector formats that machine learning models can process efficiently while preserving the sequential nature of those interactions. In some examples, the behavioral assessment component 508 can (i) capture user interaction sequences of the users, and (ii) execute a vector coding process. Capturing user interaction sequences can include tracking a series of actions that users take when interacting with software. For example, tracking user's button clicks, page views, feature usage, time spent on different screens, etc. These actions can occur in a specific order that often contains valuable information about the user's goals, preferences, and behaviors. Executing the vector encoding process can include assigning each distinct action or interaction type of the users a base representation. The sequence of these actions can then be processed to create a fixed-length vector that captures: (a) specific actions taken, (b) order in which the actions occurred, and (c) timing between actions. A vector is created using a transformer-based encoder but other implementations including types of recurrent neural network or n-gram heuristics could also be used. In some examples, the user interactions and / or sequences23IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO provide data that can serve the following applications: (a) identifying behavioral patterns that predict customer churn, (b) recognizing sequences that indicate a user is struggling with a feature, (c) detecting navigation patterns that suggest shopping or purchase intent, (d) comparing user journeys to identify similar user types or behaviors. For example, a user's navigation path through an application can include: Home —> Product Page —> Features Tab —> Pricing Page —> Account Settings —> Upgrade Options. This sequence can be encoded as a vector that captures not just which pages were visited but their order and relationship. This vector representation can then be compared by the behavioral assessment component 508 with patterns known to indicate purchase interest or compared with vectors from other users to find similar behavior profiles. In some examples, the behavioral assessment component 508 can translate temporal behavioral data into a mathematical representation which preserves sequential information, while allowing for efficient processing, comparison, and analysis by machine learning models.
[0117] Referring to FIG. 10A, a flowchart 1000 of a method for encoding user interaction sequences is shown, according to some embodiments. At step 1002, the method can include initializing encoders, e.g., if not already loaded. At a step 1004, the method can include encoding user action sequences. At a step 1006, the method can include encoding temporal patterns between action sequences. At a step 1008, the method combining sequence and temporal information. At a step 1010, calculating intervals between action sequences. At a step 1012, the method can include extracting statistical features. At a step 1014, the method can include extracting rhythm features using wavelet decomposition. FIG. 10B shows exemplary pseudocode for the method of FIG. 10A.
[0118] In some embodiments, the system 500 of FIG. 5 includes custom neural network architectures configured for behavioral sequence analysis such as (i) attentional recurrent models that combine recurrent processing with attention mechanisms to identify important patterns in interaction sequences regardless of their position in the sequence, (ii) 1 dimensional convolutional networks for extracting temporal features across multiple time scales simultaneously, and that identify characteristic patterns in usage data, (iii) selfsupervised pre-training approaches that include a pre-training methodology that can learns behavioral representations without requiring labelled data.
[0119] Referring to FIG. 11 A, a flowchart 1100 of a method for self-supervised pretraining is shown, according to some embodiments. At step 1102, the method can include generating synthetic tasks for self-supervised learning. At a step 1104, the method can include creating training data from unlabeled sequences. At a step 1106, the method can24IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO include training the model on multiple self-supervised tasks. FIG. 11B shows exemplary pseudocode for the method of FIG. 11 A.
[0120] In some embodiments, the behavioral assessment component 508 of FIG. 5 can include a periodicity detection component. In some examples, the periodicity detection component of the behavioral assessment component 508 can identify periodic patterns in user behavior using wavelet transformation techniques.
[0121] Referring to FIG. 12A, a flowchart 1200 of a method for identifying periodic patterns in user behavior is shown, according to some embodiments. At step 1202, the method can include applying continuous wavelet transform to received data. At a step 1204, the method can include finding significant coefficients from the wavelet transform. At a step 1206, the method can include grouping periodic components. FIG. 12B shows exemplary pseudocode for the method of FIG. 12 A.
[0122] In some embodiments, cross-domain assessment component 510 of FIG. 5 can combine signals from different representational spaces, e.g., behavioral and text. In some embodiments, the cross-domain assessment component 510 combines the output from the text assessment component 506 and the behavioral assessment component 508 to determine structured information including properties of the users. In an example, the cross-domain assessment component 510 can align user sequences based on behavioral data and text data. In some examples, the cross-domain assessment component 510 can implement dynamic time warping to align user sequences from different data sources with varying temporal characteristics.
[0123] Referring to FIG. 13A, a flowchart 1300 of a method for aligning user action sequences using dynamic time warping is shown, according to some embodiments. At step 1302, the method can include initializing a distance matrix. At a step 1304, the method can include calculating distances, with respect to the distance matrix, with constrain enforcement. At a step 1306, the method can include calculating a valid range based on constraints. At a step 1308, the method can include calculating distances between current points. At a step 1310, the method can include finding a minimum of three possible previous steps, e.g., including (i) an insertion, (ii) a deletion, and (iii) a match. At a step 1312, the method can include reconstructing an optimal warping path. FIG. 13B shows exemplary pseudocode for the method of FIG. 13 A.
[0124] In some embodiments, the cross-domain assessment component 510 of FIG. 5 can be used for semantic mapping between domains. In an example, the cross-domain assessment component 510 can create cross domain mappings based on behavioral data and25IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO text data. In some examples, the cross-domain assessment component 510 can create meaningful connections between textual and behavioral representations via the cross domain mappings.
[0125] Referring to FIG. 14A, a flowchart 1400 of a method for creating cross domain mapping is shown, according to some embodiments. At step 1402, the method can include initializing mapping networks. At a step 1404, the method can include creating paired examples for training. At a step 1406, the method can include training the mapping networks with symmetric constraints. At a step 1408, the method can include training text- to-behavior mapping. At a step 1410, the method can include training behavior-to-text mapping. At a step 1412, the method can include applying cycle consistency constraints. FIG. 14B shows exemplary pseudocode for the method of FIG. 14A.
[0126] In some embodiments, the cross-domain assessment component 510 of FIG. 5 can be used for signal fusion via Bayesian ensemble techniques. In an example, the crossdomain assessment component 510 can combine predictions using a Bayesian ensemble. In some examples, the cross-domain assessment component 510 can combine predictions from different models while accounting for their varying confidence levels.
[0127] Referring to FIG. 15 A, a flowchart 1500 of a method for combining predictions using a Bayesian ensemble is shown, according to some embodiments. At step 1502, the method can include calculating important weights based on determined confidences. At a step 1504, the method can include using a weighted probability combination for classification outputs. At a step 1506, the method can include using a weighted distribution combination for regression outputs. At a step 1508, the method can include using a customized rank fusion for ranking outputs. At a step 1510, the method can include converting confidences to precision parameters for beta distributions. At a step 1512, the method can include incorporating prior knowledge. At a step 1514, the method can include sampling from posterior distributions to account for uncertainty. At a step 1516, the method can include calculating weights based on expected precision. FIG. 15B shows exemplary pseudocode for the method of FIG. 15 A.
[0128] In some embodiments, the cross-domain assessment component 510 of FIG. 5 can be used for multi-objective optimization. In an example, the cross-domain assessment component 510 optimizes multiple objectives of the system 500. In some examples, the cross-domain assessment component 510 can use custom multi-objective optimization techniques to balance competing objectives like precision and recall based on the specific use cases.26IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0129] Referring to FIG. 16A, a flowchart 1600 of a method for optimizing multiple system objectives is shown, according to some embodiments. At step 1602, the method can include initializing a pareto frontier, e.g., also referred to as a pareto front. At step 1604, the method can include running an evolutionary algorithm to find pareto-optimal solutions. At step 1606, the method can include generating candidate solutions, e.g., based on the pareto front. At step 1608, the method can include evaluating each candidate solution based on one or more candidate objectives. At step 1610, the method can include updating the pareto front. At step 1612, the method can include applying constraint filtering, e.g., applying the constraint filtering to the pareto front. At step 1614, the method can include checking for convergence, e.g., based on the pareto front. At step 1616, the method can include returning a full pareto front for non-dominated solutions. At step 1618, the method can include loading an operating point selection policy for a particular use case. At step 1620, the method can include applying the operating point selection policy to select an appropriate point on the pareto front. FIG. 16B shows exemplary pseudocode for the method of FIG. 16A.
[0130] In some embodiments, the system 500 of FIG. 5 can include an adaptive processing component. In some examples, the adaptive processing component can improve on computational resource management by optimizing processing efficiency while maintaining accuracy. In an example, the adaptive processing component can optimize batch parameters based on input workload, hardware profiles, and latency requirements, among others. In some examples, the adaptive processing component can include dynamic batch optimization for accelerator hardware such as GPU / TPU hardware.
[0131] Referring to FIG. 17A, a flowchart 1700 of a method for optimizing batch parameters is shown, according to some embodiments. At step 1702, the method can include analyzing workload characteristics. At step 1704, the method can include finding optimal batch parameters through Bayesian optimization. At step 1706, the method can include simulating batch processing based on the batch parameters. At step 1708, the method can include calculating a score based on throughput and latency requirements. At a step 1710, the method can include running a Bayesian optimization to find optimal parameters. FIG. 17B shows exemplary pseudocode for the method of FIG. 17 A.
[0132] In some embodiments, the system 500 of FIG. 5 can include a feedback integration component. In an example, the feedback integration component can assess feedback signals, e.g., the feedback integration component can assess the feedback signals based on input predictions, actual outcomes, and received model parameters. In some examples, the system27IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO 500 can include a closed loop learning implementation, e.g., the system 500 can continuously improve model performance based on observed outcomes.
[0133] Referring to FIG. 18 A, a flowchart 1800 of a method for analyzing and / or assessing feedback signals is shown, according to some embodiments. At step 1802, the method can include matching predictions with actual outcomes. At step 1804, the method can include calculating performance metrics. At step 1806, the method can include detecting performance drift. At step 1808, the method can include triggering a model recalibration process. At step 1810, the method can include analyzing feature importance. At step 1812, the method can include updating feature weightings based on importance. FIG. 18B shows exemplary pseudocode for the method of FIG. 18 A.
[0134] In some embodiments, the system 500 of FIG. 5 can include a progressive data transformation pipeline. In an example, the system 500 can apply a timed data transformation, e.g., based on input data, age of the data, and an input privacy policy. In some examples, the system 500 implements data privacy preservation through automated data transformations that can become increasingly aggressive over time.
[0135] Referring to FIG. 19A, a flowchart 1900 of a method for applying a timed data transformation is shown, according to some embodiments. At step 1902, the method can include determining an applicable transformation stage based on data age. At step 1904, the method can include applying an appropriate transformation for the transformation stage based on at least one of: recent data (minimal transformation), medium-term data (increased anonymization), long-term data (aggressive anonymization), or archival data (statistical preservation). At step 1906, the method can include reducing temporal precision. At step 1908, the method can include generalizing specific identifiers. At step 1910, the method can include applying differential privacy techniques, e.g., where appropriate. FIG. 19B shows exemplary pseudocode for the method of FIG. 19A.
[0136] In some embodiments, the insight output component 512 transmits structured information including user properties to users and / or administrators. In some examples, the insight output component 512 receives the output from the cross-domain assessment component 510 and transmits output, e.g., structured information, to users and / or administrators.APPLICATION- SPECIFIC IMPLEMENTATIONS
[0137] In some embodiments, the systems 100, 500, and 600 described above can be configured for one or more applications. In some non-limiting examples, the applications28IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO can include customer support optimization, churn prevention, sales improvement, audience assessment, among other applications.CUSTOMER SUPPORT OPTIMIZATION
[0138] Existing customer support systems react to explicit support requests, missing early indicators of confusion and struggling to prioritize resources effectively. To address these limitations, the systems 100, 500, 600 can include a multi-modal configuration including (i) a user behavior detection subsystem, (ii) an issue clustering and impact prediction engine, and (iii) a support resource allocation optimizer. In some examples, the user behavior detection subsystem can detect user behavior. User behavior can include the interaction of the user with the user's digital device. Use behavior can include text data that includes contextual information that describes the user's interaction with an application on the user's device. The user's behavior can be interpreted based on text data received from the user. For example, user behavior associated with user happiness and / or frustration as interpreted through text data received from the user. In a non-limiting example, the user behavior subsystem can be referred to as a frustration early detection subsystem, among other terms. The issue clustering and impact prediction engine can cluster and prioritize user issues based on detected issues of the system, and user population. The support resource allocation optimizer can optimize support resource allocation based on known issues, available resources, and input parameters.
[0139] Referring to FIG. 20A, a flowchart 2000 of a method for detecting user behavior is shown, according to some embodiments. At step 2002, the method can include extracting linguistic markers of associated with user behavior. In some examples, extracting linguistic markers associated with user behavior can include extracting linguistic markers associated with user emotions determined through text data. User emotion can include happiness, frustration, sadness, among other emotions. Extracting linguistic markers associated with user behavior can include extracting linguistic markers based on text data received from the user, and a determined user emotion from the text data. In a non-limiting example, extracting linguistic markers associated with user behavior can include extracting linguistic markers associated with user frustration. The user behavior and / or user emotion can be determined via contextual information of the text data, emotion markers (e.g., frustration markers), among others. At step 2004, the method can include extracting behavior indicators. In some examples, the behavioral indicators can be extracted from text data, behavior data, emotion pattern data (e.g., frustration pattern data), among others. At step 2006, the method can include comparing the user's behavior against a user's baseline29IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO behavioral data. The behavioral data can be stored and / or generated from previous text data and / or behavioral data received from the user. At step 2008, the method can include calculating a behavioral probability using a calibrated machine learning model (e.g., an LLM, among others). The behavioral probability can include frustration probability, happiness, probability, sadness probability, among user emotion-based probability. FIG. 20B shows exemplary pseudocode for the method of FIG. 20A.
[0140] Referring to FIG. 21A, a flowchart 2100 of a method for clustering and prioritizing user issues is shown, according to some embodiments. At step 2102, the method can include semantic clustering of similar cluster issues. At step 2104, the method can include calculating an impact factor for each cluster. At step 2106, the method can include modeling a propagation through a user network. At step 2108, the method can include prioritizing issues based on impact and time sensitivity. FIG. 2 IB shows exemplary pseudocode for the method of FIG. 21 A.
[0141] Referring to FIG. 22A, a flowchart 2200 of a method for optimizing support resource allocation is shown, according to some embodiments. At step 2102, the method can include defining an optimization problem. In some examples, the defining an optimization problem can include maximizing support effectiveness based on resource constraints, response time, and / or user priority constraints. Defining an optimization problem can include estimating resources needed based on the impact of issues and users affected. At step 2104, the method can include solving the optimization problem. In some examples, solving the optimization problem can include returning a recommended resource allocation. FIG. 22B shows exemplary pseudocode for the method of FIG. 22A.CHURN PREVENTION
[0142] Existing churn models rely on lagging indicators and fail to identify specific causal factors that could enable effective intervention. To address these limitations, the systems 100, 500, 600 can include (i) an early warning subsystem (ii) a causal factor identification subsystem, and (iii) an intervention optimization subsystem. In some examples, the early warning subsystem can build a churn prediction model, e.g., using a survival assessment. The churn prediction model can be built based on user histories, survival times, and censoring data. The causal factor identification subsystem can identify causal factors, e.g., based on user population, treatments, and outcomes. The intervention optimization subsystem can optimize churn interventions. The churn interventions can be optimized based on user segments, interventions, and constraints.30IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0143] Referring to FIG. 23 A, a flowchart 2300 of a method for building a churn prediction model is shown, according to some embodiments. At step 2302, the method can include creating time-varying covariates from user histories. At step 2304, the method can include building an extended cox proportional hazards model. At step 2306, the method can include validating the cox model's performance. At step 2308, the method can include calculating a variable importance, e.g., based on the cox model. At step 2310, the method can include extracting current covariates for the user. At step 2312, the method can include calculating a survival probability at a specified time horizon. The survival probability can be calculated based on the current covariates and a horizon time. At step 2314, the method can include calculating a hazard ratio relative to a baseline. The hazard ratio can be calculated based on an input model (e.g., the cox model), and current covariates. In some examples, the baseline can include the input model and the current covariates. FIG. 23B shows exemplary pseudocode for the method of FIG. 23 A.
[0144] Referring to FIG. 24A, a flowchart 2400 of a method for identifying causal factors is shown, according to some embodiments. At step 2402, the method can include creating a structural causal model. At step 2404, the method can include performing a counterfactual assessment. In some examples, the counterfactual assessment can be based on a causal model, treatments, and outcomes. At step 2406, the method can include calculating average treatment effects. The average treatment effects can be calculated based on results from the counterfactual assessment. At step 2408, the method can include identifying strongest causal factors. FIG. 24B shows exemplary pseudocode for the method of FIG. 24A.
[0145] Referring to FIG. 25A, a flowchart 2500 of a method for optimizing churn interventions is shown, according to some embodiments. At step 2502, the method can include initializing a model such as a multi-armed bandit model for each segment of user segments. In some examples, initializing the model, e.g., the multi-armed bandit, for each segment of user segments can include initializing a multi-armed bandit for individual user of group of users. At step 2504, the method can include executing an optimization loop. At step 2506, the method can include selecting interventions for each segment. At step 2508, the method can include simulating outcomes. In some examples, simulating outcomes can include simulating intervention outcomes based on the selections at step 2506. At step 2510, the method can include updating bandits from the multi-armed bandits with observed results. At step 2512, the method can include extracting an optimal policy from trained bandits. FIG. 25B shows exemplary pseudocode for the method of FIG. 25 A.SALES IMPROVEMENT31IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0146] Sales teams lack reliable mechanisms to identify genuine purchase signals among noisy data, and struggle to properly time customer engagement for maximum effectiveness. To address these limitations, in some examples, the systems 100, 500, 600 detect and act on sales opportunities by including (i) a purchase intent detection subsystem, (ii) a need identification engine, and (iii) an engagement timing optimization subsystem. The purchase intent detection subsystem can detect a purchase intent of a user, e.g., based on text data from the user, user behavior, and user context. The need identification engine can identify user needs. User needs can be identified based on a user data vector, and product capability vectors. The engagement timing optimization subsystem can optimize user engagement timing, e.g., based on user profiles and user engagement history.
[0147] Referring to FIG. 26A, a flowchart 2600 of a method for detecting a purchase intent of a user is shown, according to some embodiments. At step 2602, the method can include extracting linguistic intent markers, e.g., from the user text data. At step 2604, the method can include extracting behavioral intent markers, e.g., from user behavior data. At step 2606, the method can include combining user intent signal using contextual weighting. At s. 2608, the method can include calculating a confidence interval for an intent score. FIG. 26B shows exemplary pseudocode for the method of FIG. 26A.
[0148] Referring to FIG. 27A, a flowchart 2700 of a method for identifying user needs is shown, according to some embodiments. At step 2702, the method can include projecting user data into a need-space. At step 2704, the method can include calculating similarity to product capability vectors. In some examples, the similarity to product capability vectors can be calculated based on a product capability vector map. At step 2706, the method can include identifying capability gaps, capability gaps can be identified based on a user needs vector, a product capability vector, and a user data vector. FIG. 27B shows exemplary pseudocode for the method of FIG. 27A.
[0149] Referring to FIG. 28A, a flowchart 2800 of a method for optimizing user engagement timing is shown, according to some embodiments. At step 2802, the method can include extracting receptivity patterns. In some examples, the receptivity patterns are extracted based on user profiles and user engagement history. At step 2804, the method can include predicting optimal time windows. At step 2806, the method can include calculating urgency factors. At step 2808, the method can include balancing receptivity with urgency. FIG. 28B shows exemplary pseudocode for the method of FIG. 28 A.AUDIENCE ASSESSMENT32IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0150] Traditional demographic-based audience targeting can lack precision and can fail to identify valuable users based on actual behavioral and communication patterns. To address these limitations, the systems 100, 500, 600 can extract distinguishing characteristics of high-value users and can include (i) a high-value pattern extraction subsystem, and (ii) a similarity modelling engine. The high-value pattern extraction subsystem can extract high value patterns. In some examples, the high value patterns can be extracted based on user population, and value metrics. The similarity modelling engine can build similarity models. The similarity models can be built based on reference users, and user candidate pools.
[0151] Referring to FIG. 29A, a flowchart 2900 of a method for extracting high value patterns is shown, according to some embodiments. At step 2902, the method can include segmenting users by value. At step 2904, the method can include extracting feature distributions for each user segment. At step 2906, the method can include calculating discriminative power of the determined features. At step 2908, the method can include identifying distinctive patterns of high-value users. In some examples, distinctive patterns of high-value users can be identified based on user segments and feature importance. FIG. 29B shows exemplary pseudocode for the method of FIG. 29A.
[0152] Referring to FIG. 30A, a flowchart 3000 of a method for building similarity models is shown, according to some embodiments. At step 3002, the method can include creating embeddings for reference users. At step 3004, the method can include building a vector index for efficient similarity search. The vector index can be built based on a vector search index and reference embeddings. At step 3006, the method can include defining a similarity function based on multiple criteria. The similarity function can be defined based on textual weights, behavioral weights, and temporal weights. At step 3008, the method can include finding similar users. In some examples, similar users can be found based on vector indices, a candidate pool, a similarity function, and a similarity threshold. At step 3010, the method can include calculating confidence scores. FIG. 30B shows exemplary pseudocode for the method of FIG. 30 A.
[0153] Referring to FIG. 31 to FIG. 36., exemplary detection diagrams for the systems and methods described herein are shown, according to some embodiments. In some embodiments, the FIG. 31 - FIG. 36 show exemplary detections that the systems and methods described herein can make for determining customer satisfaction levels, issue categories, upsell opportunities, and / or quantity of information shared by support agents. FIG. 31 to FIG. 36 also show exemplary results for the systems and methods described herein including structured process improvement recommendations for the customer. In33IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO some examples, FIG. 31 to FIG. 36 show exemplary results for an optimized way to improve the user's documentation to reduce workload on user support systems.IMPROVEMENTS OVER CONVENTIONAL ANALYTICS PLATFORMS
[0154] The technical field of user analytics faces several significant challenges that existing systems have failed to adequately address. Such challenges can include (a) signal integration problems: existing systems can process different data types (text, behavior, temporal patterns) in isolation, creating disconnected insights that fail to capture the complex relationships between what users say and what they do which can lead to fragmented understanding and contradictory predictions, (b) heterogeneous data normalization problems: text data exists in fundamentally different representational spaces than behavioral data, making direct comparison and integration technically challenging, where previous attempts at simple feature concatenation fail to preserve the semantic relationships between these domains, (c) computational efficiency problems: processing multiple data streams at scale can require enormous computation resources, making comprehensive multi-modal analysis prohibitively expensive for most applications using conventional architectures, (d) temporal alignment problems: user behavior and communications occur at different rates and timeframes, creating technical challenges in properly sequencing and relating events across modalities, (e) signal confidence problem: different data sources have varying levels of reliability and predictive power in different contexts, but existing systems can lack mechanisms to appropriately weigh and combine them based on their contextual reliability, and (f) privacy -utility trade-off problems: current systems either preserve raw data (creating privacy risks) or applying aggressive anonymization (reducing analytical utility), lacking a technical approach to balance these competing objectives. These technical problems have prevented existing analytics systems from fully leveraging the rich combination of text and behavioral data typically available to software providers. Furthermore, prior approaches have relied primarily on siloed analysis systems that process different data types separately, simplistic rule-based systems for combining insights, static weighting schemes that fail to adapt to context or data quality, and computational architectures that cannot scale to handle multi-modal analysis across large user populations. The need for a solution has been amplified by recent changes in the technical landscape including increased privacy restrictions limiting third-party data access, growing computational capabilities enabling more sophisticated analyses, and advances in large language model technology creating new possibilities for text understanding.34IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO
[0155] The present disclosure addresses the technical problems identified above through a multi-layered architecture with several key technical improvements including: (i) a crossdomain integration architecture: a system for integrating textual and behavioral signals through a bidirectional mapping approach that aligns their semantic spaces while preserving their unique information content, (ii) an adaptive signal processing pipeline: a computational framework that dynamically allocates processing resources based on signal characteristics, contextual requirements, and historical predictive value, (iii) a temporal relationship modelling mechanism: mechanisms for aligning and relating events across different time scales and data sources, allowing the identification of causal relationships between user communications and behaviors, (iv) a confidence-weighted signal fusion mechanism: a framework for integrating predictions from multiple models while accounting for varying levels of confidence and reliability across different contexts, and (v) a progressive data transformation system: an automated, time-triggered transformation system that gradually reduces identifiability of user data while preserving analytical utility. Furthermore, the systems and methods described herein include a bidirectional feedback mechanisms between different processing components, creating an integrated system that allows for significantly improved predictive accuracy while maintaining computational efficiency through improved optimization techniques.ADVANTAGES OVER CONVENTIONAL ANALYTICS PLATFORMS
[0156] The systems and methods described herein provide several specific technical advantages over existing systems and techniques, including (i) improved predictive accuracy: the cross-domain integration approach allows for an improvement in predictive accuracy for user behavior compared to single-modality approaches, (ii) computational efficiency: the adaptive processing architecture reduces computational resource requirements compared to native multi-modal implementations through intelligent resource allocation and model optimization techniques, (iii) temporal insight extraction: the system's temporal alignment and periodicity detection capabilities enable identification of causal relationships between user communications and actions that were previously undetectable with conventional analytics, (iv) adaptive signal weighting: the confidence-weighted fusion approach improves robustness to noisy or missing compared to fixed-weight ensemble methods.
[0157] In some embodiments, the systems and methods provide an improvement over previous systems and techniques including by being implemented one or more configurations: (i) edge-centric processing architecture: initial processing of the system and35IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO method can occur on edge devices, with aggregated insights transmitted to central systems, (ii) federated implementation: machine learning models can be trained using federated learning approaches, with no raw user data leaving local environments, (iii) real-time processing focus: a configuration optimized for immediate insight generation with reduced historical analysis capabilities but faster response times, (iv) specialized domain implementations: targeted implementations for specific industries with domain-specific analysis components, and (v) hardware-optimized variants: implementations specifically optimized for hardware accelerators like FPGAs or specialized Al chips.HARDWARE AND SOFTWARE IMPLEMENTATIONS
[0158] FIG. 37 is a block diagram of an example system 3700 that may be used in implementing the technology described in this document. As described herein, the system 3700 can also be referred to as a computer system 3700, among other terms. General- purpose computers, network appliances, mobile devices, or other electronic systems may also include at least portions of the system 3700. The system 3700 includes a processor 3702, a memory 3704, a storage device 3706, and an input / output device 3708. Each of the processors 3702, 3704, 3706, and 3708 may be interconnected, for example, using a system bus 3710. The processor 3702 is capable of processing instructions for execution within the system 3700. In some implementations, the processor 3702 is a single-threaded processor. In some implementations, the processor 3702 is a multi -threaded processor. The processor 3702 is capable of processing instructions stored in the memory 3704 or on the storage device 3706.
[0159] The memory 3704 stores information within the system 3700. In some implementations, the memory 3704 is a non-transitory computer-readable medium. In some implementations, the memory 3704 is a volatile memory unit. In some implementations, the memory 3704 is a non-volatile memory unit.
[0160] The storage device 3706 is capable of providing mass storage for the system 3700. In some implementations, the storage device 3706 is a non-transitory computer-readable medium. In various different implementations, the storage device 3706 may include, for example, a hard disk device, an optical disk device, a solid-date drive, a flash drive, or some other large capacity storage device. For example, the storage device may store long-term data (e.g., database data, file system data, etc.). The input / output device 3708 provides input / output operations for the system 3700. In some implementations, the input / output device 3708 may include one or more network interface devices, e.g., an Ethernet card, a serial communication device, e.g., an RS-232 port, and / or a wireless interface device, e.g., 36IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO an 802.11 card, a 3G wireless modem, or a 4G wireless modem. In some implementations, the input / output device may include driver devices configured to receive input data and send output data to other input / output devices, e.g., keyboard, printer and display devices 3712. In some examples, mobile computing devices, mobile communication devices, and other devices may be used.
[0161] In some implementations, at least a portion of the approaches described above may be realized by instructions that upon execution cause one or more processing devices to carry out the processes and functions described above. Such instructions may include, for example, interpreted instructions such as script instructions, or executable code, or other instructions stored in a non-transitory computer readable medium. The storage device 3706 may be implemented in a distributed way over a network, for example as a server farm or a set of widely distributed servers, or may be implemented in a single computing device.
[0162] Although an example processing system has been described in FIG. 37, embodiments of the subject matter, functional operations and processes described in this specification can be implemented in other types of digital electronic circuitry, in tangibly- embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible nonvolatile program carrier for execution by, or to control the operation of, data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0163] The term “system” may encompass all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. A processing system may include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). A processing system may include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that37IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0164] A computer program (which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0165] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0166] Computers suitable for the execution of a computer program can include, by way of example, general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. A computer generally includes a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices.
[0167] Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; and magneto optical38IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0168] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.
[0169] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0170] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0171] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can39IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO generally be integrated together in a single software product or packaged into multiple software products.
[0172] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous. Other steps or stages may be provided, or steps or stages may be eliminated, from the described processes. Accordingly, other implementations are within the scope of the following claims.TERMINOLOGY
[0173] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting.
[0174] The term “approximately”, the phrase “approximately equal to”, and other similar phrases, as used in the specification and the claims (e.g., “X has a value of approximately Y” or “X is approximately equal to Y”), should be understood to mean that one value (X) is within a predetermined range of another value (Y). The predetermined range may be plus or minus 20%, 10%, 5%, 3%, 1%, 0.1%, or less than 0.1%, unless otherwise indicated.
[0175] The indefinite articles “a” and “an,” as used in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.” The phrase “and / or,” as used in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0176] As used in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least40IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used shall only be interpreted as indicating exclusive alternatives (i.e. “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0177] As used in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a nonlimiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0178] The use of “including,” “comprising,” “having,” “containing,” “involving,” and variations thereof, is meant to encompass the items listed thereafter and additional items.
[0179] Use of ordinal terms such as “first,” “second,” “third,” etc., in the claims to modify a claim element does not by itself connote any priority, precedence, or order of one claim element over another or the temporal order in which acts of a method are performed. Ordinal terms are used merely as labels to distinguish one claim element having a certain name from another element having a same name (but for use of the ordinal term), to distinguish the claim elements.
[0180] Having thus described several aspects of at least one embodiment of this invention, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are41IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO intended to be part of this disclosure, and are intended to be within the spirit and scope of the invention. Accordingly, the foregoing description and drawings are by way of example only.42IPTS / 1289821 57.1
Claims
Attorney Docket No. TAPP-OOIWOCLAIMSWhat is claimed is:
1. A computer-implemented method, comprising: receiving, via a server, text data collected from users; generating, using the server, identifiers associated with the text data and associated with the users; receiving a prompt having instructions for a large language model (LLM) to determine structured information including user properties based on the text data; transmitting the prompt, the text data, and identifiers to the LLM; determining, using the LLM, the structured information based on the prompt, the text data and identifiers; and transmitting, using an aggregation component, the structured information and the identifiers to an administrator.
2. The computer-implemented method of claim 1, prior to receiving the text data, automatically recording, using a data gathering application, user keystrokes to generate the text data associated with the users.
3. The computer-implemented method of claim 1, further comprising: transmitting, using the server, the text data and identifiers to a text store for storage.
4. The computer-implemented method of claim 1, wherein the prompt includes instructions for the LLM to analyze, summarize, classify and assess the text data to determine the structured information.
5. The computer-implemented method of claim 1, wherein the prompt comprises a predefined prompt.
6. The computer-implemented method of claim 1, further comprising extracting the prompt from an LLM prompt store.
7. The computer-implemented method of claim 1, further comprising generating the prompt instructions for a large language model (LLM) to determine structured information based on the text data.
8. A system for generating structured information, the system comprising: a processor; and43IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO a memory storing instructions that, when executed by the processor, configure the system to: receive, via a server, text data collected from users; generate, using the server, identifiers associated with the text data and associated with the users; receive a prompt having instructions for a large language model (LLM) to determine structured information having user properties based on the text data; transmit the prompt, the text data, and identifiers to the LLM; determine, using the LLM, structured information based on the prompt, the text data and identifiers; and transmit, using an aggregation component, the structured information and the identifiers to an administrator.
9. The system of claim 8, further comprising a data gathering application that automatically records user keystrokes to generate the text data associated with the users.
10. The system of claim 8, further comprising a text store that stores the text data and identifiers.
11. The system of claim 8, wherein the prompt includes instructions for the LLM to analyze, summarize, classify and assess the text data to determine the structured information.
12. The system of claim 8, wherein the prompt comprises a pre-defined prompt.
13. The system of claim 8, further comprising an LLM prompt that stores that stores the prompt.
14. A computer-implemented method, comprising: receiving, via a data collection component, text data and user actions collected from users; transmitting the text data and user actions to a text assessment component and a behavioral assessment component, wherein the text assessment component and the behavioral assessment component each include respective large language models (LLMs); generating, using the text assessment component, a prompt based on text characteristics of the text data, the prompt including instructions for a cross-domain assessment component;44IPTS / 1289821 57.1Attorney Docket No. TAPP-OOIWO determining, using the behavioral assessment component, user behavioral patterns based on the user actions; and determining, using the cross-domain assessment component, structured information including properties of the users based on the prompt, text data, and user behavioral patterns.
15. The computer-implemented method of claim 14, further comprising storing, via a storage and pre-processing component, the text data.
16. The computer-implemented method of claim 14, wherein generating, using the text assessment component, the prompt comprises generating a parameterized prompt.
17. The computer-implemented method of claim 16, wherein generating, using the text assessment component, the parameterized prompt comprises initializing a prompt template to generate the parameterized prompt.
18. The computer-implemented method of claim 14, wherein determining, using the behavioral assessment component, user behavioral patterns comprises encoding user action sequences.
19. The computer-implemented method of claim 18, wherein determining, using the behavioral assessment component, user behavioral patterns comprises encoding temporal patterns between the user action sequences.
20. The computer-implemented method of claim 14, wherein determining, using the crossdomain assessment component, structured information including properties of the users based on the prompt, text data, and user interaction sequences comprises aligning user sequences using dynamic time warping.45IPTS / 1289821 57.1