Generative diversity system and method
Patent Information
- Application Number
- PCT/US2026/015181
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Priority Date
- 2025-02-24
- Filing Date
- 2026-02-13
- Publication Date
- 2026-08-27
Smart Images

Figure US2026015181_27082026_PF_FP_ABST
Abstract
Description
GENERATIVE DIVERSITY SYSTEM AND METHODSTATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0001] This invention was made with government support under Grant No. DRL2400781 awarded by the National Science Foundation. The government has certain rights in the invention.NOTICE OF MATERIAL SUBJECT TO COPYRIGHT PROTECTION
[0002] A portion of the material in this patent document is subject to copyright protection under the copyright laws of the United States and of other countries. The owner of the copyright rights has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the United States Patent and Trademark Office publicly available file or records, but otherwise reserves all copyright rights whatsoever. The copyright owner does not hereby waive any of its rights to have this patent document maintained in secrecy, including without limitation its rights pursuant to 37 C.F.R. § 1.14.BACKGROUNDTechnical Field
[0003] This disclosure pertains generally to systems and methods for analyzing text, and more particularly to systems and methods for assessing creativity and diversity in written content.Background Discussion
[0004] As the use of generative artificial intelligence (Al) expands rapidly, a critical concern is that these tools tend to homogenize ideation. The ideas produced by generative Al models such as large language models are often more homogeneous than ideas produced by groups of humans. Recent work has quantified the difference in idea diversity between groupsof human ideas and groups of ideas generated by Al models. This homogenization poses a substantial problem for institutions in industry, education, science and other fields that rely on generating genuinely new and different ideas to solve difficult problems and spark revolutionary innovations. It has become increasingly important for these institutions to identify potential employees, applicants, and collaborators who can use their own creativity rather than relying primarily on Al for idea generation. However, existing methods for assessing creativity and distinguishing human-generated from Al-generated content have limitations. Many approaches seek only binary classification of text as human or Al-authored, without providing deeper insights into the diversity and uniqueness of the ideas expressed. Traditional creativity assessments often rely on subjective human ratings, which can be timeconsuming and inconsistent at scale.
[0005] There is a need for improved systems and methods to quantitatively assess the diversity and uniqueness of ideas expressed verbally (e.g., oral, written, or text), in order to identify genuinely creative and original thinking. Such tools could support more effective evaluation of candidates, foster innovation, and help counteract the homogenizing effects of widespread Al use on ideation. For example, one use-case for the generative diversity system (which is also referred to herein as “the system”) is that it can be used to identify potential employees, applicants, etc. who use their own creativity rather than relying on generative Al as a primary basis for idea generation.BRIEF DESCRIPTION
[0006] To address this challenge, exemplary (i.e., example) embodiments in accordance with the present application, referred to herein as a generative diversity system, may leverage advanced natural language processing and machine learning techniques to assess the creativity and uniqueness of ideas input from text or any other verbal input (e.g., oral, written, or text). The generative diversity system may analyze 1) the diversity of individual ideas, 2) how much an idea expands a collective idea space, and 3) the diversity of trajectories through semantic embedding spaces to quantify different aspects of idea diversity and creativity. The system can analyze text using two main aspects: diversity of individual ideas and diversity of trajectory through embedding space. The system can assess how individual ideas contribute to group diversity.
[0007] In accordance with one aspect, a method is provided for assessing creativity and diversity in written content (or any transcribable content, e.g., speech-to-text). The method may comprise: receiving input text for analysis; generating embeddings for ideas in the input text; creating a heat map of idea locations in an embedding space; analyzing a trajectory of ideas through the embedding space; calculating diversity scores based onidea locations and trajectories; comparing individual text diversity to a larger group or dataset; and generating a report on generative diversity of the input text.
[0008] In accordance with another aspect, a system is provided for assessing creativity and diversity in written content. The system may comprise: a text input module configured to receive input text (or any verbal input); an embedding generation component configured to generate embeddings for ideas in the input text; a heat map generation module configured to create a heat map of idea locations in an embedding space; a trajectory analysis component configured to analyze a trajectory of ideas through the embedding space; a scoring and evaluation module configured to calculate diversity scores based on idea locations and trajectories and compare individual text diversity to a larger group or dataset; and a user interface configured to display results and generate a report on generative diversity of the input text.
[0009] Further aspects of the technology described herein will be brought out in the following portions of the specification, wherein the detailed description is for the purpose of fully disclosing preferred embodiments of the technology without placing limitations thereon.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWING(S)
[0010] The technology described herein will be more fully understood by reference to the following drawings which are for illustrative purposes only:
[0011] FIG. 1 is a block diagram illustrating an exemplary system architecture for a generative diversity system in accordance with the present disclosure.
[0012] FIG. 2 is a flow diagram illustrating an exemplary method for assessing creativity and diversity in written content in accordance with the present disclosure.
[0013] FIG. 3 depicts an exemplary scoring and evaluation module that can utilize several modules to accomplish its operations in accordance with the present disclosure.
[0014] FIG. 4 illustrates exemplary thematic profile analysis steps in accordance with the present disclosure.
[0015] FIG. 5 is a diagram illustrating novel generative diversity metrics’ correlations with human-rated creativity from a data set of creative short stories in accordance with the present disclosure.
[0016] FIG. 6 is a diagram illustrating an exemplary process in which Generative Diversity Metrics are generalized to new data in accordance with the present disclosure.
[0017] FIG. 7 is a diagram illustrating exemplary embedding space representations and heat maps indicating non-diverse thought in accordance with the present disclosure.
[0018] FIG. 8 is another diagram illustrating exemplary embedding space representations and heat maps indicating non-diverse thought in accordance with the present disclosure.DETAILED DESCRIPTION
[0019] Further aspects of the technology described herein will be brought out in the following portions of the specification, wherein the detailed description is for the purpose of disclosing preferred embodiments of the technology without placing limitations thereon.
[0020] The generative diversity system described herein may leverage two key aspects of diversity in human ideation relative to Al-generated content: diversity of individual ideas, and diversity of trajectory through embedding space.
[0021] In some exemplary embodiments, to assess diversity of individual ideas, each idea (which may be represented by sentences or words in this context) may be represented by its location in an embedding space. The embedding space may be generated using sentence embeddings or word embeddings. The generative diversity system may map the locations of ideas from a large set of texts, such as essays, responses to problem-solving prompts, stories, or other written content. This mapping may result in a “heat map” visualization where “hotter” regions indicate areas of the embedding space where ideas from the group of texts are frequently located, and “colder” regions indicate areas where ideas are less frequently located. For any individual text being assessed, such as text written by a candidate being considered for hiring, the embedding locations of ideas from that individual text may be overlaid onto the group heat map. The tool may then calculate the degree to which that text increases the diversity of ideas in the group. Ideas from the individual text that are located in relatively “cold” regions of the group heat map may yield higher generative diversity scores.
[0022] In some exemplary embodiments, to assess diversity of trajectory through embedding space, the generative diversity system may map the path taken from one idea to the next within each text. The tool may compare the trajectory of an individual text to the trajectories of other texts in a large group to determine the relative frequency or infrequency of the shape of that text's trajectory. This component may utilize a “shape uniqueness” feature as described in a separate patent application for “REPRESENTING THOUGHT TRAJECTORY THROUGH SEMANTIC SPACE AS VISIBLE GEOMETRIC FEATURES THAT ARE UPDATABLE AND RESPONSIVE,” US Patent Application No. 63 / 677,092, filed July 30, 2024, incorporated here in its entirety. Shape uniqueness may be a meta-property capturing the quantitative distinctness of a shape from other shapes, such as the distinctness of the shape generated by one individual from the shapes generated by a group of other individuals. Shape uniqueness may be calculated by combining multiple features of a shape into a single model to compare that combination of features across multiple shapes, or by using Al-based visual algorithms to identify holistic visual shape uniqueness.
[0023] As described in US Patent Application No. 63 / 677,092, a “Shape of Thought” (SoT) system can provide a developed metric that can represent a person’s thought trajectory through semantic space as visible, namable geometric features. Thought can comprise, for example, a creative performance, cognitive performance, idea progression, idea generation, and the like (e.g., as they think about a problem or generate novel ideas), as a person’s thought process happens, the person (also referred to herein as “participant” or “subject”) can provide verbal input that reflects or is representative of that thought process, to the SoT system. The verbal input can be received by the SoT system via, for example, a microphone associated with the SoT system, a touchpad, a keyboard, a stylus, or any other device that can be used to provide input. Based on the verbal input, and an analysis of the verbal input, the Shape-of-Thought system maps the trajectory of the progression through semantic space and calculates (and visually renders) the geometric features of this. Each idea within their thought process is a point along this trajectory. Ideas (points along the thought trajectory) can be mapped as the locations of word embeddings, sentence embeddings, or topic embeddings within a semantic space. For example, if a person is generating different solutions to a transportation problem involving moving a house, they might first think about physically transporting the whole house. Then they might consider 3-D printing the house in a different location. Then they might consider deconstructing the house and moving it in parts that can be recombined. Each of those three ideas has a mappable location in semantic space (i.e., the location of the sentence embeddings for each of those three sentences). Many features of the trajectory of the thought process through those ideas can be calculated and visually represented. The Shape-of-Thought system can calculate in high-dimensional spaces in which large language models are constructed, but the Shape of Thought system also provides for methods of dimension reduction that allow the Shape-of-Thought system to operate in three and / or two dimensions, such that shapes (and shape features) can be readily visualized for both measurement and interactive applications.
[0024] The Shape of Thought system can capture the spatial representation of textual embeddings in the context of assessment (e.g., creativity assessment, problem solving assessment, etc.). This process can start by extracting word, sentence, or topic embeddings from the library of all unique words, sentences, or topics that were generated by participants (subjects) that the Shape of Thought system is evaluating. Vectors are created by extracting word, sentence, or topic-embeddings (BERT, OpenAI’s Ada2, etc.). These embeddings represent the relationship of a certain text (word, sentence, or topic) to all other words, sentences or topics using hyperdimensional vectors (also called “Semantic Space”).
[0025] The Shape of Thought system can reduce the hyperdimensionality (more than three dimensions, e.g., hundreds to thousands of dimensions) of embeddings using, for example, the Uniform Manifold Approximation and Projection (UMAP) algorithm, so that every word or sentence is represented by a two or three-dimensional vector that relates to all other vectors. Each performance can be lemmatized, segmented, and tokenized, making it distinguishable in the semantic space of the embedding library being used.
[0026] Another way to reduce the complexity of a semantic space is by segmenting it into larger segments than words and sentences. To accomplish this, two methods can be used: topic modeling, and k-means clustering.
[0027] Example operations (i.e., methods) can be performed by the Shape of Thought system, which can comprise one or more computing devices (e.g., computer 900 as shown in FIG. 9 of US Patent Application No. 63 / 677,092) having a processor and memory (or other non-transitory computer-readable medium or storage device). The memory can store machine-readable and executable instructions that, when executed by the processor, facilitate performance of the example operations described here. The instructions can comprise one or more software modules (for example, modules for implementing features related to volume, fractional anisotropy, idea spacing, linearity, etc.). FIG. 1 of US Patent Application No.63 / 677,092 depicts an example of such data processing and display operations, and various features and aspects of the SoT operations. Those operations can comprise receiving a plurality of first verbal data inputs, wherein each of the plurality of first verbal data inputs are reflective of the subject’s thought, and wherein the first verbal data inputs comprise words or sentences. Based on the plurality of first verbal data inputs, the operations can comprise determining, using an analysis module executed by the SoT system, a plurality of data elements.
[0028] The data elements can comprise an idea spacing data element that relates to a semantic diversity measurement and a novelty of ideas measurement. The idea spacing data element can be derived from a calculation that is an average cosine distance between an embedding of the plurality of first verbal data inputs in a two or three-dimensional space. The idea spacing data element can be indicative of a distant and diverse idea. The idea spacing data element can also be indicative of a similar or conventional idea.
[0029] The data elements can also comprise a volume data element that relates to a semantic extent of richness of creative performance measurement. The volume data element can be derived from a calculation that is the minimum volume of a multidimensional ellipsoid that encompasses an embedding of the plurality of first verbal data inputs in the semantic space. The volume data element can also be indicative of a creative performance measurement that is semantically large and rich. The volume data element can also be indicative of the creative performance measurement that is semantically limited and narrow.
[0030] The data elements can also comprise a fractional anisotropy data element that relates to a semantic directionality and coherence of creative performance measurement. The fractional anisotropy data element can be derived from a calculation as a ratio of a largest eigenvalue to the sum of all eigenvalues of a covariance matrix of embeddings of the plurality of first data inputs in the three-dimensional space. The fractional anisotropy data element can comprise a value of “0,” which can be indicative of an isotropic semantic space that is representative of a creative performance measurement with theme and direction. The fractional anisotropy data element can also comprise a value of “1,” which can be indicative of an anisotropic semantic space that is representative of the creative performance measurement that is exploratory and diverse.
[0031] The data elements can also comprise a linearity data element that relates to a semantic progression and continuity of creative performance measurement. The linearity data element can be derived from a calculation of a correlation between an order of the plurality of first data elements and the angles between the responses, derived from embeddings in a two or three-dimensional space. The linearity data element can be indicative of a creative performance that follows a linear semantic trajectory that is representative of the creative performance being logical and consistent. The linearity data element can also be indicative of a creative performance that deviates from the linear semantic trajectory and is representative of the creative performance being surprising and unexpected.
[0032] The data elements can also comprise a semantic network nodes and edges data element related to words or sentences derived from the plurality of first data inputs. The semantic network nodes and edges data element can utilize graph theory methods (e.g., clustering coefficient, modularity) to extract additional features.
[0033] The data elements can also comprise an optimal path data element indicative of a density related to a semantic space. The optimal path data element can be calculated as an average of shortest paths between all pairs of nodes.
[0034] The data elements can also comprise a topic modeling data element comprising an arrangement of words or sentences around topics or themes based on their semantic similarity. The topic modeling data element can relate to additional features such as number of nodes / topics, themes, and semantic nodes.
[0035] The data elements can also comprise a shape uniqueness data element reflective of a quantitative distinctness of a shape from other shapes. The shape uniqueness data element can be calculated by combining multiple features of a shape into a single model to compare that combination of features across multiple shapes. The shape uniqueness data element can be calculated using Artificial Intelligence (Al) based visual algorithms to identify shape uniqueness.
[0036] The operations performed by the SoT system can comprise mapping a trajectory based on the plurality of data elements, wherein the trajectory comprises one or more geometric features displayable on a display device. The trajectory comprising the geometric features can then be displayed. As will be further detailed below, once a subject sees his or her thought in the form of geometric shapes, the subject can respond to the feedback by changing the verbal input that they submit (e.g., another set of words or sentences).
[0037] The operations performed by the SoT system can then comprise receiving, by the SoT system, a plurality of second verbal data inputs related to this response to an evaluation of the trajectory by the subject, wherein the plurality of second verbal data inputs is representative of a second set of words or sentences. Based on the second set of verbal inputs, the SoT system can then determine another set of data elements, and map a second trajectory. The subject can continue to provide inputs until satisfied with the “shape” of his or her thought, or if the subject discontinues, the operations can end. This interactive display and response, which can be considered real time or near real time, can be used to train a subject to respond in a manner consistent with a desired personality, skill, occupation, etc.
[0038] The SoT system can also be used to distinguish “human” from “Al” generated texts.
[0039] Construction of semantic spaces. As mentioned above, the shape of thought features are metrics designed to capture the spatial representation of textual embeddings in the context of creativity assessment. The first step of the process can be to generate an embedding space incorporating all the different semantic performance-sections (words and sentences). This can be performed by extracting embeddings from all the sentences or words that were generated by the participants for all the semantic tasks. Vectors are created by extracting word-or sentence-embeddings by using one of the standard embedding libraries (e.g., BERT, or OpenAl). These embeddings represent the relationship of a certain text (word or sentence) to all other words or sentences using hyperdimensional vectors. A dimension reduction method (e.g., UMAP; Mclnnes et al., 2018) is then used to reduce the hyperdimensionality of embeddings so that every word or sentence can be represented by a two- or three-dimensional vector which relates to all other vectors (for example, FIGs. 2-8 of US Patent Application No.63 / 677,092 illustrate this concept). To optimize the dimension reduction process, and to select the best performing specification, namely, the specification that retains the most information from the hyperdimensional space in the low-dimensional space, diverse specifications are calculated. The best performing specification is then chosen by correlating the similarity matrices of the low-dimensionality spaces (LD) and high-dimensionality spaces (HD). The Space with the highest correlations is chosen to be used in the next step.
[0040] The calculation of SoT features. In the next step, each performance is lemmatized, segmented, and tokenized in python. In the next step SoT features can be extracted using the HD and LD spaces using a customized python code. The SoT features include: Volume, Optimal Path, Fractional Anisotropy, Idea-Spacing, clustering of performance into nodes, and features relating to how linear is the performance.
[0041] Training and testing a classification model to recognize and predict human us. Al behavior. The final stage in this process can be to train and test a classification model to recognize and predict human and Al texts by using the SoT features. To accomplish this, the SoT features that were extracted from the creative performance can be received as input, which are then used to fine-tune a model to predict the classification of a text (human or Al). The dataset is prepared for training by the SoT system and validated. The hyper-parameters number of hidden-layers and decay are optimized, which is then tested on the validation set. The best performing hyperparameters are chosen, and the optimized model can be used to make predictions and test them on the test dataset. A confusion matrix can be used to test the performance and determine the accuracy, sensitivity, and specificity of the model, in relation to predicting human performance, Al performance, or both. Post-hoc analysis can then be performed to determine the importance or value of the predictors, which determines the importance of individual features to the classification model, with greater values indicative of larger importance. Once the finetuned model is trained, validated, and tested, the model can be used to classify new texts. To extend the output beyond a binary classification, the probability of being “human” or “Al” given the input can be used to get a more nuanced parameter of classification. The feature vectors can be used in a forward propagation in the neural network, which starts with the multiplication of the weight matrix for the first layer, and biases are added, next an activation function can be applied to introduce non-linearity. This process can be repeated for each layer. The final layer produces raw scores (logits) for each class in the problem, in our case “human” and “Al”. The logits can be converted into probabilities using the softmax function, which ensures that (1) the probabilities are non-negative, and (2) the probabilities sum to 1 across all classes. In short, P("human"|input) is calculated using the softmax function applied to the logits from the neural network's final layer.
[0042] The SoT system’s distinguishment between “human” from “Al” generated texts is further elaborated upon below:
[0043] 1. Inputs to the Model
[0044] The model can receive a feature vector (input) as input. This feature vector represents the numerical or encoded representation of the data you want to classify.
[0045] 2. Forward Propagation in the Neural Network. The neural network computes a series of transformations to predict probabilities for each class. Steps:1. Linear Transformation: The input vector is multiplied by the weight matrix for the first layer, and biases are added:z=W-input+bwhere W is the weight matrix, b is the bias vector, and z is the linear output before activation.2. Activation Function: An activation function (e.g., ReLU, tanh) is applied to introduce non-linearity:a=f(z)This process repeats for each layer until the final layer.3. Final Layer: The final layer produces raw scores (logits), one for each class:1 ogits= W out • 1 ast-1 ay er-output+boutFor a two-class problem ("human" and "Al"), there are two logits: one for each class.
[0046] 3. Softmax TransformationThe logits are converted into probabilities using the softmax function, which ensures that:• The probabilities are non-negative.• The probabilities sum to 1 across all classes.For the "human" class, the probability is computed as:P("human"linput)=exp (logit"human") / (exp (logit"human")+exp (logit" Al"))For the "Al" class:P("AI"linput)=exp (logit" AI") / (exp (logit"human")+exp (logif'AI"))Explanation:• exp (logit): Exponentiation ensures all values are positive.• The denominator normalizes the probabilities so they sum to 1.
[0047] 4. OutputThe type = "raw" option in predict() provides these softmax probabilities:• P("human"linput)• P("AI"linput)
[0048] 5. Interpretation• If P("human"linput)=0.7 the model predicts a 70% chance that the input belongs to the "human" class.• These probabilities can be thresholded (e.g., >0.5) for hard classification or used directly as confidence scores.
[0049] In summary, P("human"linput) is calculated using the softmax function applied to the logits from the neural network's final layer.
[0050] In some exemplary embodiments, the generative diversity system may calculate and utilize multiple novel metrics to quantify different aspects of idea diversity and creativity, including but not limited to: Idea Cluster Uniqueness (ICU), Diversity Polarity (DP), Diversity Polarity of Idea Cluster Uniqueness (DPICU), Thematic Profile (TP), Thematic Profile Idea Cluster Uniqueness (TPICU), Diversity Polarity of Thematic Profile Idea Cluster Uniqueness (DPTPICU), Idea Cluster Ambiguity (ICA), and Diversity Polarity of Idea Cluster Ambiguity (DPICA). These metrics, discussed further below, may provide further insights into different aspects of idea diversity and creativity.
[0051] The generative diversity system may produce visualizations of embedding spaces and generate interpretable reports and scores based on the calculated metrics. These outputs may be used to support creativity assessment, candidate evaluation, creativity training, and identification of genuinely original thinking. This tool can provide an assessment of creativity. The uniqueness of an individual’s ideation can be a core component of creativity according to widely-used standards for creativity (e.g., Runco & Jaegger, 2012). Because creativity is among the most valuable capabilities in the current and future economy (e.g., according to World Economic Forum, 2023), the insights into creativity provided by the generative diversity system can likely be valuable in their own regard.
[0052] Another application of the generative diversity system is to support institutions in identifying groups of individuals who use their own creativity to generate ideas, rather than being over-reliant on generative Al tools for ideation. Importantly, this tool is not devised to provide a binary classification of Al-generated vs. human-generated text. Rather, becausehomogenization of ideation appears to be a fundamental feature of generative Al models, higher generative diversity scores will tend to be negatively associated with the use of generative Al in ideation. Thus, selecting a group of applicants with higher generative diversity scores will tend to yield a greater number of applicants who have produced their own ideas rather than relying on ideas generated by Al. This will have value both for identifying creativity generally, and for identifying individuals who are willing and able to generate their own ideas, which can be very useful for solving difficult problems, as noted above.
[0053] As such, applications of the generative diversity system can include, for example: automated creativity assessment in both professional and educational contexts; employers and universities identifying individual differences in thought processes (not just thinking outcomes) as a means of selecting candidates, assigning responsibilities to specific individuals, etc.; individuals understanding their own thought processes; revealing the nature of changes in thought process (i.e., seeing how generative diversity changes with creativity training); educators or employers understanding their students’ or employees thought processes, which can be used as a basis to tailor educational or training curricula to those students, or changes in generative diversity can be used to evaluate the efficacy of curricula.
[0054] FIG. 1 illustrates an exemplary system architecture for a generative diversity system 100 (also referred to herein as “the system,” or “system 100”) in accordance with the present disclosure. The system 100 may comprise one or more computing devices having one or more processors 102 and memory 104. The memory 104 may store instructions that, when executed by the one or more processors 102, cause the system 100 to perform operations for assessing creativity and diversity in written content. These instructions can take the form of one or more modules, which can be computer programs, or computer routines. In example embodiments, the one or more computers can be a personal computer (such as a desktop or laptop), tablet running an application, a smartphone running and application, or the like. In example embodiments, the generative diversity system 100 can be a server, which can be accessible via a communications network 105 by other remote computer(s).
[0055] The system 100 may include a text input module 106 that can be a text preprocessing module configured to receive input text for analysis. The text can be input by an input device 107, which can be a keyboard, touchscreen, or a microphone, which can capture verbalized oral input (e.g., sound, speech, a played recording, etc.) and the system 100 can convert that input to text. The system 100 can also process files, which can be downloaded (or uploaded) to it, and process the file, extracting text.
[0056] The input text may comprise, as in an example use case, college admissions essays, job application essays, student assignments, or other written content to be evaluated for creativity and diversity.
[0057] An embedding generation module 108 may be configured to generate embeddings for ideas in the input text. The embeddings may comprise vector representations of words, sentences, or documents that capture semantic meaning. In some exemplary embodiments, the embedding generation component 108 may utilize pre-trained language models or sentence transformers to generate the embeddings.
[0058] A heat map generation module 110 may be configured to create a heat map of idea locations in an embedding space. The heat map may visually represent the frequency and distribution of ideas within the embedding space, with "hotter" regions (e.g., which can be indicated by color and shades of color, or different intensities of color, such as red for hot) indicating areas where ideas are more frequently located.
[0059] A trajectory analysis module 112 may be configured to analyze the trajectory or path taken from one idea to the next within the input text. This component may map the sequence of ideas through the embedding space to assess the diversity of thought progression, as demonstrated by shape uniqueness.
[0060] A scoring and evaluation module 114 may be configured to calculate diversity scores based on idea locations and trajectories. As described below in FIG. 3, the scoring and evaluation module 114 may implement various metrics and algorithms for assessing creativity and diversity.
[0061] The system 100 may further include a user interface 116 configured to display results and generate reports on the generative diversity of input text. The user interface can appear on, for example, a touch screen, a monitor, or some other display device. The user interface 116 may provide visualizations of embedding spaces, heat maps, idea trajectories, and various creativity and diversity metrics to aid in interpretation and decision-making. The user interface 116 may allow users to interact with the system 100, submit texts for analysis, and view results.
[0062] The system 100 may also include a database 118 for storing text corpora, embeddings, analysis results, and other relevant data. The database can be stored in memory (e g., a memory device such as a hard drive, or solid-state hard drive).
[0063] FIG. 2 illustrates an exemplary method 200 for assessing creativity and diversity in written content in accordance with the present disclosure. The method 200 may beimplemented by the exemplary generative diversity system 100 described above, or by other suitable computing devices or systems.
[0064] At step 202, input text is received for analysis. The text can be input by keyboard, touchscreen, or verbalized orally, for example, captured using a microphone or other audio input device, and converted to text. The input text may comprise one or more documents (e.g., files, or scanned in files), such as college admissions essays, to be evaluated for creativity and diversity.
[0065] At step 204, embeddings are generated for ideas in the input text. This may involve using pre-trained language models or sentence transformers to convert words, sentences, or documents into vector representations that capture semantic meaning.
[0066] At step 206, a heat map of idea locations in an embedding space is created. The heat map may visually represent the frequency and distribution of ideas within the embedding space. The heat map can use different colors, and can also use various shades and intensities of colors, to represent “hot” and “cool” areas, with red representing hot and blue representing cold.
[0067] At step 208, the trajectory of ideas through the embedding space is analyzed. This may involve mapping the sequence of ideas to assess the diversity of thought progression within the input text.
[0068] At step 210, diversity scores are calculated based on idea locations and trajectories. Various metrics and algorithms may be applied to assess creativity and diversity, including but not limited to ICU, DP, DPICU, TP, TPICU, DPTPICU, ICA, and DPICA as described in this disclosure.
[0069] At step 212, individual text diversity (diversity of an individual text) is compared to a larger group or dataset. This comparison may provide context for interpreting the creativity and diversity scores of the input text relative to a broader corpus of similar documents. As such, this comparison to a larger body of data can aid in determining diversity.
[0070] At step 214, a report is generated on the generative diversity of the input text. The report may include visualizations, metrics, and interpretations to aid in understanding thecreativity and diversity of the analyzed content. The report can be in the format of a visual representation, such as the example graph shown in FIG. 7.
[0071] As shown in FIG. 3, the scoring and evaluation module 114 may implement various metrics and algorithms for assessing creativity and diversity, including but not limited to: Idea Cluster Uniqueness (ICU), Diversity Polarity (DP), Diversity Polarity of Idea Cluster Uniqueness (DPICU), Thematic Profile (TP), Thematic Profile Idea Cluster Uniqueness (TPICU), Diversity Polarity of Thematic Profile Idea Cluster Uniqueness (DPTPICU), Idea Cluster Ambiguity (ICA), Diversity Polarity of Idea Cluster Ambiguity (DPICA), and Functional Diversity Metrics. In example embodiments, these metrics and algorithms can be executed by one or more computer-executable modules, for example, as shown in FIG. 3, the scoring an evaluation module can draw from an Idea Cluster Uniqueness (ICU) module 302, a Diversity Polarity (DP) module 304, a Diversity Polarity of Idea Cluster Uniqueness (DPICU) module 306, a Thematic Profile (TP) module 308, a Thematic Profile Idea Cluster Uniqueness (TPICU) module 310, a Diversity Polarity of Thematic Profile Idea Cluster Uniqueness (DPTPICU) 312, a Idea Cluster Ambiguity (ICA) module 314, a Diversity Polarity of Idea Cluster Ambiguity (DPICA) module 316, and Functional Diversity Metrics module 318.
[0072] Idea Cluster Uniqueness (ICU): ICU may capture how far an idea or idea set is from clusters of related ideas in a larger corpus of creative ideas. The scoring and evaluation module 114 may utilize hierarchical density-based spatial clustering of applications with noise (HDBSCAN) to identify clusters of ideas (via word, sentence, or document embeddings) that share embedding features (i.e., meaning and context). HDBSCAN can generate global-local outlier scores from hierarchies (GLOSH) that can simultaneously capture how much an idea separates itself from local clusters and the global set of ideas. When GLOSH scores are applied to creative ideas, these scores represent the uniqueness of an idea relative to clusters of related ideas - which can be termed Idea Cluster Uniqueness (ICU). As such, the higher the score, the more unique the idea, relative to clusters of related ideas. In a wellvalidated dataset of creative short stories, ICU scores correlated very highly with humanrated creativity (r = .72), nearly approaching ceiling levels of inter-rater reliability between human raters. See Appendix X for visualization of this correlation.
[0073] Diversity Polarity (DP): DP may capture how much an idea or idea set diverges from a larger corpus of ideas. The scoring and evaluation module 114 may compute cosine semantic distances between all pairwise sentences (or other embedding approach including word or document) in the corpus to get a total diversity score, then calculate difference scores by removing individual idea sets and recomputing diversity. That is, it can compute the cosine semantic distance between all pairwise sentences in the corpus to get total diversity score. Next, for each idea set (e.g., a person’s story), their sentence embeddings are removed from the larger corpus; then total diversity is recomputed on the remaining sentence embeddings. A difference score is computed by subtracting this diversity score (that excludes a single person’sideas) from the total diversity scores (contains all ideas from the entire corpus, including the single target person). This difference score can be termed Diversity Polarity (DP) because a positive score for a specific person’s ideas can indicate their ideas added diversity to the larger idea pool, whereas a negative score can indicate their ideas removed diversity from the larger idea pool. In other words, positive scores can reflect increased heterogeneity in the idea pool, whereas negative scores can reflect increased homogeneity in the idea pool, for each target idea set of interest.
[0074] Diversity Polarity of Idea Cluster Uniqueness (DPICU): DPICU may assess how the inclusion or exclusion of a specific idea affects the density of idea clusters. The scoring and evaluation module 114 can apply HDBSCAN cluster analysis to generate ICU scores for the entire corpus, then recalculate scores with individual idea sets removed to determine their impact on cluster density. As an example, if an idea increases the diversity of a corpus of ideas, HDBSCAN cluster analysis generates less dense (i.e., tightly packed, highly related) ideas. In contrast if an idea decreases the diversity of a corpus of ideas, then HDBSCAN cluster analysis can generate more densely packed ideas. To compute DPICU, each idea can first be converted to a sentence embedding. Then, HDBSCAN can be used to generate a GLOSH score (i.e., ICU score when applied to creative ideas) for each idea. The average ICU score for the entire corpus can reflect how densely packed, that is, how strongly related ideas are in the corpus. Next, each idea set can be removed from the corpus, and HDBSCAN can be conducted to generate a new set of ICU scores, that are again, averaged across the corpus (excluding the target idea set). A difference score can be computed by subtracting the average ICU score from the incomplete set (i.e., that excludes the target idea set) from the average ICU score from the complete set. A positive DPICU score can indicate the target idea set increased the diversity of ideas by decreasing the density of related ideas in clusters. A negative DPICU score can indicate the target idea set decreased the diversity of ideas by increasing the density of related ideas in clusters.
[0075] Thematic Profile (TP): TP may capture how much a target's ideas diverge from common and rare themes in a corpus of creative ideas. The scoring and evaluation module 114 may utilize topic modeling to identify themes, convert themes and target ideas into embeddings, and compute semantic distances between theme embeddings and idea embeddings to create thematic profiles. First, a large corpus of creative ideas can be reduced into a smaller, more interpretable set of themes using topic modeling (e.g., Roberts et al., 2019; Step 1, FIG. 4). These topics can then be used to identify and label common and rare themes present in the corpus. Next, each theme can be converted into an embedding (Step 2, FIG. 4). The scoring and evaluation module 114 next takes the target creative ideas it would like to explain (i.e., profile) and converts them into embeddings as well (Step 3, FIG. 4). Then, the scoring and evaluation module can compute the distance between each theme embedding and each creative idea embedding (Step 4, FIG. 7), where higher similarity values indicate the theme and target creative ideas are more semantically similar and higher semantic distance indicates more dissimilarity. The collection of distances between themes and ideas places eachidea in relative position to others’ ideas and creates the thematic profde. The thematic profile can reveal how semantically similar all ideas are to common and rare themes (see Step 4, FIG.4). The higher the semantic distances from both common and rare themes, the more the target ideas increase the diversity of the idea pool.
[0076] Thematic Profile Idea Cluster Uniqueness (TPICU): TPICU may capture the relative uniqueness of thematic profiles. The scoring and evaluation module 114 may apply HDBSCAN to thematic profiles to generate GLOSH scores representing the uniqueness of profiles relative to clusters of related profiles. Higher TPICU scores indicate the TP separates itself from clusters of related TP profiles. Each thematic profile from each target idea set contains a semantic distance score from each theme identified in the corpus (e.g., 8 themes in the creative short story data set, as shown in FIG. 8). HDBSCAN is used to capture the relative uniqueness of the TP profiles relative to clusters of related TP profiles, again, using GLOSH scores.
[0077] Diversity Polarity of Thematic Profile Idea Cluster Uniqueness (DPTPICU): DPTPICU may assess how a target idea set impacts the heterogeneity or homogeneity of thematic profiles relative to clusters of related profiles. The scoring and evaluation module 114 may compute TPICU scores for the entire corpus, then recalculate scores with individual idea sets removed to determine their impact on profile diversity. This was not computed for the creative short story data set because the thematic profile consists of 8 semantic distance values and there is a single profile for each person, which is a small number of features and rows on which to extract clusters. Positive DPTPICU scores will indicate that a target idea set increases heterogeneity of thematic profiles relative to a cluster of related thematic profiles, whereas negative scores will indicate a that a target idea set increases the homogeneity of thematic profiles relative to a cluster of related thematic profiles.
[0078] Idea Cluster Ambiguity (ICA): ICA may capture the degree to which an idea belongs to more than one cluster. The scoring and evaluation module 114 may utilize soft-clustering analysis to assign ideas to multiple clusters based on membership degree, with lower membership degrees indicating higher ambiguity and creativity. Each idea can be first converted to an embedding. Then, fuzzy or soft-clustering analysis (e.g., Ferraro, Giordani, & Serafini, 2019) including: fuzzy k-means and its variants, fuzzy c-means and its variants, and spatial fuzzy c-means, can be conducted on the idea embeddings to identify clusters of related ideas. In contrast to HDSCAN, soft-clustering assigns ideas to multiple clusters by membership degree, that is, the probability the idea belongs to each cluster. The higher the membership degree, the more likely the idea clearly belongs to a single cluster, and therefore the lower in creativity and diversity. A lower membership degree score can indicate the idea’s cluster membership is ambiguous or uncertain - that is, it doesn’t fit neatly into a single cluster of related ideas. The membership degree values can be reverse-scored so that higher ICAvalues indicate the ideas are more unique, relative to clusters of related ideas. In a well-validated dataset of creative short stories, ICA scores correlated vary highly with human-rated creativity (r = .59). See FIG. 8 for visualization of this correlation.
[0079] Diversity Polarity of Idea Cluster Ambiguity (DPICA): DPICA may capture how an idea or idea set impacts the ambiguity of cluster assignments for the entire corpus. The scoring and evaluation module 114 may compute average ICA scores for the corpus, then recalculate scores with individual idea sets removed to determine their impact on overall ambiguity. For example, if an idea increases the ambiguity of cluster assignments, then it does not fit neatly into a single cluster, and is likely more creative and increases the diversity of ideas in a corpus of ideas. In contrast, if an idea decreases the ambiguity of cluster assignments, then it does neatly fit into a single cluster, and is likely less creative and decreases the diversity of ideas in a corpus of ideas. To compute DPICA, each idea can first be converted to a sentence embedding. Then, soft-clustering can be applied to the embeddings and used to assign ideas to multiple clusters based on membership degree. The average ICA score for the entire corpus reflects how neatly each idea fits into single clusters of related ideas. Next, each idea set can be removed from the corpus, and soft-clustering is conducted to generate a new set of ICA scores, that are again, averaged across the corpus (excluding the target idea set). A difference score can be computed by subtracting the average ICA score from the incomplete set (i.e., that excludes the target idea set) from the average ICA score from the complete set. A positive DPICA score can indicate the target idea set increased the diversity of ideas by decreasing the probability each idea fits neatly into single clusters of related ideas in clusters. A negative DPICA score can indicate the target idea set decreased the diversity of ideas by increasing the probability each idea fits neatly into single clusters of related ideas in clusters.
[0080] Functional diversity metrics: these can include functional richness, functional specialization, and functional originality, described further below.
[0081] FIG. 4 illustrates exemplary thematic profile analysis steps in accordance with the present disclosure. The figure shows the process of identifying common and creative themes using topic modeling, converting themes and target ideas into embeddings, and computing semantic distances to derive thematic profiles.
[0082] FIG. 5 is a diagram illustrating novel generative diversity metrics’ correlations with human-rated creativity from a data set of creative short stories. Each thematic profile from each target idea set contains a semantic distance score from each theme identified in the corpus (e.g., 8 themes in the creative short story data set, as shown in FIG. 8). In a well-validated dataset of creative short stories, ICA scores correlated vary highly with human-rated creativity (r = .59); FIG. 8 provides a visualization of this correlation.
[0083] As the correlation matrix in FIG. 5 demonstrates, the new Generative Diversity metrics can predict human creativity in the moderate to very strong range. Critically, the correlations between each of the new metrics were in the low to moderate range, indicating each is capturing something unique about the creativity of the text. This suggests not only that these metrics are distinct from each other, but also that they can be combined for performance that exceeds any one metric alone.
[0084] FIG. 6 describes a process in which GenDiv Metrics is generalized to new data. This can be a computational method that allows each metric from modules that can implement the process depicted in FIG. 6 to generate scores for new data without undergoing, unsupervised machine learning or analysis with a large corpus of ideas. An example of this method can be seen in FIG. 6, as applied to ICU scores. If one relied solely on HDBSCAN cluster analysis to derive Idea Cluster Uniqueness (ICU) scores, then a new cluster analysis would most likely need to be conducted, and new ICU scores generated for the entire corpus for each new data point. In many applications, this may not be feasible, so this developed approach relies on a novel and automated method of hyperparameter tuning and ensemble machine learning to allow the generation of ICU scores for new data efficiently and reliably. As FIG. 6 shows, after converting ideas into embeddings, a method of hyperparameter tuning can be applied to HDSCAN, where HDBSCAN can be conducted on a full range of hyperparameter values (i . e. , the minimum cluster size parameter) and our approach selects the HDBSCAN model that minimizes the percent of maximum GLOSH score assignments (i.e., GLOSH value = 1) in order to generate the more optimal continuous distribution (i.e., minimizes skew), instead of a GLOSH distribution that often skews to a high percent of maximum score assignments. While HDBSCAN is an unsupervised machine learning approach, this method of hyperparameter tuning can be done in an automated fashion without the need for human curation.
[0085] As FIG. 6 Step 5 shows, next, the GLOSH scores can be extracted from the HDBSCAN model with a hyperparameter setting tuned to minimize the GLOSH maximal values. One can refer to these GLOSH scores as ICU scores from HDBSCAN, and EDGE score more broadly for the collection of Generative Diversity metrics to reflect they are designed to capture how much an idea expands divergent generative expressions (i.e., EDGE). In Step 6 of FIG. 6, ensemble machine learning can be applied using gradient boosted machine learning (GBM) using new EDGE scores as labels and idea embeddings and predictors. This solves a critical problem, which is how to apply an unsupervised learning approach (i.e., HDBSCAN) to new data as seamlessly as one can apply new data to a model derived from supervised machine learning. Now, using the GBM trained to generate EDGE scores, the system can generate EDGE scores on new data unseen by HDBSCAN and the GBM.
[0086] Using the short stories data set as described above with respect to FIG. 5, the validity of this approach can be demonstrated by showing that the GBM-derived ICU scores applied to new data, unseen by both the cluster analysis and GBM ensemble model, strongly correlate with human creativity and reliably across 200 bootstrapped samples of 80% of the data, with a r = .55, 95% bootstrapped Confidence Interval [.32, .69],
[0087] The 6-step method as shown in FIG. 6 can be applied to all generative diversity metrics with some modification for each metric. For example, Diversity Polarity (DP) does not employ cluster analysis, so Steps 3-4 are unnecessary. For ICA and DPICA, a different form of hyperparameter tuning would be needed as they employ soft-clustering instead of HDBSCAN. Instead, an automated hyperparameter tuning that selects the cluster model that maximizes the variance of membership degree values would be used. So, Steps 3 and 4 can be modified, so that the ICA DPICA scores come from a model with hyperparameters that maximize the variance in membership degree. The remaining steps remain unchanged. However, a key and shared innovation for all generative diversity metrics, is that a GBM ensemble learning model is trained to generate each individual generative diversity metric without the process of unsupervised learning (i.e., cluster analysis, topic modeling).Increasing Reliability of Generative Diversity Metrics
[0088] Multiple techniques will be used to increase the reliability and generalizability of all generative diversity metrics. One approach that can be used is to compute all generative diversity metrics with at least three different large language models, to get better coverage conceptual idea space, as determined by different LLM architectures (e.g., BERT, LLAMA, multilingual e5), and then standardize the scores, and then either average them together or use them separately in a more complex model. Another approach for the clustering-based generative diversity metrics is to generate each generative diversity metric using at least two different criteria to tune hyperparameters that could include tuning to maximize the variance, minimizing the percent of maximum score assignments, and minimizing the percent of minimum scoring assignments. Scores derived each of these three hyperparameter tuning criteria will reflect different facets of generative diversity, which can again be averaged together or used separately in a more complex model. Application of both multiple LLMs and hyperparameter settings will increase the reliability of generative diversity and generalizability across domains of application.Expanding Divergent Generative Expression (EDGE) Types
[0089] As FIG. 7 depicts, to facilitate the application of generative diversity metrics to education and industry settings, we developed a method that partitions the continuous generative diversity metrics into 4 categories: 1) Central EDGE Type, 2) Peripheral EDGE Type, 3) Multidimensional EDGE Type, and 4) Singular EDGE Type. Lower scores on an individual generative diversity metric can reflect an idea aligns closely with a common cluster or theme (i.e. Central type), whereas moderate scores reflect that an idea is approaching the periphery of a cluster / theme (i.e., Peripheral type). Moderately high scores reflect that an idea falls in the in-between spaces, so not in a specific cluster / theme, but rather share elements of multiple themes (i.e., Multidimensional type). The highest scores reflect that an idea separatesitself from all themes and most other ideas (i.e., Singular type). This categorization scheme builds on the power of visualizing ideas in semantic space relative to clusters / themes and others’ ideas and adds a name to help convey the meaning of being in a particular location in idea space.
[0090] Nearly all the above metrics already derive scores by comparing to a larger body of text, and index how much each idea separates itself from common clusters / themes and other ideas generally
[0091] FIG. 8 illustrates another exemplary embedding space representations and heat maps in accordance with the present disclosure. FIG. 8 depicts how the new generative diversity scores (termed EDGE scores here) places homogenized Al ideas in common / central spaces and human ideas at the edges of the collective idea space.
[0092] The systems and methods described herein may provide several advantages over existing approaches for assessing creativity and diversity in written content. By leveraging advanced natural language processing techniques and novel metrics, the generative diversity system can provide more nuanced and comprehensive evaluations of creativity that go beyond simple binary classifications of human-generated versus Algenerated text.
[0093] The system's ability to analyze both the diversity of individual ideas and the diversity of thought trajectories offers insights into multiple dimensions of creativity. This multi-faceted approach can help institutions identify individuals who demonstrate genuine creativity and original thinking, rather than relying heavily on generative Al tools for ideation.
[0094] Furthermore, the interpretable nature of metrics like Thematic Profile and Idea Cluster Uniqueness provides actionable insights that can be used to understand and foster creativity in educational and professional contexts. The system 100's flexibility allows for application across various domains, from college admissions to employee selection and beyond.ADDITIONAL TECHNIQUES AND IMPLEMENTATIONSVisual Embedding Space Representation
[0095] The method of Visual Embedding Space Representation can provide a systematic approach to distinguishing unique essays within a pool by leveraging advanced natural language processing and deep learning techniques. This method can have several steps designed to convert textual data into visual formats that can be analyzed using Convolutional Neural Networks (CNNs) for pattern detection.• a. Conversion of Essays into Embeddings: Each essay in the pool can be transformed into a high-dimensional vector representation that encapsulates its semantic content. This transformation can be achieved by utilizing a highperforming sentence embedding model, such as the multilingual sentencetransformer referenced in the prior metrics. The embedding model processes the textual data of each essay and outputs a vector in a high-dimensional space where semantically similar essays are positioned closer together, and dissimilar essays are farther apart.• b. Dimensionality Reduction: Given that the embeddings reside in a highdimensional space, direct visualization and analysis are computationally challenging. To address this, dimensionality reduction techniques can be applied to project the high-dimensional embeddings into a lower-dimensional space, typically 2D or 3D. Methods such as Uniform Manifold Approximation and Projection (UMAP) or t-distributed Stochastic Neighbor Embedding (t-SNE) can be employed for this purpose. These techniques are adept at preserving the local and global structure of the data, and can facilitate ensuring that the relative distances between essays in the reduced space better reflect their semantic similarities and differences accurately.• c. Visual Representation: With the embeddings projected into a 2D or 3D space, visual representations such as scatter plots or heatmaps are created. In these visualizations, each point corresponds to an essay, positioned based on its reduced embedding coordinates. Additional dimensions of information can be incorporated by color-coding or size-coding the points using metrics like the Idea Cluster Uniqueness (ICU) scores. For instance, essays with higher ICU scores — indicating greater uniqueness — can be highlighted in a distinct color or assigned larger marker sizes. This visual augmentation aids in intuitively identifying essays that stand out from the rest.• d. Application of Convolutional Neural Networks for Pattern Detection: To automate and enhance the detection of unique essays, the visual representations canbe processed using Convolutional Neural Networks (CNNs), which are proficient in identifying patterns within images. The process can involve two substeps:Substep 1: Image Generation: The scatter plots or heatmaps generated in the previous step can be converted into image formats suitable for processing by CNNs. These images capture the spatial distribution of essays in the reduced embedding space, along with any additional encoded metrics.Substep 2: CNN Analysis: A CNN model can be trained on these images to detect clusters and outliers. The training process can involve feeding the CNN with labeled images where regions of high density correspond to common themes, and sparse regions represent unique essays. The CNN can be operative to learn to classify these regions based on visual features such as the concentration and dispersion of points. Once trained, the CNN can efficiently process new images to identify and classify essays that are potential outliers due to their semantic uniqueness. This automated detection can significantly enhance the scalability and objectivity of the uniqueness assessment. CNNs are adept at extracting hierarchical features from images, making them suitable for identifying complex patterns in visualized data. By transforming textual data into embeddings and then into visual representations, the system 100 can leverage both NLP and computer vision techniques. This method can be effective for distinguishing essay uniqueness as it captures semantic content and visual patterns indicative of uniqueness.Outlier Detection with Autoencoders
[0096] The Outlier Detection with Autoencoders method leverages the capabilities of neural networks to learn data representations and identify anomalies within a dataset of essays. Autoencoders, comprising of encoder and decoder networks, are trained to reconstruct input data, and the reconstruction error can be used as an indicator of how well an essay conforms to the learned patterns.
[0097] Training Autoencoders: The process can begin by using neural network-based autoencoders to learn compressed representations of the idea embeddings derived from the essays. Each essay can be converted by the system 100 into an embedding using a suitable language model, capturing its semantic content in a high-dimensional vector. The autoencoder can then be trained on these embeddings, learning to encode each essay into a lowerdimensional latent space and decode it back to its original form. The objective can be to minimize the reconstruction error across all essays during training.
[0098] Anomaly Detection: After the autoencoder is trained, it can be used by the system 100 to reconstruct each essay embedding. The reconstruction error — the difference between the original embedding and the reconstructed embedding — can be calculated for each essay. Essays that are poorly reconstructed, exhibiting high reconstruction errors, can be classified as anomalies or outliers. These essays deviate significantly from the patterns learned by the autoencoder, indicating that they possess unique characteristics not prevalent in the majority of the dataset.
[0099] Alternative to GLOSH Scores: This method can provide an alternative to using Global-Local Outlier Scores from Hierarchies (GLOSH) for detecting uniqueness. While GLOSH scores rely on clustering algorithms to identify outliers, autoencoders focus on reconstructing data and identifying deviations based on reconstruction errors. This neural network-based approach can capture complex, non-linear relationships within the data, potentially uncovering unique essays that may not be detected by traditional clustering methods.
[0100] Technical Background and Justification: Autoencoders can be used to capture the most prevalent patterns in the data during training. Essays that do not conform to these patterns result in higher reconstruction errors, effectively highlighting uniqueness. This method may be suitable for identifying unique essays without requiring labeled data.Applying Graph Neural Networks (GNNs)
[0101] The application of Graph Neural Networks (GNNs) by the system 100 uses an approach to modeling and analyzing the relationships between essays by representing them as graphs. In this framework, essays can be represented as nodes within a graph, and edges represent semantic similarities between pairs of essays.
[0102] Idea Networks: to construct the idea network, each essay can first be converted by the system 100 into an embedding that captures its semantic content. The semantic similarities between essays are then calculated, often using measures like, for example, cosine similarity. Edges are established between nodes (essays) based on these similarities, possibly with weights reflecting the strength of the semantic connection. The resulting graph can embody the relational structure of the essay pool, highlighting how essays are interconnected through shared themes or ideas.
[0103] GNN Analysis: Graph Neural Networks can be designed to operate directly on graph structures, learning representations of nodes, edges, or entire graphs that capture both the features of individual nodes and the topology of the graph. By applying GNNs to the idea network, structural properties can be analyzed such as community structures, centrality measures, and connectivity patterns.
[0104] Uncovering New Metrics of Diversity: Through the system 100’s use of GNN analysis, it can be possible to uncover new metrics of diversity based on network topology. For example, essays that are peripheral in the network or have low connectivity might represent unique ideas not shared by other essays. Central essays with high connectivity might cover common themes. Additionally, structural anomalies detected by the GNN could indicate essays that bridge disparate clusters, representing new connections between different ideas.Temporal Analysis with Recurrent Neural Networks (RNNs)
[0105] Temporal Analysis with Recurrent Neural Networks (RNNs) focuses on capturing the sequential or temporal dynamics within essays, and can be particularly useful if the essays contain a progression of ideas or narratives.
[0106] Sequential Patterns: Essays often have an inherent structure, with ideas unfolding over the course of the text. By treating an essay as a sequence of sentences or paragraphs, each represented by embeddings, RNNs or Transformer models can be used by the system 100 to model the progression of ideas. These models can be designed to handle sequential data, capturing dependencies and patterns over time.
[0107] Uniqueness Over Time: Analyzing how the uniqueness of ideas evolves throughout an essay provides a dynamic view of creativity. By assessing the semantic novelty at each point in the essay, one can identify sections where the author of the essay introduces particularly unique concepts or deviates from common themes. This temporal analysis can reveal the structure of creativity within an essay, highlighting moments of high originality.
[0108] The integration of advanced neural network techniques by the system 100 such as CNNs, autoencoders, GNNs, and RNNs offers a multifaceted approach to assessing essay uniqueness within a large pool. Each method provides unique advantages: Visual Embedding Space Representation leverages visualization and CNNs for intuitive and automated detection of unique essays based on spatial distribution; Outlier Detection with Autoencoders can use reconstruction errors to identify essays that deviate from learned patterns, capturing non-linear forms of uniqueness; Applying Graph Neural Networks models essays as interconnected nodes, revealing structural properties and new diversity metrics based on network analysis;Temporal Analysis with RNNs examines the evolution of ideas within essays.Functional Diversity Metrics
[0109] In Ecology, functional diversity refers to the range and value of functional attributes present within a biological community. These attributes influence how species interact with their environment and each other, contributing to ecosystem processes and resilience. Instead of merely categorizing species (taxonomic diversity), functional diversity evaluates the differences in species' roles, such as variations in feeding strategies, habitat use, and reproductive methods. Ecosystems rich in functional diversity tend to be more adaptable and resilient, better equipped to response to environmental changes and disturbances.
[0110] When analyzing an idea, an analogy can be drawn between ecosystems and text. Each text functions as an ecosystem, and the words or sentences within it are similar to species inhabiting that system. Just as species contribute different functions to an ecosystem, words or sentences convey various themes and ideas within a text. Assessing the functional diversity of a text can involve examining the range and uniqueness of these themes and ideas, providing deeper insights into the content's creativity and the writer's breadth of thought.
[0111] Below are examples of functional diversity metrics, adapted from ecology, that can be used by the system 100 to assess creativity. While the adaptation of functional diversity metrics at the sentence level can be illustrated for clarity, this approach can also be applied at other levels, such as word or theme. We also explain how each metric can be adapted when the primary interest is to assess how an idea or text compares to Algenerated ideas or texts.
[0112] Functional Richness: Measures the convex hull volume of the embedding space occupied by sentences in a text. A larger volume indicates that the text covers a broader range of ideas and concepts, reflecting higher functional richness. When assessing against Algenerated ideas, functional richness can also indicate the extent to which the text includes ideas positioned outside a set (cloud) of Al ideas generated in response to the same task or prompt.
[0113] Functional specialization: Calculates the average distance between each sentence and the centroid (average position) of all sentences in the embedding space. A higher value suggests that the sentences are more specialized, focusing on specific themes that are distinct from the overall average. When assessing against Al-generated ideas, functional specialization can also be adapted to measure the distance from the centroid of a large set (cloud) of Al ideas generated in response to the same task or prompt.
[0114] Functional originality: Assesses the average distance from each sentence to its nearest neighbor within the global pool of sentence (e g., all sentences in a corpus). Sentences that are farther from their nearest neighbors are considered more original, contributing unique ideas to the corpus. This concept is similar to measuring how unique a species is within a global species pool. When assessing against Al -generated ideas, functional originality can also be adapted to identify sentences that are positioned outside a set (cloud) of Al ideas generated in response to the same task or prompt.Academic Domain and Personal Attribute-Based Text Embedding (ADAPT-E):Explainable Embeddings through Academic and Value-Based Representations
[0115] While semantic embeddings from existing models have advanced assessing creativity through measures like semantic distance, they often lack interpretability. For example, although these embeddings can quantify how far apart two ideas are in a semantic space, they do not provide sufficient insights into the underlying reasons for that distance. This limitation makes it difficult to fully appreciate how ideas diverge in meaningful ways, especially when assessing complex constructs such as creativity in college admissions essays.
[0116] This challenge by introducing ADAPT-E, an embedding approach that can be used by system 100 that maps ideas into a unified representational space defined interpretable embedding domains such as academic domains and / or personal attributes. Similar to traditional embedding algorithms like BERT, ADAPT-E generates vector representations (embeddings) for each idea. However, unlike traditional models, ADAPTE’s embeddings are designed to be human-interpretable, with meaningful dimensions corresponding to domains such as academic fields and personal attributes. This enhanced explainability can enable more intuitive comparison and analysis.
[0117] Each ADAPT-E embedding can be utilized by the system 100 not only for established creativity measures such as the DSI but also for GLOSH-based metrics, including ICU, DP, and DPICU. ADAPT-E can enhance conventional models by preserving the functional advantages of traditional embedding methods and providing improved transparency and interpretability. In addition to assessing creativity, ADAPTE's embeddings create comprehensive profiles of applicants, capturing aspects including, but not limited to, academic interests and personal attributes. These profiles can provide deeper insights into each applicant’s unique perspectives and highlight how their ideas diverge in meaningful ways.
[0118] Construction of the ADAPT-E Embedding Space: The basics of the ADAPT-E approach can be implemented by the system 100 in several steps. Although only sentence-level embeddings are mentioned as an example, word-level embeddings are also possible with system 100. Additionally, while academic domains and personal attributes have beendemonstrated by the system 100, embeddings for various other domains are equally feasible. Similarly, in addition to demonstrating application to college admission essays, evaluations by the system 100 of various writings in different contexts are possible — for example, in an employment scenario or educational instructional or course settings.
[0119] Here, the ADAPT-E method implemented by the system 100 for sentence-level embedding is described. The initial step can involve separating each sentence of an essay, then generating embeddings for each academic domain and personal attribute. The detailed procedures for each are described below.
[0120] Academic Domain Embedding: Construction of he academic domain embeddings can be used to map sentences onto key academic domains relevant to higher education (e.g., Geography, Biology, Engineering). We utilize the Academic Research Domains retrieved from Clarivate’s Web of Science, which include 147 subdomains categorized under, for example, five main academic domains: Arts & Humanities, Social Sciences, Life Sciences & Biomedicine, Physical Sciences, and Technology. Alternative categorizations or domain lists may also be employed, depending on context and relevance. The approach remains adaptable and is not limited to any specific set of academic domains, thereby ensuring broad applicability across various frameworks of academic classification.
[0121] LLM-Based Embedding Generation’. A Large Language Model (LLM) can be employed by the system 100 to generate a vector representing the probability that each sentence relates to each academic domain in our list. These resulting vectors constitute the academic domain embeddings. Techniques used by the system 100 such as finetuning or retrieval-augmented generation (RAG) can be applied to LLMs to enhance their performance in providing accurate probabilities. This approach can leverage the LLM's understanding of language and context to assess the relevance of a sentence to various academic domains.
[0122] Keyword-Based Approach with TF-IDF: Alternatively, a keyword-based approach can be employed by the system 100 using a keyword-domain frequency table. To build this table, keywords from academic papers published in academic journals are collected and aggregated across different academic domains. Each keyword can be assigned a weight by the system 100 based on Term Frequency-Inverse Document Frequency (TF-IDF):o Step 1 - Term Frequency (TF): Measures how frequently a keyword appears in papers within a specific academic domain.o Step 2 - Inverse Document Frequency (IDF): Assesses the rarity of the keyword across all domains by calculating the logarithm of the total number of domains divided by the number of domains where the keyword appears.o Step 3 - TF-IDF Weighting: The TF and IDF values are multiplied to obtain the TF-IDF weight for each keyword within each domain. A higher TF-IDF score can indicate that the keyword is significant to that domain and less common in others. In our context, TF-IDF can be used by the system 100 to identify the importance ofa keyword to a particular domain by considering its frequency in that domain and its rarity across other domains.
[0123] For each sentence in an admissions essay, for example, embeddings are computed by the system 100 using the following steps (Steps 4 to 6):o Step 4 - Keyword Extraction: Identifying the keywords present in the sentence. o Step 5 - Mapping to Domain Weights: Matching these keywords to their corresponding TF-IDF weights in the keyword-domain table.o Step 6 - Sentence-Level Embedding: Aggregating the domain weights of the keywords (e.g., summing or averaging) to create a vector that represents the sentence's relevance to each academic domain. This embedding reflects the degree to which the sentence pertains to each academic domain based on the weighted presence of domain-specific keywords.
[0124] Personal Attribute Embedding: Construction of the personal attribute embeddings is used to map sentences onto key personal attributes that might be significant in the college admissions context. 10 personal attributes highlighted in recent literature can be selected, including Growth Mindset, Emotional Intelligence, Resilience, Community Engagement, among others. Alternative arrays of personal attributes with a larger number of attributes can be employed, depending on context and relevance. The approach remains adaptable and is not limited to any specific set of personal attributes, thereby ensuring broad applicability across various frameworks of psychological constructs and assessment models.
[0125] Below is a step-by-step description of a method for generating personal attribute embeddings used by system 100:o Step 1 - Generation of Example Sentences Using LLM: For each personal attribute, the system 100 can use a Large Language Model (LLM) to generate multiple example sentences that exemplify that attribute. These sentences serve as prototypical representations of each attribute.o Step 2 - Computing Embeddings with SBERT: The system 100 can compute the 768-dimensional Sentence-BERT (SBERT) embeddings for each sentence in the essay and for the example sentences generated for each attribute. The system 100 can use SBERT to produce semantic embeddings of sentences, but other embedding models, such as OpenAFs text-embedding-3 -large, can also be employed. o Step 3 - Calculating Semantic Distances: For each sentence in the essay, the system 100 can calculate the semantic distance to the example sentences of each attribute using cosine similarity. Specifically, the system 100 can compute the cosine similarity between the essay sentence embedding and each of the example embeddings for an attribute, then average these similarities to obtain a single similarity score for that attribute.o Step 4 - Constructing the Attribute Embedding: The set of averaged similarity scores for all attributes can form the attribute embedding for that sentence. Using this embedding by the system 100 can effectively capture the degree to which the sentence reflects each personal attribute.
[0126] Alternatively, established psychometric instruments such as the Big Five personality traits or the VIA Character Strengths Survey can be directly utilized by the system 100. In this scenario, semantic distances are computed using the individual items from these scales instead of sample sentences generated by LLMs for each attribute. This approach can allow for a more standardized assessment based on well-validated psychological constructs, enhancing the robustness and applicability of the attribute embeddings.
[0127] This pipeline allows the system 100 to transform qualitative textual data into quantitative embeddings that can be analyzed and compared by the system 100. By leveraging the semantic understanding capabilities of LLMs and embedding models or established psychometric measures, the system 100 can assess how closely an applicant's essay reflects desired personal attributes, enhancing the explainability of the embeddings and providing deeper insights into the applicant's unique perspectives.
[0128] Application of ADAPT-E: By mapping ideas into the ADAPT-E space, the system 100 can use its embeddings to compute creativity measures based on semantic distances, while making resulting outcomes more interpretable. These distances reflect meaningful dimensions including, but not limited to, academic domains and personal attributes. For instance, larger cosine distances between sentences in an essay, when using only the academic domain embedding, suggest diverse academic interests. On the other hand, using personal attribute embeddings reveals how broadly or narrowly an applicant covers personal qualities like leadership or resilience. Unlike traditional semantic based approaches, use by the system 100 of ADAPT-E provides explainable reasons for the (dis)similarity. For example, it can indicate that two ideas are dissimilar because they belong to different academic domains or vary in personal attributes such as leadership and resilience.
[0129] When these two embeddings are combined, the system 100 using ADAPT-E can highlight how academic interests and personal attributes intersect. This combined approach allows the system 100 to identify whether an applicant discusses different personal qualities within the same academic domain or connects their focused academic interests in one domain to a variety of personal attributes. As a result, the combined embeddings offer a more nuanced, albeit less straightforward, understanding of the applicant's ideas and perspectives. This flexibility allows prioritization of analysis on academic diversity, personal traits, or the relationship between the two based on a chosen focus.
[0130] Below are brief examples demonstrating how aforementioned ICU, DP, and DPICU metrics can be applied using ADAPT-E embeddings across different configurations for ICU, DP, and DPICU:
[0131] Idea Cluster Uniqueness (ICU): ICU measures how distinct an idea is relative to clusters of similar ideas in a corpus.• cademic Domain Only: A student who writes about interdisciplinary research spaAnning biology and engineering may receive a high ICU score. This is because few essays focus on these intersecting domains, making their essay unique in its academic focus.• Personal Attribute Only: A student’s essay that emphasizes both resilience and emotional intelligence in navigating challenges may have a high ICU score if other essays in the corpus mostly highlight only a single attribute, making this combination stand out.• Combined: An essay that discusses resilience in the context of pursuing innovations in environmental science (biology + engineering) could receive an even higher ICU score, as it is unique both in its academic and personal dimension combination.
[0132] Diversity Polarity (DP): DP can be used by system 100 to capture how an idea either adds to or detracts from the overall diversity of the idea pool.• Academic Domain Only: If most essays focus on humanities, but one student writes about computer science and robotics, their DP score can be be positive, indicating that they have added diversity to the academic focus of the essay pool.• Personal Attribute Only: If most essays focus on community engagement and teamwork, but one essay centers on personal growth and resilience, the DP score can be positive, indicating that the student’s essay increased diversity in terms of personal attributes.• Combined: A student who writes about applying resilience (personal attribute) while studying data science and artificial intelligence (academic domains) in a humanities-heavy corpus would yield a positive DP score for contributing to both academic and personal diversity.
[0133] Diversity Polarity of Idea Cluster Uniqueness (DPICU): DPICU can be used to assess how the inclusion or exclusion of a specific idea affects the density of idea clusters.• Academic Domain Only: If a student’s essay on synthetic biology (a rare academic domain) is removed from a corpus of essays, and HDBSCAN reveals that the remaining essays are more tightly clustered, the student’s DPICU score can be be positive, indicating that their essay previously dispersed the idea clusters and added diversity.• Personal Attribute Only: If removing an essay focusing on leadership, empathy, and growth mindset causes the remaining essays to cluster more tightly around a single attribute like teamwork, the removed essay can have have a positive DPICU score.• Combined: If removing an essay discussing the use of leadership (attribute) in managing technological innovations in Al (academic domain) makes the remaining essays more densely packed around a few domains and attributes, the DPICU score can be positive, showing the essay had previously contributed to spreading out the clusters.
[0134] Functional Diversity Metrics — Examples Illustrating the Approach: Consider an essay where a student discusses experiences in both astrophysics and classical music, intertwining themes of innovation and perseverance. Using ADAPT-E embeddings, this essay would occupy a unique and expansive position in the embedding space, contributing to functional richness by covering diverse academic domains and personal attributes. The sentences might be distant from the centroid, indicating functional specialization, and far from other sentences in the corpus, reflecting functional originality. In contrast, an essay focused solely on common themes within a single domain — such as teamwork in business management — would exhibit lower functional diversity metrics. Its sentences would cluster closely together in the embedding space, indicating less diversity in the functions they represent.Profile Analysis of Target Attributes (PATA)
[0135] NUMBERED P: PATA emphasizes interpretability and flexibility of which human or text characteristics one is interested in examining. PATA can be used by the system 100 to capture how much a target attribute or attribute profile increases or decreases the diversity of a larger pool by:• Step 1 - Identify target attributes and convert to sentence embeddings using an LLM: Target attributes could be an academic domain, personal attribute like empathy, or mission-related objective like “helping people thrive”. These target attributes are converted into a sentence embedding.• Step 2 -Convert target text to sentence embeddings using an LLM: Target text could be a college essay, cover letter, or ideas from a creativity task. Each idea or sentence is converted to a sentence embedding.• Step 3 -Compute painvise semantic distances between each target attribute and target idea: Now each target idea has an average distance between it and the target attribute, indicating how much the target idea is expanding or shrinking the attribute space.• Step 4 —Create profile of attribute scope expansion or reduction: This is an optional step if there is also an available outcome of interest. For example, ifcollege GPA is an outcome of interest, each idea-attribute’s semantic distance score from Step 3 can be used to predict the outcome and determine which attributes are most predictive. The profile indicates how much each applicant / document increases or decreases attribute scope while simultaneously indicating how important it is to the target outcome.• Step 5 —Identify uniqueness of attribute profile using cluster analysis: Take the profile of attribute semantic distances and conduct both HDSCAN and soft-cluster analysis to generate ICU and ICA scores for the attribute profiles. This is another way to capture not just how the applicants influences each individual attribute’s scope, but how their entire attribute scope profile either increases or decreases the scope of attribute profiles.hiEDGE Approach (human idea Enabled Divergent Generative Expression)
[0136] hiEDGE metrics can be used to integrate combinations of our new generative diversity metrics in a way that simultaneously optimizes these metrics’ ability to capture how any idea increases or decreases the diversity of ideas and how much human ideas uniquely expand the idea pool beyond Al-generated ideas. The system 100 can be operable to perform these exemplary steps:o Step 1 - Integrate Multiple Generative Diversity Metrics Using Ensemble Machine Learning: Using ensemble machine learning (e.g., XGBoost), scores for each of our new generative diversity metrics are used as predictors of a primary target outcome. This could be human-rated creativity, human-rated personal attributes, college GPA, or human vs. Al-generated classification. This allows for complex, non-linear, and many interactive combinations of generative diversity to predict our outcome of interest, while also addressing the multicollinearity that will exist in our predictors. After fine-tuning this model (e.g., interaction depth, number of trees, shrinkage), a predictive model that results can takes all the new generative diversity scores as inputs and then outputs a single score optimized to a selected outcome.o Step 2 - Train an LLM to Predict Optimized and Integrated Generative Diversity Scores: Using the latest high performing LLM (e.g., LLAMA 3 open source), the system 100 can train a new model that uses each idea (e.g., sentences from essays) as input and the optimized combined diversity score from Step 1 as labels. A different model can be created by the system 100 for each target outcome, which could be human creativity, human performance metric (e.g., college GPA), or critically, a dichotomous human vs. Al-generated classification. In all cases, a generative diversity single score is the label for training the LLM, the difference is which kind of target outcome the labels have been trained on (i.e., human-rated creativity, human vs. Al classification).o Step 3 - Generate New Generative Diversity Scores, Optimized to Target Outcome: 3 separate LLM-derived models from Step 2 can be used, whose labelswere optimized at Step 1, to predict human-rated creativity, college GPA, or human vs. Al-generated classification. Now, a set of new ideas (e.g., new essays, new cover letters) can be input in the LLM-trained model and receive optimized generative diversity scores, including a score that captures how much each idea expands the idea space relative to Al-generated ideas.
[0137] Although the description above contains many details, these should not be construed as limiting the scope of the disclosure but as merely providing illustrations of some of the presently preferred embodiments. Therefore, it will be appreciated that the scope of the disclosure fully encompasses other embodiments which may become obvious to those skilled in the art.
[0138] In the claims, reference to an element in the singular is not intended to mean “one and only one” unless explicitly so stated, but rather “one or more.” All structural and functional equivalents to the elements of the disclosed embodiments that are known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the present claims. Furthermore, no element, component, or method step in the present disclosure is intended to be dedicated to the public regardless of whether the element, component, or method step is explicitly recited in the claims. No claim element herein is to be construed as a “means plus function” element unless the element is expressly recited using the phrase “means for”. No claim element herein is to be construed as a “step plus function” element unless the element is expressly recited using the phrase “step for.”
Claims
CLAIMSWhat is claimed is:
1. A system for assessing creativity and diversity in written content, comprising:one or more computers having one or more processors and memory storing instructions that, when executed by the one or more processors, cause the one or more computers to:receive input text for analysis;generate embeddings for ideas in the input text;create a heat map of idea locations in an embedding space;analyze a trajectory of ideas through the embedding space;calculate diversity scores based on idea locations and trajectories;compare individual text diversity to a larger group or dataset; and generate a report on generative diversity of the input text.
2. The system of claim 1, wherein calculating diversity scores comprises use of one or more of the following modules: Idea Cluster Uniqueness (ICU) module, a Diversity Polarity (DP) module, a Diversity Polarity of Idea Cluster Uniqueness (DPICU) module, a Thematic Profile (TP) module, a Thematic Profile Idea Cluster Uniqueness (TPICU) module, a Diversity Polarity of Thematic Profile Idea Cluster Uniqueness (DPTPICU), an Idea Cluster Ambiguity (ICA) module, a Diversity Polarity of Idea Cluster Ambiguity (DPICA) module, and a Functional Diversity Metrics module, calculating an Idea Cluster Uniqueness (ICU) score that captures how far an idea or idea set is from clusters of related ideas in a larger corpus.
3. The system of claim 1, wherein the calculating diversity scores comprises calculating an Idea Cluster Uniqueness (ICU) score that captures how far an idea or idea set is from clusters of related ideas in a larger corpus, the calculating diversity scores further comprising:applying hierarchical density-based spatial clustering of applications with noise (HDBSCAN) to identify clusters of ideas that share embedding features; and generating global-local outlier scores from hierarchies (GLOSH) to represent uniqueness of an idea relative to clusters of related ideas.
4. The system of claim 1, wherein calculating diversity scores comprises calculating a Diversity Polarity (DP) score that captures how much an idea or idea set diverges from a larger corpus of ideas.
5. The system of claim 4, wherein calculating the DP score comprises:computing cosine semantic distances between all pairwise sentences in a corpus to get a total diversity score; removing individual idea sets and recomputing diversity; and calculating difference scores to determine impact of individual idea sets on overall diversity.
6. The system of claim 1, wherein calculating diversity scores comprises calculating a Diversity Polarity of Idea Cluster Uniqueness (DPICU) score that assesses how inclusion or exclusion of a specific idea affects density of idea clusters.
7. The system of claim 1, wherein calculating diversity scores comprises calculating a Thematic Profile (TP) score that captures how much a target's ideas diverge from common and rare themes in a corpus of creative ideas.
8. The system of claim 7, wherein calculating the TP score comprises:utilizing topic modeling to identify themes in a corpus; convertingthemes and target ideas into embeddings; andcomputing semantic distances between theme embeddings and idea embeddings to create thematic profiles.
9. The system of claim 1, wherein calculating diversity scores comprises calculating a Thematic Profile Idea Cluster Uniqueness (TPICU) score that captures relative uniqueness of thematic profiles.
10. The system of claim 1, wherein calculating diversity scores comprises calculating a Diversity Polarity of Thematic Profile Idea Cluster Uniqueness (DPTPICU) score that assesses how a target idea set impacts heterogeneity or homogeneity of thematic profiles relative to clusters of related profiles.
11. The system of claim 1, wherein calculating diversity scores comprises calculating an Idea Cluster Ambiguity (ICA) score that captures a degree to which an idea belongs to more than one cluster.
12. The system of claim 11, wherein calculating the ICA score comprises:utilizing soft-clustering analysis to assign ideas to multiple clusters based on membership degree; and interpreting lower membership degrees as indicating higher ambiguity and creativity.
13. The system of claim 1, wherein calculating diversity scores comprises calculating a Diversity Polarity of Idea Cluster Ambiguity (DPICA) score that captures how an idea or idea set impacts ambiguity of cluster assignments for an entire corpus.
14. The system of claim 1, wherein the instructions further cause the one or more computers to:apply convolutional neural networks (CNNs) to visual representations of embedding spaces to detect patterns and outliers in idea distributions.
15. The system of claim 1, wherein the instructions further cause the one or more computers to:utilize autoencoders to detect outliers based on reconstruction errors when encoding and decoding idea embeddings.
16. The system of claim 1, wherein the instructions further cause the one or more computers to:apply graph neural networks (GNNs) to model relationships between ideas represented as nodes in a graph.
17. The system of claim 1, wherein the instructions further cause the one or more computers to:utilize recurrent neural networks (RNNs) to analyze temporal patterns in idea sequences within the input text.
18. The system of claim 1, wherein the instructions further cause the one or more computers to:calculate functional diversity metrics adapted from ecological measures to assess range and uniqueness of ideas within the input text.
19. The system of claim 1, wherein generating embeddings for ideas comprises: utilizing Academic Domain and Personal Attribute-Based Text Embedding (ADAPT-E) to create interpretable embeddings based on academic domains and personal attributes.
20. A method for assessing creativity and diversity in written content, comprising:receiving, by one or more computers, input text for analysis; generating, by the one or more computers, embeddings for ideas in the input text;creating, by the one or more computers, a heat map of idea locations in an embedding space;analyzing, by the one or more computers, a trajectory of ideas through the embedding space;calculating, by the one or more computers, diversity scores based on idea locations and trajectories;comparing, by the one or more computers, individual text diversity to a larger group or dataset; andgenerating, by the one or more computers, a report on generative diversity of the input text.