Systems and methods for the analogical modeling of psychological structure

Analogical modeling using protein interactions and psychological structures addresses the lack of a physical framework for self-consciousness, enabling a physical interpretation of self-report data for mental health and educational applications.

WO2025226818A1PCT designated stage Publication Date: 2025-10-30SEATTLE CHILDRENS HOSPITAL (DBA SEATTLE CHILDRENS RES INST)
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/025974
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-28
Filing Date
2025-04-23
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing self-report questionnaires lack a tangible, physical framework to understand self-consciousness, failing to anchor abstract subjective experiences to a physical structure.

Method used

Analogical modeling is employed, drawing parallels between protein interactions and psychological self-structures to analyze and visualize patterns in self-report questionnaire data, leveraging variations in response patterns across individuals to map the topology of self-consciousness, using structural biology, artificial intelligence, and psychology/psychiatry.

Benefits of technology

This approach effectively captures the psychological relevance of self-report data, enabling a physical interpretation of self-consciousness, facilitating whole psyche analysis and bridging the subjective and physical domains for mental health and educational applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025025974_30102025_PF_FP_ABST
    Figure US2025025974_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods that integrate structural biology, artificial intelligence, and psychology / psychiatry to explore the physical boundaries of self-consciousness are described. Parallels between protein interactions and psychological self-structures are drawn to analyze and visualize patterns in self-report questionnaire data. Variation in self-report response patterns across individuals can be leveraged to map the topology of self-consciousness boundary, akin to genetic linkage mapping. It is thus possible to conduct whole psyche analysis and leverage variations in the response patterns across individuals to solve the physical topology of the self-boundary.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FORTHE ANALOGICAL MODELING OF PSYCHOLOGICAL STRUCTURECROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 637,761 filed on April 23, 2024, and U.S. Provisional Patent Application No. 63 / 688, 139 filed on August 28, 2024, the entire contents of each of which are incorporated by reference herein in their entirety.FIELD OF THE DISCLOSURE

[0002] The current disclosure provides systems and methods that integrate structural biology, artificial intelligence, and psychology / psychiatry to explore the physical boundaries of self-consciousness. Parallels between protein interactions and psychological self-structures are drawn to analyze and visualize patterns in self-report questionnaire data. Variation in selfreport response patterns across individuals can be leveraged to map the topology of self-consciousness boundary, akin to genetic linkage mapping. It is thus possible to conduct whole psyche analysis and leverage variations in the response patterns across individuals to solve the physical topology of the self-boundary.BACKGROUND OF THE DISCLOSURE

[0003] Analogical modeling has been pivotal in scientific discovery, stimulating new hypotheses upon comparing unfamiliar phenomena to well-understood physical systems (A. Wegener, (Friedrich Vieweg & Sohn, Braunschweig, 1915); N. Wiener, (The MIT Press, 2019); N. Bohr, Philosophical Magazine 26, 1-25 (1913); Watson and Crick, Nature 171, 737— 738 (1953)). A significant unresolved question is the physical boundary of self-consciousness. Self-report questionnaires probe “self’ experiences in consciousness and generate substantial data but fall short of offering a tangible, physical framework for understanding self-consciousness (D. J. Chalmers, (New York : Oxford University Press, 1997)). The core challenge lies in anchoring abstract subjective experiences to a physical structure, presumably the brain (F. Crick, (Scribners, 1994)). This may be facilitated by a concrete physical analogy that could render the observed patterns in selfreport data more physically interpretable. Before the current disclosure, such a physical analogy was absent.SUMMARY OF THE DISCLOSURE

[0004] The current disclosure provides a novel framework and methodological approaches that integrate structural biology, artificial intelligence, and psychology / psychiatry to explore the physical boundaries of self-consciousness. By drawing parallels between protein interactions and psychological self-structures (e.g., an individual’s experiences, knowledge, values, beliefs, and the organization thereof), an approach to analyze and visualize patterns in self-report questionnaire data was developed, akin to visualizing protein interfaces. The disclosed findings reveal that semantic representations by language models effectively capture the psychological relevance of self-report data and that variation in self-report response patterns across individuals may be leveraged to map the topology of self-consciousness boundary,akin to genetic linkage mapping. These findings demonstrate that one can conduct whole psyche analysis and leverage variations in the response patterns across individuals to solve the physical topology of the self-boundary. Thus, this work opens the door to understanding self-consciousness by bridging the subjective and the physical, with implications for basic sciences, mental health, social and educational applications.BRIEF DESCRIPTION OF THE FIGURES

[0005] Some of the drawings submitted herein may be better understood in color. Applicant considers the color versions of the drawings as part of the original submission and reserves the right to present color images of the drawings in later proceedings.

[0006] FIG. 1 illustrates an example environment for analogical modeling and improving the psychological relevance of semantic representations.

[0007] FIGS. 2A and 2B illustrate example processes and according to implementations described herein. (2A) Example process for using a model to determine a relationship of data to a reference set. In various implementations, the data includes at least one prompt and / or at least one response of an individual to the prompt. The model may be configured to generate an output indicative of the relationship. In some implementations, the output may include a visual three- dimensional model. (2B) Example process for generating and training a model described herein. In various implementations, the model may be generated and trained using training data. The performance of the model may be evaluated.

[0008] FIG. 3 shows an example computer architecture for a computer capable of executing program components for implementing the functionality described herein.

[0009] FIGs. 4A-4J illustrate analogical modeling of self-consciousness boundary using protein interactions. (4A) Schematic depicting consciousness versus protein interactions. (4B) Analogy between self-report questionnaire items and protein residues organization. (4C) Analogy between human self-report response patterns and self-to-other protein interaction interface patterns. (4D) Nanobody-antigen interactions as a particularly suitable analogical system for modeling diverse human self-response patterns to psychological questionnaires. PDB Structure codes, 13G9A, 2-3K1 K. (4E) Distance matrices in psychological questionnaires and protein structures. (4F) Clusters of distance inter-relatedness reflect patches of surfaces on a physical structure. Example of 3 clusters is shown on PDB nanobody 3K1 K. (4G) Strategy to test whether order of component sequence and preserved inter-relatedness is important for generating same amino acid groupings as the original data. (4H) Inter-relatedness, but not sequence order, is important for maintaining the correct grouping of amino acids in relation to the actual structural information. Non-parametric Friedman test showed significant difference across groups (Friedman statistic: Pooled 285.3, p<1.0E-15). Post hoc Dunn’s multiple comparisons test revealed significantly lower adjusted Rand index (ARI) shuffling inter-relatedness than shuffling order (Nanobodies and Antigens: p<1.0E-15). n = 107 nanobodies, 107 antigens. (4I) Multidimensional Scaling (MDS) effectively reconstitute correct alpha carbon backbone structure over other dimension reduction methods. PDB nanobody 3K1 K. (4J) Quantification of RMSD fit between actual and reconstituted structures across 214 structures, including nanobody (n=107)and antigen (n=107) structures. Non-parametric Friedman test showed significant difference across groups (Friedman statistic: Nanobody 289.6, p<1.OE-15; Antigen 225.4, p<1.OE-15). Post hoc Dunn's multiple comparisons test revealed significantly lower RMSD for MDS versus all other conditions (Nanobodies: p<1 .OE-7, Antigens: p<1. OE-15). ****, p < 0.0001.

[0010] FIGs. 5A-5C illustrate visualization of semantic structures of questionnaires using MDS. (5A) Visualization of semantic structures across psychology and psychiatry questionnaires of varying sizes with MDS pipeline. Rotation is along vertical axis. Distances are approximate and not drawn to scale. Colored clusters derived from affinity propagation clustering of semantic distance matrices. (5B) Combined items from FIG. 5A questionnaires. (5C) Combined items from 30 questionnaires. In FIGs. 5B, 5C, the left corresponds to structure labeled by the identity of the questionnaire, and the right corresponds to same structure labeled by the identity of semantic interrelatedness clusters.

[0011] FIGs. 6A-6C illustrate visualization of semantic structures of psychiatric questionnaires using MDS. (6A) Visualization of semantic structures across DSM-V psychiatry questionnaires across adult response groups with MDS pipeline. Rotation is along vertical axis. Distances are approximate and not drawn to scale. Colored clusters derived from affinity propagation clustering of semantic distance matrices. (6B) Structures for adolescent group. (6C) Structures for parent / guardian of child group. The left corresponds to structure labeled by the identity of the questionnaire, and the right corresponds to same structure labeled by the identity of semantic interrelatedness clusters.

[0012] FIG. 7 illustrates visualization of massive items database using MDS. Visualization of semantic structures across 7 large items databases with MDS. The left corresponds to structure labeled by the identity of the questionnaire, and the right corresponds to same structure labeled by the identity of semantic interrelatedness clusters.

[0013] FIG. 8 illustrates various visualization schemes possible with psychological structures generated by MDS. Various visualization modes for visualizing semantic structures across 7 large items databases with MDS.

[0014] FIG. 9 illustrates effect of item size on local and global measures of structure integrity. Average Jaccard similarity was used to measure the overlap in closest neighbor sets in the high dimensional versus MDS 3D space. Spearman’s rank correlation was used to measure stability of overall relational structure. Item variables are presented in logarithmic scale. Linear regression was used, and 95% confidence bounds around each line are shown as dashed lines.

[0015] FIG. 10 illustrates that semantic inter-relatedness in language models facilitates grouping of items in a psychologically relevant manner. More recent language models are beginning to encode psychologically relevant groupings in its semantic representation of text.

[0016] FIGs. 11A-11C illustrate generalization of the finding that language models facilitate groupings of items in a psychologically relevant manner. (11 A) Evaluation of language models' ability to group self-report items into psychologically relevant groupings. 31 questionnaires from 30 studies were studied, including 637 items and 110 empirically defined facets. Publications of questionnaire facet data range from 1997 to 2023. (11 B) Large screen of 50 language models for the ARI measure. Median level shown with bars. Controls indicate shuffled embeddings as input for semantic processing. The X- axis, from 1 to 50 is: (1) ada2, (2) sentence-transformers_sentence-t5-xl, (3) intfloat_e5-base-v2, (4) hkunlpjnstructor- large, (5) thenlper_gte-small, (6) angle_llama_13b_nli, (7) angle_llama_7b_nli_v2, (8) Ilmrails_ember-v1 , (9) BAAI_bge-Iarge-en-v1.5, (10) sentence-transformers_sentence-t5-base, (11) avsolatorio_GIST-Embedding-vO, (12) sentence- transformers_all-reoberta-large-v1, (13) intfloat_e5-small, (14) sentence-transformers_sentence-t5-large, (15) sentence- transformers_gtr-t5-large, (16) hkunlpjnstructor-base, (17) sentence-transformers_gtr-t5-xl, (18) thenlper_gte-base, (19) intfloat_e5-small-v2, (20) intfloat_e5-large, (21) sentence-transformers jjaraphrase-multilingual-mpnet-base-v2, (22) hkunlpjnstructor-xl, (23) jamesgpt1_sf_model_e5, (24) sentence-transformers_all-mpnet-base-v2, (25) jinaaijina- embedding-l-en-v1 , (26) sentence-transformers_all-MiniLM-L12-v2, (27) sentence-transformers_paraphrase-rnultilingual- MiniLM-L12-v2, (28) sentence-transformers_multi-qa-mpnet-base-dot-v1, (29) sentence-transformers_gtr-t5-base, (30) sentence-transformers_distiluse-base-multilingual-cased-v1, (31) intfloat_multilingual-e5-base, (32) Intfloat_e5-large-v2, (33) jinaaijina-embedding-b-en-v1, (34) sentence-transformers_sentence-t5-xxl, (35) intfloat_multilingual-e5-small, (36) sentence-transformers_multi-qa-distilbert-cos-v1, (37) sentence-transformers_distiluse-base-multilingual-cased-v2, (38) sentence-transformers_all-MiniLM-L6-v2, (39) sentence-transformers_paraphrase-albert-small-v2, (40) thenlper_gte- large, (41) symanto_sn-xlm-roberta-base-snli-mnli-anli-xnli, (42) sentence-transformers_all-distilroberta-v1 , (43) jinaaijina-embedding-s-en-v1 , (44) sentence-transformers_LaBSE, (45) sentence-transformers_multi-qa-MiniLM-L6-cos- v1, (46) sentence-transformers_paraphrase-MiniLM-L3-v2, (47) clips_mfaq, (48) NeuML_pubmedbert-base-embeddings, (49) sentence-transformers_gtr-t5-xxl, (50) and thenlper_gte-large-zh. (11C) Trained language models show significantly increased ability (boxes marked with a triangle; alpha value corrected by Bonferroni method for multiple comparisons) above control models without training in the grouping of items into psychologically relevant groupings. The X-axis, from 1 to 32, is: (1) bert-base-uncased, (2) roberta-base, (3) distilbert-base-uncased, (4) sentence-transformers / all-mpnet-base- v2, (5) sentence-transformers / multi-qa-mpnet-base-dot-v1, (6) sentence-transformers / all-distilroberta-v1 , (7) sentence- transformers / all-MiniLM-L12-v2, (8) sentence-transformers / multi-qa-distilbert-cos-v1, (9) sentence-transformers / all- MiniLM-L6-v2, (10) sentence-transformers / multi-qa-MiniLM-L6-cos-v1, (11) sentence-transformers / paraphrase- multilingual-mpnet-base-v2, (12) sentence-transformers / paraphrase-albert-small-v2, (13) sentence- transformers / paraphrase-multilingual-MiniLM-L12-v2, (14) sentence-transformers / paraphrase-MiniLM-L3-v2, (15) sentence-transformers / distiluse-base-multilingual-cased-v1 , (16) ada2, (17) bert-base-uncased_control, (18) roberta- base_control, (19) distilbert-base-uncased_control, (20) sentence-transformers / all-mpnet-base-v2_control, (21) sentence- transformers / multi-qa-mpnet-base-dot-v1_control, (22) sentence-transformers / all-distilroberta-v1_control, (23) sentence- transformers / all-MiniLM-L12-v2_control, (24) sentence-transformers / multi-qa-distilbert-cos-v1_control, (25) sentence- transformers / all-MiniLM-L6-v2_control, (26) sentence-transformers / multi-qa-MiniLM-L6-cos-v1_control, (27) sentence- transformers / paraphrase-multilingual-mpnet-base-v2_control, (28) sentence-transformers / paraphrase-albert-small- v2_control, (29) sentence-transformers / paraphrase-multilingual-MiniLM-L12-v2_control, (30) sentence- transformers / paraphrase-MiniLM-L3-v2_control, and (31) ada2_control.

[0017] FIGs. 12A-12C illustrate that language models facilitate groupings of DSM-V items in a psychiatrically relevant manner. (12A) Evaluation of language models’ ability to group DSM-V self-report items into psychiatrically relevant groupings. 4 large questionnaires designed for adults were studied, including 410 items and 54 total facets. Questionnaires were PID-5, TR Self-Rated Level 1 Cross-Cutting Symptoms Measure, Level 2 Cross-cutting symptom measures, Disorder-Specific Severity Measures. The X-axis, from 1 to 48, is: (1) ada2, (2) sentence-transformers_all-roberta-large-v1, (3) hkunlpjnstructor-xl, (4) sentence-transformers_sentence-t5-xl, (5) avsolatorio_GIST-Embedding-vO, (6) sentence- transformers_all-mpnet-base-v2, (7) hkunlpjnstructor-large, (8) BAAI_bge-large-en-v1.5, (9) sentence- transformers_multi-qa-mpnet-base-dot-v1, (10) hkunlpjnstructor-base, (11) thenlper_gte-small, (12) sentence- transformers_multi-qa-distilbert-cos-v1 , (13) sentence-transformers_sentence-t5-large, (14) jamesgpt1_sf_model_e5, (15) sentence-transformers_gtr-t5-xl, (16) sentence-transformers_paraphrase-multilingual-mpnet-base-v2, (17) sentence- transformers_all-distilroberta-v1, (18) intfloat_e5-base-v2, (19) sentence-transformers_all-MiniLM-L6-v2, (20) sentence- transformers_gtr-t5-base, (21) thenlper_gte-base, (22) Ilmrails_ember-v1, (23) sentence-transformers_sentence-t5-base, (24) sentence-transformers_all-MiniLM-L12-v2, (25) sentence-transformers_gtr-t5-large, (26) sentence- transformers_multi-qa-MiniLM-L6-cos-v1, (27) intfloat_e5-large, (28) jinaaijina-embedding-l-en-v1, (29) sentence- transformers_gtr-t5-xxl, (30) sentence-transformers jjistiluse-base-multilingual-cased-v2, (31) sentence- transformers_distiluse-base-multilingual-cased-v1, (32) symanto_sn-xlm-roberta-base-snli-mnli-anli-xnli, (33) intfloat_e5- small, (34) intfloat_multilingual-e5-small, (35) thenlper_gte-large, (36) intfloat_e5-large-v2, (37) intfloat_e5-small-v2, (38) sentence-transformers_sentence-t5-xxl, (39) jinaaijina-embedding-b-en-v1 , (40) intfloat_multilingual-e5-base, (41) clips_mfaq, (42) NeuML_pubmedbert-base-embeddings, (43) sentence-transformers_LaBSE, (44) sentence- transformers_paraphrase-albert-small-v2, (45) sentence-transformers_paraphrase-multilingual-MiniLM-L12-v2, (46) sentence-transformers_paraphrase-MiniLM-L3-v2, (47) jinaaijina-embedding-s-en-v1 , and (48) thenlper_gte-large-zh. (12B) 3 large questionnaires designed for children ages 11-17 were studied, including 195 items and 28 empirically defined facets. Questionnaires were Level 2 Cross-Cutting Symptom Measures, TR Parent / Guardian-Rated Level 1 Cross-Cutting Symptom Measures, Disorder-specific severity measures. The X-axis, from 1 to 49, is: (1) sentence- transformers_sentence-t5-xxl, (2) avsolatorio_GIST-Embedding-vO, (3) sentence-transformers_sentence-t5-large, (4) angle_llama_13b_nli, (5) sentence-transformers_sentence-t5-xl, (6) intfloat_e5-large-v2, (7) jinaai Jina-embedding-l-en- v1, (8) hkunlpjnstructor-xl, (9) intfloat_e5-base-v2, (10) sentence-transformers_all-mpnet-base-v2, (11) sentence- transformers_gtr-t5-xl, (12) anglejlama_7b_nli_v2, (13) intfloat_e5-small, (14) intfloat_e5-large, (15) jamesgpt1_sf_model_e5, (16) hkunlpjnstructor-large, (17) thenlper_gte-base, (18) sentence-transformers_gtr-t5-large, (19) sentence-transformers_all-roberta-large-v1, (20) sentence-transformers_gtr-t5-base, (21) BAAI_bge-large-en-v1.5, (22) Intfloat_e5-small-v2, (23) thenlper_gte-small, (24) sentence-transformers_LaBSE, (25) hkunlpjnstructor-base, (26) sentence-transformers_paraphrase-multilingual-MiniLM-L12-v2, (27) sentence-transformers_paraphrase-albert-small-v2, (28) sentence-transformers_multi-qa-mpnet-base-dot-v1 , (29) sentence-transformers_sentence-t5-base, (30) sentence- transformers_all-distilroberta-v1, (31) Ilmrails_ember-v1, (32) thenlper_gte-large, (33) jinaaijina-embedding-b-en-v1, (34) symanto_sn-xlm-roberta-base-snli-mnli-anli-xnli, (35) sentence-transformers_distiluse-base-multilingual-cased-v2, (36) sentence-transformers_multi-qa-MiniLM-L6-cos-v1, (37) sentence-transformers_allMiniLM-L12-v2, (38) Intfloat_multilingual-e5-small, (39) Intfloat_multilingual-e5-base, (40) sentence-transformers_all-MiniLM-L6-v2, (41) sentence-transformers_multi-qa-distilbert-cos-v1, (42) sentence-transformers_distiluse-base-multilingual-cased-v1, (43) sentence-transformers_paraphrase-multilingual-mpnet-base-v2, (44) clips_mfaq, (45) sentence-transformers_paraphrase-MiniLM-L3-v2, (46) thenlper_gte-large-zh, (47) jinaai Jina-embedding-s-en-v1, (48) sentence-transformers_gtr-t5-xxl, and (49) NeuML_pubmedbert-base-embedding. (12C) 2 large questionnaires designed for parents / guardians of children age 6 and above were studied, including 128 items and 20 facets. Questionnaires were Level 2 Cross-Cutting Symptom Measures, TR Parent / Guardian-Rated Level 1 Cross-Cutting Symptom Measures. Facets are defined based on unique name within a questionnaire. Facets with identical or similar names across questionnaires were kept in separate groupings. The X-axis, from 1 to 35, is: (1) sentence-transformers_sentence-t5-xxl, (2) intfloat_multilingual-e5-base, (3) sentence- transformers_all-distilroberta-v1, (4) sentence-transformers_sentence-t5-xl, (5) angle_llama_13b_nli, (6) avsolatorio_GIST-Embedding-vO, (7) sentence-transformers_sentence-t5-large, (8) intfloat_e5-base-v2, (9) hkunlpjnstructor-base, (10) angle_llama_7b_nli_v2, (11) intfloat_e5-small-v2, (12) sentence-transformers_sentence-t5- base, (13) sentence-transformers_multi-qa-mpnet-base-dot-v1, (14) jinaai Jina-embedding-b-en-v1, (15) hkunlpjnstructor-large, (16) thenlper_gte-small, (17) intfloat_e5-small, (18) jinaai Jina-embedding-l-en-v1, (19) sentence- transformers_LaBSE, (20) thenlper_gte-base, (21) sentence-transformers_all-mpnet-base-v2, (22) intfloat_multilingual-e5- small, (23) clips_mfaq, (24) sentence-transformers_gtr-t5-base, (25) sentence-transformers_gtr-t5-large, (26) sentence- transformers_all-MiniLM-L12-v2, (27) sentence-transformers_paraphrase-albert-small-v2, (28) sentence- transformers_multi-qa-MiniLM-L6-cos-v1, (29) jinaai Jina-embedding-s-en-v1, (30) hkunlpjnstructor-xl, (31) intfloat_e5- large, (32) intfloat_e5-large-v2, (33) sentence-transformers_all-roberta-large-v1, (34) sentence-transformers_all-MiniLM- L6-v2, (35) entence-transformers_paraphrase-multilingual-mpnet-base-v2, (36) sentence-transformers_gtr-t5-xl, (37) jamesgpt1_sf_model_e5, (38) BAAI_bge-large-en-v1.5, (39) sentence-transformers_paraphrase-multilingual-MiniLM-L12- v2, (40) NeuML_pubmedbert-base-embeddings, (41) sentence-transformers_distiluse-base-multilingual-cased-v2, (42) sentence-transformers_paraphrase-MiniLM-L3-v2, (43) symanto_sn-xlm-roberta-base-snli-mnli-anli-xnli, (44) thenlper_gte-large, (45) sentence-transformers_multi-qa-distilbert-cos-v1, (46) Ilmrails_ember-v1 , (47) sentence- transformers_distiluse-base-multilingual-cased-v1, (48) sentence-transformers_gtr-t5-xxl, and (49) thenlper_gte-large-zh.

[0018] FIG. 13 illustrates applying analogical modeling to visualize development of language models. MDS-generated semantic structure of the PID-5 questionnaire in untrained (left, right) and trained (right, overlaid onto untrained model) SBERT model. This illustrates the potential of the analogical modeling to provide diagnostic and intuitive visual analysis of neural network development. Shading indicates clusters of semantic inter-relatedness (shades are repeated for different clusters).

[0019] FIG. 14 illustrates that the semantic-interrelatedness space created by MDS explains psychological factor loading data patterns. Narcissism Personality Inventory is 40 items in size. Language model is MiniLM-L6-v2. MDS is applied on pairwise cosine distance matrix containing the narcissism items. Mesh volume is depicted around location of each item (x) in MDS 3D space. Shading is based on empirically derived facet groupings of each item. The partitioning of Entitlement items in this particular MDS space explains each items’ factor loadings in relation to Authority facet, suggesting that semantic inter-relatedness explains psychological patterns.

[0020] FIG. 15 illustrates that psychological factor loading values are related to the proximity between items in MDS space. Narcissism Personality Inventory is 40 items in size. MDS is applied on pairwise cosine distance matrix containingthe narcissism items. Mesh volume is depicted around location of each item (x) in MDS 3D space. Shading is based on empirically derived factor loadings corresponding to the specified facet above each structure. Different shades indicate high, medium and low factor loading values corresponding to the specified narcissism factor.

[0021] Fl G. 16 illustrates that geometric closeness in MDS space is predictive of empirically derived factor loadings for a self-report questionnaire. Narcissism Personality Inventory is 40 items in size and is divided into 7 facets (also called "group” here). The angular closeness (cosine distance) of an item is calculated against the facet / group's centroid position in MDS space, and Spearman correlation coefficient is calculated to test whether geometric closeness is correlated with the item's factor loading specific to specific facet / group. The results indicate that MDS provides the most predictive structural information relative to other dimension reduction methods, PCA / t-SNE / UMAP.

[0022] FIGs. 17A and 17B illustrate that neural networks can be trained to predict factor loadings non-randomly. (17A) A large language model (LLM), sentence-t5-xl, was coupled with a multilayer perceptron (MLP) and trained on real or shuffled data documenting the associations between narcissism items and their corresponding factor loading values relative to each narcissism facet (n = 40 items, 7 facets). This information was trained on the models in a 10-fold cross validation manner. Mean-squared error (MSE) was used as loss function to assess the ability of the system to learn the patterns in a nonrandom way. Each condition was repeated in 10 independent runs. The MSE of data from 4000 shuffles of data was tested as a control. This data was not trained on neural networks. Results indicate statistically significant MSE reduction specifically for the combination of trained LLM + trained MLP + real data, indicating that neural networks can be trained to predict factor loadings from text input alone. Statistics: On log-transformed data. Brown-Forsythe ANOVA test showing difference across groups (F(Numerator Degrees of Freedom (DFn), Denominator Degrees of Freedom (DFd)): 118.7 (8.000, 60.47); p<0.0001). Welch’s ANOVA test (W(DFn, DFd): 161.0 (8.000, 32.66); p-value < 0.0001)). Dunnett’s T3 multiple comparisons test. **** indicates p<0.0001. Plots are mean+ / - SEM. (17B) Extension of model to make predictions about item membership across 11 questionnaires with complete factor loading tables to train and validate models on. Shown is an evaluation of accuracy performance across different thresholds for factor loading scores of an item to be considered for a particular factor. It can be seen that Trained Encoder (MLP) outperforms Untrained Encoder, Trained Encoder fed Shuffled Input Data and random guessing levels, indicating that the system is capable of making non-random assignment of factor membership to unseen items.

[0023] FIGs. 18A-18D illustrate fine-tuning language models to improve classification of questionnaire items into appropriate psychological groupings. (18A-18C) 3 models shown to have variable performance on the psychological item grouping task (FIGs. 11 , 12) were fine-tuned to identify appropriate psychological facet grouping for each item. Training was performed on 31 self-report questionnaires (30 studies) randomly sampled from the web, with 638 items grouped into 110 facets. Results indicate that all 3 models could be fine-tuned to improve their classification of items into appropriate psychological groupings. Training and Validation losses indicate generalizable learning in a manner that is specific to real data. Accuracy parameter indicate improve performance in classification task above random chance. (18D) Final model accuracy on unseen test set after all epochs are completed.

[0024] FIG. 19 illustrates an MDS- and LLM-generated structure of a psychiatric symptom questionnaire - MID-60.Arrows demonstrate closeness in positioning of items with similar meanings. Shading represents clusters grouped by semantic inter-relatedness.

[0025] FIGs. 20A and 20B illustrate human subject response patterns mapped onto MDS-generated structure of a psychiatric symptom questionnaire - MID-60, analogously to visualizing protein-protein interface. (20A) Semantic structure of MID-60 in mesh surface view. Shading indicates clusters of semantically related items. (20B) Clusters of human subject response patterns on MID-60. Varying in shading gradient indicates agreement with item. N = 314 subjects.

[0026] FIGs. 21A-21C illustrate solving the physical topology of a nanobody protein using variation of interface signals across nanobodies. (21A) A nanobody interface database (n=70) was analyzed by first constructing a signal similarity matrix, followed by MDS to solve the structural topology of the amino acids. Shading represents amino acids clustering by their signal similarity patterns. (21 B) Actual conserved nanobody structure showing actual locations of amino acids colored in the same way in FIG. 18A. Note the striking preservation in physical topology despite distortions to the structure. (21C) Procrustes analysis measuring disparity. Alignment of actual structure in FIG. 21 B to reconstituted structure in FIG. 21 A shows non-random alignment, as revealed by comparison to the value seen when aligning a structure with shuffled data to the actual structure. This experiment illustrates the ability to solve a physically relevant topology from signal variations in a dataset that bears similar data structures as those datasets reporting psychological self-report response patterns.

[0027] FIG. 22 illustrates application of signal variation analysis on the human subject response patterns of a psychiatric symptom questionnaire - MID-60, resulting in a putative physical topological map representing the boundary of selfconsciousness. Human response patterns on the MID-60 are processed identically to the processing of protein interface dataset in order to derive a putative physical topology of the self-boundary. Shading indicates clusters grouped by signal similarity.

[0028] FIG. 23 illustrates that language models that perform better on the psychological I semantic grouping tasks give MDS-generated semantic structures that align better to MDS-generated structures from signal variation analysis. MID-60 questionnaire items were processed to create MDS-generated semantic structures for 47 language models, and the performance of each language model on the psychological I semantic grouping overlap task (FIGs. 13, 14 using median ARI metric) correlate with better alignment (reduced disparity from Procrustes analysis) to the MDS structure generated from MID-60 human response variation patterns (FIG. 22). Linear regression R-squared - 0.097, Slope Significantly nonzero: p = 0.03. These results support the concept that MDS semantic structures of various language models are psychologically relevant and serve as models of a psychologically, and potentially physically relevant self-structure. These results justify the use of these structures as models for sampling and generating psychological items with the aim of rapidly and comprehensively conducting whole psyche analysis.

[0029] FIGs. 24A and 24B illustrate one form of whole psyche analysis based on content of this work. (24A) MDS- generated semantic structure is suitable for global cluster sampling in large items databases (>100), and suitable for refined, psychologically relevant item sampling in small items databases (<100). (24B) Pipeline for whole psyche analysis and potential applications.

[0030] FIGs. 25A and 25B illustrate developing psychologically relevant language models (LLMs) by utilizing structuralalignment between semantic and psychological structures. (25A) Training LLMs or equivalent on structural alignment to psychological structure in original dimensions. (25B) Training LLMs on structural alignment to psychological structure in different dimensions. This method distinguishes from traditional neural network I LLM training paradigm by using a structural alignment approach to train models based on many inter-relatedness between items, thus rearranging embedding space of LLMs in a psychologically relevant way.

[0031] FIG. 26 illustrates the concept of analogical modeling of the physical structure of awareness.

[0032] FIGs. 27A-27D illustrate analogical modeling (referred to as “SELF-MAP”) validation and application. (27A) Validation on ground-truth nanobody-to-target structures to reconstruct shared “Self” nanobody structure is shown in FIGs. 4C (right), 21 A, and 21 B. As shown in FIG. 27A, disparity value from real data-derived structure in falls outside entire range of values from 10,000 shuffled datasets. (27B) Application on MID-60 self-report data, with reference to the self-reports response patterns shown in FIG. 4C (left). (27C) SELF-MAP models constructed from 10,000 subsampled datasets from a mental health population (n=2010) were aligned to the model created from a college sample (n=314). The two MID-60 sample populations are non-randomly and significantly aligned in SELF-MAP topologies. Unpaired t-test. ****; p<0.0001. (27D) Analysis of average inter-item similarity across 1 ,000 sub-samples (mental health samples) of SELF-MAP structures reveal items with less structural stability.

[0033] FIGs. 28A-28F illustrate aligning SELF-MAP model to brain structure. (28A) Published functional magnetic resonance imaging (fMRI) study by Lebois, et al. (2022) co-administered MID60 items. (28B) The analysis of the published fMRI coordinates / unique variances reveal Frontal-Posterior (F-P) bias in positioning between Partially Dissociation Intrusions (PDI) / Depersonalization / Derealization (DPDR) groups of items. Kruskal-Wallis Multiple Comparisons Test, ** p = 0.0012, n = 15-16 brain regions. (28C) fMRI data-guided alignment of SELF-MAP structures to the brain's F-P axis. (28D) Alignment using 40% of PDI / DPDR items, followed by testing of F-P positioning of the unseen PDI / DPDR items, reveal a significant F-P segregation between the two groups, as predicted. n=200 subsamples, p<0.0001. Sidak multiple comparisons. One-Way ANOVA between conditions (F (2.023, 402.5) = 56.55, p<0.0001). (28E) SELF-MAP F-P models constructed from 2 MID-60 datasets correlate, but not with LLM-derived SELF-MAP structure. n=27 items. Spearman Rank Correlation - Bonferroni-corrected P<0.0001. (28F) Predictions of SELF-MAP model.

[0034] FIG. 29 illustrates that LLM training using human response data improves “physical” relevance but not “psychological” relevance of fMRI-aligned SELF-MAP model.DETAILED DESCRIPTION

[0035] Analogical modeling has been pivotal in scientific discovery, stimulating new hypotheses upon comparing unfamiliar phenomena to well-understood physical systems (A. Wegener, (Friedrich Vieweg & Sohn, Braunschweig, 1915); N. Wiener, (The MIT Press, 2019); N. Bohr, Philosophical Magazine 26, 1-25 (1913); Watson and Crick, Nature 171, 737- 738 (1953)). A significant unresolved question is the physical boundary of self-consciousness. Self-report questionnaires probe “self’ experiences in consciousness and generate substantial data but fall short of offering a tangible, physical framework for understanding self-consciousness (D. J. Chalmers, (New York : Oxford University Press, 1997)). The corechallenge lies in anchoring abstract subjective experiences to a physical structure, presumably the brain (F. Crick, (Scribners, 1994)). This may be facilitated by a concrete physical analogy that could render the observed patterns in selfreport data more physically interpretable. Before the current disclosure, such a physical analogy was absent.

[0036] Recent advances in artificial intelligence (Al), particularly in neural network architectures (A. Vaswani, et al., Adv Neural Inf Process Syst, (2017) vol. 30; Radford and Narasimhan, “Improving Language Understanding by Generative PreTraining” (2018); J. Devlin, et al., Proceedings of the 2019 NAACL-HLT Conference, Volume 1 (2019), pp. 4171-4186), present a promising avenue for rigorously exploring the organization of subjective experiences. Language models, which have become increasingly sophisticated (T. Brown, et al., Adv Neural Inf Process Syst, (2020) vol. 33, pp. 1877-1901 ; J. Ni, et al., Findings of the Assoc, for Comp. Ling. 2022, (2022), pp. 1864-1874; T. Wolf, et al., Proceedings of the 2020 EMNLP Conference: System Demonstrations (2020)), offer tools to quantify the meaning and semantic similarities in text (D. Cer, et al., Proceedings of the 2018 EMNLP Conference: System Demonstrations, (2018), pp. 169-174; Reimers and Gurevych, arXiv e-prints, (2019)), potentially bridging the gap between linguistic analyses of subjective experiences and the physical structures underlying the self.

[0037] The current disclosure provides a novel framework and methodological approaches that integrate structural biology, artificial intelligence, and psychology / psychiatry to explore the physical boundaries of self-consciousness. By drawing parallels between protein interactions and psychological self-structures, an approach to analyze and visualize patterns in self-report questionnaire data was developed, akin to visualizing protein interfaces. In this context, proteinprotein interactions, notably the binding mechanisms of camelid antibodies (S. Muyldermans, Annu. Rev. Biochem. 82, 775-797 (2013)) (nanobodies) to antigens, emerge as an intriguing analogical model. Nanobodies, with their robustly conserved yet diverse antigen-interacting patterns, mirror the varied responses individuals exhibit in self-report questionnaires concerning self-perception. Although protein interactions are not necessarily the only suitable analogical model, they provide a rich source of empirically validated structures and data patterns to ground the interpretation of similar patterns in self-report data (H. M. Berman, et al, Nucleic Acids Res. 28, 235-242 (2000); Krissinel and Henrick, J. Mol. Biol. 372, 774-797 (2007)). These advantages can guide the interpretation of complex self-report data patterns and abstract psychological concepts in a physically relevant way. This disclosure introduces an interdisciplinary approach that leverages the structural characteristics of protein interactions to develop a framework and methodologies for understanding the physical boundary of self-structure.

[0038] The disclosed findings reveal that semantic representations by language models effectively capture the psychological relevance of self-report data and that variation in self-report response patterns across individuals may be leveraged to map the topology of self-consciousness boundary, akin to genetic linkage mapping. These findings demonstrate that one can conduct whole psyche analysis and leverage variations in the response patterns across individuals to solve the physical topology of the self-boundary. Thus, this disclosure opens the door to understanding selfconsciousness by bridging the subjective and the physical, with implications for basic sciences, mental health, social, and educational applications.

[0039] FIG. 1 illustrates an example environment 100 for analogical modeling and improving the psychological relevanceof semantic representations. As shown, a prediction system 102 includes a trainer 104, a predictive model 106, and a visualizer 108. The prediction system 102, for instance, is embodied in hardware, software, or a combination thereof. In various implementations, the predictive model 106 includes one or more machine learning (ML) model(s) 110 and a classifier 112. While FIG. 1 illustrates the ML model(s) 110 and the classifier 112 separately, implementations of the present disclosure are not so limited. In various implementations, the predictive model 106 includes the ML model(s) 110 or the classifier 112. In some examples, the functionality of the ML model(s) 110 is included in the classifier 112 or the functionality of the classifier 112 is included in the ML model(s) 110.

[0040] In various implementations, the ML model(s) 110 and / or the classifier 112 include a language model, an encoder model, an encoder-decoder model, an embedding model, a transformer model, a neural network, a topic model, a natural language processing model , a deep learning model, or another ML model known in the art, that are defined according to one or more parameters. In various implementations, the predictive model 106 (e.g., the ML model(s) 110 and / or the classifier 112) includes one or more deep learning models, such as neural networks (NNs) or multilayer perceptrons (MLPs). In various implementations, the predictive model 106 includes one or more language models, such large language models (LLMs). The ML model(s) 110 and / or the classifier 112 may de generated by selecting a machine learning algorithm, a feature extraction algorithm, an artificial intelligence algorithm, a Bayesian algorithm, a statistical analysis algorithm, a topic modeling algorithm, a clustering algorithm, or another suitable algorithm. The trainer 104 is configured to optimize the parameter(s) of the predictive model 106 based on training data 114.

[0041] The training data 114 includes training prompts and / or training responses 116. In various implementations, the training prompts and / or training responses 116 include prompts and / or responses from a database (e.g., a database including psychological surveys, a database including ongoing or published studies, a database of self-report data, a database of marketing research, etc.), prompts and / or responses from experimental data, prompts and / or responses from records (e.g., patient records, clinical records, marketing research records, etc.), or the like. In various implementations, the training prompt includes text, a visual prompt, or a sound. For example, the text may include a question, a description of a scenario, or the like. The visual prompt, for example, may include a thematic apperception test, a Rorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, a captcha, or the like. The training prompt, in some implementations, may be an open-ended prompt or a close-ended prompt (e.g., a prompt with a finite number of responses). The training response, in some examples, may be a response to an open-ended prompt. In various implementations, the training response includes a response to a close-ended prompt. For instance, the training response may include a text response, a selection of a value from a scale, a selection of a choice from a multiple-choice question, a verbal response, a physiological response, or a behavior.

[0042] In some implementations, the training prompts and / or training responses 116 are derived from psychological surveys. For instance, the training prompts and / or training responses 116 may include a prompt and / or a response to a prompt from a Narcissism Personality Inventory, a McLean Screening Instrument, a Multidimensional Inventory of Dissociation 60-item (MID-60), a Personality Inventory for Diagnostic and Statistical Manual of Mental Disorders (DSM)-5 (PID-5), a Level 1 Cross-Cutting Symptom Measure, a Level 2 Cross-Cutting Symptom Measure, a Disorder-SpecificSeverity Measure, a Levenson's Self-Report Psychopathy Test, a thematic apperception test, a Rorschach test, a Draw- A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, a captcha, or the like. In some implementations, the training prompts and / or training responses 116 may include a prompt and / or a response to a prompt from a Beck Depression Inventory (BDI), a Generalized Anxiety Disorder (GAD-7), a Depression Anxiety Stress Scale (DASS), a Brief-COPE, a Positive and Negative Affect Schedule (PANAS), a State Trait Anxiety Inventory (STAI), a modified version of a Russel Mood Circumplex, a modified version of a NIH Sleep Diary, a 36 Item Short Form Health Survey (SF-36), a 5D Altered state of Consciousness Scale (5d-ASC), a daily sleep diary questionnaire, a Hamilton Rating Scale for Depression, a Hamilton Anxiety Rating Scale, a mood survey, a self-reporting survey, a sleep diary, or any other survey known in the art.

[0043] In various implementations, the training prompts and / or training responses 116 are associated with at least one personal attribute. Examples of personal attributes include, but are not limited to, an emotional state, a preference (e.g., a product preference, a lifestyle preference, a religious affiliation, or the like), a demographic characteristic (e.g., gender, age, ethnicity, socioeconomic status, or the like), a diagnosis of a psychological condition or a lack thereof, or the like. An emotional state may include, for example, depression, anxiety, stress, coping, mood, attention, phobias, demoralization, or rumination. In some examples, the personal attribute includes a comprehensive psychological assessment of a subject (e.g., the subject providing the training response). For instance, a training prompt may be configured to determine a personal attribute of a subject answering the training prompt. A training response, in some examples, is from a subject having a personal attribute. In various implementations, a training response may be a benchmark training response determined by an expert, a trained professional, or the like. The benchmark training response, for instance, may be configured to assist to determine a personal attribute associated with a response.

[0044] The training data 114, in some examples, further includes labels 118. The labels 118 may indicate whether a training prompt or a training response is associated with a personal attribute. In some examples, the labels 118 are generated by at least one expert (e.g., a physician with specialized training for identifying one or more personal attributes, a psychologist or therapist with specialized training for identifying one or more personal attributes, or the like). In some examples, the labels 118 are generated computationally (e.g., by a data scientist, statistician, researcher, or by one or more processors). For instance, data about a subject associated with a training response may be analyzed, and one or more personal attributes associated with the training response or a training prompt associated with the training response may be determined.

[0045] The trainer 104 may optimize the parameters of the ML model(s) 110 and the classifier 112 in the predictive model 106 based on the training data 114. The ML model(s) 110 and / or classifier 112 may include one or more NNs. The term "Neural Network (NN)," and its equivalents, may refer to a model with multiple hidden layers, wherein the model receives an input (e.g., an image) and transforms the input by performing operations via the hidden layers. An individual hidden layer may include multiple “neurons,” each of which may be disconnected from other neurons in the layer. An individual neuron within a particular layer may be connected to multiple (e.g., all) of the neurons in the previous layer. A NN may further include at least one fully connected layer that receives a feature map output by the hidden layers and transformsthe feature map into the output of the NN.

[0046] The trainer 104 may optimize the classifier 112, such that the classifier accurately outputs the labels 118 in response to receiving the training prompts and / or training responses 116 as inputs. Examples of the classifier 112 include, for example, a deep-learning aided classifier, such as a psychological classifier (e.g., configured to identify a psychological condition, an emotional state, etc.), a behavioral classifier (e.g., configured to identify a behavior), a demographic classifier (e.g., configured to identify a demographic classifier), or another type of personal attribute classifier.

[0047] The trainer 104 can perform various techniques to train (e.g., optimize the parameters of) the classifier 112 using the training prompts and / or training responses 116. For instance, the trainer 104 may perform a training technique utilizing stochastic gradient descent with backpropagation, or any other machine learning training technique known to those of skill in the art. The trainer 104 may be configured to identify predictive features of the training prompts and / or training responses 116 that are indicative of the labels 118. The predictive features, for instance, may include factor loadings. The trainer 104, in various examples, determines semantic structures of the training prompts and / or training responses 116. In various implementations, an accuracy of factor loadings from a trained ML model is higher than an accuracy of factor loadings from an untrained ML model.

[0048] In various implementations, the trainer 104 may be configured to train the predictive model 106 by optimizing various parameters within the predictive model 106 (e.g., within the ML model(s) 110 and / or the classifier 112) based on the training prompts and / or training responses 116. For example, the trainer 104 may input the training prompts and / or training responses 116 into the predictive model 106 and compare outputs of the predictive model 106 to the labels 118. The trainer 104 may further modify various parameters of the predictive model 106 (e.g., filters in the ML model(s) 110 and / or the classifier 112) in order to ensure that the outputs of the predictive model 106 are sufficiently similar and / or identical to the labels 118. For instance, the trainer 104 may identify values of the parameters that result in a minimum of loss between the outputs of the predictive model 106 and the labels 118.

[0049] In various implementations, the trainer 104 may be configured to perform unsupervised training that does not utilize the labels. For instance, the trainer 104 may be configured to determine similarities or patterns between each of the training prompts and / or responses 116. In some examples, the trainer 104 determines a pattern in the semantic structures of the training prompts and / or responses 116.

[0050] Once the ML model(s) 110 and the classifier 112 are trained, the predictive model 106 can be utilized to analyze new prompts and / or responses. For example, a prompt and / or a response 120 may be generated or collected (e.g., by a researcher, a data scientist, a psychologist, a clinician, a trained user, etc.). In some examples, the prompt and / or the response 120 are collected from an input source 122, such as a database, an experimental result (e.g., a self-report survey, a clinical study, etc.), a record, or the like. The response, for instance, may be collected from a subject 124. The subject 124 may be a human, a non-human primate, a mammal, or another organism. In some examples, the subject 124 is an adult, an adolescent, a child. In some examples, the subject 124 is a parent or guardian of a child. In various examples, the prompt and / or the response 120 is different than the training prompts and / or training responses 116 used to train the predictive model 106. For instance, the subject 124 providing the response may be different than the subjects providingthe training responses.

[0051] The prompt may include text, a visual prompt, a sound, or the like. In some examples, the prompt includes a question, a description of a scenario, or the like. The prompt, for example, may include a thematic apperception test, a Rorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, a captcha, or the like. The prompt, in some implementations, may be an open-ended prompt or a close-ended prompt (e.g., a prompt with a finite number of answers). The response, in some examples, may be a response to an open-ended prompt. In various implementations, the response includes a response to a close-ended prompt. For instance, the response may include a text response, a selection of a value from a scale, a selection of a choice from a multiple-choice question, a verbal response, a physiological response, or a behavior. In various examples, the prompt and / or response 120 include a prompt and / or a response from a psychological survey or other source disclosed herein.

[0052] The prompt and / or response 120 may have the same or similar format to the training prompts and / or training responses 116. For instance, the prompt and the training prompts may include visual prompts. In some examples, the response and the training responses may include behavioral responses.

[0053] In some examples, the ML model(s) 110 accept the prompt and / or the response 120 as an input. The ML model(s) 110 generate a vector, an embedding, a matrix, or the like based on the prompt and / or the response 120 . For instance, the ML model(s) 110 may include a language model configured to generate an embedding indicative of the semantic structure of the prompt and / or the response 120. For instance, the ML model(s) 110 may be configured to accept two or more different prompts and determine a similarity of the semantic structure of the different prompts. Accordingly, the embeddings associated with each of the different prompts may be similar or identical.

[0054] In various implementations, the classifier 112 accepts the prompt and / or the response 120 as an input. In some examples, the classifier 112 accepts an output of the ML model(s) 110 (e.g., the vector, the embedding, the matrix, etc.) as an input. The predictive model 106 may output a predicted label 126 (e.g., a predicted personal attribute associated with the prompt and / or the response 120) to one or more end user device(s) 128. In some implementations, the predictive model 106 may output the prompt and / or the response 120 to the end user device(s) 128. The end user device(s) 128 may include, for instance, a device used by a doctor, a therapist, a coach, a teacher, a researcher, the subject 124 providing the response, or the like.

[0055] In various implementations, the end user device(s) 128 are configured to output an indication of the predicted label 126 (e.g., the predicted personal attribute). In some examples, the indication may include a message on a graphical user interface (e.g., a visual display). The indication may include a personalized message, a recommendation, the predicted label, or the like. In some implementations, the personalized message includes a marketing message, and the marketing message may be output, for instance, to the subject 124 providing the response via an email, social media, a mobile application, or the like. In various examples, the recommendation includes a treatment regimen, such as a drug (e.g., a prescribed medication, an over-the-counter drug, a dietary supplement, or the like), a behavior modification program, or any other suitable treatment regimen. The behavior modification program, in some cases, is configured to change a substance intake (e.g., food intake, alcohol intake, drug intake, etc.), physical exercise, sleep schedule, performance ofself-defeating behaviors, social interactions, compulsive behaviors, or the like of the subject 124 providing the response. In some examples, a first end user device may be configured to transmit the indication to a second end user device.

[0056] In some implementations, the predictive model 106 may output the prompt and / or the response 120, the output of the ML model(s) 110 (e.g., the vector, the embedding, the matrix, etc.), the output of the classifier 112, or the predicted label 126 to the visualizer 108.

[0057] In various implementations, it may be beneficial to visualize the output of the predictive model 106. For instance, the predictive model 106 may be applied each prompt of a psychological survey, and it may be beneficial to understand the relationship between each prompt and the personal attributes associated with the psychological survey. In some examples, the predictive model 106 may be applied to an individual's responses to each prompt, and it may be beneficial to determine a comprehensive psychological assessment of the individual. These issues can be addressed, for example, by generating a multidimensional model 130 that represents the output of the predictive model 106. The multidimensional model 130, in various implementations, enables analysis of and pattern identification in the output of the predictive model 106. In some implementations, multidimensional models may be generated for multiple psychological surveys, or for multiple individuals' responses to one or more psychological surveys. The multidimensional models can be analyzed (e.g., visually, computationally, etc.) to improve understanding of, for instance, the relationship between semantic structure and personal attributes.

[0058] The visualizer 108, in various examples, is configured to generate the multidimensional model 130 that is indicative of a semantic structure or a classification associated with the prompt and / or the response 120. For instance, a trained user may analyze the multidimensional model 130 to determine a personal attribute associated with the prompt and / or the response 120. The multidimensional model 130 may include a two-dimensional model, a three-dimensional model, or a model of another dimensionality. The visualizer 108, in various examples, is configured to apply a clustering technique, a dimensionality reduction technique, a data visualization technique, a structural alignment, or the like to an output of the predictive model 106. For instance, the visualizer 108 may apply multidimensional scaling (MDS), principal component analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), or the like to an output of the ML model(s). Examples of clustering techniques include, but are not limited to, affinity propagation clustering, spectral clustering, hierarchical clustering, k-means clustering, or mean shift clustering. The clustering technique may include an objective function configured to optimize a spatial distribution of one or more clusters. Examples of structural alignment include, but are not limited to, Procrustes analysis, root-mean-square deviation measurement, hotspot analysis, or kernel density estimation.

[0059] The visualizer 108, in some examples, may be configured to determine an Adjusted Rand Index (ARI) or a Normalized Mutual Information (NMI) to evaluate a result of at least one of the clustering technique, the dimensionality reduction technique, the data visualization technique, or the structural alignment. In some examples, applying MDS includes determining a cosine similarity, a Euclidean distance, a Manhattan distance, or the like. The visualizer 108, in various instances, performs scaling. For instance, the visualizer may be configured to perform non-linear scaling if the prompt and / or response 120 include less than 61 prompts and / or responses. The visualizer may perform linear scaling ifthe prompt and / or response 120 include more than 61 prompts and / or responses. In various implementations, the predictive model 106 and / or the visualizer 108 are configured to receive between 1 to 1 million prompts and / or responses as an input.

[0060] While FIG. 1 describes the ML model(s) 110, the classifier 112, and the visualizer 108 separately, implementations of the present disclosure are not so limited. In some examples, the visualizer 108 and / or the classifier 112 are part of the ML model(s) 110. In some examples, techniques performed by the ML model(s) 110, the classifier 112, or the visualizer 108 may be performed by a different component (e.g. , of the prediction system 102).

[0061] In various implementations, the visualizer 108 is configured to determine a position of a data point in a multidimensional space that is indicative of a semantic structure of the prompt and / or the response 120. In some implementations, the multidimensional model 130 illustrates a relationship between a plurality of prompts and / or responses. In various implementations, the visualizer 108 is configured to determine a continuous multidimensional structure based on the distribution of data points. For instance, the multidimensional model 130 may be configured to illustrate a distribution of data points associated with the semantic structure of reference prompts and / or responses. The reference prompts and / or responses may include the training prompts and / or the training responses 116. In some examples, the reference prompts and / or responses are different that the training prompts and / or training responses 116. For instance, the predictive model 106 may be applied to the reference prompts and / or responses, and the visualizer 108 may determine a multidimensional distribution of data points associated with the output, from the predictive model 106, for each of the reference prompts and / or responses. Based on the multidimensional distribution of data points, the visualizer may generate a three- dimensional model that includes at least one data point that is indicative of the semantic structure of the prompt and / or response 120. In some examples, the multidimensional model 130 includes a distribution of data points that is indicative a relationship between each of the prompt and / or the response 120. For instance, a distance between each of the data points (e.g., data points associated with the reference prompts and / or responses, the prompt and / or response 120, or the training prompts and / or responses 116) may correspond to a similarity of semantic structure of each of the data points. The terms “relationship,” “similarity," and their equivalents, as used herein, refer to a measure of how alike two items are. For instance, a relationship of or a similarity between two items may be measured by semantic similarity (e.g., semantic interrelatedness), similarity in sound (e.g., frequency, tone, amplitude, etc.), similarity in behavior (e.g., speed, type, action, etc.), visual similarity (e.g., color, intensity, shape, style, etc.), or another metric that is applied to the two items. This distance, in various examples, includes a cosine similarity, a Euclidean distance, a Manhattan distance, or another suitable metric.

[0062] The relationship may be represented on the surface of the multidimensional model 130, by a volume (e.g., a mesh volume) of the multidimensional model 130, or by another characteristic of the multidimensional model 130 (e.g., distance, color, area, structure, texture, etc.). In some implementations, the multidimensional model 130 includes a surface model, a wire model, a wireframe model, a space-filling model, a cartoon model, a ribbon model, a point cloud model, or any other suitable model. According to various implementations, the multidimensional model 130 is independent of the order that a plurality of prompts and / or responses 120 are inputted into the visualizer 108 or the predictive model 106.

[0063] In various implementations, the visualizer 108 is configured to apply a structural biology tool. For instance, the visualizer 108 may apply a structural biology tool to the relationship between one or more prompt(s) and / or response(s)120. In some examples, the visualizer 108 applies the structural biology tool to the output of the ML model(s) 110, to the output of the classifier 112, or any other metric described herein. The structural biology tool may include PyMOL, ChimeraX, Visual Molecular Dynamics, or the like. In various implementations, determining the multidimensional model 130 is analogous to determining a three-dimensional structure of a protein based on an interrelationship of amino acids.

[0064] In various implementations, one or more multidimensional model(s) 130 can be analyzed (e.g., visually, computationally, etc ). For example, two or more multidimensional models may be compared to determine a similarity between the corresponding prompts and / or responses. Accordingly, the multidimensional models can be utilized to determine a similarity between a survey including the prompts, an individual providing the responses, or the like. In some implementations, a local analysis may be performed on a multidimensional model 130 to determine a relationship (e.g., a similarity) between two or more data points). In various examples, a global analysis may be performed on a multidimensional model 130 to determine a personal attribute associated with the distribution of data points of the multidimensional model 130.

[0065] In some examples, the output of the predictive model 106 or the visualizer 108 may be compared to biological data. For instance, a multidimensional model 130 associated with the subject 124 providing a response may be compared to clinical information (e.g., medical diagnosis, physiological data, anatomical information, family and medical history, health assessment, biomarker level, genetic information, etc.). In some examples, the biological data includes a brain structure.

[0066] In various implementations, a questionnaire may be generated based on the output of the predictive model 106 and / or the visualizer 108. For instance, the end user device(s) 128 and / or a trained user may generate a psychological questionnaire, a marketing survey, or the like. In some examples, the questionnaire is configured to be administered to the subject 124 providing the response. For instance, the questionnaire may be configured to determine a different personal attribute than the predicted label 126. In some examples, the questionnaire may be configured to determine a comprehensive psychological assessment of the subject 124.

[0067] FIG. 2A illustrates an example process 200 for using a model to determine a relationship of data to a reference set. The process 200 may be performed by an entity, such as at least one of the prediction system 102, the trainer 104, the predictive model 106, the visualizer 108, the end user device(s), or one or more processors.

[0068] At 202, the entity determines a relationship of data to a reference set by inputting the data into a model. In various implementations, the data includes a prompt and / or a response (e.g., the prompt and / or response 120) generated or collected from a subject (e.g., the subject 124), an input source (e.g., the input source 122), or another source. The reference set may include training data (e.g., the training prompts and / or training responses 116) used to train the model. In some examples, the reference set includes prompts and / or responses that are different that the data and the training data. In some example, the data includes a plurality of prompts and / or responses, and the model is configured to determine a relationship between each of the plurality of prompts and / or responses. The model may include a predictive model (e.g., the predictive model 106) that includes one or more ML model(s) (e.g., ML model(s) 110) and / or a classifier (e.g., the classifier 112). The predictive model, for instance, is configured to predict a personal attribute associated with the data (e.g., the prompt and / or response). In various implementations, the output of the ML model(s) and the output of the classifierare independent of the order in which the data is inputted into the model.

[0069] At 204, the entity generates an output indicative of the relationship of the data to the reference set. The output may include a multidimensional model (e.g., the multidimensional model 130). In various implementations, the model further includes a visualizer (e.g., the visualizer 108) that is configured to generate the multidimensional model based on the data, the output of the ML model(s), or the output of the classifier. In some examples, the output includes a predicted personal attribute (e.g., the predicted label 126). In various implementations, the output is transmitted to one or more end user device(s) (e.g., the end user device(s) 128). In some examples, the multidimensional model and / or the predicted personal attribute is analyzed. For instance, the entity may generate the multidimensional model, analyze the multidimensional model to determine a pattern in the data, and output the pattern to the end user device(s). In various examples, the relationship of the data to the reference set is analyzed to determine a pattern in the data, a personalized message or recommendation, or another output. In some examples, the relationship of the data to the reference set is analyzed to generate a questionnaire, for instance, to determine a comprehensive psychological assessment of the subject providing the response.

[0070] FIG. 2B illustrates an example process 206 for generating and training a model described herein. The process 206 may be performed by an entity, such as at least one of the prediction system 102, the trainer 104, the predictive model 106, the visualizer 108, the end user device(s), or one or more processors, or a trained user, such as a researcher, a data scientist, a data analyst, a software engineer, or the like.

[0071] At 208, the entity generates a model configured to determine a relationship of data (e.g., the prompt and / or response 120) to a reference set. In various implementations, generating the model (e.g., the predictive model 106, the ML model(s) 110, the classifier 112, or any other model described herein) includes selecting a machine learning algorithm, a feature extraction algorithm, an artificial intelligence algorithm, a Bayesian algorithm, a statistical analysis algorithm, a topic modeling algorithm, a clustering algorithm, or another suitable algorithm.

[0072] At 210, the entity trains the model based on training data (e.g., training data 114). In various examples, the training data includes training prompts and / or training responses (e.g., the training prompts and / or training responses 116). In some cases, the training data includes labels (e.g., the labels 118). The training data may be inputted into a trainer (e.g., the trainer 104) configured to train the model to predict personal attributes associated with the data. In some examples, the trainer trains each component (e.g., the ML model(s) 110, the classifier 112) of the model (e.g., the predictive model 106). In various examples, the trainer may use supervised learning by using the labels to improve the predictions of the model. In some cases, the trainer may use unsupervised learning by identifying patterns in the training data without using the labels. Based on training the model, the entity, in some examples, inputs the data into the model to determine predicted labels associated with the data.

[0073] FIG. 3 shows an example computer architecture for a computer 300 capable of executing program components for implementing the functionality described herein. The computer architecture shown in FIG. 3 illustrates a conventional computer, workstation, desktop computer, laptop, tablet, network appliance, e-reader, smartphone, or other computing device, and can be utilized to execute any of the processes (e.g. , process 200, process 206, or any other process describedherein) presented herein.

[0074] The computer 300 includes a baseboard 302, or "motherboard,” which is a printed circuit board to which a multitude of components or devices can be connected by way of a system bus or other electrical communication paths. In one illustrative configuration, one or more processing units, such as (“CPUs”) 304, GPUs, TPUs, ASICs, FPGAs, or the like, and / or threads, kernels, cores, and / or the like thereof, that may operate in conjunction with a chipset 306. The CPUs 304 can be standard programmable processors that perform arithmetic and logical operations necessary for the operation of the computer 300.

[0075] The CPUs 304 perform operations by transitioning from one discrete, physical state to the next through the manipulation of switching elements that differentiate between and change these states. Switching elements generally include electronic circuits that maintain one of two binary states, such as flip-flops, and electronic circuits that provide an output state based on the logical combination of the states of one or more other switching elements, such as logic gates. These basic switching elements can be combined to create more complex logic circuits, including registers, adders- subtractors, arithmetic logic units, floating-point units, and the like.

[0076] The chipset 306 provides an interface between the CPUs 304 and the remainder of the components and devices on the baseboard 302. The chipset 306 can provide an interface to a random-access memory (RAM) 308 or any other suitable form of memory, used as the main memory in the computer 300. The chipset 306 can further provide an interface to a computer-readable storage medium such as a read-only memory (ROM) 310 or non-volatile RAM (NVRAM) for storing basic routines that help to startup the computer 300 and to transfer information between the various components and devices. The ROM 310 or NVRAM can also store other software components necessary for the operation of the computer 300 in accordance with the configurations described herein.

[0077] The computer 300 can operate in a networked environment using logical connections to remote computing devices and computer systems through a network, such as the network 324. The chipset 306 can include functionality for providing network connectivity through a network interface controller (NIC) 312, such as a gigabit Ethernet adapter. The NIC 312 is capable of connecting the computer 300 to other computing devices over the network 324. It should be appreciated that multiple NICs 312 can be present in the computer 300, connecting the computer 300 to other types of networks and remote computer systems. In some instances, the NICs 312 may include at least on ingress port and / or at least one egress port.

[0078] The computer 300 can include an input / output (I / O), such as a controller sufficient to transmit processorexecutable instructions to or receive processor-executable instructions from a device. For example, the I / O controller include or interface with one or more user interface devices (e.g. , a display, speaker, a keyboard, a mouse, a trackpad, a touchscreen), one or more servers, laboratory equipment (e.g., DNA / RNA sequencing device(s), staining and / or stain imaging device(s), probe hybridization / imaging / decoding device(s)), and / or the like. Interfacing with any of these devices may additionally or alternatively be executed by the network interface controller 312.

[0079] The computer 300 can be connected to a storage device 318 that provides non-volatile storage for the computer. The storage device 318 can store an operating system 320, programs 322, and data, which have been described in greater detail herein. The storage device 318 can be connected to the computer 300 through a storage controller 314 connectedto the chipset 306. The storage device 318 can consist of one or more physical storage units. The storage controller 314 can interface with the physical storage units through a serial attached small computer system interface (SCSI) (SAS) interface, a serial advanced technology attachment (SATA) interface, a fiber channel (FC) interface, or other type of interface for physically connecting and transferring data between computers and physical storage units.

[0080] The computer 300 can store data on the storage device 318 by transforming the physical state of the physical storage units to reflect the information being stored. The specific transformation of physical state can depend on various factors, in different embodiments of this description. Examples of such factors can include, but are not limited to, the technology used to implement the physical storage units, whether the storage device 318 is characterized as primary or secondary storage, and the like.

[0081] For example, the computer 300 can store information to the storage device 318 by issuing instructions through the storage controller 314 to alter the magnetic characteristics of a particular location within a magnetic disk drive unit, the reflective or refractive characteristics of a particular location in an optical storage unit, or the electrical characteristics of a particular capacitor, transistor, or other discrete component in a solid-state storage unit. Other transformations of physical media are possible without departing from the scope and spirit of the present description, with the foregoing examples provided only to facilitate this description. The computer 300 can further read information from the storage device 318 by detecting the physical states or characteristics of one or more particular locations within the physical storage units.

[0082] In addition to the storage device 318 described above, the computer 300 can have access to other computer- readable storage media to store and retrieve information, such as program modules, data structures, or other data. It should be appreciated by those skilled in the art that computer-readable storage media is any available media that provides for the non-transitory storage of data and that can be accessed by the computer 300. In some examples, the operations performed by any network node described herein may be supported by one or more devices similar to computer 300. Stated otherwise, some or all of the operations performed by a network node may be performed by one or more computer devices 300 operating in a cloud-based arrangement.

[0083] By way of example, and not limitation, computer-readable storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology. Computer-readable storage media includes, but is not limited to, RAM, ROM, erasable programmable ROM ("EPROM”), electrically-erasable programmable ROM (“EEPROM”), flash memory or other solid-state memory technology, compact disc ROM ("CD-ROM”), digital versatile disk ("DVD”), high definition DVD (“HD-DVD"), BLU-RAY, or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information in a non-transitory fashion.

[0084] As mentioned briefly above, the storage device 318 can store an operating system 320 utilized to control the operation of the computer 300. According to one embodiment, the operating system 320 includes the LINUX® operating system. According to another embodiment, the operating system 320 includes the WINDOWS SERVER® operating system from MICROSOFT Corporation of Redmond, Washington. According to further embodiments, the operating system can include the UNIX® operating system or one of its variants. It should be appreciated that other operating systems can alsobe utilized. The storage device 318 can store other system or application programs and data utilized by the computer 300.

[0085] In one embodiment, the storage device 318 or other computer-readable storage media includes one or more programs 322. The programs 322, for example, include computer-executable instructions which, when loaded into the computer 300, transform the computer from a general-purpose computing system into a special-purpose computer capable of implementing the embodiments described herein. These computer-executable instructions transform the computer 300 by specifying how the CPUs 304 transition between states, as described above. According to one embodiment, the computer 300 has access to computer-readable storage media storing computer-executable instructions which, when executed by the computer 300, perform the various processes described herein. The computer 300 can also include computer-readable storage media having instructions stored thereupon for performing any of the other computer- implemented operations described herein. The program(s) 322, for example, include one or more processes. The process(es) may include instructions that, when executed by the CPU(s) 304, cause the computer 300 and / or the CPU(s) 304 to perform one or more operations.

[0086] The computer 300 can also include one or more input / output controllers 326 for receiving and processing input from a number of input devices, such as a keyboard, a mouse, a touchpad, a touch screen, an electronic stylus, or other type of input device. Similarly, an input / output controller 326 can provide output to a display, such as a computer monitor, a flat-panel display, a digital projector, a printer, or other type of output device. It will be appreciated that the computer 300 might not include all of the components shown in FIG. 3, can include other components that are not explicitly shown in FIG. 3, or might utilize an architecture completely different than that shown in FIG. 3.

[0087] In some instances, one or more components may be referred to herein as "configured to,” “configurable to,” “operable / operative to,” “adapted / adaptable,” “able to,” “conformable / conformed to," etc. Those skilled in the art will recognize that such terms (e.g., “configured to”) can generally encompass active-state components and / or inactive-state components and / or standby-state components, unless context requires otherwise.

[0088] Exemplary Embodiments.1. A method including: inputting data including: (i) a prompt and / or (ii) a response of an individual to the prompt into a model configured to determine a relationship of the prompt and / or the response to a reference set and generating an output indicative of the relationship.2. The method of embodiment 1 , wherein the prompt and the response are indicative of a personal attribute of the individual.3. The method of embodiment 1 or 2, wherein the prompt includes at least one of text, a visual prompt, or a sound.4. The method of embodiment 3, wherein the text includes a question and / or a description of a scenario.5. The method of embodiment 4, wherein the question is an open-ended question.6. The method of any of embodiments 3-5, wherein the visual prompt includes at least one of a thematic apperception test, a Rorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, or a captcha.7. The method of any of embodiments 1-6, wherein the response includes at least one of a text response, a selectionof a value from a scale, a selection of a choice from a multiple-choice question, a verbal response, a physiological response, or a behavior.8. The method of any of embodiments 1-7, wherein the relationship includes a similarity between: a semantic structure of the reference set; and a semantic structure of the prompt and / or the response.9. The method of any of embodiments 1-8, wherein the data includes the response, and wherein the method further includes determining, based on the relationship, a personal attribute of the individual.10. The method of any of embodiments 1-9, wherein the individual is an adult, an adolescent, or a child.11 . The method of embodiment 10, wherein the adult is a parent or guardian of a child.12. The method of any of embodiments 9-11 , wherein the personal attribute includes at least one of an emotional state of the individual, a preference of the individual, a demographic characteristic of the individual, or a diagnosis of a psychological condition or a lack thereof of the individual.13. The method of embodiment 12, wherein the emotional state includes at least one of depression, anxiety, stress, coping, mood, attention, phobias, demoralization, or rumination.14. The method of any of embodiments 9-13, further including: outputting, based on the personal attribute, a personalized message or a recommendation.15. The method of embodiment 14, wherein the personalized message includes a marketing message.16. The method of embodiment 15, wherein outputting the personalized message includes outputting the marketing message via an email, social media, or a mobile application.17. The method of any of embodiments 14-16, wherein the recommendation includes a treatment regimen.18. The method of embodiment 17, wherein the treatment regimen indicates at least one of a drug or a behavior modification program.19. The method of embodiment 18, wherein the drug includes at least one of a prescribed medication, an over-the- counter drug, or a dietary supplement.20. The method of embodiment 18 or 19, wherein the behavior modification program is configured to change at least one of a substance intake, physical exercise, sleep schedule, performance of self-defeating behaviors, social interactions, or compulsive behaviors of the individual.21 . The method of embodiment 20, wherein the substance intake includes at least one of food intake, alcohol intake, or drug intake.22. The method of any of embodiments 14-21, wherein outputting the personalized message includes displaying the recommendation on a graphical user interface.23. The method of any of embodiments 14-22, wherein outputting the personalized message includes outputting the recommendation to a third party.24. The method of embodiment 23, wherein the third party is at least one of a doctor, a therapist, a coach, a teacher, or a researcher.25. The method of any of embodiments 2-24, wherein the data includes the response, and the personal attributeincludes a comprehensive psychological assessment of the individual.26. The method of any of embodiments 1-25, wherein the output includes a visual model of the relationship.27. The method of embodiment 26, wherein generating the output indicative of the relationship includes generating the visual model by applying a structural biology visualization tool to the relationship.28. The method of embodiment 27, wherein the structural biology visualization tool includes PyMOL, ChimeraX, or Visual Molecular Dynamics.29. The method of any of embodiments 26-28, wherein the visual model includes a multi-dimensional model.30. The method of any of embodiments 26-29, wherein the visual model includes a three-dimensional model.31. The method of embodiment 30, wherein the data includes multiple fields, each of the multiple fields includes a prompt and / or response, and wherein the three-dimensional model is configured to represent an interrelationship between semantic structures of the multiple fields.32. The method of embodiment 30 or 31 , wherein the data is inputted in an order, and wherein the three-dimensional model is independent of the order.33. The method of any of embodiments 30-32, wherein the three-dimensional model presents the relationship on a surface of a structure.34. The method of embodiment 33, wherein the three-dimensional model includes: at least one first data point corresponding to a semantic structure of the prompt and / or response.35. The method of embodiment 34, wherein the three-dimensional model further includes at least one second data point corresponding to a semantic structure of the reference set.36. The method of embodiment 34 or 35, wherein the surface includes the at least one first data point.37. The method of any of embodiments 34-36, wherein the data includes multiple fields, each of the multiple fields includes a prompt and / or response, wherein the at least one first data point includes multiple first data points respectively corresponding to semantic structures of the multiple fields, and wherein a distance between the multiple first data points correspond to similarities between the semantic structures.38. The method of embodiment 37, wherein the distance includes a cosine similarity, a Euclidean distance, or a Manhattan distance.39. The method of embodiment 37 or 38, wherein generating the three-dimensional model includes determining a distribution of the multiple first data points based on an interrelationship of the semantic structures of the multiple fields.40. The method of embodiment 39, wherein determining the distribution is analogous to determining a three- dimensional structure of a protein based on an interrelationship of amino acids.41 . The method of any of embodiments 37-40, wherein the data includes more than one prompt from a psychological questionnaire, and wherein the model includes a three-dimensional representation of the psychological questionnaire.42. The method of any of embodiments 33-41 , further including comparing the three-dimensional model to biological data.43. The method of embodiment 42, wherein the biological data includes a brain structure.44. The method of any of embodiments 30-44, wherein the three-dimensional model includes at least one of a surface model, a wire model, a wireframe model, a space-filling model, a cartoon model, a ribbon model, or a point cloud model.45. The method of embodiment 44, wherein the three-dimensional model presents the relationship using at least one of distance, color, or volume.46. The method of embodiment 45, wherein the volume is a mesh volume.47. The method of any of embodiments 26-46, wherein the data includes a response, and wherein the method further includes evaluating the visual model to identify a personal attribute of the individual.48. The method of any of embodiments 1-47, wherein the data includes a first prompt or a first response from a first individual, and the relationship is a first relationship, the method further including: inputting data including a second prompt or a second response from a second individual into the model; and generating a second output indicative of a second relationship of the second response and the reference set.49. The method of embodiment 48, further including: determining that the first response is different than the second response; and determining that the first relationship is the same as the second relationship.50. The method of embodiment 49, wherein the second output is a second three-dimensional model, and wherein a first three-dimensional model corresponding to the first response is the same as a second three-dimensional model corresponding to the second response.51 . The method of any of embodiments 48-50, further including: determining that the first response is different than the second response; and determining that the first relationship is different than the second relationship.52. The method of any of embodiments 1-51 , wherein the model includes a machine learning (ML) model.53. The method of embodiment 52, wherein the ML model includes a language model, an encoder model, an encoderdecoder model, an embedding model, a transformer model, a neural network, a topic model, natural language processing, or deep learning.54. The method of any of embodiments 1-53, wherein the model includes a language model and a neural network.55. The method of embodiment 54, wherein the language model includes a large language model (LLM) and the neural network includes a Multilayer Perceptron (MLP).56. The method of any of embodiments 52-55, further including: training the ML model by optimizing parameters of the ML model based on training data, the training data including example prompts and / or example responses identified from example samples.57. The method of embodiment 56, wherein the example samples include: the example prompts from at least one of a Narcissism Personality Inventory, a McLean Screening Instrument, a Multidimensional Inventory of Dissociation 60-item (MID-60), a Personality Inventory for Diagnostic and Statistical Manual of Mental Disorders (DSM)-5 (PID-5), a Level 1 Cross-Cutting Symptom Measure, a Level 2 Cross-Cutting Symptom Measure, a Disorder-Specific Severity Measure, a Levenson's Self-Report Psychopathy Test, a thematic apperception test, a Rorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, or a captcha and / or the example responses to a prompt from at least one of the Narcissism Personality Inventory, McLean Screening Instrument, MID-60, PID-5, Level 1Cross-Cutting Symptom Measure, Level 2 Cross-Cutting Symptom Measure, Disorder-Specific Severity Measure, Levenson's Self-Report Psychopathy Test, a thematic apperception test, a Rorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, or a captcha.58. The method of embodiment 56 or 57, wherein the example samples include: the example prompts from at least one of a Beck Depression Inventory (BDI), a Generalized Anxiety Disorder (GAD-7), a Depression Anxiety Stress Scale (DASS), a Brief-COPE, a Positive and Negative Affect Schedule (PANAS), a State Trait Anxiety Inventory (STAI), a modified version of a Russel Mood Circumplex, a modified version of a NIH Sleep Diary, a 36 Item Short Form Health Survey (SF-36), a 5D Altered state of Consciousness Scale (5d-ASC), a daily sleep diary questionnaire, a Hamilton Rating Scale for Depression, a Hamilton Anxiety Rating Scale, a mood survey, a self-reporting survey, or a sleep diary, and / or the example responses to the example prompts from at least one of the BDI, GAD-7, DASS, Brief-COPE, PANAS, STAI, modified version of the Russel Mood Circumplex, modified version of the NIH Sleep Diary, SF-36, 5d-ASC, daily sleep diary questionnaire, Hamilton Rating Scale for Depression, Hamilton Anxiety Rating Scale, mood survey, self-reporting survey, or sleep diary.59. The method of any of embodiments 56-58, wherein the example samples include at least one of a database, an experimental result, or a record.60. The method of any of embodiments 56-59, wherein the example samples include multiple example fields, each of the multiple example fields including an example prompt and / or example response, wherein the training data further includes labels indicating whether the example samples are associated with a personal attribute, and wherein training the ML model includes identifying, using supervised ML based on the labels, predictive features of the example fields that are indicative of the personal attribute.61 . The method of embodiment 60, wherein the predictive features include factor loadings.62. The method of embodiment 61 , wherein the ML model includes a neural network, the method further including: based on training the ML model: determining that an accuracy of the factor loadings of the trained ML model is higher than an accuracy of factor loadings of an untrained ML model; or determining that the accuracy of the factor loadings of the trained ML model is higher than a threshold.63. The method of any of embodiments 56-62, wherein training the ML model includes determining a similarity of semantic structures of the example prompts and / or the example responses.64. The method of embodiment 63, further including identifying a pattern in the semantic structures of the example prompts and / or the example responses.65. The method of any of embodiments 1-64, wherein the data includes more than one prompt and / or response, and wherein the model is configured to represent a semantic interrelatedness between each of the prompt and / or the response.66. The method of any of embodiments 3-65, wherein the model is configured to generate at least one of a vector, an embedding, or a matrix, and wherein the vector, the embedding, or the matrix is indicative of a semantic structure of the text.67. The method of embodiment 66, wherein the model is configured to apply at least one of a clustering technique, adimensionality reduction technique, a data visualization technique, or a structural alignment to the vector, the embedding, or the matrix.68. The method of embodiment 67, wherein the dimensionality reduction technique includes at least one of multidimensional scaling (MDS), principal component analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t- SNE), or Uniform Manifold Approximation and Projection (UMAP).69. The method of embodiment 68, wherein applying the MDS includes determining at least one of a cosine similarity, a Euclidean distance, or a Manhattan distance.70. The method of embodiment 68 or 69, wherein the model is configured to generate multiple encodings, each of the multiple encodings including a vector, an embedding, or a matrix, the method further including based on applying the dimensionality reduction technique, determining a distribution of data points in a multi-dimensional space, wherein the data points correspond to each of the multiple encodings.71. The method of embodiment 70, wherein the multi-dimensional space is a three-dimensional space, the method further including determining a continuous three-dimensional structure based on the distribution of data points.72. The method of embodiment 71, further including: performing a local analysis of a region of the continuous three- dimensional structure; and determining a relationship between two or more data points.73. The method of embodiment 71 or 72, further including: performing a global analysis of the continuous three- dimensional structure; and determining a personal attribute associated with the distribution of data points.74. The method of any of embodiments 67-73, wherein the clustering technique includes at least one of affinity propagation clustering, spectral clustering, hierarchical clustering, k-means clustering, or mean shift clustering.75. The method of any of embodiments 67-74, wherein the clustering technique includes an objective function configured to optimize a spatial distribution of one or more clusters.76. The method of any of embodiments 67-75, wherein the model is further configured to determine an Adjusted Rand Index (ARI) or a Normalized Mutual Information (NMI) to evaluate a result of at least one of the clustering technique, the dimensionality reduction technique, the data visualization technique, or the structural alignment.77. The method of any of embodiments 1-76, further including: based on the output, generating a questionnaire.78. The method of embodiment 77, wherein the questionnaire includes a question indicative of a personal attribute of an individual.79. The method of embodiment 78, wherein the personal attribute includes a comprehensive psychological assessment.80. The method of any of embodiments 1 -79, wherein the prompt or the response include 1 to 1 million prompts and / or 1 to 1 million responses.81 . The method of any of embodiments 1-80, wherein the prompt or the response include less than 61 prompts and / or responses, and wherein the model is configured to perform non-linear scaling.82. The method of any of embodiments 1-81 , wherein the prompt or the response include more than 60 prompts and / or responses, and wherein the model is configured to perform linear scaling.83. The method of any of embodiments 1-82, further including generating the model.84. The method of embodiment 83, wherein generating the model includes selecting at least one of a machine learning algorithm, a feature extraction algorithm, an artificial intelligence algorithm, a Bayesian algorithm, a statistical analysis algorithm, a topic modeling algorithm, or a clustering algorithm.85. A method, including: generating a machine learning (ML) model configured to determine a relationship of a prompt and / or a response to a reference set; training the ML model by optimizing parameters of the ML model based on training data, the training data including example prompts and / or example responses identified from example samples; and generating a report indicating a performance of the ML model.86. The method of embodiment 85, wherein generating the ML model includes selecting at least one of a machine learning algorithm, a feature extraction algorithm, an artificial intelligence algorithm, a Bayesian algorithm, a statistical analysis algorithm, a topic modeling algorithm, or a clustering algorithm.87. The method of embodiment 85 or 86, wherein the example prompts include at least one of text, a visual prompt, or a sound.88. The method of embodiment 87, wherein the text includes a question and / or a description of a scenario.89. The method of embodiment 88, wherein the question is an open-ended question.90. The method of any of embodiments 87-89, wherein the visual prompts include at least one of a thematic apperception test, a Rorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, or a captcha.91 . The method of any of embodiments 85-90, wherein the example responses include at least one of a text response, a selection of a value from a scale, a selection of a choice from a multiple-choice question, a verbal response, a physiological response, or a behavior.92. The method of any of embodiments 85-91 , wherein the example prompts and / or example responses include text, wherein the ML model is configured to generate at least one of a vector, an embedding, or a matrix; andwherein the vector, the embedding, or the matrix is indicative of a semantic structure of the text.93. The method of embodiment 92, wherein the ML model is configured to apply at least one of a clustering technique, a dimensionality reduction technique, a data visualization technique, or structural alignment to the vector, the embedding, or the matrix.94. The method of embodiment 93, wherein the dimensionality reduction technique includes at least one of multidimensional scaling (MDS), principal component analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t- SNE), or Uniform Manifold Approximation and Projection (UMAP).95. The method of embodiment 94, wherein applying the MDS includes determining at least one of a cosine similarity, a Euclidean distance, or a Manhattan distance.96. The method of any of embodiments 93-95, wherein the clustering technique includes at least one of affinity propagation clustering, spectral clustering, hierarchical clustering, k-means clustering, or mean shift clustering.97. The method of any of embodiments 93-96, wherein the clustering technique includes an objective functionconfigured to optimize a spatial distribution of one or more clusters.98. The method of any of embodiments 93-97, further including determining an Adjusted Rand Index (ARI) or a Normalized Mutual Information (NMI) to evaluate a result of the clustering technique.99. The method of any of embodiments 93-98, wherein the ML model is configured to generate, based on the applying, a three-dimensional structure indicative of the semantic structure of the text.100. The method of embodiment 99, wherein the three-dimensional structure is a continuous three-dimensional structure.101. The method of any of embodiments 93-100, wherein the structural alignment includes at least one of Procrustes analysis, root-mean-square deviation measurement, hotspot analysis, or kernel density estimation.102. The method of any of embodiments 93-101 , wherein the structural alignment includes comparing the semantic structure of the text to a personal attribute associated with text.103. The method of any of embodiments 85-102, wherein the example samples include at least one of a database, an experimental result, or a record.104. The method of any of embodiments 85-103, wherein the training data further includes labels indicating whether the example samples are associated with a personal attribute, and wherein training the ML model includes identifying, using supervised ML based on the labels, predictive attributes of the example prompts and / or example responses that are indicative of the labels.105. The method of embodiment 104, wherein the predictive attributes include factor loadings.106. The method of any of embodiments 92-105, wherein training the ML model includes determining a similarity of semantic structures of the text.107. The method of any of embodiments 92-106, further including identifying a pattern in semantic structures of the text.108. The method of any of embodiments 85-107, wherein the performance includes a structural integrity of the ML model.109. A system, including: a processor; and memory storing instructions that, when executed by the processor, cause the processor to perform operations including: inputting data including: (I) a prompt and / or (ii) a response of an individual to the prompt into a model configured to determine a relationship of the prompt and / or the response to a reference set; and generating an output indicative of the relationship.110. The system of embodiment 109, wherein the prompt is configured to determine a personal attribute of an individual.111. The system of embodiment 109 or 110, further including a transceiver configured to receive a communication signal indicative of the prompt and / or response.112. The system of any of embodiments 109-111 , further including a transceiver configured to transmit, to an external device, a communication signal indicative of the output.113. The system of any of embodiments 109-112, further including a display configured to visually present the output.114. The system of embodiment 113, wherein the output includes a visual model of the relationship.115. The system of embodiment 114, wherein the visual model includes a multi-dimensional model.116. The system of embodiment 114 or 115, wherein the visual model includes a three-dimensional model.117. The system of embodiment 116, wherein the three-dimensional model presents the relationship on a surface of a structure.118. The system of embodiment 116 or 117, wherein the three-dimensional model includes at least one of a surface model, a wire model, a wireframe model, a space-filling model, a cartoon model, a ribbon model, or a point cloud model.119. The system of any of embodiments 116-118, wherein the three-dimensional model presents the relationship using at least one of distance, color, or volume.120. The system of embodiment 119, wherein the volume is a mesh volume.121 . A non-transitory computer readable medium storing instructions for performing operations including: inputting data including: (i) a prompt and / or (ii) a response of an individual to the prompt into a model configured to determine a relationship of the prompt and / or the response to a reference set; and generating an output indicative of the relationship.

[0089] Experimental Example 1. Methods. Human data: Human response data from MID-60 questionnaires were obtained from Southern Cross University, Australia and the University of New England, Australia. The data was exempt due to being de-identified and did not require ethics approval.

[0090] Protein Data Bank (PDB): For test of iterative arrangement on molecular structure, the nanobody from 3K1 K was used as a general template to map all nanobody interface information. To facilitate easier processing of data, a new PDB was created re-numbering problematic residue numbers 52A, 82A, 82B, 82C and breakages in numbering. This resulted in a structure that no longer showed breakages in two parts of the structure (original residue numbering had 98 skipped to 101 , new numbering system made the numbering continuous).

[0091] PDB chain data: To gather chain data, PDB codes were requested from PDB under the search query, “Nanobody”, and interfaces of interest were identified from PDBePISA (E. Krissinel and K. Henrick (2007). J. Mol. Biol. 372, 774—797) for these PDB entries. To pick a particular interface from PDBePISA, interfaces which did not include one nanobody and one protein were discarded, and interfaces where a nanobody was observed binding to a ligand were also discarded. By using Python’s Bio.pairwise2's global alignments, removing leading and trailing dashes in the aligned strings, and calculating the percent identity between the resulting strings, nanobodies can be effectively identified using a cutoff value of 48% for the percent identity (FIGs. 4A-4J). Using Scikit’s logistic regression model trained on 505 interfaces to identify biologically significant interfaces, 94 of which were labeled for high complementarity determining region (CDR) binding, the interface with the highest probability was chosen to be downloaded from the remaining set of interfaces. The parameters involved in the model for each interface were total buried surface area (BSA), the number of occupied residues on the nanobody chain, the portion of BSA on the CDR loops, the sum of BSA on CDR1 and CDR2 plus FR2, and PDBePISA's complexation significance score (CSS) score. BSA on each residue was obtained through PDBePISA’s values, and the number of occupied residues were counted by counting the number of residues with nonzero BSA. To identify the regions on each nanobody, alignments of the nanobody sequence to a sample sequence consisting of only the framework regionswere performed, using penalties of 0 for extended gaps penalties and penalties of 2 and 3 for single gaps. A successful alignment showing the locations of the regions had all of the residues in each framework region in contiguous substrings in the aligned framework sequence string. PDB files of the nanobody and protein were downloaded by downloading the PDB file for the entry from PDB and extracting the chains by their chain IDs.

[0092] Root mean square deviation (RMSD) analysis: 107 nanobody and antigen structures were respectively employed for a total of 214 structures The coordinates of their alpha carbon (CA) atoms were utilized to construct pairwise distance matrices. Employing principal component analysis (PCA), t-distributed stochastic neighbor embedding (tSNE),u manifold approximation and projection (UMAP), and multidimensional scaling (MDS), each matrix was reduced to three dimensions. Aligning the reduced structures with true 3D counterparts using RMSD, average RMSD values were computed across all nanobodies and antigens. Results are presented in a bar plot, with each bar representing a reduction method, offering insights into the effectiveness of dimensionality reduction techniques for both nanobodies and antigens.

[0093] Distance matrix shuffling experiment: To explore how alterations in the spatial arrangement and interrelatedness of CA atoms affect clustering, two types of shuffling were performed on 107 nanobodies and antigens. Initially, the positions of the CA atoms were shuffled within each structure, and then affinity propagation clustering was applied to see how this shuffling impacts the clustering assignment compared to the original, unshuffled arrangement, using the Adjusted Rand Index (ARI) for comparison. Subsequently, these shuffled positions were reassigned back to their original sequence order— effectively altering their spatial interrelatedness without changing the sequence order— to examine the clustering outcome under these modified conditions. By using this approach, the effects of changing physical locations from alterations in sequence order on the structural clustering could be disentangled. The outcomes, reflected in ARI scores, were visualized through a bar plot, contrasting the clustering results of shuffling both order and interrelatedness against the original structure's clustering.

[0094] Rendering of psychological structure: Protein visualization software PyMOL (Schrodinger, LLC, Version 1.8 (2015)) was repurposed for visualization of psychological structures. Self-report questionnaires varying in item size (Mclean Screening Instrument for Mclean Screening Instrument for Borderline Personality Disorder (M. C. Zanarini, et al., J. Pers. Disord. 17, 568-573 (2003)) - 10 items, Narcissism Personality Inventory (R. Raskin, H. Terry, J. Pers. Soc. Psychol. 54, 890-902 (1988)) - 40 items, Multi-inventory of Dissociation (MID-60) (M.-A. Kate, et al., J. Trauma Dissociation 22, 265— 287 (2021))- 60 items and Personality Inventory for DSM-5 (PID-5) (R. F. Krueger, et al., Psychol. Med. 42, 1879-1890 (2012))- 220 items) were chosen to test the utility for visualization across a wide range of self-report questionnaires.

[0095] Data Processing and Embedding Generation: Questionnaire responses, including open-ended textual descriptions, were transformed into quantifiable data using the 'sentence-t5-xl' Sentence Transformer model. This model was selected based on the comparisons of language model performances to group psychological items. The model converts each response into a high-dimensional (768-dimensional) vector. This process enabled a computational analysis of the semantic relationships within the questionnaire data, irrespective of the questionnaire's focus area.

[0096] Dimensionality Reduction via Multidimensional Scaling (MDS): To facilitate a comprehensible analysis of the semantic space, MDS was applied with a three-dimensional target space, using cosine dissimilarity (1 - cosine similarity)as the metric. This reduction aimed to maintain the semantic distances among the responses, enabling an interpretable visualization and clustering in a reduced space. The choice of three dimensions was made to balance between preserving data complexity and visualization feasibility.

[0097] Clustering: Affinity Propagation was employed to categorize the reduced vectors into clusters, eliminating the need to pre-specify the number of clusters. This algorithm was chosen for its ability to determine the number of clusters based on the properties of the data itself.

[0098] Distance adjustments: In the initial tests with PyMOL of psychological structures, it was noted that the clarity of structural patterns was obscured by increasing number of items. To accommodate variations in questionnaire size, an adjustment factor designed to modulate target distances for intra- and inter-cluster dispersion was introduced.

[0099] For datasets with less than or equal to 60 responses, a specific scaling approach utilizing a non-linear adjustment through the use of an exponent was applied. Specifically, for sample sizes below or equal to the base size of 60, the adjustment factor scales target distances by employing an exponent of 0.5. This results in a calculated scale factor equal to the minimum scale plus the increase afforded by the ratio of the sample size to the base size, raised to the power of 0.5, multiplied by the difference between one and the minimum scale. This nuanced approach ensures a less steep adjustment curve for smaller datasets.

[0100] For sample sizes exceeding 60, the adjustment methodology was refined to include a linear scaling section between the base size of 60 and a target size of 200, with a reduced slope to mitigate excessive scaling. This linear scaling ensures a more gradual increase in the scaling factor within this range, providing a balanced distribution of points and preventing undue clumping. Beyond 200, the adjustment employs a more complex formula involving an exponent of 3. The scale factor for these larger datasets was derived from raising the ratio of the sample size to the target size to the power of 3. To further refine the scaling, a logarithmic adjustment was applied, based on the logarithm of one plus the ratio of the sample size minus the base size to the target size minus the base size. This was then scaled by 0.5 and added to one, ensuring that the scale factor remains between a minimum of 0.5 and a maximum of 2.5, thus offering adaptability across a broad range of dataset sizes.

[0101] In various implementations of the present disclosure, the non-linear adjustment was applied for datasets with less than or equal to less than 30, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 300, 500, or greater than 500 responses. In various implementations, the linear scaling section was applied for data sets between a base size of less than 30, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 300, 500, or greater than 500 responses and a target size of less than 50, 50, 100, 150, 200, 250, 300, 350, 400, 600, 800, or greater than 800 responses. In various implementations, it can be determined experimentally whether to apply the non-linear adjustment, the linear scaling section, or a different scaling approach. Various models and algorithms according to the present disclosure may utilize different scaling approaches.

[0102] Optimization of Median Scaling Factor: A scaling factor optimization aimed to align the median intra- and intercluster distances with specified targets, utilizing a numerical minimization process. This optimization ensured the clusters’ spatial distribution accurately reflected the underlying data structure, adjusting for dataset size. The objective function minimized the absolute deviation from target median distances, set at 4 for intra-cluster and 10 for inter-cluster distances,across the sample spectrum. The adjustment factor is applied to these values to obtain the actual distances to optimize for scaling.

[0103] Structural Representation and Visualization: The optimized three-dimensional embeddings were used to construct a PDB format representation, assigning each questionnaire response to a 'residue' within a virtual 'protein structure’. This approach facilitated the use of molecular visualization tools (e.g., PyMOL) for detailed analysis of data clustering. Clusters were color-coded using a predefined palette for the top eight clusters based on membership size, with remaining clusters depicted in grey, enhancing visual differentiation. This was chosen to demonstrate the formation of surface patches by semantic clusters in these structures without overwhelming the reader with excessive coloring.

[0104] Local versus Global structural integrity analysis: Average Jaccard similarity was used to measure the overlap in closest neighbor sets in the high dimensional versus MDS 3D space. Spearman’s rank correlation was used to measure stability of overall relational structure. Items variable was presented on logarithmic scale. Linear regression was applied to visualize and extrapolate from the trend.

[0105] Initial tests: Initial exploratory tests involved the Levenson's Self-Report Psychopathy Test (M. R. Levenson, et al., J. Pers. Soc. Psychol. 68, 151-158 (1995)) and Narcissism instruments. These explorations led to formulation of the hypothesis that semantic similarities produced by some embedding-producing models can relate to psychological factor membership. The instruments were excluded from the hypothesis testing phase.

[0106] Creation of items database for screening embedding models: To determine whether semantic similarity relates to psychological factor membership, the literature was objectively sampled for a comprehensive range of psychological questionnaires. A strategy analogous to the snowball sampling method was employed. First, a search with the search prompt "principal components analysis factor inventory” was performed on December 21 , 2023, on Pubmed and filtered for Free Full-Texts. 452 search results were returned. Many results did not correspond to psychological questionnaires. Each publication was manually reviewed, and publications that utilized PCA to derive factor loadings for each item of the questionnaire and published the results were filtered for. Many fitting results correspond to evaluation of non-English versions of a questionnaire. When these cases were encountered, another paper detailing English version of the same questionnaire was searched for, or if the questionnaire was a revision of another questionnaire available in English - the original questionnaire being revised. If an article showing PCA factor loadings for the English items were found, the results for the English version were input into the database. For some items, the instrument was set up such that a prompting statement was followed by items to complete the sentence (ex. I self-harm to....(item)). The prompting statement was fused to the item to form the strings for analysis. Factor membership may sometimes be ambiguous due to above criterion factor loadings for an item across factor groups. In these cases, factor membership was assigned according to which factor does the item hold the highest factor loading. In total, 30 questionnaires from the database were sampled.

[0107] Evaluation of the relationship between semantic similarity and factor membership: To evaluate whether semantic similarity is related to psychological factor membership, it was determined whether factor membership can be predicted by the inter-relatedness between psychological items. Cosine distance was used as a measure of semantic similarity. The number of factors per questionnaire was determined, and that value was used to perform k-means clustering to producean equivalent group of clusters defined by the pairwise similarity matrix. It was then measured how well factor membership overlapped with cluster membership, using the ARI and normalized mutual information (NMI) measures.

[0108] Impact of sentence manipulation on MDS positioning: Sentences were chosen based on an earlier study on the stability of sentence embeddings in UMAP space. The algorithm identified sentences with highly stable embedding positions regardless of new UMAP initializations.

[0109] Language models: To compile a comprehensive list of language models which output embeddings for evaluation, models from Hugging Face and Open-AI were gathered. In February 2024, an online search was performed at Huggingface (huggingface.co / models) for the top models filtered according to Most Liked, Most Downloads and Trending filters. Filters for sentence-transformers library and English language were further specified. The first page of 30 models from each filtered results were collected, and a unique list of language models was derived from the combination of these lists. The sentence-t5 (J. Ni, et al. Proceedings of the 2022 EMNLP, (2022), pp. 9844-9855), gtr-t5 {Id.), jina (M. Gunther, et al., Proceedings of the 3rd Workshop for NLP-OSS 2023, (2023), pp. 8-18) and AnglE (X. Li, J. Li, arXiv preprint, (2023)) series of models were added due to their reports of improved performances. Lastly, models scoring high on Semantic Task Search (sf_model_e5, ember-v1 , bge-large-en-v1.5, gte-large-quant) were added from the Multilingual Text Embedding Benchmark (MTEB) leadership based on early indications the Semantic Task Search was correlated with increased performances for grouping psychological factor performances.[O1 1O] The language model sentence-t5-xl was used to provide text embeddings for calculating example distance matrix shown in FIG. 4E.

[0111] Language model parameters: To evaluate correlations between ARI performance on the psychological item grouping task with model parameters, a structured assessment using the MTEB was conducted.

[0112] The evaluation process was designed to ensure comprehensive coverage while acknowledging computational limitations. As such, certain high-capacity models— specifically sentence-transformers / sentence-t5-xxl, sentence- transformers / gtr-t5-xxl, text-embedding-ada-002, SeanLee97 / angle_llama_7b_nli_v2, and SeanLee97 / angle_llama_13b_nli— were not directly evaluated due to memory constraints. Instead, performance data for these models on MTEB tasks were sourced from publicly available repositories and documentation, including the MTEB leaderboard and model-specific pages on the HuggingFace platform.

[0113] For models within the scope of the available computational resources, a custom Python script was utilized to systematically load each model from the HuggingFace repository, execute a series of benchmark tasks as defined by MTEB, and record the outcomes. The results of these benchmarks were recorded in a JSON file to facilitate a standardized and reproducible approach to data collection. Subsequent to the evaluation phase, another Python script was employed to process the collected JSON files. This script was tasked with parsing the results, aggregating key performance metrics, and compiling the data into concise and informative CSV files for further analysis.

[0114] Evaluating semantic clustering amongst psychopathy questionnaire items: To enable distinct visualization of language model performances, a custom colormap was generated, utilizing Matplotlib for Python. This colormap includes five segments, each corresponding to specific performance intervals identified by average baseline cosine similarity values.The segments are defined by distinct colors: Olive Green (RGB: 128, 179, 76), Dark Blue (RGB: 51, 51, 153), Orange (RGB: 230, 153, 0), Beige (RGB: 230, 204, 128), and White (RGB: 255, 255, 255), ensuring visual distinction across the performance spectrum. The colormap employs a resolution of 100 bins to achieve smooth transitions. This methodological approach facilitates nuanced comparison of language model performances by encapsulating multi-layer distinctions within the specified color range.

[0115] DSM-V items database: To create a comprehensive items database of psychiatric questionnaires, all items were extracted from the American Psychiatric Association - DSM-V Text Revision online assessment resources (www.psychiatry.org / psychiatrists / practice / dsm / educational-resources / assessment-measures). Facet groupings were organized as described in the assessments.

[0116] Prediction of Factor Loadings: It was an aim to show the predictability of psychological facets by embeddings produced by large language models (LLMs). To demonstrate the feasibility of such, a multilayer perceptron (MLP) that can predict factor loadings of psychological questions in "A Principal-Components Analysis of the Narcissistic Personality Inventory and Further Evidence of Its Construct Validity” was designed. To perform prediction of the factor loadings of a given psychological survey, two components were used, a sentence transformer model with t5-base to produce embeddings and an MLP to go from embeddings to a factor loading. A separate MLP for predicting each factor loading column in the paper was trained. K-fold cross-validation was performed with 10 folds, and the mean-squared error (MSE) was taken from the development set and averaged across each iteration in K-fold cross-validation. A baseline was computed by taking the MSE across 4000 random shufflings of the factor loadings as predictions. It was also shown that the sentence transformer and MLP are necessary for prediction by computing loss if one or the other is untrained.

[0117] The sentence transformer t5-base model is the encoder-only version described in “Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models.”

[0118] Specifically, the MLP was defined as:$\begin{verbatim} self.linearl = nn.Linear(input_size, 20) self.linear2 = nn.Linear(20, 20) self.linear3 = nn.Linear(20, 20) self.linear4 = nn.Linear(20, output_size) self.tanh = nn.Tanh() \end{verbatim}$

[0119] In this case, input_size = embedding size, which was 768 for the disclosed sentence transformer model, and output_size = 1 as one factor loading from the embedding in each MLP was predicted. It was found found that this simple architecture was adequate for achieving significant improvements in prediction than that of the computed baselines.

[0120] It was also an aim to utilize the disclosed model to predict class membership of a survey question in a paper. Enhancements in the model architecture were implemented by increasing its size and fine-tuning the hyper-parameters as follows:$\begin{verbatim} seif.linearl = nn.Linear(input_size, 1000) self.linear2 = nn.Linear(1000, 500) self.linear3 = nn.Linear(500, 100) self.linear4 = nn.Linear(100, 50) self.linear5 = nn.Linear(50, output_size) self.tanh = nn.Tanh()\end{verbatim}$.

[0121] The foundational sentence encoder remains the unchanged ‘sentence-t5 base'. The selection of these hyperparameters was informed by trials conducted on the previously mentioned study on narcissism. Subsequently, the refined model was evaluated across tables in ten separate studies, each including eleven unique tables related to survey questions and various psychological dimensions.

[0122] For each evaluation, a 4-fold cross-validation was conducted after shuffling the dataset. In this process, one fold was split in half into validation and test subsets. The validation subset aided in choosing the optimal MLP during the training phase based on validation MSE loss which was later used for testing. The test subset facilitated the calculation of performance metrics, including accuracy, recall, precision, and F1 score. To ensure the integrity of the evaluation, each MLP was initialized randomly before being applied to a specific table, thereby preventing the transfer of information between tables. These metrics were computed for three models with the same architecture, one with an untrained encoder, one with a trained encoder, and one with a trained encoder with an MLP trained on a shuffled mismatched training set. Each 4-fold cross-validation was run 3 times on each table and the final graphs represent overall overages across all iterations and all tables.

[0123] Signal variation analysis: To explore if biological data can inform the mapping of physical topology in self-report constructs, the nanobody interface dataset was analyzed to determine whether the variation of interface signals across individuals, mapped onto the highly conserved nanobody structure, may be exploited to solve for a physically relevant topology of the nanobody's amino acid relative positioning in 3D space. Interface values across individuals at each residue position in the conserved nanobody structure were treated as a vector. Pairwise Euclidean distance was calculated for each residue vector, generating a signal similarity matrix. Next, MDS was applied on the signal similarity matrix to represent the signal inter-relatedness in 3D space. A PDB file was constructed with the MDS 3D coordinates, and the resulting structure was visualized on PyMOL. Labeling for specific groups of amino acids was facilitated by Affinity Propagation Clustering of the signal similarity matrix. To evaluate whether the MDS output contains physically relevant information, Procrustes analysis was performed to evaluate level of alignment. As a control, the identity of the amino acids in the conserved nanobody structure was shuffled, and the Procrustes value was compared to that derived from the original structure. This procedure was repeated on self-report questionnaires response data to construct a putative physical topology representation of subjective experiences, with individual response variation to each self-report items being treated as a vector for pairwise distance calculations.

[0124] Statistical Analysis. Standard statistical analyses were performed on Prism (Versions 7, 9, 10; GraphPad Software, Inc.; RRID:SCR_002798). To determine appropriate tests for comparisons, datasets were assessed for normality using Anderson-Darling, D’Agostino & Pearson, Shapiro-Wilk and / or Kolmogorov-Smirnov tests whenever applicable. Datasets were also visualized for normality using QQ plots and assessed for equal variance by examining the Residual plot (Residuals versus Predicted Y). Parametric or non-parametric tests were chosen based on the combination of these analyses. Data were transformed logarithmically (with or without addition of a constant prior to transformation) whenever it was appropriate to promote normality and equal variance. Unless specified, sphericity was not assumed, and Geisser- Greenhouse correction was applied in all ANOVA tests. The appropriate post hoc multiple comparisons tests were applied to compare between the means of specific conditions wherever applicable. Significance was set at alpha = 0.05. Bonferroni p adjustment was used to account for multiple comparisons in this case. When Graphpad Prism does not output exact p- value, Excel (version 16.78.3; Microsoft) was used with the ANOVA-specific FDIST(F, DFN, DFD) where DFN is the Numerator Degrees of Freedom and DFD is the Denominator Degrees of Freedom. For detailed description of statistical procedures please refer to Supplementary Information.

[0125] Detailed Statistical Procedures. Nanobody Order and Inter-relatedness Shuffling analysis: Non-parametric Friedman test was performed to test the hypothesis that ARI measure varied across groups. Post hoc Dunn's multiple comparisons test was performed to test the hypothesis that ARI score differs between sequence order shuffling versus inter-relatedness shuffling.

[0126] Nanobody RMSD analysis - Non-parametric Friedman test was performed to test the hypothesis that RMSD measures varied across groups. Post hoc Dunn's multiple comparisons test was performed to test the hypothesis that MDS-associated RMSD values were significantly different from those of PCA, tSNE, or UMAP.

[0127] Results. Representing Semantic Structure by solving Semantic inter-relatedness from Language Models. It was investigated whether a system with known physical connections to data patterns can illuminate the structural properties of self-consciousness. Intriguing parallels between the patterns in self-report questionnaires and those in proteins were identified. First, just as self-report questionnaires consist of individual text items, proteins are made up of amino acids (FIG. 4B). Second, the way humans respond to questionnaire items mirrors the interaction patterns seen among individual proteins of a specific class, notably camelid antibodies (nanobodies) (FIGs. 4C, 4D). Despite a highly conserved structure, nanobodies exhibit diverse interaction patterns due to variations that affect how they interact with other proteins. This diversity is similar to how individuals’ responses to self-report items vary (FIG. 40).

[0128] It was noticed that distance matrices can be constructed to represent the inter-relatedness between components in both systems (FIG. 4E). For proteins, amino acids that are share similar inter-relationships with other amino acids tend to form clusters, which correspond to surfaces on the protein (FIG. 4F). Importantly, the clustering is based on the patten of inter-relatedness rather than the sequence order of amino acids, suggesting that one can also understand the structure behind questionnaire items' inter-relatedness without considering their order. To validate this analogy, the order of components or the inter-relatedness were shuffled within individual proteins (FIG. 4G). Shuffling the order did not affect cluster formation based on distance matrices, as shown by stable Adjusted Rand Index (ARI) scores (L. Hubert, P. Arabie,J. Classif. 2, 193-218 (1985)) (FIG. 4H). However, when inter-relatedness is shuffled, the overlap between original and shuffled cluster groupings was significantly disrupted, demonstrating the importance of inter-relatedness in maintaining the system’s integrity (FIG. 4H). These findings indicate that, as demonstrated with protein structures, the order of questionnaire items is not important for the analogy, since inter-relatedness relates to surfaces of a self-structure, and this boundary is what is being modeled between the two systems.

[0129] It was tested whether it is possible to visually represent measures of inter-relatedness by solving these distance matrices (FIGs. 4I, 4J). It was found that multidimensional scaling (MDS) (W. S. Torgerson, Psychometrika 17, 401-419 (1952)), a technique designed to preserve distances between datapoints in a specified dimensional space, could indeed be applied to solve the distance matrices across proteins of a dataset (FIGs. 4I, 4J). In this dataset, nanobodies (n=107) have a fairly homogenous group of structures, and antigens (n=107) have widely divergent structures. Remarkably, MDS was able to reconstitute structures across the dataset. MDS significantly outperformed UMAP, PGA, and T-Sne in faithfully representing the structure of the protein, with little deviation from original structure in terms of RMSD, despite the efforts to optimize RMSD by varying scaling factors to find the lowest RMSD values for UMAP, PCA, T-Sne.

[0130] To test whether it is possible to extend this approach to self-report questionnaires, a language model, sentence- t5-xl, was used to generate semantic distance matrices of various psychological and psychiatric questionnaires and MDS was applied on the matrices to visualize the information (FIG. 5A). Dynamic scaling of MDS distances was applied depending on questionnaire size to pull items together into a visually intuitive structure (Methods). In all cases, clusters of semantic inter-relatedness formed patches of surfaces on 3D structures analogous to those seen on physical protein structures, albeit with occasional overlapping patches suggestive of expected distortion of high dimensional data (FIG. 5A).

[0131] To further test the possibility of visualizing multiple questionnaires in a single structure, the four questionnaires from FIG. 5A and 30 questionnaires sampled from the web were combined (FIG. 5C). A new structure was observed that resulted in fewer clusters than the expected number based on the assumption that each questionnaires had nonoverlapping items. This indicated that some items between questionnaires have overlapping semantic-interrelatedness. Indeed, this was visually confirmed by the presence of items that cluster into the same group but were derived from different questionnaires. Additionally, inspection of items' actual texts indicated that intermingled items from different questionnaires do indeed share similar semantics. While the distortion of high-dimensional data into a 3D space presents a challenge, visual inspections indicate the clusters of semantic interrelatedness are robustly seen across structures over broad ranges (10 to 637 items, from 1 to 30 questionnaires). The visualization approach was tested on various psychiatric questionnaires as well, resulting in confirmation of similar patterns across psychological and psychiatric questionnaires, and across response groups ranging from adults to adolescent (children ages 11-17) and parent / guardians of children age 6 and up (FIGs. 6A-6C). To further test the possibility of visualizing a massive dataset, the ability to visualize a 9,999 items database pooled from text from various questionnaires and text databases was demonstrated, with similar distribution of clustered items into patches along the surface of the structure (FIG. 7). Finally, by repurposing protein visualization tools, psychological structures from all types of perspectives familiar to structural biologists were visualized, ranging from mesh to surface to sticks, etc. (FIG. 8). These results comprehensively establish the utility of integrating MDS and structuralbiological visualization methods to analyze psychological structures.

[0132] As with all dimension reduction techniques like MDS, distortion of high-dimensional information in 3D space is unavoidable. Investigating this issue, it was found that local structures as measured by Average Jaccard Similarity, is mostly preserved at items size below 100 but rapidly deteriorates as items size increases further. On the other hand, global structures as measured by Spearman rank distance correlations is highly robust to item size increase (FIG. 9). Extrapolation from linear regression of the Spearman values suggests that the global structure could remain moderately retained with items datasets of as large as 1 million items. This is a notable finding because the overall global structure amongst areas of the psychological structure is retained despite large variation in items size. These data indicate that MDS is suitable for analyzing both local and global structures at low items size and is appropriate for global level analysis at large items size. Overall, molecular analogies can inform new ways to model semantic structure of self-report questionnaires of varying sizes. Further, the ability to generate and visualize semantic structures of questionnaires and identify localized 3D spaces of inter-mingled self-report items from different questionnaires makes possible novel questionnaire generation strategies that do away with traditional questionnaire boundaries between items. Lastly, these results establish the feasibility of leveraging MDS-generated structures to conduct global sampling of massive databases of many items for whole psyche analysis.

[0133] Psychological Relevance of Semantic Representation by Language Models. Given the ability of MDS to produce a semantic structure based on semantic inter-relatedness, it was determined whether the semantic representations by language models are psychologically relevant. Language models were leveraged to measure similarity in meaning between self-report items analogously to physical distance between amino acids of protein structures. Language models transform text into numerical vectors, or embeddings, allowing semantic similarities to be computed between them. The popular cosine similarity measure was used to calculate semantic distance between texts. The process of grouping self-report statements in psychological research often begin by analyzing human response patterns with principal component analyses and related methods. The factor loading values from these analyses inform the categorization of items into specific psychological groupings. Self-report questionnaires with published factor loading matrices and measured pairwise semantic similarity (cosine method) across all self-report items. Pairwise semantic similarity patterns were more likely to cluster between items of the same empirically determined psychological factor in specific embedding types - specifically those from more recent and larger embedding model (FIG. 10). Further, there is variation in visible clustering patterns amongst the models tested, indicating that some models may be more effective at capturing psychological relevance than others.

[0134] To thoroughly investigate whether language models encode psychologically relevant groupings, 50 models were selected, with most of them identified through filtering for the most popular, liked, downloaded, or recent language models from HuggingFace (T. Wolf, et al., Proceedings of the 2020 EMNLP Conference: System Demonstrations (2020); N. Reimers, I. Gurevych, arXiv e-prints, (2019)). The ada-002 embedding model from OpenAI was further included, along with several models reported to have enhanced performances (J. Ni, et al., Findings of the Assoc, for Comp. Ling. 2022, (2022), pp. 1864-1874; J. Ni, et al., Proceedings of the 2022 EMNLP, (2022), pp. 9844-9855; X. Li, J. Li, arXiv preprint, (2023)).To evaluate model performances on an objectively and comprehensively sampled list of self-report questionnaires, a search was performed for a new set of peer-reviewed articles from Pubmed mentioning principal component analysis to derive factor loadings and used a strategy analogous to the snowball method to extensively collect factor grouping information (Methods). The ARI (Hubert and Arable, J. Classif. 2, 193-218 (1985)) was used as a method of overlap between the groupings created by semantic interrelatedness from language models versus the groupings reported by psychological experiments (FIG. 11 A). Language model showed a spectrum of performances on this task, with the most recent and larger language models sitting at the top of the performance list (FIG. 11 B). As controls, the performances of pretrained language models were confirmed to be significantly above untrained language models (FIG. 11 B), and that shuffling of semantic embeddings was found to obliterate significant ARI performances for each model (FIG. 11 A). Further confirming the generality of this finding, language models were found to group psychiatric items from psychiatric DSM-V self-report measures from adults, adolescents, and parents of child into the defined facets significantly above chance, indicating broad generalizability (FIGs. 12A-12C). Thus, semantic inter-relatedness generated by language models can be leveraged to categorize psychological and psychiatric items into similar groupings as those obtained from self-report experiments, with room for improvement.

[0135] Development of language models. Interestingly, visualization of untrained and trained models revealed clear differences in semantic structures (FIG. 13). The structural evolution from an undifferentiated ball of tightly overlapping items to a larger, differentiated structure bearing patches of semantically inter-related items is reminiscent of embryonic development, and suggest that analogical modeling may be leveraged to track and analyze neural network development over training.

[0136] Semantic self-structural properties are predictive of psychological data. Next, it was determined whether the reconstituted MDS space provides psychologically relevant information. The narcissism personality inventory questionnaire was analyzed, for which extensive psychological data has been collected (Raskin and Terry, J. Pers. Soc. Psychol. 54, 890-902 (1988)). Factor loadings are values obtained from factor analysis of psychological response patterns. The value of these loadings informs the grouping of psychological items, which then become known as psychological concepts called facets or factors. It was noticed that item placement in MDS space appeared to correlate their factor loading values to a particular narcissism facet. For example, 3 of 6 Entitlement items that show weakly positive factor loading values, or association, to the Authority facet, were positioned on one side of a cluster of Authority items in MDS space (FIG. 14). This contrasts with the positioning of the other 3 of 6 Entitlement items that show negative or zero factor loading values, or association, to the Authority facet (FIG. 14). These observations indicate that MDS space, built to represent semantic interrelatedness between psychological items, can explain empirically derived psychological response patterns that had no tangible explanations.

[0137] For further evaluation, the factor loadings of narcissism personality inventory were examined across its 7 facets by overlaying the values for each item and facet onto the MDS-generated narcissism semantic structure. A qualitative trend was noted, whereby items that are closer together in MDS space share similar factor loading values, with a rough gradient of high to low values as items move away from the highest value items (FIG. 15). To rigorously test whether this observationis quantitatively true, it was assessed whether geometric properties of the narcissism structure in MDS-space are predictive of factor loadings. The MDS coordinates of each narcissism facet's centroid were identified, and it was determined whether angular closeness in MDS space, as measured by cosine similarity between the centroid and item of interest, is related to the item’s empirically measured factor loading score for each facet. Indeed, significant (0.3-07) correlation values were found across the facets, and MDS consistently outperformed PCA / T-Sne / UMAP in these metrics (FIG. 16). These results indicate that MDS-generated semantic structures are psychologically relevant, encoding psychological information that used to require experimentation.

[0138] MDS-generated semantic structures derive from the embeddings of language models. It was determined whether it is possible to train neural networks to improve their ability to predict factor loadings from text alone. The ability to improve factor loading predictions by feeding neural networks psychological experimental data would demonstrate proof-of-concept of the ability to develop psychologically relevant tools for diverse applications. A pipeline of text analysis was developed whereby a large language model, sentence-t5-xl, fed embeddings into an MLP trained to produce factor loading predictions. Using 10-fold cross validation, the narcissism inventory's factor loading scores for all items across all 7 facets were trained and the model showed significant, non-random improvement in loss function compared to models paired with untrained MLP, or when MLP was trained with untrained language models, or when the input data was shuffled (FIG. 17A). It was further discovered that the model developed above-chance level prediction of facet grouping across 11 questionnaires after training, but not with shuffled input data or with untrained MLP (FIG. 17B). These findings indicated the advantages and practicality of using MDS-derived 3D semantic structures to predict psychological responses, without experimentation.

[0139] To further ask whether it is possible to improve language model embedding space itself, language models were fine-tuned using the questionnaire items dataset sampled from the web. Language models were fine-tuned to recognize the facets associated with each item, and demonstrated improved accuracy in validation and test sets (FIG. 18). These findings indicated the ability to improve language models for psychological applications, as measured by the metric in FIGs. 11 and 12.

[0140] Psychological Self structure of a Dissociative Experiences self-report questionnaire that may underly selfconsciousness boundary experiences. To apply the described framework and methodologies to the study of consciousness and mental health, the MID-60, a questionnaire developed to study dissociative experiences, was assessed. The MID-60 is well known to be increased in individuals exposed to early life trauma and amongst psychiatric patients. A closer inspection of the MID-60 semantic structure revealed that clusters grouped by semantic inter-relatedness tended to form patches of surfaces as observed in the protein analogy. Importantly, the meaning of the items that form clusters tend to be very similar (FIG. 19). To further evaluate the utility of this structure for understanding dissociative disorders, response patterns from 314 individuals were overlaid, and affinity propagation clustering was applied to parse individuals by their response patterns. 11 clusters of response patterns that were shared across individuals were detected. Intriguingly, hotspots of response across items were observed, indicating correlated responses related to semantic closeness. Further, experiences were related to the response categories, with traumatized individuals showing strong agreement response to most items on the MID-60, while individuals with little to no traumatic self-reports tended to score low on the MID-60 items(FIGs. 20A, 20B). These observations indicated that the psychological self-structure could be useful for rapid and intuitive self-report assessments in the clinic and beyond.

[0141] Relationship between subjective experiences and physical organization of experiences. The hard problem of consciousness concerns how subjective experience is translated at the physical level. While brain activities can be associated with subjective experiences, there is not a model of how subjective experiences perceived by the self may be organized in a physically relevant topology. Imagine the physical self-structures are physically exposed to inputs from different modalities of consciousness, and increased intensity of those inputs translate to perceptible experience. Over time, the self develops language to describe different types of subjective experiences, but the relationship between the semantics of those experiences and the physical organization of the inputs that generate those experiences are not clear. As an analogy, when a self-molecule interacts with other molecule at an interface, that interaction may lead to a specific combination of experiences. If the self-molecule is able to respond to questionnaires about its experience, the agreement with combinations of particular experiential statements may reveal the physical relationship between statements, due to the variation in response patterns across individuals. It is possible to infer the physical topology of subjective selfexperiences in a physically relevant manner if one can account for these patterns.

[0142] The semantic structures described earlier are indicated as partially correct solutions to the physical organization of the self-structure in regard to subjective experiences. It is indicated that, like in genetic linkage mapping, it is the estimation of the strength of co-occurrence between psychological item response strength that provides a way to solve the physically relevant topology of subjective experiences on a physical substrate, presumably the brain.

[0143] The interface matrix documenting the interaction between specific portions of nanobodies binding to their antigens can be thought to be analogous to psychological self-report response matrix, with each item being analogous to portions of nanobodies and the strength of agreement to a psychological item analogous to strength of a particular type of physical interaction with other consciousness. Variation in interface signals amongst nanobodies may be exploited to piece together a physically relevant topology of the conserved nanobody structure. If there exists such a method, it can be applied to solve the physical topology of self-consciousness by exploiting the variation of self-report response patterns across individuals.

[0144] This problem was studied by first constructing signal similarity matrixes comparing the Euclidean distance between interface signals at particularly conserved amino acid positions of the nanobody across individuals. The matrix was then clustered to identify amino acid positions that share similar interface signal patterns. Next, MDS was applied to construct a 3D representation of the signal similarities. When cluster label identity was overlaid onto both the actual nanobody conserved structure and compared the relative 3D positioning of the signal cluster of amino acids relative to other clusters, it was found that the physical topology created by MDS is aligned with that of the actual positioning of the signal cluster amino acids on the conserved nanobody structure (FIGs. 21A-21C). Procrustes analysis comparing the MDS output of the positioning of the signal clusters with the actual positioning on the nanobody structure showed a non-random significant alignment between the MDA output and actual physical placement of the signal cluster amino acid (FIG. 21 C). This establishes the feasibility of exploiting variation in interface signal patterns across individuals to produce a physically relevant topology reflecting actual structure.

[0145] This principle was extended to psychological response patterns of the MID-60 questionnaire and obtained a psychological structure that represents the putative physical topology of self-structure boundary (FIG. 22). Based on actual MID-60 self-report data, the self-topology seems to divide into several notable spatial groupings. These results provide a map that, according to the analysis of nanobody-antigen interface signals, encodes physically relevant information about the positioning of physical substrates underlying agreement with these items. If this is true, the psychological structure formed by the MID-60 response data should align better to semantic structures derived from language models that are more psychologically relevant. Indeed, the topology of this structure is better aligned to those from language models that perform better on the ARI item grouping task (FIG. 23), indicating that language models reflect various hypothesis about the self-structure, with signal-MDS structure deriving from actual self-report response data holding the most information regarding the physical topology of the self.

[0146] Discussion. In this Example, molecular complexes were adopted as analogical models for self versus other consciousness in order to derive physically relevant connection between the subjective report of experiences by selfconsciousness and physical reality. This work bridges disparate fields of structural biology, psychology / psychiatry and artificial intelligence to model self-structure and provides a novel framework and methodologies encouraging cross- interdisciplinary collaborations across these and related fields.

[0147] The commonality between protein and psychological situation is the boundary of a self in contact with non-self. Describing the topological boundary in a physically relevant way opens the door to finding neurobiological architecture that best explains the topology. The disclosed physical topology mapping approach is reminiscent of the genetic linkage maps conceptualized to represent the physical closeness of phenotypes on a line, which ultimately informed the targeting of actual chromosomal locations and genes in genetics (T. H. Morgan, et al., (Holt, Oxford, England, 1915)). The self-structure boundary is anticipated to require a higher order structure in order to describe its topology, but the physical topology can be mapped with the response pattern of individuals and may eventually be fully represented by language models. The concept of the self remains an elusive target in psychological and neuroscientific research. The disclosed approach, while utilizing static models to analyze self-report questionnaire data, is predicated on the assumption that these models can offer a window into the complex dynamics underlying self-consciousness. Analogously, proteins are inherently dynamic, yet static models have provided widely-validated insights, giving confidence to the framework developed here. The disclosed work proposes a method through which its contours may be explored and understood. The advantage of MDS in visualizing self-report questionnaire data lies in its ability to preserve overall topology in a visually intuitive 3D space. While distortion of high dimensional data is a real issue and demands innovative solutions (Ding and Regev, Nat. Comm. 12, 2554 (2021)), the application of MDS for the goals here is appropriate. Nevertheless, this study provides a concrete framework and testable models regarding the relative spatial organization of physical entities that gives rise to specific combinations of subjective experiences.

[0148] Genetic linkage maps are useful tools for studying the physical basis of inheritance of traits even when limited traits are available. The self-topological maps created by specific questionnaires can similarly yield useful insight into particular aspects of the psyche. At the same time, a systematic expansion of analysis to saturate the range of humansubjective experiences in single individuals ill be important to allow for efforts to fully deduce the physical topology of selfboundary from variation in individual response patterns. The system described allows for such mapping to be performed while also providing useful tools for psyche analysis even in limited items size. Questionnaire-specific self-topological maps allow for tracking of physical structure changes that may help follow the appearance of hidden experiences linked physically. Finally, as genetic patterns measured by phenotyping is not entirely predictable due to the influence of epigenetic mechanisms, the uncanny similarity in questionnaire response patterns influenced by factors such as inattention, dishonesty and social desirability is noted (D. L. Paulhus, (Academic Press, 1991) Vol. 1., pp. 17-59), indicating a novel perspective to dissect these influences inspired by the epigenetics field.

[0149] Applications with demonstrated methods. The demonstration here makes possible multiple applications across broad areas of society where self-report questionnaires are applied. First, the ability to delineate semantic inter-relatedness clusters and generate semantic structure for visualization allows intuitive understanding of relationship between items. This facilitates various applications such as education, psychological / psychiatric / social sciences research. Second, the discovery of psychological relevance of language models and their ability to preserve semantic and psychological relevance in global and local ways, depending on the dataset size, allow for development of strategies to systematically sample questionnaire items from the semantic MDS space (FIGs. 24A, 24B). This is not necessarily confined to 3D but may be operated at higher or lower dimensions. The advantage of operating at higher dimension of MDS space, in some examples, may be to improve psychological relevance of language models at larger items datasets in terms of local neighborhoods, whereas operating at lower dimensions, in some examples, allows for visualization (in 2D and 3D) and could save computational costs while allowing for a more global understanding of semantic inter-relatedness between clusters of items. Third, the ability to delineate the items space allows for one to develop generative model creating new items for rapid and efficient generation of questionnaires for broad applications. This would save time and resources in the gathering of large- scale data about the whole psyche. Fourth, the demonstration that structural alignment between MDS-generated semantic and psychological structures makes possible a new approach to training neural networks or language models. Traditional language models are typically trained on a single text or comparing texts. However, training in this way may, in some cases, continuously disrupt proper learning from earlier trials because of a lack of consideration of local and global embedding space changes that are unseen from each feedback update to the model through methods like gradient descent. However, the relevance of a language model for a particular application may be more properly trained by consistently imposing constraints for optimization that consider the larger portions of the embedding space, or in its entirety (FIGs. 25A, 25B). Dimension reduction methods may allow for this by providing structural coordinates that facilitate alignment analysis, such as using Procrustes analysis, across a larger portion of embedding space. Lastly, although language models are specified here, the ideas here may be extended to other types of neural networks with similar architectures as language models.

[0150] The advent of advanced language models and the discovery of their psychological relevance offer unprecedented opportunities for whole psyche analysis, with applications in psychological and psychiatry. While language models can significantly enhance understanding of the psyche's semantic architecture, the ultimate goal should be to integrate these insights with response pattern analyses, as the fine-tuning of models on real-world response patterns will provide a morerealistic representation of the psyche and result in more rapid and efficient ways to probe individual psyche, resulting in more effective and personalized diagnostic and therapeutic interventions, as well as development of novel educational tools. By leveraging the full potential of language models in conjunction with response pattern analysis in this Example, these results indicate that the entirety of the psyche can be mapped, understood, and, when necessary, therapeutically navigated.

[0151] Experimental Example 2. This Example describes the development of SELF-MAP, which utilizes relative selfreport ratings that can encode a similarly tangible "Self' interface (FIGs. 26, 4C, and 4D). SELF-MAP reflects proteinprotein interactions, where correlated interaction patterns between protein parts reflect participation in a physical interface. Each survey item was treated as a “position” on a hypothesized self structure, so that inter-item response patterns across individuals model potential structural proximities in the brain.

[0152] Methods. Signal variation analysis. To explore if biological data can inform the mapping of physical topology in self-report constructs, the nanobody interface dataset was analyzed to determine whether the variation of interface signals across individuals, mapped onto the highly conserved nanobody structure, could be used to solve for a physically relevant topology of the nanobody’s amino acid relative positioning in 3D space. The interface data was first processed for normalization. Each residue position (columns in FIG. 40) in the dataset underwent a logarithmic transformation to mitigate skewness from exponential distributions (CDR loops which summed up buried surface area values across variable number of residues). To accommodate zero values in the data, a small constant (e.g., 0.01) was added to each parameter value before applying the transformation in this Example. Following the logarithmic conversion, the data was normalized to a 0- 1 scale by subtracting the minimum value of the transformed data and dividing by the range (e.g., maximum minus minimum) of that data. The normalized values were then rescaled to a 0-4 range to align with the analytical framework. To finalize, the scaled values were adjusted into a 1-5 range, suitable for further analysis, by adding 1 to each and rounding to the nearest integer.

[0153] Interface values across individuals at each residue position in the conserved nanobody structure were treated as a vector. Pairwise cosine distance was calculated for each residue vector, generating a signal similarity matrix. Next, MDS was applied on the signal similarity matrix to represent the signal inter-relatedness in 3D space. A PDB file was constructed with the MDS 3D coordinates, and the resulting structure was visualized on PyMOL. Labeling for specific groups of amino acids was facilitated by Affinity Propagation Clustering of the signal similarity matrix. To evaluate whether the MDS output contains physically relevant information, Procrustes analysis was performed to evaluate level of alignment. As a control, the identities of the amino acids in the conserved nanobody structure were shuffled, and the Procrustes value was compared to that derived from the original structure. This procedure was repeated on self-report questionnaires response data to construct a putative physical topology representation of subjective experiences, with individual response variation to each self-report items being treated as a vector for pairwise distance calculations.

[0154] Procrustes alignment between Language models and signal-MDS structures. To determine whether language models that are more psychologically relevant tend to align better in their 3D MDS structure with 3D structures derived from signal-MDS analysis of self-report data on the MID-60 questionnaire, Procrustes analysis was performed. The signal-MDS 3D structure was derived from using cosine distance as the basis for distance matrix construction between item signal variation across individuals. To derive the 3D MDS semantic structure, MID60 text items were fed into each language model to obtain embeddings, and cosine distance matrix was constructed based cosine similarity comparisons between text items. MDS was applied to solve the 3D coordinates of each structure. The 3D coordinates of the semantic structures were then aligned to the 3D MDS coordinates of the signal-variation analysis of MID-60 response data. As a control, each semantic MDS coordinate was paired with a shuffled 3D coordinates made from the signal-variation-MDS data. Linear regression analysis was then applied to determine whether there was a non-zero linear relationship between median ARI values measuring overlap between semantic groupings by language models and psychological groupings.

[0155] Item SELF-MAP Topological Reproducibility Analysis. To further assess the reproducibility of topological relationships between items, an analysis was developed to assess closeness of inter-item relationships across subsamples. The larger MID-60 dataset (mental health population = 2010) was used to subsample 1000 smaller subsamples (n = 314, equivalent to the college population sample size). Each dataset was processed to create a SELF-MAP structure. The inter-relationship between each MID-60 item in each subsampled SELF-MAP structure was calculated with Euclidean distance or cosine similarity distance, and the averages across item-specific distance vector (against all other ID60 items) were calculated. Upon visualization of a histogram plotting distance value bins against item frequency within each bin, it was found that items regarding conversion symptoms were consistently least reproducible, while items regarding derealization / depersonalization were consistently most reproducible. To identify items consistently positioned on the less stable half of the distribution, all items landing to the right of the histogram peak in the Euclidean distance case were assessed. It was determined which item also mapped onto the lower end of the cosine similarity case. Based on this logic, 13 items were identified. To cut off 20% of MID60 items, 12 of 13 items were cut off, leaving the borderline item “When you are angry, doing....” item in the retained list.

[0156] Brain alignment. Next, it was assessed whether the topology of the putative self-structure was useful for deducing its spatial organization of self-interface in the brain. A previous study reported fMRI data associating different subsets of MID items with brain regions, surveyed on an independent psychiatric population. MID and MID60 share items that fall within the Partially Dissociation Intrusions (PDI) and Depersonalization / Derealization (DPDR) sub-groups. The MID + fMRI dataset was leveraged to assess a possible strategy to deduce the self-structure organization in the brain. Using t- statistics from the MID fMRI study, the relative bias of unique variance contribution from MID PDI or MID DPDR towards the MID PDI or MID DPDR full variance model was determined for specific brain regions. That subtracted score was multiplied by each dimension of the brain coordinates to assess whether there was bias of MID DPI or MID DPDR contribution towards different brain networks / dimensions. A statistically significant bias of the MID DPI contribution was detected in the frontal half of the brain relative to MID DPDR, specifically when differential CEN and Salience Network associations were considered with these items. This was consistent with a frontal bias of PDI relative to DPDR if only frontal regions were considered, and a lack of significant bias when only posterior regions were considered. In this Example, there was a lack of bias on the medial-lateral nor dorsal-ventral dimensions. This data indicated that MID PDI and MID DPDR show specific differential contribution to brain activities along the frontal-posterior axis. This finding was leveraged to alignthe questionnaire data-derived topology to the brain. Roy's greatest root method was applied to test whether the maximal eigenvector that parses out MID60 PDI and MID60 DPDR items in 3D space was statistically significant, indicating that this was true for both MID60 questionnaire datasets. Then Linear Discrimination Analysis (LDA) was used to derive the eigenvector explaining most of variance between MID60 DPI and MID60 DPDR groups. All MID60 items from the original 3D topology space were projected onto the new 3D space alignment items in one dimension based on greatest separation. In this Example, this projection appeared to align with the brain frontal-posterior axis as explained above. To evaluate predictive power of this alignment, the distribution of items was then analyzed for trends of non-MID60 DPI or non-MID60 DPDR items to uncover consistent trends that could help further distinguish the question of whether the MID60 items were localized to frontal only, frontal-posterior spread, or posterior only. Spearman correlation showed that the F-P order of selftopology deduced

[0157] Results. SELF-MAP was validated using ground truth systems. “Self-structures” were first simulated in a 2D environment: SELF-MAP accurately recovered the underlying self-structure from analyzing snapshots of self-to-other interactions (data not shown). The same principle was then applied to real protein complexes (n = 389 nanobodies-to- target structures) whose ground-truth 3D configurations are known (Dingus, J. G., et al. eLife 11 , e68253 (2022); Tang, J. C. et al. Elife 5, (2016)). SELF-MAP successfully recovered actual spatial relationships between shared protein residues, thereby reconstructing a shared Self structure (Methods; FIGs. 4G, 21 A, 21 B, 27A). FIGs. 21 A and 21 B indicate the SELFMAP corresponding to FIG. 40 (right). FIG. 21A illustrates the re-construction (e.g. , the shared “self nanobody structure), while FIG. 21 B illustrates the ground truth (e.g., the x-ray solved nanobody).

[0158] To test whether SELF-MAP can uncover physical information regarding Self from self-report data, a dissociative symptom survey (MID-60) administered to both college and clinic populations was utilized (FIGs. 4C and 27B-27D) (Kate, M.-A., et al. Journal of Trauma & Dissociation 22, 265-287 (2021)). FIG. 27B indicates the SELF-MAP corresponding to FIG. 4C (left). FIG. 27B illustrates the re-construction (e.g., the shared “self” structure (MID-60)). Dissociative experiences are diverse and reflect disconnection of “Self to various “Other” facets of consciousness. Thus, MID-60 was found to cover a range of subjective experiences potentially reflecting distinct physical boundaries of the Self.

[0159] Confabulation concerns were minimized by repeatedly reconstructing item-level maps using different subsamples, filtering items with high positional instability, and confirming that random permutations of item-response data failed to recreate the coherent structure (Methods; FIGs. 27C and 27D). Focusing on relative self-report magnitudes reduced biases linked to overall response style, while convergent topologies in both populations reinforced that SELF-MAP captures biologically meaningful rather than idiosyncratic signal.

[0160] The physical relevance of the SELF-MAP model was tested by aligning it to brain structure and assessing fit to known patterns. Published brain coordinates and metrics associated with MID60 items in a fMRI study were utilized (Lebois, L. A. M. et al. Neuropsychopharmacology 47, 2261-2270 (2022)). The analyses revealed a Frontal-Posterior (F-P) distinction between Partially Dissociated Intrusions (PDI) and Depersonalization / Derealization (DPDR) subgroups of dissociative experiences in their associations in the brain (FIGs. 28A and 28B). The SELF-MAP structure was aligned to the F-P brain axis, using LDA (Methods; FIG. 28C). In a test of physical relevance, only 40% of items from each dissociativesubgroup, PDI and DPDR, were subsampled. It was verified that the derived frontal-posterior (F-P) axis still separated these subgroups. (FIG. 28D). Despite demographic contrasts, a robust topological pattern emerged: items involving “parts” of self clustered in what appears to be a frontal zone, whereas memory and DPDR, localized posteriorly. Spearman rank correlations between the F-P positions across MID-60 items not used for alignment showed a strong agreement in positioning between the college and clinical datasets’ item positions. Conversely, a purely semantic “SELF-MAP” derived from large language model (LLM) embeddings of the MID-60 items failed to replicate the same F-P structure or correlate meaningfully with the item positions in either dataset (not shown). Combining lesion, stimulation, neuroimaging and structural fiber analyses, the predicted locations of MID-60 items aligned with cingulate cortices, indicating a specific anatomical axis for these dissociative experiences (not shown).

[0161] In this Example, it was shown that LLMs can be fine-tuned with psychological data, using a pairwise comparison of item cosine similarity, using the signal variation vector specific to the pair of items, to improve the “physical relevance” of the LDA-mediated, fMRI-aligned SELF-MAP model (FIG. 29), while “psychological” relevance of the model, defined by the ability of LLM semantic-embedding-mediated clustering to produce overlapping groups with psychologist-defined factor groupings, remained unaltered (FIG. 29). This Example indicates that psychological human self-report data, at least from dissociative experiences, encode information that is more tuned to the physical structure of the brain.

[0162] Analysis to Test Physical Relevance of Self-Report Response Data. To investigate the physical relevance of selfreport data in characterizing subjective experiences an LDA approach was implemented to align self-report derived structure to the frontal-posterior axis of the brain. Self-report data was divided into three groups based on the co-occurrence of MID60 "PDI" and "DPDR" criteria. For each group, 40% or 60% of the data were subsampled multiple times (e.g., up to 200 permutations), and the remaining 60% or 40% were treated as held-out test data, respectively.

[0163] The subsampled training sets were used to align data within a shared discriminative subspace using LDA, trained on 3D coordinates extracted from SELF-MAP structures (e.g., spatial coordinates "X", "Y", "Z"). To control for directional ambiguity in the latent subspace, the sign of the LDA discriminant axis was determined based on the relative alignment of MID60 PDI and MID60 DPDR items within the training sets, ensuring interpretability across permutations. Critically, this sign flip was applied consistently across all items, including those used in training and held-out test data.

[0164] The second phase of the analysis evaluated the generalizability of this alignment to held-out data. By isolating rows flagged as held-out in each permutation, the discriminant scores (LDA1) were extracted for MID60 PDI and MID60 DPDR items not used for alignment. Statistical comparisons of average LDA1 scores between these groups were performed to determine whether the self-report data encoded physically relevant differences in subjective experience. Additionally, a null distribution was generated by randomly shuffling group labels relative to LDA1 scores for the held-out data, allowing for a direct assessment of whether the observed discriminative patterns exceeded those expected by chance.

[0165] This framework was found to be rigorous in separating training and test data while leveraging multiple permutations to quantify the robustness of group distinctions. By comparing LDA1 score differences in the original and shuffled data, a systematic test to determine whether self-report-derived structure contains physically meaningful information that can be aligned to the brain, beyond random noise or statistical artifacts, is presented in this Example. Thisapproach rigorously evaluates the alignment of self-reported subjective experiences with brain structures.

[0166] Discussion. According to this Example, item-level patterns of subjective variation, similar to correlated interface patterns in proteins, can reveal a physically relevant map of consciousness. By situating MID-60 item clusters along an axis reminiscent of known frontal-posterior functional divisions, SELF-MAP uncovered a structural underpinning for subjective experiences in this Example.

[0167] Closing Paragraphs. As will be understood by one of ordinary skill in the art, each implementation disclosed herein can comprise, consist essentially of or consist of its particular stated element, step, or component. Thus, the terms “include” or “including” should be interpreted to recite: “comprise, consist of, or consist essentially of." The transition term “comprise” or “comprises" means has, but is not limited to, and allows for the inclusion of unspecified elements, steps, ingredients, or components, even in major amounts. The transitional phrase “consisting of1excludes any element, step, ingredient or component not specified. The transition phrase “consisting essentially of’ limits the scope of the implementation to the specified elements, steps, ingredients or components and to those that do not materially affect the implementation. As used herein, the term “based on” is equivalent to “based at least partly on,” unless otherwise specified.

[0168] Unless otherwise indicated, all numbers expressing quantities, properties, conditions, and so forth used in the specification and claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. When further clarity is required, the term “about” has the meaning reasonably ascribed to it by a person skilled in the art when used in conjunction with a stated numerical value or range, i.e. denoting somewhat more or somewhat less than the stated value or range, to within a range of ±20% of the stated value; ±19% of the stated value; ±18% of the stated value; ±17% of the stated value; ±16% of the stated value; ±15% of the stated value; ±14% of the stated value; ±13% of the stated value; ±12% of the stated value; ±11 % of the stated value; ±10% of the stated value; ±9% of the stated value; ±8% of the stated value; ±7% of the stated value; ±6% of the stated value; ±5% of the stated value; ±4% of the stated value; ±3% of the stated value; ±2% of the stated value; or ±1% of the stated value.

[0169] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. Any numerical value, however, inherently contains certain errors necessarily resulting from the standard deviation found in their respective testing measurements.

[0170] The terms “a," “an,” “the” and similar referents used in the context of describing implementations (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described hereincan be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein is intended merely to better illuminate implementations of the disclosure and does not pose a limitation on the scope of the disclosure. No language in the specification should be construed as indicating any non-claimed element essential to the practice of implementations of the disclosure.

[0171] Groupings of alternative elements or implementations disclosed herein are not to be construed as limitations. Each group member may be referred to and claimed individually or in any combination with other members of the group or other elements found herein. It is anticipated that one or more members of a group may be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.

[0172] Certain implementations are described herein, including the best mode known to the inventors for carrying out implementations of the disclosure. Of course, variations on these described implementations will become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventors expect skilled artisans to employ such variations as appropriate, and the inventors intend for implementations to be practiced otherwise than specifically described herein. Accordingly, the scope of this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by implementations of the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.

Claims

CLAIMSWhat is claimed is:

1. A method comprising: identifying data comprising a set of responses of an individual to a set of prompts; inputting the data into a model configured to determine a similarity between a semantic structure of the set of responses to a semantic structure of a reference set; generating a visual model indicative of the similarity by determining a distribution of multiple data points based on an interrelationship of the semantic structure of the set of responses and the semantic structure of the reference set; and determining, based on the visual model, a personal attribute of the individual.

2. A method comprising: inputting data comprising:(i) a prompt and / or(ii) a response of an individual to the prompt into a model configured to determine a relationship of the prompt and / or the response to a reference set; and generating an output indicative of the relationship.

3. The method of claim 2, wherein the prompt and the response are indicative of a personal attribute of the individual.

4. The method of claim 2, wherein the prompt comprises at least one of text, a visual prompt, or a sound.

5. The method of claim 4, wherein the text comprises a question and / or a description of a scenario.

6. The method of claim 5, wherein the question is an open-ended question.

7. The method of claim 4, wherein the visual prompt comprises at least one of a thematic apperception test, aRorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, or a captcha.

8. The method of claim 2, wherein the response comprises at least one of a text response, a selection of a value from a scale, a selection of a choice from a multiple-choice question, a verbal response, a physiological response, or a behavior.

9. The method of claim 2, wherein the relationship comprises a similarity between: a semantic structure of the reference set; and a semantic structure of the prompt and / or the response.

10. The method of claim 2, wherein the data comprises the response, and wherein the method further comprises determining, based on the relationship, a personal attribute of the individual.11 . The method of claim 2, wherein the individual is an adult, an adolescent, or a child.

12. The method of claim 11 , wherein the adult is a parent or guardian of a child.

13. The method of claim 10, wherein the personal attribute comprises at least one of an emotional state of theindividual, a preference of the individual, a demographic characteristic of the individual, or a diagnosis of a psychological condition or a lack thereof of the individual.

14. The method of claim 13, wherein the emotional state comprises at least one of depression, anxiety, stress, coping, mood, attention, phobias, demoralization, or rumination.

15. The method of claim 10, further comprising: outputting, based on the personal attribute, a personalized message or a recommendation.

16. The method of claim 15, wherein the personalized message comprises a marketing message.

17. The method of claim 16, wherein outputting the personalized message comprises outputting the marketing message via an email, social media, or a mobile application.

18. The method of claim 15, wherein the recommendation comprises a treatment regimen.

19. The method of claim 18, wherein the treatment regimen indicates at least one of a drug or a behavior modification program.

20. The method of claim 19, wherein the drug comprises at least one of a prescribed medication, an over-the- counter drug, or a dietary supplement.21 . The method of claim 19, wherein the behavior modification program is configured to change at least one of a substance intake, physical exercise, sleep schedule, performance of self-defeating behaviors, social interactions, or compulsive behaviors of the individual.

22. The method of claim 21 , wherein the substance intake comprises at least one of food intake, alcohol intake, or drug intake.

23. The method of claim 15, wherein outputting the personalized message comprises displaying the recommendation on a graphical user interface.

24. The method of claim 15, wherein outputting the personalized message comprises outputting the recommendation to a third party.

25. The method of claim 24, wherein the third party is at least one of a doctor, a therapist, a coach, a teacher, or a researcher.

26. The method of claim 3, wherein the data comprises the response, and the personal attribute comprises a comprehensive psychological assessment of the individual.

27. The method of claim 2, wherein the output comprises a visual model of the relationship.

28. The method of claim 27, wherein generating the output indicative of the relationship comprises generating the visual model by applying a structural biology visualization tool to the relationship.

29. The method of claim 28, wherein the structural biology visualization tool comprises PyMOL, ChimeraX, or Visual Molecular Dynamics.

30. The method of claim 27, wherein the visual model comprises a multi-dimensional model.31 . The method of claim 27, wherein the visual model comprises a three-dimensional model.

32. The method of claim 31 , wherein the data comprises multiple fields, each of the multiple fields comprises aprompt and / or response, and wherein the three-dimensional model is configured to represent an interrelationship between semantic structures of the multiple fields.

33. The method of claim 31 , wherein the data is inputted in an order, and wherein the three-dimensional model is independent of the order.

34. The method of claim 31 , wherein the three-dimensional model presents the relationship on a surface of a structure.

35. The method of claim 34, wherein the three-dimensional model comprises: at least one first data point corresponding to a semantic structure of the prompt and / or response.

36. The method of claim 35, wherein the three-dimensional model further comprises at least one second data point corresponding to a semantic structure of the reference set.

37. The method of claim 35, wherein the surface comprises the at least one first data point.

38. The method of claim 35, wherein the data comprises multiple fields, each of the multiple fields comprises a prompt and / or response, wherein the at least one first data point comprises multiple first data points respectively corresponding to semantic structures of the multiple fields, and wherein a distance between the multiple first data points correspond to similarities between the semantic structures.

39. The method of claim 38, wherein the distance comprises a cosine similarity, a Euclidean distance, or a Manhattan distance.

40. The method of claim 38, wherein generating the three-dimensional model comprises determining a distribution of the multiple first data points based on an interrelationship of the semantic structures of the multiple fields.41 . The method of claim 40, wherein determining the distribution is analogous to determining a three-dimensional structure of a protein based on an interrelationship of amino acids.

42. The method of claim 38, wherein the data comprises more than one prompt from a psychological questionnaire, and wherein the model comprises a three-dimensional representation of the psychological questionnaire.

43. The method of claim 34, further comprising comparing the three-dimensional model to biological data.

44. The method of claim 43, wherein the biological data comprises a brain structure.

45. The method of claim 31 , wherein the three-dimensional model comprises at least one of a surface model, a wire model, a wireframe model, a space-filling model, a cartoon model, a ribbon model, or a point cloud model.

46. The method of claim 45, wherein the three-dimensional model presents the relationship using at least one of distance, color, or volume.

47. The method of claim 46, wherein the volume is a mesh volume.

48. The method of claim 27, wherein the data comprises a response, and wherein the method further comprises evaluating the visual model to identify a personal attribute of the individual.

49. The method of claim 2, wherein the data comprises a first prompt or a first response from a first individual, and the relationship is a first relationship, the method further comprising: inputting data comprising a second prompt or a second response from a second individual into the model; and generating a second output indicative of a second relationship of the second response and the reference set.

50. The method of claim 49, further comprising: determining that the first response is different than the second response; and determining that the first relationship is the same as the second relationship.51 . The method of claim 50, wherein the second output is a second three-dimensional model, and wherein a first three-dimensional model corresponding to the first response is the same as a second three-dimensional model corresponding to the second response.

52. The method of claim 49, further comprising: determining that the first response is different than the second response; and determining that the first relationship is different than the second relationship.

53. The method of claim 2, wherein the model comprises a machine learning (ML) model.

54. The method of claim 53, wherein the ML model comprises a language model, an encoder model, an encoderdecoder model, an embedding model, a transformer model, a neural network, a topic model, natural language processing, or deep learning.

55. The method of claim 2, wherein the model comprises a language model and a neural network.

56. The method of claim 55, wherein the language model comprises a large language model (LLM) and the neural network comprises a Multilayer Perceptron (MLP).

57. The method of claim 53, further comprising: training the ML model by optimizing parameters of the ML model based on training data, the training data comprising example prompts and / or example responses identified from example samples.

58. The method of claim 57, wherein the example samples comprise: the example prompts from at least one of a Narcissism Personality Inventory, a McLean Screening Instrument, a Multidimensional Inventory of Dissociation 60-item (MID-60), a Personality Inventory for Diagnostic and Statistical Manual of Mental Disorders (DSM)-5 (PID-5), a Level 1 Cross-Cutting Symptom Measure, a Level 2 Cross-Cutting Symptom Measure, a Disorder-Specific Severity Measure, a Levenson’s Self-Report Psychopathy Test, a thematic apperception test, a Rorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, or a captcha and / or the example responses to a prompt from at least one of the Narcissism Personality Inventory, McLean Screening Instrument, MID-60, PID-5, Level 1 Cross-Cutting Symptom Measure, Level 2 Cross-Cutting Symptom Measure, Disorder- Specific Severity Measure, Levenson’s Self-Report Psychopathy Test, a thematic apperception test, a Rorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, or a captcha.

59. The method of claim 57, wherein the example samples comprise:the example prompts from at least one of a Beck Depression Inventory (BDI), a Generalized Anxiety Disorder (GAD-7), a Depression Anxiety Stress Scale (DASS), a Brief-COPE, a Positive and Negative Affect Schedule (PANAS), a State Trait Anxiety Inventory (STAI), a modified version of a Russel Mood Circumplex, a modified version of a NIH Sleep Diary, a 36 Item Short Form Health Survey (SF-36), a 5D Altered state of Consciousness Scale (5d-ASC), a daily sleep diary questionnaire, a Hamilton Rating Scale for Depression, a Hamilton Anxiety Rating Scale, a mood survey, a selfreporting survey, or a sleep diary, and / or the example responses to the example prompts from at least one of the BDI, GAD-7, DASS, Brief-COPE, PANAS, STAI, modified version of the Russel Mood Circumplex, modified version of the NIH Sleep Diary, SF-36, 5d-ASC, daily sleep diary questionnaire, Hamilton Rating Scale for Depression, Hamilton Anxiety Rating Scale, mood survey, selfreporting survey, or sleep diary.

60. The method of claim 57, wherein the example samples comprise at least one of a database, an experimental result, or a record.61 . The method of claim 57, wherein the example samples comprise multiple example fields, each of the multiple example fields comprising an example prompt and / or example response, wherein the training data further comprises labels indicating whether the example samples are associated with a personal attribute, and wherein training the ML model comprises identifying, using supervised ML based on the labels, predictive features of the example fields that are indicative of the personal attribute.

62. The method of claim 61 , wherein the predictive features comprise factor loadings.

63. The method of claim 62, wherein the ML model comprises a neural network, the method further comprising: based on training the ML model: determining that an accuracy of the factor loadings of the trained ML model is higher than an accuracy of factor loadings of an untrained ML model; or determining that the accuracy of the factor loadings of the trained ML model is higher than a threshold.

64. The method of claim 57, wherein training the ML model comprises determining a similarity of semantic structures of the example prompts and / or the example responses.

65. The method of claim 64, further comprising identifying a pattern in the semantic structures of the example prompts and / or the example responses.

66. The method of claim 2, wherein the data comprises more than one prompt and / or response, and wherein the model is configured to represent a semantic interrelatedness between each of the prompt and / or the response67. The method of claim 4, wherein the model is configured to generate at least one of a vector, an embedding, or a matrix, and wherein the vector, the embedding, or the matrix is indicative of a semantic structure of the text.

68. The method of claim 67, wherein the model is configured to apply at least one of a clustering technique, adimensionality reduction technique, a data visualization technique, or a structural alignment to the vector, the embedding, or the matrix.

69. The method of claim 68, wherein the dimensionality reduction technique comprises at least one of multidimensional scaling (MDS), principal component analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t- SNE), or Uniform Manifold Approximation and Projection (UMAP).

70. The method of claim 69, wherein applying the MDS comprises determining at least one of a cosine similarity, a Euclidean distance, or a Manhattan distance.

71. The method of claim 69, wherein the model is configured to generate multiple encodings, each of the multiple encodings comprising a vector, an embedding, or a matrix, the method further comprising based on applying the dimensionality reduction technique, determining a distribution of data points in a multi-dimensional space, wherein the data points correspond to each of the multiple encodings.

72. The method of claim 71 , wherein the multi-dimensional space is a three-dimensional space, the method further comprising determining a continuous three-dimensional structure based on the distribution of data points.

73. The method of claim 72, further comprising: performing a local analysis of a region of the continuous three-dimensional structure; and determining a relationship between two or more data points.

74. The method of claim 72, further comprising: performing a global analysis of the continuous three-dimensional structure; and determining a personal attribute associated with the distribution of data points.

75. The method of claim 68, wherein the clustering technique comprises at least one of affinity propagation clustering, spectral clustering, hierarchical clustering, k-means clustering, or mean shift clustering.

76. The method of claim 68, wherein the clustering technique comprises an objective function configured to optimize a spatial distribution of one or more clusters.

77. The method of claim 68, wherein the model is further configured to determine an Adjusted Rand Index (ARI) or a Normalized Mutual Information (NMI) to evaluate a result of at least one of the clustering technique, the dimensionality reduction technique, the data visualization technique, or the structural alignment.

78. The method of claim 2, further comprising: based on the output, generating a questionnaire.

79. The method of claim 78, wherein the questionnaire comprises a question indicative of a personal attribute of an individual.

80. The method of claim 79, wherein the personal attribute comprises a comprehensive psychological assessment.81 . The method of claim 2, wherein the prompt or the response comprise 1 to 1 million prompts and / or 1 to 1 million responses.

82. The method of claim 2, wherein the prompt or the response comprise less than 61 prompts and / or responses, andwherein the model is configured to perform non-linear scaling.

83. The method of claim 2, wherein the prompt or the response comprise more than 60 prompts and / or responses, and wherein the model is configured to perform linear scaling.

84. The method of claim 2, further comprising generating the model.

85. The method of claim 84, wherein generating the model comprises selecting at least one of a machine learning algorithm, a feature extraction algorithm, an artificial intelligence algorithm, a Bayesian algorithm, a statistical analysis algorithm, a topic modeling algorithm, or a clustering algorithm.

86. A method, comprising: generating a machine learning (ML) model configured to determine a relationship of a prompt and / or a response to a reference set; training the ML model by optimizing parameters of the ML model based on training data, the training data comprising example prompts and / or example responses identified from example samples; and generating a report indicating a performance of the ML model.

87. The method of claim 86, wherein generating the ML model comprises selecting at least one of a machine learning algorithm, a feature extraction algorithm, an artificial intelligence algorithm, a Bayesian algorithm, a statistical analysis algorithm, a topic modeling algorithm, or a clustering algorithm.

88. The method of claim 86, wherein the example prompts comprise at least one of text, a visual prompt, or a sound.

89. The method of claim 88, wherein the text comprises a question and / or a description of a scenario.

90. The method of claim 89, wherein the question is an open-ended question.91 . The method of claim 88, wherein the visual prompts comprise at least one of a thematic apperception test, aRorschach test, a Draw-A-Person Test, a House-Tree-Person test, a Bender-Gestalt Test, a Kinetic Family Drawing, or a captcha.

92. The method of claim 86, wherein the example responses comprise at least one of a text response, a selection of a value from a scale, a selection of a choice from a multiple-choice question, a verbal response, a physiological response, or a behavior.

93. The method of claim 86, wherein the example prompts and / or example responses comprise text, wherein the ML model is configured to generate at least one of a vector, an embedding, or a matrix; and wherein the vector, the embedding, or the matrix is indicative of a semantic structure of the text.

94. The method of claim 93, wherein the ML model is configured to apply at least one of a clustering technique, a dimensionality reduction technique, a data visualization technique, or structural alignment to the vector, the embedding, or the matrix.

95. The method of claim 94, wherein the dimensionality reduction technique comprises at least one ofmultidimensional scaling (MDS), principal component analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t- SNE), or Uniform Manifold Approximation and Projection (UMAP).

96. The method of claim 95, wherein applying the MDS comprises determining at least one of a cosine similarity, a Euclidean distance, or a Manhattan distance.

97. The method of claim 94, wherein the clustering technique comprises at least one of affinity propagation clustering, spectral clustering, hierarchical clustering, k-means clustering, or mean shift clustering98. The method of claim 94, wherein the clustering technique comprises an objective function configured to optimize a spatial distribution of one or more clusters.

99. The method of claim 94, further comprising determining an Adjusted Rand Index (ARI) or a Normalized Mutual Information (NMI) to evaluate a result of the clustering technique.

100. The method of claim 94, wherein the ML model is configured to generate, based on the applying, a three- dimensional structure indicative of the semantic structure of the text.

101. The method of claim 100, wherein the three-dimensional structure is a continuous three-dimensional structure.

102. The method of claim 94, wherein the structural alignment comprises at least one of Procrustes analysis, rootmean-square deviation measurement, hotspot analysis, or kernel density estimation.

103. The method of claim 94, wherein the structural alignment comprises comparing the semantic structure of the text to a personal attribute associated with text.

104. The method of claim 86, wherein the example samples comprise at least one of a database, an experimental result, or a record.

105. The method of claim 86, wherein the training data further comprises labels indicating whether the example samples are associated with a personal attribute, and wherein training the ML model comprises identifying, using supervised ML based on the labels, predictive attributes of the example prompts and / or example responses that are indicative of the labels.

106. The method of claim 105, wherein the predictive attributes comprise factor loadings.

107. The method of claim 93, wherein training the ML model comprises determining a similarity of semantic structures of the text.

108. The method of claim 93, further comprising identifying a pattern in semantic structures of the text.

109. The method of claim 86, wherein the performance comprises a structural integrity of the ML model.

110. A system, comprising: a processor; and memory storing instructions that, when executed by the processor, cause the processor to perform operations comprising: inputting data comprising:(i) a prompt and / or(ii) a response of an individual to the promptinto a model configured to determine a relationship of the prompt and / or the response to a reference set; and generating an output indicative of the relationship.

111. The system of claim 110, wherein the prompt is configured to determine a personal attribute of an individual.

112. The system of claim 110, further comprising a transceiver configured to receive a communication signal indicative of the prompt and / or response.

113. The system of claim 110, further comprising a transceiver configured to transmit, to an external device, a communication signal indicative of the output.

114. The system of claim 110, further comprising a display configured to visually present the output.

115. The system of claim 114, wherein the output comprises a visual model of the relationship.

116. The system of claim 115, wherein the visual model comprises a multi-dimensional model.

117. The system of claim 115, wherein the visual model comprises a three-dimensional model.

118. The system of claim 117, wherein the three-dimensional model presents the relationship on a surface of a structure.

119. The system of claim 117, wherein the three-dimensional model comprises at least one of a surface model, a wire model, a wireframe model, a space-filling model, a cartoon model, a ribbon model, or a point cloud model.

120. The system of claim 117, wherein the three-dimensional model presents the relationship using at least one of distance, color, or volume.

121. The system of claim 120, wherein the volume is a mesh volume.

122. A non-transitory computer readable medium storing instructions for performing operations comprising: inputting data comprising:(i) a prompt and / or(ii) a response of an individual to the prompt into a model configured to determine a relationship of the prompt and / or the response to a reference set; and generating an output indicative of the relationship.

Citation Information

Patent Citations

  • Mind-body correlation data evaluation apparatus and method of evaluating mind-body correlation data

    US20070167690A1

  • System and method for providing a user with recommendations indicating a fitness level of one of more topical skin products with a personal care device

    US20190080385A1

  • Adapting a sequence model for use in predicting future device interactions with a computing system

    US20210004682A1

  • Methods and compositions for screening and treating developmental disorders

    US20210017601A1

  • Method for determining a representation of a subjective state of an individual with vectorial semantic approach

    US20210056267A1