A computer automated system and method for randomized dataset creation and evaluation of language models in understanding basic human emotions
A computer automated system aggregates, standardizes, and evaluates LLMs' emotional understanding using statistical methods, addressing the need for rigorous testing and ensuring empathetic responses for improved human-machine interactions.
Patent Information
- Application Number
- PCT/IB2024/062729
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-16
- Publication Date
- 2025-07-03
AI Technical Summary
There is a need for comprehensive and rigorous testing and evaluation methodologies to assess the ability of large language models (LLMs) to comprehend and interpret basic human emotions, particularly for applications in mental health support and assisting people with disabilities, to ensure empathetic and appropriate responses.
A computer automated system aggregates user data from diverse sources, standardizes and normalizes it, estimates emotional states, triangulates the data, and estimates errors using statistical methods like Chi-Squared Goodness of Fit and p-values, while evaluating performance metrics such as precision, recall, and F1 score through confusion matrices and AUC-ROC/AUC-PR.
Enables accurate evaluation of LLMs' emotional understanding capabilities, ensuring they can handle complex human situations and provide empathetic responses, thereby improving human-machine interactions.
Smart Images

Figure IB2024062729_03072025_PF_FP_ABST
Abstract
Description
A COMPUTER AUTOMATED SYSTEM AND METHOD FOR RANDOMIZED DATASET CREATION AND EVALUATION OF LANGUAGE MODELS IN UNDERSTANDING BASIC HUMAN EMOTIONSBACKGROUNDField
[0001] Embodiments disclosed are in the field of regenerative artificial intelligence based computer automated systems and methods for randomized dataset creation and evaluation of language models in understanding basic human emotions.Related Art
[0002] The recent emergence and widespread adoption of large language models (LLMs) have highlighted the need for a comprehensive and rigorous testing and evaluation methodology. There remains a need to assess the ability of LLMs to comprehend and interpret basic human emotions in an analytical and statistical manner. Such an approach is crucial for assessing the inherent capacity of language models and for providing a solid foundation for their deployment in various applications.
[0003] Yet additionally, there remains a need for evaluation systems and methods to facilitate the development of LLMs that can address deeper human issues, including mental health support and assisting people with disabilities. By testing the ability of these models to understand human emotions, researchers and developers can ensure that they are equipped to handle complex situations and provide appropriate responses.
[0004] Yet additionally, there remains a need for computer automated systems and methods that enable and help foster more empathetic responses from LLMs. By training these models to understand human emotions in a more nuanced and sophisticated way, we can create a more human-like interaction with these systems. This is particularly important in situationswhere users may be vulnerable, such as those seeking mental health support or assistance with disabilities.
[0005] The development and implementation of a detailed test and evaluation system and methodology for LLMs is a critical step in ensuring the continued advancement and application of these models. By leveraging the power of these systems and methods to understand and respond to human emotions, we can address a wide range of important issues and improve the overall quality of human-machine interactions.
[0006] Embodiments disclosed address precisely the aforementioned challenges.SUMMARY
[0007] A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions. One general aspect includes a computer automated system for randomized dataset creation and evaluation of language models in understanding basic human emotions. According to an embodiment, the computer automated system is caused to aggregate user data from a plurality of data sources. The neural architecture is configured to standardize the aggregated user data from the plurality of data sources. Embodiments disclosed then normalize the standardized, aggregated user data from the plurality of data sources, and are configured to estimate an emotional state or states of the aggregated user data based on predefined criteria. A preferred embodiment is further configured to triangulate each of the estimated emotional state or states of the aggregated user data, to estimate user data sample distribution and to estimate an error or errors in pre-labelled user data. Other embodiments ofthis aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0008] One general aspect includes a computer implemented method for randomized dataset creation and evaluation of language models in understanding basic human emotions. An embodiment includes a computer implemented method comprising aggregating user data from a plurality of data sources. The neural architecture is configured for standardizing the aggregated user data from the plurality of data sources. An embodiment of the computer implemented method includes normalizing the standardized, aggregated user data from the plurality of data sources, and is further configured for estimating an emotional state or states of the aggregated user data based on pre -defined criteria. A preferred embodiment of the computer implemented method includes triangulating each of the estimated emotional state or states of the aggregated user data, estimating data sample distribution, and estimating an error or errors in pre-labelled data. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To further clarify the advantages and features of the present disclosure, a more particular description of the disclosure will be rendered by reference to specific embodiments thereof, which is illustrated in the appended drawings. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered limiting of its scope. The disclosure will be described and explained with additional specificity and detail with the accompanying drawings.
[0010] The subject matter that is regarded as the invention is particularly pointed out and distinctly claimed in the claims at the conclusion of the specification. The foregoing and other aspects, features, and advantages of the invention are apparent from the following detailed description taken in conjunction with the accompanying drawings in which:
[0011] FIG. 1 illustrates the computer automated system according to an embodiment.
[0012] TABLE 1 illustrates data collection according to the embodiment depicted in FIG. 1.
[0013] FIG. 2 illustrates data sample collection in detail, according to an embodiment.
[0014] FIG. 3 FIG. 3 illustrates data labeling in further detail, according to an embodiment of the system and method.
[0015] FIG. 4 illustrates testing of the Machine Learning and Predictive Analytics in further detail, according to an embodiment of the system and method.
[0016] FIG. 5 illustrates the output of a model according to an embodiment.
[0017] FIG. 6 illustrates the computer implemented method according to an embodiment.DETAILED DESCRIPTION
[0018] For the purpose of promoting an understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings, and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended. Such alterations and further modifications in the illustrated system, and such further applications of the principles of the invention as illustrated therein would be contemplated as would normally occur to one skilled in the art to which the invention relates. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skillin the art. The system, methods, and examples provided herein are illustrative only and are not intended to be limiting.
[0019] Embodiments disclosed include a computer automated system for randomized dataset creation and evaluation of language models in understanding basic human emotions. According to an embodiment, the computer automated system is caused to aggregate user data from a plurality of data sources. According to an embodiment, the plurality of data sources comprise diverse individuals in a population, and the data is aggregated based on age, gender, nationality, language, culture, race, financial profile, academic profile, and marital status, amongst other things. The neural architecture is configured to standardize the aggregated user data from the plurality of data sources. Embodiments disclosed then normalize the standardized, aggregated user data from the plurality of data sources, and are configured to estimate an emotional state or states of the aggregated user data based on predefined criteria. A preferred embodiment is further configured to triangulate each of the estimated emotional state or states of the aggregated user data, to estimate user data sample distribution and to estimate an error or errors in pre-labelled user data. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.
[0020] According to an embodiment, the computer automated system is further caused to aggregate the user data from a pre-defined population based on at least one of age, gender, ethnicity, language, culture, race, financial profile, academic profile, and marital status.
[0021] According to an embodiment, the computer automated system is further caused to: in estimating user data sample distribution, estimating the sample distribution by Chi-Squared Goodness of Fit test i.e., a statistical test used to determine how well a sample data set matches the distribution of a theoretical probability distribution. In other words, it assesseswhether the observed data follows a specified distribution, often comparing it to an expected or theoretical distribution.
[0022] Additionally, in estimating an error or errors in pre-labelled user data, the computer automated system calculates a p-value that shows the likelihood of observing a difference between a plurality of groups in the data, assuming there is no real difference between the group; and wherein if the p-value is less than a pre-defined significance level (usually set at 0.05), then the difference between the groups is considered statistically significant such that a low p-value indicates that there is strong evidence to suggest that there is a real difference between the groups being compared. The p-value is a measure used in statistical hypothesis testing to determine the evidence against a null hypothesis. It represents the probability of obtaining results as extreme as, or more extreme than, the observed results under the assumption that the null hypothesis is true. A lower p-value suggests stronger evidence against the null hypothesis, often leading to its rejection. Commonly, a significance level (alpha) is set (e.g., 0.05), and if the p-value is less than alpha, the null hypothesis is rejected. It should be noted however, that a low p-value does not prove a specific alternative hypothesis; it only indicates that the observed results are unlikely under the assumption of the null hypothesis.
[0023] According to an additional embodiment, the computer automated system is caused to evaluate the performance of a machine-learning model by comparing a predicted user label data by the machine learning model with the pre -labelled user data. Further, and preferably, in evaluating the performance of the machine-learning model, the computer automated system is caused to return at least one of a true positive, a true negative, a false positive and a false negative result, and display the result on a user interface.
[0024] According to an ideal embodiment, based on the returned result, the computer automated system is caused to calculate at least one of a precision, a recall and an Fl score.The precision comprises the proportion of true positive predictions from all of the plurality of positive predictions made by the machine-learning model. The recall comprises the proportion of true positive predictions by the machine-learning model from all of the prelabelled user data. And the Fl score comprises a harmonic mean of precision and recall.
[0025] FIG. 1 illustrates the computer automated system according to an embodiment. The computer automated system 100 comprises central processing unit (CPU) 102 operatively coupled to memory element 104 having instructions encoded thereon. The computer automated system 100 is caused to aggregate user data via a data aggregation engine 110 from a plurality of data sources 122 through network 120. According to an embodiment, the plurality of data sources comprise diverse individuals in a population, and the data is aggregated based on age, gender, nationality, language, culture, race, financial profile, academic profile, and marital status, amongst other things. Alternatively, the data sources comprises a dataset labelled for the presence of emotions classified by trained clinicians, psychologists, psychiatrists and linguistics to ensure that the dataset used for the research is of high quality and that the labels provided are accurate. To reduce bias in human labeling, the data sources provide input from a diverse group of labelers, use random sampling techniques, and validate the labeling process, before being further validated and processed by labelling engine 114. Labelling engine 114 is further configured to estimate an emotional state or states of the aggregated user data based on pre -defined criteria. Data triangulation engine 112 is configured to triangulate each of the estimated emotional state or states of the aggregated user data, to estimate user data sample distribution and to estimate an error or errors in pre -labelled user data. Randomization engine 116 randomizes and picks samples from triangulated data for further accuracy. According to one embodiment, sample selection in randomization is 10%. Other variations are possible, and may even be preferable, as would be apparent to a person having ordinary skill in the art. Analysis Of Variance(ANOVA) and Sample Label Error Estimation Engine 118 assumes that the data within each group are normally distributed and have equal variances. At the time of labeling, if there are outliers or errors present within a group, this can cause problems for the analysis. ANOVA assumes that the residuals (differences between actual and predicted values) are normally distributed and have equal variances across groups. It's crucial to check whether these assumptions hold true and to use other methods suitable for the data if they don't. In other words, if the data do not meet these assumptions, then it is necessary to consider other statistical methods that are more appropriate for the specific situation. Analysis Of Variance or ANOVA engine 118 calculates a p-value that shows the likelihood of observing a difference between groups in the data, assuming there is no real difference between the groups. If the p-value is less than a significance level (usually set at 0.05), then the difference between the groups is considered statistically significant. In other words, a low p-value indicates that there is strong evidence to suggest that there is a real difference between the groups being compared. Machine learning engine 106 is configured to generate a model of predicted labels in emotion detection. Confusion Matrix engine 108 evaluates the performance of the machine-learning model by comparing the model's predicted labels to the actual labels. In the context of emotion detection, the predicted labels are the model's classification of which emotion is present in natural language data, while the actual labels are clinician-labeled annotations indicating the presence or absence of, for example, anger. The confusion matrix provides four possible results: true positive, true negative, false positive, and false negative. These results can be used to calculate other metrics like precision, recall, and Fl score, providing a detailed breakdown of the LLM's performance.
[0026] TABLE 1 illustrates data collection according to the embodiment depicted in FIG. 1 and includes aggregation of a population, wherein the data is aggregated based on age,gender, nationality, language, culture, race, financial profile, academic profile, and marital status, amongst other things.
[0027] FIG. 2 illustrates data sample collection in detail, according to an embodiment. According to the illustrated embodiment, first an emotion is selected, secondly the selected emotion is then defined. Thirdly data is collected from a pre-defined population based on age, gender, ethnicity, language, culture, race, financial profile, academic profile, and marital status, according to the illustrated example embodiment. Fourthly, the defined emotion is triangulated for the defined population. Fifth, emotions other than the selected emotion are triangulated for the define population. And sixth, movies, news articles, tweets, songs, reels, social media, parts, etc. that depict both the triangulated selected emotion and emotions other than the selected emotion or emotions, are selected. And finally, in the seventh step, randomization comprising picking samples from triangulated data for further accuracy is performed. According to one embodiment, sample selection in randomization is 10%. Other variations are possible, and may even be preferable, as would be apparent to a person having ordinary skill in the art.
[0028] Let's split down the supplied prompt into smaller components to explain it more thoroughly. Choose an emotion to study: The first step is to choose an emotion to study. Emotions are personal experiences that elicit a variety of physiological and psychological responses. Happiness, sadness, wrath, fear, contempt, and surprise are all examples of emotions.
[0029] Define emotion: Once we've decided on an emotion, we must define it. Depending on the theoretical framework and context in which they are researched, emotions can be characterized in a variety of ways. Emotions are complex psychological states that include a subjective feeling, physiological changes, and behavioral responses, according to one generally accepted definition.
[0030] Triangulate the population's specified emotion: Triangulation is the use of various methodologies or sources of data to verify a conclusion or hypothesis. In the context of emotion research, triangulation entails gathering data from several sources and employing various methods to quantify the defined emotion for a specific population. For example, we may triangulate the experience of a defined emotion in a specific group using self -report questionnaires, physiological markers such as heart rate or skin conductance, and behavioral observations.
[0031] Other emotions other than the defined emotion should be triangulated for the populations: We can triangulate other feelings for the population in addition to the defined emotion. This entails picking various emotions to examine, defining them, and measuring them using a variety of methodologies. We can acquire a better knowledge of how different emotions are experienced and expressed in diverse settings by studying several emotions in the same population.
[0032] FIG. 3 illustrates data labeling in further detail, according to an embodiment of the system and method. Step 302 entails measuring the sample distribution estimation before sending the data for pre-processing (step 304). The data is cleaned in step 306 and any anomalies are removed or corrected. Data is partitioned into subsets in step 308 to make the task more manageable and efficient. Step 310 entails preparing the data for labeling which essentially comprises appropriately formatting the data. Step 312 entails sample labeling . And after sample label randomization (step 314), step 316 involves sample label error estimation. This step loops back to sample labeling step 312 which re-labels the data based on the result of the error estimation.
[0033] FIG. 4 illustrates testing of the Machine Learning and Predictive Analytics in further detail, according to an embodiment of the system and method. Step 402 entails creating high quality prompts to obtain accurate predictions from large language models Language ModelAPIs. Step 404 comprises an interaction with the language model via a Large LanguageModel API to make predictions (step 406). Step 408 includes evaluation of the predictions based on which the language model is modified until desired predictions are obtained. Evaluation further comprises evaluating the performance of a machine-learning model by comparing the model's predicted labels to the actual labels via a confusion matrix (step 408A). The confusion matrix provides four possible results: true positive, true negative, false positive, and false negative. These results can be used to calculate other metrics like precision, recall, and Fl score (step 408B), providing a detailed breakdown of the LLM's performance.
[0034] Precision measures the proportion of true positive predictions among all positive predictions made by the model. In other words, it measures how often the model correctly identified a positive case. Precision is calculated as:
[0035] Recall, also known as sensitivity or true positive rate, measures the proportion of true positive predictions among all actual positive cases in the dataset. In other words, it measures how often the model correctly identified a positive case out of all the positive cases in the dataset. Recall is calculated as:
[0036] Fl score is the harmonic mean of precision and recall. It is a balanced measure that takes both precision and recall into account. The Fl score is calculated as:
[0037] A higher precision indicates that the model makes fewer false positive predictions. A higher recall indicates that the model makes fewer false negative predictions. A high Fl score indicates that the model has both high precision and high recall.
[0038] Step 408C includes Area Under Curve-Receiver Operating Characteristis (AUC-ROC) and Area Under Curve-Precision Recall (AUC-PR). The AUC-ROC and AUC-PR metrics are commonly used to evaluate the performance of Large Language Models (LLMs) in various classification tasks, including emotion detection. AUC-ROC measures a model's ability to differentiate between positive and negative examples by calculating the area under the ROC curve, while AUC-PR measures a model's ability to rank positive examples higher than negative examples by calculating the area under the precision-recall (PR) curve. Higher AUC- ROC and AUC-PR values indicate better performance of the model, but these metrics should be used in conjunction with other evaluation metrics to fully assess the model's performance.
[0039] Step 410 includes error estimation to determine an error type, i.e. type 1 or type 2 (step 410A). A Type 1 error occurs when a null hypothesis that is true is rejected, resulting in a false positive outcome in language models. A Type 2 error happens when a null hypothesis that is false is not rejected, leading to a false negative result. Estimation of these errors in language models depends on the application and metrics used, such as precision and recall scores in text classification. Ongoing monitoring and improvement are necessary to improve estimation methods as language models evolve.
[0040] FIG. 5 illustrates the output of a model according to an embodiment. A Softmax function can be used to convert the output of an LLM (Language Model) into a probability distribution over the set of possible output tokens. Here's how to do it: Obtain the output vector from the LLM model for a given input sequence (step 502). This output vector will typically be a vector of logits, which are real-valued numbers that indicate the strength of evidence foreach possible output token. Apply the Softmax function to the output vector to obtain a probability distribution (step 504). The Softmax function is defined as:
[0041] where z is the vector of logits, i is the index of the current output token, and j ranges over all possible output tokens. Exponentiation (step 506) and normalization (step 508) lead to step 510 wherein the Softmax function maps each logit to a probability value that is between 0 and 1 and ensures that the sum of all probabilities is equal to 1. The result is a set of possible lables or categories (step 512).
[0042] An embodiment includes a computer implemented method comprising aggregating user data from a plurality of data sources, the plurality of data sources may comprise diverse individuals in a population, and the data is aggregated based on age, gender, nationality, language, culture, race, financial profile, academic profile, and marital status, amongst other things. An embodiment of the computer implemented method includes standardizing the aggregated user data from the plurality of data sources. Preferably, the computer implemented method further includes normalizing the standardized, aggregated user data from the plurality of data sources, and additionally, estimating an emotional state or states of the aggregated user data based on pre -defined criteria. Additional embodiments of the computer implemented method comprise triangulating each of the estimated emotional state or states of the aggregated user data, estimating data sample distribution, and estimating an error or errors in pre-labelled data.
[0043] According to an embodiment of the computer implemented method, aggregating user data comprises aggregating the user data from a pre-defined population based on at least one of age, gender, ethnicity, language, culture, race, financial profile, academic profile, and marital status.
[0044] According to an additional embodiment of the computer implemented method, estimating user data sample distribution comprises estimating the sample distribution by Chi- Squared Goodness of Fit test which as explained above, is used to determine how well a sample data set matches the distribution of a theoretical probability distribution, wherein it assesses whether the observed data follows a specified distribution, often comparing it to an expected or theoretical distribution.
[0045] According to an embodiment of the computer implemented method, estimating an error or errors in pre-labelled user data comprises calculating a p-value that shows the likelihood of observing a difference between a plurality of groups in the data, assuming there is no real difference between the group. And preferably, if the p-value is less than a predefined significance level (usually set at 0.05), then the difference between the groups is considered statistically significant such that a low p-value indicates that there is strong evidence to suggest that there is a real difference between the groups being compared. The p- value, as explained above is a measure used in statistical hypothesis testing to determine the evidence against a null hypothesis and represents the probability of obtaining results as extreme as, or more extreme than, the observed results under the assumption that the null hypothesis is true. A lower p-value suggests stronger evidence against the null hypothesis, often leading to its rejection. According to an embodiment, a significance level (alpha) is set (at, say for example, at 0.05), and if the p-value is less than alpha, the null hypothesis is rejected. It should be noted that while a low p-value does not prove a specific alternative hypothesis, it indicates that the observed results are unlikely under the assumption of the null hypothesis.
[0046] According to a preferred embodiment, the computer implemented method further comprises evaluating the performance of a machine-learning model by comparing a predicted user label data by the machine learning model with the pre -labelled user data. And ideally, inevaluating the performance of the machine-learning model, the method comprises returning at least one of a true positive, a true negative, a false positive and a false negative result.
[0047] According to a preferred embodiment, the computer implemented method further comprises, based on the returned result, calculating at least one of a precision, a recall and an Fl score. The precision comprises the proportion of true positive predictions from all of the plurality of positive predictions made by the machine-learning model. The recall comprises the proportion of true positive predictions by the machine -learning model from all of the prelabelled user data. And the Fl score comprises a harmonic mean of precision and recall.
[0048] FIG. 6 illustrates the computer implemented method according to an embodiment. The computer implemented method 600 begins with a data aggregation step 602 wherein data is aggregated from the plurality of sources. According to the illustrated embodiment of the method, data aggregation is followed by a data triangulation step 604 wherein triangulation is based on a labelled first emotion collected from the plurality of data sources. Step 606 entails triangulation based on a second one or more emotions collected from the plurality of data sources. Step 608 includes data labeling followed by ANOVA and sample label error estimation in step 610. Step 612 comprises generating a model of predicted labels in emotion detection via machine learning and predictive analytics. Step 614 includes evaluating the performance of the machine-learning model by comparing the model's predicted labels to the actual labels via a confusion matrix. Step 616 comprises AUC-ROC and AUC-PR. AUC- ROC measures a model's ability to differentiate between positive and negative examples by calculating the area under the ROC curve, while AUC-PR measures a model's ability to rank positive examples higher than negative examples by calculating the area under the precisionrecall curve. Higher AUC-ROC and AUC-PR values indicate better performance of the model, but these metrics should be used in conjunction with other evaluation metrics to fullyassess the model's performance. And step 618 includes normalization, softmax and exponentiation. Step 620 concludes the process.
[0049] Normalization is a process used to scale and transform data to a standard range, typically between 0 and 1 or -1 and 1. It helps in bringing different features to a similar scale, preventing some features from dominating others in machine learning algorithms.
[0050] Softmax is a mathematical function that takes a vector of arbitrary real-valued scores and converts them into probabilities. It is often used in the output layer of a neural network for multi-class classification problems. Softmax exponentiates the input values and normalizes them to obtain a probability distribution, where the sum of probabilities equals 1. Exponentiation is the mathematical operation of raising a number to a power. In the context of softmax, it involves raising the scores or logits (raw model outputs) to the power of the base of the natural logarithm (e), which helps in transforming the scores into a probability distribution. The exponentiation is followed by normalization to ensure that the resulting probabilities add up to 1.
[0051] Embodiments disclosed include systems and methods that enable comprehensive and rigorous testing and evaluation of large language models (LLMs). Embodiments disclosed include systems and methods that enable assessment of the ability of LLMs to comprehend and interpret basic human emotions in an analytical and statistical manner. Embodiments disclosed thus enable assessing of the inherent capacity of language models and their eligibility for deployment in various applications.
[0052] Embodiments disclosed include systems and methods that accurately evaluate, in order to facilitate the development and deployment of LLMs that can address deeper human issues, including mental health support and assisting people with disabilities . By testing the ability of these models to understand human emotions, researchers and developers can ensure that they are equipped to handle complex situations and provide appropriate responses.
[0053] Embodiments disclosed include computer automated systems and methods that enable more empathetic responses from LLMs. And finally, embodiments disclosed include systems and methods that can address a wide range of important issues and improve the overall quality of human-machine interactions.
[0054] Since various possible embodiments might be made of the above invention, and since various changes might be made in the embodiments above set forth, it is to be understood that all matter herein described or shown in the accompanying drawings is to be interpreted as illustrative and not to be considered in a limiting sense. Thus, it will be understood by those skilled in the art of computer automated artificial intelligence-based systems and methods, and more particularly systems and methods for randomized dataset creation and evaluation of language models in understanding basic human emotions, that although the preferred and alternate embodiments have been shown and described in accordance with the Patent Statutes, the invention is not limited thereto or thereby.
[0055] The figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. It should also be noted that, in some alternative implementations, the functions noted / illu strated may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.
[0056] The terminology used herein is for the purpose of describing embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations,elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0057] In general, the routines executed to implement the embodiments of the invention, may be part of an operating system or a specific application, component, program, module, object, or sequence of instructions. The computer program of the present invention typically is comprised of a multitude of instructions that will be translated by the native computer into a machine-accessible format and hence executable instructions. Also, programs are comprised of variables and data structures that either reside locally to the program or are found in memory or on storage devices. In addition, various programs described hereinafter may be identified based upon the application for which they are implemented in a specific embodiment of the invention. However, it should be appreciated that any program nomenclature that follows is used merely for convenience, and thus the invention should not be limited to use solely in any specific application identified and / or implied by such nomenclature.
[0058] The present invention and some of its advantages have been described in detail for some embodiments. It should be understood that although the system and process are described with reference to systems and methods for randomized dataset creation and evaluation of language models in understanding basic human emotions, the system and method is highly reconfigurable, and may be used in other systems as well. Portions of the embodiment may be used to support artificial intelligence -based applications completely unrelated to language models in understanding basic human emotions. Modifications of the embodiments may be used in machine-to-machine interactions that could potentially replace human intervention. It should also be understood that various changes, substitutions, and alterations can be made herein without departing from the spirit and scope of the invention as defined by the appended claims. An embodiment of the invention may achieve multipleobjectives, but not every embodiment falling within the scope of the attached claims will achieve every objective. Moreover, the scope of the present application is not intended to be limited to the embodiments of the process, machine, manufacture, composition of matter, means, methods and steps described in the specification. A person having ordinary skill in the art will readily appreciate from the disclosure of the present invention that processes, machines, manufacture, compositions of matter, means, methods, or steps, presently existing or later to be developed are equivalent to, and fall within the scope of, what is claimed. Accordingly, the appended claims are intended to include within their scope such processes, machines, manufacture, compositions of matter, means, methods, or steps.
Claims
WE CLAIM:
1. A computer automated system comprising a processing unit coupled to a memory element and having instructions encoded therein, which instructions when implemented by the processing unit, cause the computer automated system to: aggregate user data from a plurality of data sources; standardize the aggregated user data from the plurality of data sources; normalize the standardized, aggregated user data from the plurality of data sources; estimate an emotional state or states of the aggregated user data based on pre-defined criteria; triangulate each of the estimated emotional state or states of the aggregated user data; estimate user data sample distribution; and estimate an error or errors in pre-labelled user data.
2. The computer automated system of claim 1 wherein the computer automated system is further caused to: aggregate the user data from a pre -defined population based on at least one of age, gender, ethnicity, language, culture, race, financial profile, academic profile, and marital status.
3. The computer automated system of claim 1 wherein the computer automated system is further caused to: in estimating user data sample distribution, determine how well a sample data set matches the distribution of a theoretical probability distribution; wherein the determining comprises assessing whether the observed data follows a specified distribution; and wherein the assessing comprises comparing the observed data to an expected or theoretical distribution.
4. The computer automated system of claim 1 wherein the computer automated system is further caused to: in estimating an error or errors in pre-labelled user data, calculating a p-value that shows the likelihood of observing a difference between a plurality of groups in the data, assuming there is no real difference between the group; and wherein if the p-value is less than a pre-defined significance level (usually set at 0.05), then the difference between the groups is considered statistically significant such that a low p-value indicates that there is strong evidence to suggest that there is a real difference between the groups being compared.
5. The computer automated system of claim 1 wherein the computer automated system is further caused to: evaluate the performance of a machine-learning model by comparing a predicted user label data by the machine learning model with the pre-labelled user data.
6. The computer automated system of claim 4 wherein the computer automated system is further caused to: in evaluating the performance of the machine-learning model, return at least one of a true positive, a true negative, a false positive and a false negative result, and display the result on a user interface.
7. The computer automated system of claim 5 wherein the computer automated system is further caused to: based on the returned result, calculate at least one of a precision, a recall and an Fl score; wherein the precision comprises the proportion of true positive predictions from all of the plurality of positive predictions made by the machine-learning model; wherein the recall comprises the proportion of true positive predictions by the machine-learning model from all of the pre-labelled user data; and wherein the Fl score comprises a harmonic mean of precision and recall.
8. A computer implemented method comprising: aggregating user data from a plurality of data sources; standardizing the aggregated user data from the plurality of data sources;normalizing the standardized, aggregated user data from the plurality of data sources; estimating an emotional state or states of the aggregated user data based on predefined criteria; triangulating each of the estimated emotional state or states of the aggregated user data; estimating data sample distribution; and estimating an error or errors in pre-labelled data.
9. The computer implemented method of claim 8 wherein aggregating user data comprises aggregating the user data from a pre-defined population based on at least one of age, gender, ethnicity, language, culture, race, financial profile, academic profile, and marital status.
10. The computer implemented method of claim 8 wherein: in estimating user data sample distribution, estimating the sample distribution; wherein the estimating comprises determining how well a sample data set matches the distribution of a theoretical probability distribution; wherein the determining comprises assessing whether the observed data follows a specified distribution; and wherein the assessing comprises comparing the observed data to an expected or theoretical distribution.
11. The computer implemented method of claim 8 wherein: in estimating an error or errors in prelabelled user data, calculating a p-value that shows the likelihood of observing a difference between a plurality of groups in the data, assuming there is no real difference between the group; and wherein if the p-value is less than a pre-defined significance level (usually set at 0.05), then the difference between the groups is considered statistically significant such that a low p-value indicates that there is strong evidence to suggest that there is a real difference between the groups being compared.
12. The computer implemented method of claim 8 further comprising: evaluating the performance of a machine-learning model by comparing a predicted user label data by the machine learning model with the pre-labelled user data.
13. The computer implemented method of claim 12 further comprising: in evaluating the performance of the machine-learning model, returning at least one of a true 15 positive, a true negative, a false positive and a false negative result.
14. The computer implemented method of claim 13 further comprising: based on the returned result, calculating at least one of a precision, a recall and an Fl score; wherein the precision comprises the proportion of true positive predictions 20 from all of the plurality of positive predictions made by the machine -learning model; wherein the recall comprises the proportion of true positive predictions by the machine-learning model from all of the pre-labelled user data; and wherein the Fl score comprises a harmonic mean of precision and recall.
Citation Information
Patent Citations
Data evaluation system, data evaluation method, and data evaluation program
US20170323013A1
Distributed analysis for cognitive state metrics
US20200342979A1