System and method of analyzing physical attributes

The method enhances emotion and thought inference by using machine learning and computer vision to analyze video data, providing a more accurate and individualized assessment of physical attributes.

WO2025111266A1PCT designated stage expired Publication Date: 2025-05-30AMGI ANIMATION LLC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/056524
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-11-19
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing methods for analyzing physical attributes to infer emotions or thoughts are crude and simplistic, failing to account for individualized patterns and other non-facial indicators.

Method used

A method that involves identifying features in video data using computer vision techniques and machine learning algorithms, translating these features into feature data over time, comparing them to previous data, and scoring the comparison to determine if it exceeds a threshold.

Benefits of technology

This method provides a more nuanced and individualized analysis of physical attributes, improving the accuracy of emotion and thought inference by accounting for patterns and deviations over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024056524_30052025_PF_FP_ABST
    Figure US2024056524_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The method may identify features of a video, select a feature in the video and translate the feature using a video transformation algorithm into feature data at a plurality of points in time. The method may store the feature data, compare the feature data to previous feature data using a comparison algorithm, scoring the comparison using a scoring algorithm and determine if the score is over a threshold.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD OF ANALYZING PHYSICAL ATTRIBUTESCross Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 602,206, entitled “SYSTEM AND METHOD OF ANALYZING PHYSICAL ATTRIBUTES”, filed November 22, 2023, reference of which is hereby incorporated herein in its entirety.Background

[0002] Many industries would benefit from analyzing physical attributes for clues on the emotions a person is feeling and the thoughts a person is having. Past attempts to find emotions or thoughts in images or videos have been somewhat crude with cameras looking for a smile or a frown and noting the emotions or thoughts normally tied to these facial actions. These analysis attempts are not individualized and are simplistic as they fail to account for other indications of emotions or thoughts.Summary

[0003] A method of analyzing physical attributes such as from video into data may be disclosed. The method may identify features of a video, select a feature in the video and translate the feature using a video transformation algorithm into feature data at a plurality of points in time. The method may store the feature data, compare the feature data to previous feature data using a comparison algorithm, score the comparison using a scoring algorithm and determine if the score is over a threshold.Brief Description of the Drawings

[0004] Fig. 1 may be an illustration of computer based learning system;

[0005] Fig. 2 may be an illustration of a method in accordance with the claims;

[0006] Fig. 3 may be an illustration of a convolutional neural network; and

[0007] Fig. 4 may be an illustration of a computer that may be physically transformed to execute the method.

[0008] Persons of ordinary skill in the art will appreciate that elements in the figures are illustrated for simplicity and clarity so not all connections and options have been shown to avoid obscuring the inventive aspects. For example, common but well-understood elements that are useful or necessary in a commercially feasible embodiment are not often depicted in order to facilitate a less obstructed view of these various embodiments of thepresent disclosure. It will be further appreciated that certain actions and / or steps may be described or depicted in a particular order of occurrence while those skilled in the art will understand that such specificity with respect to sequence is not actually required. It will also be understood that the terms and expressions used herein are to be defined with respect to their corresponding respective areas of inquiry and study except where specific meanings have otherwise been set forth herein. All dimensions specified in this disclosure may be by way of example only and are not intended to be limiting. Further, the proportions shown in these Figures may not be necessarily to scale. As will be understood, the actual dimensions and proportions of any system, any device or part of a system or device disclosed in this disclosure may be determined by its intended use.Specification

[0009] Many industries would benefit from analyzing human attributes for clues on the emotions a person is feeling or thoughts a person is having. Past attempts to find emotions or thoughts in images or videos or other sensors have been somewhat crude with cameras looking for a smile or a frown and noting the emotions normally tied to these facial actions. These analysis attempts are not individualized and are simplistic as they fail to account for other indications of emotions.

[0010] A method of analyzing physical sensor information into data may be disclosed. The method may identify features of a data from a sensor, select a feature in the data and translate the features at a plurality of points in time. The method may store the feature data, compare the feature data to previous feature data using a comparison algorithm, score the comparison using a scoring algorithm and determine if the score is over a threshold. In an example which will be used throughout the description, the sensor may be an image sensor and features may be identified in the output video. The method may translate the features using a video transformation algorithm into feature data at a plurality of points in time. The method may store the feature data, compare the feature data to previous feature data using a comparison algorithm, score the comparison using a scoring algorithm and determine if the score is over a threshold.

[0011] Methods and devices that may implement the embodiments of the various features of the invention will now be described with reference to the drawings. The drawings and the associated descriptions may be provided to illustrate embodiments of the invention and not to limit the scope of the invention. Reference in the specification to “oneembodiment” or “an embodiment” may be intended to indicate that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least an embodiment of the invention. The appearances of the phrase “in one embodiment” or “an embodiment” in various places in the specification may not necessarily be referring to the same embodiment.

[0012] Throughout the drawings, reference numbers may be re-used to indicate correspondence between referenced elements. As used in this disclosure, except where the context requires otherwise, the term “comprise” and variations of the term, such as “comprising”, “comprises” and “comprised” may not be intended to exclude other additives, components, integers or steps.

[0013] In the following description, specific details may be given to provide a thorough understanding of the embodiments. However, it may be understood by one of ordinary skill in the art that the embodiments may be practiced without these specific details. Well-known circuits, structures and techniques may not be shown in detail in order not to obscure the embodiments. For example, circuits may be shown in block diagrams in order not to obscure the embodiments in unnecessary detail.

[0014] Also, it is noted that the embodiments may be described as a process that is depicted as a flowchart, a flow diagram, a structure diagram, or a block diagram. The flowcharts and block diagrams in the figures may illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer programs according to various embodiments disclosed. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, that may include one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures.

[0015] Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. A process may be terminated when its operations are completed. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function. Additionally, each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, may be implemented by special purposehardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0016] Moreover, a storage may represent one or more devices for storing data, including read-only memory (ROM), random access memory (RAM), magnetic disk storage mediums, optical storage mediums, flash memory devices and / or other non-transitory machine readable mediums for storing information. The term "machine readable medium" may include, but is not limited to portable or fixed storage devices, optical storage devices, wireless channels and various other non-transitory mediums capable of storing, comprising, containing, executing or carrying instruction(s) and / or data.

[0017] Furthermore, embodiments may be implemented by hardware, software, firmware, middleware, microcode, or a combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine-readable medium such as a storage medium or other storage(s). One or more than one processor may perform the necessary tasks in series, distributed, concurrently or in parallel. A code segment may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or a combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted through a suitable means including memory sharing, message passing, token passing, network transmission, etc. and are also referred to as an interface, where the interface is the point of interaction with software, or computer hardware, or with peripheral devices.

[0018] Figure 2 may disclose a method of analyzing physical attributes for thought data that may be analyzed. For example, the method may analyze motion in video for thought data such as emotional indications and patterns to be analyzed. At a very high level, physical response, whether intentional or unintentional, may provide insight into what a person or animal is thinking or planning or feeling. The physical response may be in many forms, such as a rise in temperature, a smell of sweat, a sound of pain or suffering or a change of taste. In a very simple example, humans may scratch their head if they are confused. By studying physical movements using machine learning or artificial intelligence, patterns of behavior may emerge. If a human is displaying a pattern of behavior, an insight into their mind may be determined. Similarly, if a human deviates from a pattern of behavior, additional insights into the thoughts of the human may be determined. Finally, the analysismay be specific to individuals as each individual may have different mannerism which may provide insight into the thoughts of the individual.

[0019] The technical problems in the past have been many. Determining which features matter to indicate a thought process has been hit or miss. Similarly, determining patterns of behavior based on the features has been challenging and has had many false positives as the systems do not improve over time or take feedback into account. By using artificial intelligence systems specifically designed and to determine features, the ability to use features has improved. Additionally, artificial intelligence has greatly improved the ability of the system and method to analyze the large amount of data to determine patterns.

[0020] The uses of the system and method are many. As an example and not limitation, determining if a person in an education session understands a concept may be greatly improved by analyzing features of the person. Identification verification may be possible by studying a broad range of feature. Determining criminal intent may be possible by studying features. The ability to sell goods or services may be improved by studying features. The examples are many and varied and should not be limited by these examples.

[0021] In some embodiments, the sensor data such as a video may be converted into a plurality of images for analysis. In some embodiments, the video may have a frame rate and the frame rate may be used to determine the number of images to use per second. Obviously, more or less frames per second may also be used. In some embodiments, the images may be normalized so that comparisons across videos may be consistent. For example, the size, brightness or contrast of the images that make up the video may be adjusted to be consistent across the images being used as comparisons.

[0022] The physical sensor may take on many forms. It may be an image sensor, a voice sensor, a scent sensor, a touch sensor, a taste sensor, a magnetic sensor or a temperature sensor. The output data may be analyzed for features relevant to the physical sensor. In addition, the physical sensor may be a plurality of physical sensors. The data from the plurality of physical sensors may be analyzed together to create an even more detailed set of data to be analyzed to attempt to determine the thoughts of a person or animal. Throughout the description, a video sensor and video data will be used for simplicity and is not meant to be limiting in any sense.

[0023] At block 200, features of an output from a sensor such as video may be identified. Some sample features may include edges, colors, shapes, or key points. There may be a plurality of features that may be analyzed alone of together. For example, a face may have 50 or 100 features that may be analyzed. Logically, a combination of features maybe used. As an example, edges of a face and points on a nose may be used together as the features to be analyzed. The features also may be points related to parts of a human body.

[0024] The features may be extracted in a variety of ways. In one embodiment, the features are extracted using computer vision techniques and a variety of computer vision techniques are possible. In other embodiments, features are extracted using pre-trained machine learning models. For example, the pre-trained machine learning model may be a convolutional neural network (CNN). The CNN in the may be trained on millions of images of people and may have learned to understand the thoughts from the photos. This CNN may be novel because it has been created and trained on known images only. Logically, other types of learning algorithms in the estimator may be used. For example, the learning algorithm may be a fully connected neural network (FCN) in one embodiment.

[0025] Turning the images into data may entail taking measurements of different points on the object. The points may be compared to baseline of measurements for the object and the changes may be noted. The system may then analyze the changes to determine the extent of movement.

[0026] More specifically, referring to Fig. 3, the learning algorithm may include a convolutional neural network 510 (CNN) and a transformer 320. In one embodiment, the CNN 310 may determine one or more features 351-354 in each photo 341-344. In one example, the CNN may determine the features 351-354 which may be a set of numbers but the amount of features 351-354 may be varied up or down depending on many factors.

[0027] The CNN may be trained on millions of images of people and may have learned to understand the thoughts of the person from the photos. This CNN may be novel because it has been created and trained on known images. Logically, other types of learning algorithms in the may be used. For example, the learning algorithm may be a fully connected neural network (FCN) in one embodiment. The analysis of the features may indicate the changes to the physical appearance.

[0028] In training, the transformer 320 may take the features 351-354 of multiple images 341-344 of the same person (the outputs of the CNN) as well as additional data such as the stated thought process of the individual in the photo 360 to create a model. Once the model is trained, the transformer may generate predictions of the thought process of the individual 370. In some embodiments, the estimation of the thought process 370 may be in real time. The transformer 320 used in this invention may be trained on a dataset specifically created for predicting thoughts 370.

[0029] The trained model which may be in the transformer 320 may take the features of multiple images 341-344 of the same object as well as outside information in order to predict the thought process of the object. The learning algorithm also may analyze other relevant information about the object.

[0030] At block 205, a feature in the video may be selected to be analyzed. Different features may be observed for different purposes. For example, attempting to determine if a person is preparing to commit a crime may involve looking at one set of features while determining if a person is enjoying a comedy film may involve looking at a different set of feature.

[0031] Machine learning may be used to assist in determining the feature or features that should be studied for different scenarios. As mentioned previously, past films of people preparing to commit crimes to determine features that may need to be observed. Similarly, past films of people enjoying a comedy film may be studied to determine features that may need to be observed. As with any machine learning algorithm, additional data may be communicated to the model over time to improve the model.

[0032] In yet another embodiment, the features that are studied may be specific to an individual. Over time, a variety of features may be captured and analyzed to determine what the features represent. For example, a student may be monitored and the feature may be stored to determine the amount of understanding of a student. Scratching a head may be a sign of confusion for some students but may be a sign of complete mastery in other students. My personalizing the analysis, better results and more personalized results may be possible.

[0033] At block 210, the features may be translated using a transformation algorithm such as a video transformation algorithm into feature data at a plurality of points in time. A variety of video transformation algorithms may be used. In one embodiment, the video transformation algorithm is a Fast Fourier Transform (FFT) algorithm. Logically, other video transformation algorithms are possible and are contemplated. The video transformation algorithm may make it easier to analyze the feature data as the feature data may be converted into numbers which are easier to compare rather than photos or videos.

[0034] The features may be converted using a conversion algorithm into a representative format. The format may be consistent over the analysis and over time such that comparisons or videos may be based on comparing the same formats. In some embodiments, the conversion algorithm creates a data structure to store data representing movement in the video. For example, the data structure may include a feature vector or a plurality of feature vectors.

[0035] At block 215, the feature data may be stored in a memory. The memory may be selected based on the type and amount of data. For example, a large amount of data may require a different type of memory than a small amount of data. In addition, the memory may be local or remote or in the cloud depending on the amount of data, the need for speed, etc.

[0036] At block 220, the feature data may be compared to previous feature data using a comparison algorithm. A variety of comparison algorithms are possible and are contemplated. At a high level, the comparison algorithm may compare the video under analysis to videos in the video database. More specifically, the images that make up the various videos may be compared to images in the image database. Even more specifically, the data representing the image may be compared to the data representing the images to data representing previously analyzed images. In one embodiment, the comparison algorithm may include a similarity metric selected from a group of Euclidean distance, cosine similarity, or dynamic time warping for temporal patterns. Of course, other similarity measures are possible and are contemplated.

[0037] At block 225, the comparison may be scored using a scoring algorithm. The scoring algorithm may be an indication of how similar the images or videos (or features in the images or videos) may be at a point in time. For example, if the video images are converted using a Fast Fourier Transform (FFT), the similarity metric may compare the FFT results which is easier than comparing images. At a high level, the FFT converts a signal into individual spectral components and thereby provides frequency information about the signal. As an example, the similarity metric may determine the Euclidean distance between FFT results of the image in question and a scored image from the past. In another embodiment, the features may be in the form of grayscale or color intensity values for the pixel(s) representing the chosen point which may be the feature. The grayscale or color intensity values may be easy to compare in a mathematical approach.

[0038] In a more detailed approach, the pixel intensity values for a point on a face at each time point may be analyzed as a time series and FFT may be performed on both time series to convert them into the frequency domain which will result in the amplitude and phase information of the frequencies present in the intensity variations at the point on the face. A comparison metric may be defined to evaluate the difference or similarity between the FFT results at the two time points wherein the comparison metric may be based on the magnitude and phase of the dominant frequency components. Logically, a scoring mechanism may be used to determine when the point has changed significantly or remained relatively stable between the two time points.

[0039] The database of patterns may include a significant number of videos used to create a database of patterns or feature representations from those videos. The database may be organized for efficient retrieval and comparison. The format of the data and the database may take on many forms depending on the quantity of data, the speed of retrieval needed and the data type.Machine Learning

[0040] Machine learning may be used to recognize patterns. The machine learning model may be trained on a model on an existing dataset and using the model to predict whether the movement in the new video matches known patterns. The machine learning model may be used to predict future actions based on past pattern recognition. The machine learning model may also be used to determine pattern deviation. Logically, pattern deviation may be used to determine future actions.

[0041] A framework for machine learning algorithm like a large language model may involve a combination of one or more components, sometimes three components: (1) representation, (2) evaluation, and (3) optimization components. Representation components refer to computing units that perform steps to represent knowledge in different ways, including but not limited to as one or more decision trees, sets of rules, instances, graphical models, neural networks, support vector machines, model ensembles, and / or others.Evaluation components refer to computing units that perform steps to represent the way hypotheses (e.g., candidate programs) are evaluated, including but not limited to as accuracy, prediction and recall, squared error, likelihood, posterior probability, cost, margin, entropy k- L divergence, and / or others. Optimization components refer to computing units that perform steps that generate candidate programs in different ways, including but not limited to combinatorial optimization, convex optimization, constrained optimization, and / or others. In some embodiments, other components and / or sub-components of the aforementioned components may be present in the system to further enhance and supplement the aforementioned machine learning functionality.

[0042] Machine learning algorithms sometimes rely on unique computing system structures. Machine learning algorithms may leverage neural networks, which are systems that approximate biological neural networks (e.g., the human brain). Such structures, while significantly more complex than conventional computer systems, are beneficial in implementing machine learning. For example, an artificial neural network may be comprised of a large set of nodes which, like neurons in the brain, may be dynamically configured to effectuate learning and decision-making.

[0043] Machine learning tasks are sometimes broadly categorized as either unsupervised learning or supervised learning. In unsupervised learning, a machine learning algorithm is left to generate any output (e.g., to label as desired) without feedback. The machine learning algorithm may teach itself (e.g., observe past output), but otherwise operates without (or mostly without) feedback from, for example, a human administrator. Meanwhile, in supervised learning, a machine learning algorithm is provided feedback on its output. Feedback may be provided in a variety of ways, including via active learning, semisupervised learning, and / or reinforcement learning. In active learning, a machine learning algorithm is allowed to query answers from an administrator. For example, the machine learning algorithm may make a guess in a face detection algorithm, ask an administrator to identify the photo in the picture, and compare the guess and the administrator's response. In semi-supervised learning, a machine learning algorithm is provided a set of example labels along with unlabeled data. For example, the machine learning algorithm may be provided a data set of 100 photos with labeled human faces and 10,000 random, unlabeled photos. In reinforcement learning, a machine learning algorithm is rewarded for correct labels, allowing it to iteratively observe conditions until rewards are consistently earned. For example, for every face correctly identified, the machine learning algorithm may be given a point and / or a score (e.g., “75% correct”). An embodiment involving supervised machine learning is described herein.

[0044] As elaborated herein, in practice, machine learning systems and their underlying components are tuned by data scientists to perform numerous steps to perfect machine learning systems. The process is sometimes iterative and may entail looping through a series of steps: (1) understanding the domain, prior knowledge, and goals; (2) data integration, selection, cleaning, and pre-processing; (3) learning models; (4) interpreting results; and / or (5) consolidating and deploying discovered knowledge. This may further include conferring with domain experts to refine the goals and make the goals more clear, given the nearly infinite number of variables that can possible be optimized in the machine learning system. Meanwhile, one or more of data integration, selection, cleaning, and / or preprocessing steps can sometimes be the most time consuming because the old adage, “garbage in, garbage out,” also reigns true in machine learning systems.

[0045] By way of example, FIG. 1 illustrates a simplified example of an artificial neural network 100 on which a machine learning algorithm may be executed. FIG. 1 is merely an example of nonlinear processing using an artificial neural network; other forms ofnonlinear processing may be used to implement a machine learning algorithm in accordance with features described herein.

[0046] In FIG. 1, each of input nodes 110 a-n is connected to a first set of processing nodes 120 a-n. Each of the first set of processing nodes 120 a-n is connected to each of a second set of processing nodes 130 a-n. Each of the second set of processing nodes 130 a-n is connected to each of output nodes 140 a-n. Though only two sets of processing nodes are shown, any number of processing nodes may be implemented. Similarly, though only four input nodes, five processing nodes, and two output nodes per set are shown in FIG. 1, any number of nodes may be implemented per set. Data flows in FIG. 1 are depicted from left to right: data may be input into an input node, may flow through one or more processing nodes, and may be output by an output node. Input into the input nodes 110 a-n may originate from an external source 160. Output may be sent to a feedback system 150 and / or to storage 170. The feedback system 150 may send output to the input nodes 110 a-n for successive processing iterations with the same or different input data.

[0047] In one illustrative method using feedback system 150, the system may use machine learning to determine an output. The output may include anomaly scores, heat scores / values, confidence values, and / or classification output. The system may use any machine learning model including xgboosted decision trees, auto-encoders, perceptron, decision trees, support vector machines, regression, and / or a neural network. The neural network may be any type of neural network including a feed forward network, radial basis network, recurrent neural network, long / short term memory, gated recurrent unit, auto encoder, variational autoencoder, convolutional network, residual network, Kohonen network, and / or other type. In one example, the output data in the machine learning system may be represented as multi-dimensional arrays, an extension of two-dimensional tables (such as matrices) to data with higher dimensionality.

[0048] The neural network may include an input layer, a number of intermediate layers, and an output layer. Each layer may have its own weights. The input layer may be configured to receive as input one or more feature vectors described herein. The intermediate layers may be convolutional layers, pooling layers, dense (fully connected) layers, and / or other types. The input layer may pass inputs to the intermediate layers. In one example, each intermediate layer may process the output from the previous layer and then pass output to the next intermediate layer. The output layer may be configured to output a classification or a real value. In one example, the layers in the neural network may use an activation function such as a sigmoid function, a Tan h function, a ReLu function, and / or other functions.Moreover, the neural network may include a loss function. A loss function may, in some examples, measure a number of missed positives; alternatively, it may also measure a number of false positives. The loss function may be used to determine error when comparing an output value and a target value. For example, when training the neural network the output of the output layer may be used as a prediction and may be compared with a target value of a training instance to determine an error. The error may be used to update weights in each layer of the neural network.

[0049] In one example, the neural network may include a technique for updating the weights in one or more of the layers based on the error. The neural network may use gradient descent to update weights. Alternatively, the neural network may use an optimizer to update weights in each layer. For example, the optimizer may use various techniques, or combination of techniques, to update weights in each layer. When appropriate, the neural network may include a mechanism to prevent overfitting — regularization (such as LI or L2), dropout, and / or other techniques. The neural network may also increase the amount of training data used to prevent overfitting.

[0050] Once data for machine learning has been created, an optimization process may be used to transform the machine learning model. The optimization process may include (1) training the data to predict an outcome, (2) defining a loss function that serves as an accurate measure to evaluate the machine learning model's performance, (3) minimizing the loss function, such as through a gradient descent algorithm or other algorithms, and / or (4) optimizing a sampling method, such as using a stochastic gradient descent (SGD) method where instead of feeding an entire dataset to the machine learning algorithm for the computation of each step, a subset of data is sampled sequentially. In one example, optimization comprises minimizing the number of false positives to maximize a user's experience. Alternatively, an optimization function may minimize the number of missed positives to optimize minimization of losses from exploits.

[0051] In one example, FIG. 1 depicts nodes that may perform various types of processing, such as discrete computations, computer programs, and / or mathematical functions implemented by a computing device. For example, the input nodes 110 a-n may comprise logical inputs of different data sources, such as one or more data servers. The processing nodes 120 a-n may comprise parallel processes executing on multiple servers in a data center. And, the output nodes 140 a-n may be the logical outputs that ultimately are stored in results data stores, such as the same or different data servers as for the input nodes 110 a-n. Notably, the nodes need not be distinct. For example, two nodes in any two sets mayperform the exact same processing. The same node may be repeated for the same or different sets.

[0052] Each of the nodes may be connected to one or more other nodes. The connections may connect the output of a node to the input of another node. A connection may be correlated with a weighting value. For example, one connection may be weighted as more important or significant than another, thereby influencing the degree of further processing as input traverses across the artificial neural network. Such connections may be modified such that the artificial neural network 100 may learn and / or be dynamically reconfigured. Though nodes are depicted as having connections only to successive nodes in FIG. 1, connections may be formed between any nodes. For example, one processing node may be configured to send output to a previous processing node.

[0053] Input received in the input nodes 110 a-n may be processed through processing nodes, such as the first set of processing nodes 120 a-n and the second set of processing nodes 130 a-n. The processing may result in output in output nodes 140 a-n. As depicted by the connections from the first set of processing nodes 120 a-n and the second set of processing nodes 130 a-n, processing may comprise multiple steps or sequences. For example, the first set of processing nodes 120 a-n may be a rough data filter, whereas the second set of processing nodes 130 a-n may be a more detailed data filter.

[0054] The artificial neural network 100 may be configured to effectuate decisionmaking. As a simplified example for the purposes of explanation, the artificial neural network 100 may be configured to detect faces in photographs. The input nodes 110 a-n may be provided with a digital copy of a photograph. The first set of processing nodes 120 a-n may be each configured to perform specific steps to remove non-facial content, such as large contiguous sections of the color red. The second set of processing nodes 130 a-n may be each configured to look for rough approximations of faces, such as facial shapes and skin tones. Multiple subsequent sets may further refine this processing, each looking for further more specific tasks, with each node performing some form of processing which need not necessarily operate in the furtherance of that task. The artificial neural network 100 may then predict the location on the face. The prediction may be correct or incorrect.

[0055] The feedback system 150 may be configured to determine whether or not the artificial neural network 100 made a correct decision. Feedback may comprise an indication of a correct answer and / or an indication of an incorrect answer and / or a degree of correctness (e.g., a percentage). For example, in the facial recognition example provided above, the feedback system 150 may be configured to determine if the face was correctly identified and,if so, what percentage of the face was correctly identified. The feedback system 150 may already know a correct answer, such that the feedback system may train the artificial neural network 100 by indicating whether it made a correct decision. The feedback system 150 may comprise human input, such as an administrator telling the artificial neural network 100 whether it made a correct decision. The feedback system may provide feedback (e.g., an indication of whether the previous output was correct or incorrect) to the artificial neural network 100 via input nodes 110 a-n or may transmit such information to one or more nodes. The feedback system 150 may additionally or alternatively be coupled to the storage 170 such that output is stored. The feedback system may not have correct answers at all, but instead base feedback on further processing: for example, the feedback system may comprise a system programmed to identify faces, such that the feedback allows the artificial neural network 100 to compare its results to that of a manually programmed system.

[0056] The artificial neural network 100 may be dynamically modified to learn and provide better input. Based on, for example, previous input and output and feedback from the feedback system 150, the artificial neural network 100 may modify itself. For example, processing in nodes may change and / or connections may be weighted differently. Following on the example provided previously, the facial prediction may have been incorrect because the photos provided to the algorithm were tinted in a manner which made all faces look red. As such, the node which excluded sections of photos containing large contiguous sections of the color red could be considered unreliable, and the connections to that node may be weighted significantly less. Additionally or alternatively, the node may be reconfigured to process photos differently. The modifications may be predictions and / or guesses by the artificial neural network 100, such that the artificial neural network 100 may vary its nodes and connections to test hypotheses.

[0057] The artificial neural network 100 need not have a set number of processing nodes or number of sets of processing nodes, but may increase or decrease its complexity. For example, the artificial neural network 100 may determine that one or more processing nodes are unnecessary or should be repurposed, and either discard or reconfigure the processing nodes on that basis. As another example, the artificial neural network 100 may determine that further processing of all or part of the input is required and add additional processing nodes and / or sets of processing nodes on that basis.

[0058] The feedback provided by the feedback system 150 may be mere reinforcement (e.g., providing an indication that output is correct or incorrect, awarding the machine learning algorithm a number of points, or the like) or may be specific (e.g.,providing the correct output). For example, the machine learning algorithm 100 may be asked to detect faces in photographs. Based on an output, the feedback system 150 may indicate a score (e.g., 75% accuracy, an indication that the guess was accurate, or the like) or a specific response (e.g., specifically identifying where the face was located).

[0059] The artificial neural network 100 may be supported or replaced by other forms of machine learning. For example, one or more of the nodes of artificial neural network 100 may implement a decision tree, associational rule set, logic programming, regression model, cluster analysis mechanisms, Bayesian network, propositional formulae, generative models, and / or other algorithms or forms of decision-making. The artificial neural network 100 may effectuate deep learning.

[0060] A large language model may be a language model characterized by its large size. Their size is enabled by Al accelerators, which are able to process vast amounts of text data, mostly scraped from the Internet. The artificial neural networks which are built can contain from tens of millions and up to billions of weights and are (pre-)trained using selfsupervised learning and semi-supervised learning. Transformer architecture contributed to faster training.

[0061] As language models, they work by taking an input text and repeatedly predicting the next token or word. Up to 2020, fine tuning was the only way a model could be adapted to be able to accomplish specific tasks. Larger sized models, such as GPT-3, however, can be prompt-engineered to achieve similar results. They are thought to acquire embodied knowledge about syntax, semantics and "ontology" inherent in human language corpora large language models are trained using self-supervised learning or semi-supervised learning. This means that they are trained on large amounts of unlabeled text. Large language models can adjust their internal parameters and learn from new inputs from users over time.

[0062] Large language models are trained to predict the next word in a sentence based on the previous input sentence. This is a self-supervised learning task because you are not defining separate output labels. The process is repeated until the model reaches an acceptable level of accuracy. Some large language models, like InstructGPT and ChatGPT, use both supervised learning and reinforcement learning. The combination of the two is crucial for optimal performance.

[0063] Referring again to Fig. 2, at block 230, the method may determine if the score is over a threshold. The similarity score may be compared to a threshold. The threshold may be set initially and may be adjusted over time to reflect the accuracy of the model.

[0064] At block 235, in response to the score being over the threshold, the physical movements may be classified and the classification may related to what the subject is thinking. For example, a head scratch may be determined to be confusion.Verification and Refinement

[0065] Additional verification steps may be used to reduce false positives of the comparison. More specifically, the comparison analysis may be refined based on feedback and new information to improve the accuracy of the pattern recognition algorithm. For example, for many people, a head scratch indicates confusion. In a smaller percentage of people, a head scratch may indicate confidence in an answer. The feedback for the individual may be used to improve the system and method for the individual. Similarly, some movements may be affected by media or current events and the meaning of movements or other physical manifestations of thoughts may change over time.

[0066] Computing devices are used through the method and system. As shown in Fig. 4, the computing device 401 that executes the method may include a processor 402 that is coupled to an interconnection bus. The processor 402 may include a register set or register space 404, which is depicted in Fig. 4 as being entirely on-chip, but which could alternatively be located entirely or partially off-chip and directly coupled to the processor 402 via dedicated electrical connections and / or via the interconnection bus. The processor 402 may be any suitable processor, processing unit or microprocessor. Although not shown in Fig. 4, the computing device 401 may be a multi-processor device and, thus, may include one or more additional processors that are identical or similar to the processor 402 and that are communicatively coupled to the interconnection bus.

[0067] The processor 402 of Fig. 4 may be coupled to a chipset 406, which includes a memory controller 408 and a peripheral input / output (I / O) controller 410. As is well known, a chipset may typically provide I / O and memory management functions as well as a plurality of general purpose and / or special purpose registers, timers, etc. that are accessible or used by one or more processors coupled to the chipset 406. The memory controller 408 may perform functions that enable the processor 402 (or processors if there are multiple processors) to access a system memory 412 and a mass storage memory 414, that may include either or both of an in-memory cache (e.g., a cache within the memory 412) or an on-disk cache (e.g., a cache within the mass storage memory 414).

[0068] The system memory 412 may include any desired type of volatile and / or nonvolatile memory such as, for example, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, read-only memory (ROM), etc. The massstorage memory 414 may include any desired type of mass storage device. For example, the computing device 401 may be used to implement a module 416 (e.g., the various modules as herein described). The mass storage memory 414 may include a hard disk drive, an optical drive, a tape storage device, a solid-state memory (e.g., a flash memory, a RAM memory, etc.), a magnetic memory (e.g., a hard drive), or any other memory suitable for mass storage. As used herein, the terms module, block, function, operation, procedure, routine, step, and method refer to tangible computer program logic or tangible computer executable instructions that provide the specified functionality to the computing device 401, the systems and methods described herein. Thus, a module, block, function, operation, procedure, routine, step, and method can be implemented in hardware, firmware, and / or software.

[0069] In one embodiment, program modules and routines may be stored in mass storage memory 414, loaded into system memory 412, and executed by a processor 402 or may be provided from computer program products that are stored in tangible computer- readable storage mediums (e.g. RAM, hard disk, optical / magnetic media, etc.).

[0070] The peripheral I / O controller 410 may perform functions that enable the processor 402 to communicate with a peripheral input / output (I / O) device 424, a network interface 426, a local network transceiver 428, (via the network interface 426) via a peripheral I / O bus. The I / O device 424 may be any desired type of I / O device such as, for example, a keyboard, a display (e.g., a liquid crystal display (LCD), a cathode ray tube (CRT) display, etc.), a navigation device (e.g., a mouse, a trackball, a capacitive touch pad, a joystick, etc.), etc. The I / O device 424 may be used with the module 416, etc., to receive data from the transceiver 428, send the data to the components of the system 100, and perform any operations related to the methods as described herein. The local network transceiver 428 may include support for a Wi-Fi network, Bluetooth, Infrared, cellular, or other wireless data transmission protocols. In other embodiments, one element may simultaneously support each of the various wireless protocols employed by the computing device 401. For example, a software-defined radio may be able to support multiple protocols via downloadable instructions. In operation, the computing device 401 may be able to periodically poll for visible wireless network transmitters (both cellular and local network) on a periodic basis. Such polling may be possible even while normal wireless traffic is being supported on the computing device 401. The network interface 426 may be, for example, an Ethernet device, an asynchronous transfer mode (ATM) device, an 802.11 wireless interface device, a DSL modem, a cable modem, a cellular modem, etc., that enables the system 100 to communicatewith another computer system having at least the elements described in relation to the system 100.

[0071] While the memory controller 408 and the I / O controller 410 are depicted in Fig. 4 as separate functional blocks within the chipset 406, the functions performed by these blocks may be integrated within a single integrated circuit or may be implemented using two or more separate integrated circuits. The computing environment 400 may also implement the module 416 on a remote computing device 430. The remote computing device 430 may communicate with the computing device 401 over an Ethernet link 432. In some embodiments, the module 416 may be retrieved by the computing device 401 from a cloud computing server 434 via the Internet 436. When using the cloud computing server 434, the retrieved module 416 may be programmatically linked with the computing device 401. The module 416 may be a collection of various software playgrounds including artificial intelligence software and document creation software or may also be a Java® applet executing within a Java® Virtual Machine (JVM) environment resident in the computing device 401 or the remote computing device 430. The module 416 may also be a “plug-in” adapted to execute in a web-browser located on the computing devices 401 and 430. In some embodiments, the module 416 may communicate with back end components 438 via the Internet 436.

[0072] The system 400 may include but is not limited to any combination of a LAN, a MAN, a WAN, a mobile, a wired or wireless network, a private network, or a virtual private network. Moreover, while only one remote computing device 430 is illustrated in Fig. 6 to simplify and clarify the description, it is understood that any number of client computers may be supported and may be in communication within the system 400.

[0073] Additionally, certain embodiments may be described herein as including logic or a number of components, modules, blocks, or mechanisms. Modules and method blocks may constitute either software modules (e.g., code or instructions embodied on a machine- readable medium or in a transmission signal, wherein the code is executed by a processor) or hardware modules. A hardware module may be a tangible unit capable of performing certain operations and may be configured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware module that operates to perform certain operations as described herein.

[0074] In various embodiments, a hardware module may be implemented mechanically or electronically. For example, a hardware module may comprise dedicated circuitry or logic that is permanently configured (e.g., as a special-purpose processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) to perform certain operations. A hardware module may also comprise programmable logic or circuitry (e.g., as encompassed within a processor or other programmable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that the decision to implement a hardware module mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations.

[0075] Accordingly, the term “hardware module” may be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. As used herein, “hardware-implemented module” may refer to a hardware module. Considering embodiments in which hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any one instance in time. For example, where the hardware modules include a processor configured using software, the processor may be configured as respective different hardware modules at different times. Software may accordingly configure a processor, for example, to constitute a particular hardware module at one instance of time and to constitute a different hardware module at a different instance of time.

[0076] Hardware modules may provide information to, and receive information from, other hardware modules. Accordingly, the described hardware modules may be regarded as being communicatively coupled. Where multiple of such hardware modules exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the hardware modules. In embodiments in which multiple hardware modules are configured or instantiated at different times, communications between such hardware modules may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware modules have access. For example, one hardware module may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further hardware module may then, at a later time, access the memory device to retrieve and processthe stored output. Hardware modules may also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information).

[0077] The various operations of example methods described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor- implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example embodiments, comprise processor- implemented modules.

[0078] The methods or routines described herein may be at least partially processor- implemented. For example, at least some of the operations of a method may be performed by one or processors or processor-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processor or processors may be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other embodiments the processors may be distributed across a number of locations.

[0079] The one or more processors may also operate to support performance of the relevant operations in a “cloud computing” environment or as a “software as a service” (SaaS). For example, at least some of the operations may be performed by a group of computers (as examples of machines including processors), these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., application program interfaces (APIs).)

[0080] The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the one or more processors or processor-implemented modules may be located in a single geographic location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the one or more processors or processor-implemented modules may be distributed across a number of geographic locations.

[0081] Some portions of this specification may be presented in terms of algorithms or symbolic representations of operations on data stored as bits or binary digital signals within a machine memory (e.g., a computer memory). These algorithms or symbolic representations may be examples of techniques used by those of ordinary skill in the data processing arts toconvey the substance of their work to others skilled in the art. As used herein, an “algorithm” may be a self-consi stent sequence of operations or similar processing leading to a desired result. In this context, algorithms and operations may involve physical manipulation of physical quantities. Typically, but not necessarily, such quantities may take the form of electrical, magnetic, or optical signals capable of being stored, accessed, transferred, combined, compared, or otherwise manipulated by a machine. It is convenient at times, principally for reasons of common usage, to refer to such signals using words such as “data,” “content,” “bits,” “values,” “elements,” “symbols,” “characters,” “terms,” “numbers,” “numerals,” or the like. These words, however, may be merely convenient labels and are to be associated with appropriate physical quantities.

[0082] Unless specifically stated otherwise, discussions herein using words such as “processing,” “computing,” “calculating,” “determining,” “presenting,” “displaying,” or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0083] As used herein any reference to “embodiments,” “some embodiments” or “an embodiment” or “teaching” may mean that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase “in some embodiments” or “teachings” in various places in the specification may not necessarily all be referring to the same embodiment.

[0084] Some embodiments may be described using the expression “coupled” and “connected” along with their derivatives. For example, some embodiments may be described using the term “coupled” to indicate that two or more elements are in direct physical or electrical contact. The term “coupled,” however, may also mean that two or more elements are not in direct contact with each other, but yet still co-operate or interact with each other. The embodiments may not be limited in this context.

[0085] Further, the figures depict preferred embodiments for purposes of illustration only. One skilled in the art may be readily recognize from the following discussion that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein.

[0086] Upon reading this disclosure, those of skill in the art may appreciate still additional alternative structural and functional designs for the systems and methods described herein through the disclosed principles herein. Thus, while particular embodiments and applications have been illustrated and described, it is to be understood that the disclosed embodiments may not be limited to the precise construction and components disclosed herein. Various modifications, changes and variations, which may be apparent to those skilled in the art, may be made in the arrangement, operation and details of the systems and methods disclosed herein without departing from the spirit and scope defined in any appended claims.

Claims

Claims1. A method of analyzing indications of physical attributes for thought indications and patterns to be analyzed comprising: identifying features from a physical sensor; selecting a feature from the physical sensor;) translating the feature using a transformation algorithm into feature data at a plurality of points in time; storing the feature data; comparing the feature data to previous feature data using a comparison algorithm; scoring the comparison using a scoring algorithm; determining if the score is over a threshold; and in response to the score being over a threshold, classifying the results as thoughts.

2. The method of claim 1, wherein the physical sensor comprises at least one from a group comprising: an image sensor; a voice sensor; a scent sensor; a touch sensor; a taste sensor; a magnetic sensor; a temperature sensor.

3. The method of claim 2, further comprising if the physical sensor is an image sensor, a scoring mechanism to determine when the point has changed significantly or remained relatively stable between the two time points.

4. The method of claim 2 wherein if the physical sensor is an image sensor, the video transformation algorithm is a Fast Fourier Transform (FFT) algorithm.

5. The method of claim 2, wherein if the physical sensor is an image sensor, the feature comprises at least one of the group comprising: edges, colors, shapes, or key points.

6. The method of claim 2, further comprising if the physical sensor is an image sensor, analyzing a significant amount of videos to create a database of patterns or feature representations from those videos.

7. The method of claim 2, wherein if the physical sensor is an image sensor, a comparison algorithm compares the video to videos in the video database8. The method of claim 7, wherein the comparison algorithm comprises a similarity metric comprising at least one of Euclidean distance, cosine similarity, or dynamic time warping for temporal patterns.

9. The method of claim 2, wherein if the physical sensor is an image sensor, a machine learning model is trained on a model on an existing dataset and using the model to predict whether the movement in the new video matches known patterns.

10. The method of claim 9, wherein the machine learning model is used to predict future actions based on past pattern recognition.

11. The method of claim 9, wherein the machine learning model is used to determine pattern deviation and wherein the pattern deviation is used to determine future actions.

12. The method of claim 4, further comprising determining a similarity metric to compare the FFT results.

13. The method of claim 12, further comprising using the Euclidean distance between FFT results.

14. The method of claim 13, further comprising: scoring the similarity; and determining if the similarity is over a threshold.

15. The method of claim 2, further comprising if the physical sensor is an image sensor,: analyzing the pixel intensity values at each time point as a time series;performing FFT on both time series to convert them into the frequency domain which will result in the amplitude and phase information of the frequencies present in the intensity variations at that point on the face.

Citation Information

Patent Citations

  • Adaptive brain training computer system and method

    US20150351655A1

  • Image Processing System for Extracting a Behavioral Profile from Images of an Individual Specific to an Event

    US20210326586A1

  • Mental state monitoring system

    US20220071535A1