Method and system for filtering text display
Patent Information
- Application Number
- PCT/EP2026/058940
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026058940_01102026_PF_FP_ABST
Abstract
Description
Method and System for Filtering Text DisplayTechnical Field
[0001] The invention is directed to a method, device, system and computer readable programmable medium for identifying an intent in a real time conversation.Background Art
[0002] Identifying intents of certain types of behaviour in online conversations is a task of increasing importance when a considerable number of applications for computing devices (such as WhatsApp®, TikTok®, Instagram®, Facebook® or Telegram®) are used. Such intents may be indicative of cyberbullying, self-harm or online grooming, among other types of undesirable behaviour. This is of particular concern especially where applications are accessible to minors.
[0003] However, identifying intents can be a challenging task for the applications in question and for the devices in which they are installed. For example, some solutions can be based on evaluating the intent of a current message without considering the context of this message (such as the content of previous messages in the conversation). These solutions may result in an inaccurate identification of a certain behaviour. Other approaches may incorporate previous messages in the conversation into a large language model. However, the length that those large language models require may lead to high computational costs. Moreover, handling an increased number of messages that may belong to different ongoing conversations may raise concerns related to privacy and cyber-security. It would therefore be desirable to provide an intent identification method and device that outputs their results in a more reliable and efficient manner.Summary of invention
[0004] The invention is described herein with reference to the appended claims.
[0005] A method for identifying an intent in a real time conversation performed by a computing device is provided. The method comprises:inputting, at a language model, LM, a first text message;inputting, at the language model, a conversational state history, CSH, vector, wherein the CSH vector comprises an array of numbers and wherein the CSH vector is representative of text messages received at the LM prior to the first text message;outputting, by the language model, a measurement of the likelihood the first text message comprises an intent.
[0006] Since the method comprises inputting a CSH vector that is representative of messages received at LM prior to a first message, the LM is able to output a measurement of the likelihood the first text message comprises an intent that not only takes into account the content of the first text message, but also that of the previous messages. Put another way, the present method, by means of the CSH vector, considers representation of the previous messages to assist in evaluating the intent of the current message. Otherwise, the current message would have to be evaluated without any previous context and, therefore, the ability to predict the intent in the current message may fail in most cases.
[0007] The first text message and the text messages received at the LM prior to the first text message may belong to a conversation. As used herein, a "conversation" should be construed as a chat (such as a chat on an application) established between two or more people. In an example, the start of a conversation is the first message among the people taking part in the conversation. In another example, a conversation may be defined based on the time between a message and its response against an average response time to a message. Applying the method to a certain conversation may be advantageous to determine or estimate the context of the conversation, relying on the CSH of this conversation, and outputting the measurement of the likelihood the first text message comprises an intent according to such context.
[0008] The method may further comprise updating the CSH vector with the first text message.Hence, the CSH vector may adapt the context, such as the context of a conversation or the context of any other group of messages, to include the content of the first text message. This way, the updated CSH vector may be more efficiently used as input together with messages received at the LM subsequently to the first text message.
[0009] The method may further comprise retaining a copy of the CSH vector in a vector database of the computing device. Since each CSH vector may be representative of a conversation at a certain time period, storing a copy of the CSH vector may allow that there is a representation of each conversation on the computing device. Moreover, since the vector database is on the computing device, there may be no need for computational efforts from other devices, such as an external effort. Put another way, the measurement of the likelihood the first text message comprises an intent can be estimated locally at thecomputing device, which may minimise the need for server-side storage to maintain the context of each user.
[0010] The array of numbers may have a predetermined length. An update of the CSH vector may replace at least a previous number of the array of numbers with an updated number of the array of numbers. The predetermined length may be advantageously adjusted such that, the larger the length is, the more conversational and context information can be stored, but the more computational effort will be required to output the measurement. In an example, the update of the array of numbers may be carried out as follows: all_numbers_of_updated_array = all_numbers_of_old_array + updating_rate * all_numbers_of_a_new_information_array. The updating_rate may be a floating number between 0 and 1.0. The new information array may be defined as an embedding vector of the first text message.
[0011] The array of numbers may be stored in a recurrent architecture, such as a long short-term memory, LSTM, or a Recurring Memory Transformer. This may be an efficient manner of creating vectors representative of previous text messages. For example, the array of numbers may represent a cell state of the recurrent architecture.
[0012] LSTM may rely on one or more of an input gate, an output gate or a forget gate, for instance by means of a BERT, Bidirectional Encoder Representations from Transformers. These gates may be relied upon to control how much new information should be integrated into the CSH vector and how much old information will be discarded and / or retained. In an example, the CSH vector may retain all the text information from the start of a conversation using this technique. Accordingly, the CSH vector may be used as a conversational context for classification of the first text message.
[0013] The method may further comprise sending, to an external server, the CSH vector, wherein the CSH vector is updated with the first text message and wherein the CSH vector is representative of text messages received at the LM up to the first text message; receiving, at the LM and after the first text message, a second text message; sending, to the external server, the second text message; and receiving, from the external server, a server measurement of the likelihood the second text message comprises an intent.
[0014] In some scenarios, the computing device may not be able to carry out the operations to output a further measurement of the likelihood a text message, such as the second text message, comprises an intent. For example, this may happen when the computing devicehas low battery. A method in which the external server is used to obtain a server measurement of the likelihood the second text message comprises an intent may be advantageous in such scenarios.
[0015] The method may further comprise establishing a session between the computing device and the external server. Typically, the context of a conversation may not change abruptly in a short time window. Therefore, even in situations in which it may be useful to rely on the external server to provide a new measurement of likelihood a text message comprises an intent, sending information (such as the text message, the CSH vector or the measurement) back and forth between the computing device and the server for each message may be redundant. A session between the server and the computing device may thus be set up at specific periods. This may reduce the burden of the network transmission.
[0016] The intent may comprise one or more of: cyberbullying intent, self-harm intent or online grooming intent. Those are categories of particular concern in which intent identification may be particularly beneficial.
[0017] The computing device may be a user equipment, UE. The UE may be a mobile UE. The mobile UE may be a mobile phone, a tablet, a laptop computer, an intelligent watch, a gaming console or any other UE.
[0018] The LM may be a small LM, SLM.
[0019] The first text message may be associated with a chat application installed in the computing device.
[0020] There is also provided a computer readable programmable medium carrying a computer programme stored thereon which, when executed by a processor, implements any of the methods disclosed herein. The computer readable programmable medium may be embodied, for instance, on a record medium, carrier signal or read-only memory.
[0021] There is also provided a computing device comprising:a memory; andone or more processors operatively coupled to the memory, the one or more processors configured to:input, at a language model, LM, a first text message;input, at the LM, a conversational state history, CSH, vector, wherein the CSH vector comprises an array of numbers and wherein the CSH vector is representative of text messages received at the LM prior to the first text message;output, by the language model, a measurement of the likelihood the first text message comprises a harmful intent.
[0022] The one or more processors, when comparing the inbound message with the set of content stored in the computing device, may be further configured to: send, to an external server, the CSH vector, wherein the CSH vector is updated with the first text message and wherein the CSH vector is representative of text messages received at the LM up to the first text message; receive, at the LM and after the first text message, a second text message;send, to the external server, the second text message; and receive, from the external server, a server measurement of the likelihood the second text message comprises an intent.
[0023] The computing devices disclosed herein may be advantageous for the same reasons as set forth for or derivable from the methods performed by computing devices disclosed herein. Likewise, the one or more processors of the computing device may be configured to carry out any of the actions disclosed herein for the methods performed by computing devices.
[0024] There is also provided a computing system comprising any of the computing devices disclosed herein and an external server, wherein the external server comprises:a server memory; andone or more server processors operatively coupled to the server memory, the one or more server processors configured to:upon receiving, by the external server, the CSH vector and the second text message, output, by a server LM, the server measurement of the likelihood the second text message comprises a harmful intent.
[0025] The computing systems disclosed herein may be advantageous for the same reasons as set forth for or derivable from the methods disclosed herein. Likewise, the one or more processors and unit processors of the computing system may be configured to carry out any of the actions disclosed herein for the methods performed by the computing system.
[0026] It is appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. It should also be appreciated that terminology explicitly employed herein that also may appear in anydisclosure incorporated by reference should be accorded a meaning consistent with the particular concepts disclosed herein.Detailed description
[0027] Aspects of the present invention and certain features, advantages, and details thereof are explained more fully below with reference to the non-limiting examples illustrated in the accompanying drawings. Descriptions of well-known processing techniques, systems, components, etc. are omitted so as to not unnecessarily obscure the invention in detail. It should be understood that the detailed description and the specific examples, while indicating aspects of the invention, are given by way of illustration only, and not by way of limitation. Various substitutions, modifications, additions, and / or arrangements, within the spirit and / or scope of the underlying inventive concepts will be apparent to those skilled in the art from this disclosure. Note further that numerous inventive aspects and features are disclosed herein, and unless inconsistent, each disclosed aspect or feature is combinable with any other disclosed aspect or feature as desired for a particular embodiment of the concepts disclosed herein.
[0028] Unless described or implied as exclusive alternatives, features throughout the drawings and descriptions should be taken as cumulative, such that features expressly associated with some particular embodiments can be combined with other embodiments.
[0029] While certain exemplary embodiments have been described and shown in the accompanying drawings, it is to be understood that such embodiments are merely illustrative of, and not restrictive on, the broad invention, and that this invention not be limited to the specific constructions and arrangements shown and described, since various other changes, combinations, omissions, modifications and substitutions, in addition to those set forth in the above paragraphs, are possible. Those skilled in the art will appreciate that various adaptations, modifications, and combinations of the herein described embodiments can be configured without departing from the scope and spirit of the invention. Therefore, it is to be understood that, within the scope of the included claims, the invention may be practiced other than as specifically described herein.
[0030] Additionally, illustrative embodiments are described below using specific code, designs, architectures, protocols, layouts, schematics, or tools only as examples, and not by way of limitation. Furthermore, the illustrative embodiments are described in certain instancesusing particular software, tools, or data processing environments only as example for clarity of description. The illustrative embodiments can be used in conjunction with other comparable or similarly purposed structures, systems, applications, or architectures. One or more aspects of an illustrative embodiment can be implemented in hardware, software, or a combination thereof.
[0031] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of computer-implemented methods and computing systems according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions that may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus (the term "apparatus" includes systems and computer program products). The processor may execute the computer readable program instructions thereby creating a means for implementing the actions specified in the flowchart illustrations and / or block diagrams. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the actions specified in the flowchart illustrations and / or block diagrams. In particular, the computer readable program instructions may be used to produce a computer-implemented method by executing the instructions to implement the actions specified in the flowchart illustrations and / or block diagrams.
[0032] The computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions, which implement the function / act specified in the flowchart and / or block diagram block or blocks.
[0033] The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions, which execute on the computer or otherprogrammable apparatus, provide steps for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. Alternatively, computer program implemented steps or acts may be combined with operator or human implemented steps or acts in order to carry out an embodiment of the invention.
[0034] In the flowchart illustrations and / or block diagrams disclosed herein, each block in the flowchart / diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.
[0035] Computer program instructions are configured to carry out operations of the present invention and may be or may incorporate assembler instructions, instruction-set- architecture ("ISA") instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, source code, and / or object code written in any combination of one or more programming languages.
[0036] An application program may be deployed by providing computer infrastructure operable to perform one or more embodiments disclosed herein by integrating computer readable code into a computing system thereby performing the computer-implemented methods disclosed herein. Although various computing environments are described above, these are only examples that can be used to incorporate and use one or more embodiments. Many variations are possible.
[0037] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprise" (and any form of comprise, such as "comprises" and "comprising"), "have" (and any form of have, such as "has" and "having"), "include" (and any form of include, such as "includes" and "including"), and "contain" (and any form contain, such as "contains" and "containing") are open-ended linking verbs. As a result, a method or device that "comprises," "has," "includes" or "contains" one or more steps or elements possesses those one or more steps or elements, but is not limited to possessing only those one or more steps or elements. Likewise, a stepof a method or an element of a device that "comprises," "has," "includes" or "contains" one or more features possesses those one or more features, but is not limited to possessing only those one or more features. Furthermore, a device or structure that is configured in a certain way is configured in at least that way, but may also be configured in ways that are not listed.
[0038] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below, if any, are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiment was chosen and described in order to best explain the principles of one or more aspects of the invention and the practical application, and to enable others of ordinary skill in the art to understand one or more aspects of the invention for various embodiments with various modifications as are suited to the particular use contemplated.
[0039] It will be understood that relative terms are intended to encompass different orientations or sequences in addition to the orientations and sequences depicted in the drawings and described herein. Relative terminology, such as "substantially" or "about," describe the specified devices, materials, transmissions, steps, parameters, or ranges as well as those that do not materially affect the basic and novel characteristics of the claimed inventions as whole (as would be appreciated by one of ordinary skill in the art).
[0040] The terms "coupled," "fixed," "attached to," "communicatively coupled to," "operatively coupled to," and the like refer to both: (i) direct connecting, coupling, fixing, attaching, communicatively coupling; and (ii) indirect connecting coupling, fixing, attaching, communicatively coupling via one or more intermediate components or features, unless otherwise specified herein. "Communicatively coupled to" and "operatively coupled to" can refer to physically and / or electrically related components.Brief description of drawings
[0041] These and other features and advantages of the invention will become more evident in the light of the following detailed description of preferred embodiments, given only by way of illustrative and non-limiting example, in reference to the attached figures:
[0042] Figure 1 describes a computing system comprising a computing device and an external server.
[0043] Figure 2A is a diagram of a feedforward network, according to at least one embodiment, utilized in machine learning.
[0044] Figure 2B is a diagram of a convolution neural network, according to at least one embodiment, utilized in machine learning.
[0045] Figure 2C is a diagram of a portion of the convolution neural network of FIG. 2B, according to at least one embodiment, illustrating assigned weights at connections or neurons.
[0046] Figure 3 is a diagram representing an exemplary weighted sum computation in a node in an artificial neural network.
[0047] Figure 4 is a diagram of a Recurrent Neural Network RNN, according to at least one embodiment, utilized in machine learning.
[0048] Figure 5 is a schematic logic diagram of an artificial intelligence program including a frontend and a back-end algorithm.
[0049] Figure 6 is a flow chart representing a method, according to at least one embodiment, of model development and deployment by machine learning.
[0050] Figure 7 includes a flow-chart illustrating an example method performed by a computing device according to this disclosure.Description of embodiments
[0051] In the embodiment of figure 1, a computing system 1 comprising a computing device 10 and an external server 20 is shown. The computing device 10 comprises a memory 11 and a processor 12 operatively coupled to the memory. The processor 12 is configured to perform any of the methods disclosed herein in connection with the computing device, such as the methods of figure 7. For example, the processor 12 may be configured to perform the actions of elements 102 to 118 of figure 7. The memory 11 may be configured to store a vector database. The processor 12 of figure 1 is configured to run a language model, LM. Afirst text message may be input into the LM. Likewise, a conversational state history, CSH, vector may also be input into the LM. Given that the CSH vector is representative of text messages received at the LM prior to the first text message, the LM may derive, for example, the context of a conversation that the first text message is part of. Accordingly, the LM may output a measurement of the likelihood the first text message comprises an intent which may be more accurate than message than estimation of intents that do not consider the context of the analysed text message.
[0052] The server 20 of the computing system 1 of figure 1 is external to the computing device 10 and comprises a server memory 21 and a server processor 22 operatively coupled to the server memory 21. When the computing device is able (or estimates that it is not advisable) to carry out the operations to output a further measurement of the likelihood a text message, such as the second text message, comprises an intent (for example, because the battery of the computing device is low), the CSH vector updated with the first text message may be sent to the external server 20. The external server may also receive a second text message. The server processor 22, upon receiving the CSH vector and the second text message, may output a server measurement of the likelihood the second text message comprises a harmful intent. The external server 20 may then send, to the computing device 10, the server measurement.
[0053] The processor 12 of the embodiment of figure 1 is configured to display the text messages, for example when it is estimated that they do not comprise an intent, in a user interface (Ul) screen 13.
[0054] The processor 12, the server processor 22, and any other processors described herein (which are comprised in the one or more processors disclosed herein), generally include circuitry for implementing communication and / or logic functions of the computing device, the computing system or the external server. For example, the one or more processors may include a digital signal processor, a microprocessor and various analog to digital converters, digital to analog converters and / or other support circuits. Control and signal processing functions of the computing device, computing system and external server are allocated between these devices according to their respective capabilities. The one or more processors thus may also include the functionality to encode and interleave messages and data prior to modulation and transmission. The one or more processors can additionally include an internal data modem. Further, the one or more processors may includefunctionality to operate one or more software programs, which may be stored in the memory 11 or server memory 21. For example, the one or more processors may be capable of operating a connectivity program, such as a web browser application. The web browser application may then allow the computing device to transmit and receive web content, such as, for example, location-based content and / or other web page content, according to a Wireless Application Protocol (WAP), Hypertext Transfer Protocol (HTTP), and / or the like.
[0055] The memory 11, server memory 21 or any other storage unit can each also store any of a number of pieces of information, and data, used by the computing device or the external server and the applications and devices that facilitate functions of the computing device or external server, or are in communication with them, to implement the functions described herein and others not expressly described. For example, the memory 11 and the server memory 21 may include such data as user authentication information, etc.
[0056] The one or more processors, in various examples, can operatively perform calculations, can process instructions for execution, and can manipulate information. The one or more processors can execute machine-executable instructions stored in the memory 11 or the server memory 21 to thereby perform methods and functions as described or implied herein, for example by one or more corresponding flow charts expressly provided or implied as would be understood by one of ordinary skill in the art to which the subject matters of these descriptions pertain. The one or more processors can be or can include, as nonlimiting examples, a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a digital signal processor (DSP), a field programmable gate array (FPGA), a state machine, a controller, gated or transistor logic, discrete physical hardware components and combinations thereof. In some embodiments, particular portions or steps of methods and functions described herein are performed in whole or in part by way of the one or more processors, while in other embodiments methods and functions described herein include cloud-based computing in whole or in part such that the one or more processors facilitates local operations including, as non-limiting examples, communication, data transfer, and user inputs and outputs such as receiving commands from and providing displays to the user.Artificial Intelligence Technology
[0057] As used herein, an artificial intelligence system, artificial intelligence algorithm, artificial intelligence module, program, and the like, generally refer to computer implemented programs that are suitable to simulate intelligent behavior (i.e., intelligent human behavior) and / or computer systems and associated programs suitable to perform tasks that typically require a human to perform, such as tasks requiring visual perception, speech recognition, decision-making, translation, and the like. An artificial intelligence system may include, for example, at least one of a series of associated if-then logic statements, a statistical model suitable to map raw sensory data into symbolic categories and the like, or a machine learning program. A machine learning program, machine learning algorithm, or machine learning module, as used herein, is generally a type of artificial intelligence including one or more algorithms that can learn and / or adjust parameters based on input data provided to the algorithm. In some instances, machine learning programs, algorithms, and modules are used at least in part in implementing artificial intelligence ("Al") functions, systems, and methods.
[0058] Artificial Intelligence and / or machine learning programs may be associated with or conducted by one or more processors, memory devices, and / or storage devices of a computing system or device. It should be appreciated that the Al algorithm or program may be incorporated within the existing system architecture or be configured as a standalone modular component, controller, or the like communicatively coupled to the system. An Al program and / or machine learning program may generally be configured to perform methods and functions as described or implied herein, for example by one or more corresponding flow charts expressly provided or implied as would be understood by one of ordinary skill in the art to which the subjects matters of these descriptions pertain.
[0059] A machine learning program may be configured to use various analytical tools (e.g., algorithmic applications) to leverage data to make predictions or decisions. Machine learning programs may be configured to implement various algorithmic processes and learning approaches including, for example, decision tree learning, association rule learning, artificial neural networks, recurrent artificial neural networks, long short term memory networks, inductive logic programming, support vector machines, clustering, Bayesian networks, reinforcement learning, representation learning, similarity and metric learning, sparse dictionary learning, genetic algorithms, k-nearest neighbor ("KNN"), and the like. In someembodiments, the machine learning algorithm may include one or more image recognition algorithms suitable to determine one or more categories to which an input, such as data communicated from a visual sensor or a file in JPEG, PNG or other format, representing an image or portion thereof, belongs. Additionally or alternatively, the machine learning algorithm may include one or more regression algorithms configured to output a numerical value given an input. Further, the machine learning may include one or more pattern recognition algorithms, e.g., a module, subroutine or the like capable of translating text or string characters and / or a speech recognition module or subroutine. In various embodiments, the machine learning module may include a machine learning acceleration logic, e.g., a fixed function matrix multiplication logic, in order to implement the stored processes and / or optimize the machine learning logic training and interface.
[0060] Machine learning models are trained using various data inputs and techniques. Example training methods may include, for example, supervised learning, (e.g., decision tree learning, support vector machines, similarity and metric learning, etc.), unsupervised learning, (e.g., association rule learning, clustering, etc.), reinforcement learning, semi-supervised learning, self-supervised learning, multi-instance learning, inductive learning, deductive inference, transductive learning, sparse dictionary learning and the like. Example clustering algorithms used in unsupervised learning may include, for example, k-means clustering, density based special clustering of applications with noise ("DBSCAN"), mean shift clustering, expectation maximization ("EM") clustering using Gaussian mixture models ("GMM"), agglomerative hierarchical clustering, or the like. According to one embodiment, clustering of data may be performed using a cluster model to group data points based on certain similarities using unlabeled data. Example cluster models may include, for example, connectivity models, centroid models, distribution models, density models, group models, graph based models, neural models and the like.
[0061] One subfield of machine learning includes neural networks, which take inspiration from biological neural networks. In machine learning, a neural network includes interconnected units that process information by responding to external inputs to find connections and derive meaning from undefined data. A neural network can, in a sense, learn to perform tasks by interpreting numerical patterns that take the shape of vectors and by categorizing data based on similarities, without being programmed with any task-specific rules. A neural network generally includes connected units, neurons, or nodes (e.g., connected by synapses) and mayallow for the machine learning program to improve performance. A neural network may define a network of functions, which have a graphical relationship. Various neural networks that implement machine learning exist including, for example, feedforward artificial neural networks, perceptron and multilayer perceptron neural networks, radial basis function artificial neural networks, recurrent artificial neural networks, modular neural networks, long short term memory networks, as well as various other neural networks.
[0062] Neural networks may perform a supervised learning process where known inputs and known outputs are utilized to categorize, classify, or predict a quality of a future input. However, additional or alternative embodiments of the machine learning program may be trained utilizing unsupervised or semi-supervised training, where none of the outputs or some of the outputs are unknown, respectively. Typically, a machine learning algorithm is trained (e.g., utilizing a training data set) prior to modeling the problem with which the algorithm is associated. Supervised training of the neural network may include choosing a network topology suitable for the problem being modeled by the network and providing a set of training data representative of the problem. Generally, the machine learning algorithm may adjust the weight coefficients until any error in the output data generated by the algorithm is less than a predetermined, acceptable level. For instance, the training process may include comparing the generated output produced by the network in response to the training data with a desired or correct output. An associated erroramount may then be determined forthe generated output data, such as for each output data point generated in the output layer. The associated error amount may be communicated back through the system as an error signal, where the weight coefficients assigned in the hidden layer are adjusted based on the error signal. For instance, the associated error amount (e.g., a value between -1 and 1) may be used to modify the previous coefficient, e.g., a propagated value. The machine learning algorithm may be considered sufficiently trained when the associated error amount for the output data is less than the predetermined, acceptable level (e.g., each data point within the output layer includes an error amount less than the predetermined, acceptable level). Thus, the parameters determined from the training process can be utilized with new input data to categorize, classify, and / or predict other values based on the new input data.
[0063] An artificial neural network ("ANN"), also known as a feedforward network, may be utilized, e.g., an acyclic graph with nodes arranged in layers. A feedforward network (see, e.g., feedforward network 260 referenced in FIG. 2A) may include a topography with a hidden layer264 between an input layer 262 and an output layer 266. The input layer 262, having nodes commonly referenced in FIG. 2A as input nodes 272 for convenience, communicates input data, variables, matrices, or the like to the hidden layer 264, having nodes 274. The hidden layer 264 generates a representation and / or transformation of the input data into a form that is suitable for generating output data. Adjacent layers of the topography are connected at the edges of the nodes of the respective layers, but nodes within a layer typically are not separated by an edge. In at least one embodiment of such a feedforward network, data is communicated to the nodes 272 of the input layer, which then communicates the data to the hidden layer 264. The hidden layer 264 may be configured to determine the state of the nodes in the respective layers and assign weight coefficients or parameters of the nodes based on the edges separating each of the layers, e.g., an activation function implemented between the input data communicated from the input layer 262 and the output data communicated to the nodes 276 of the output layer 266. It should be appreciated that the form of the output from the neural network may generally depend on the type of model represented by the algorithm. Although the feedforward network 260 of FIG. 2A expressly includes a single hidden layer 264, other embodiments of feedforward networks within the scope of the descriptions can include any number of hidden layers. The hidden layers are intermediate the input and output layers and are generally where all or most of the computation is done.
[0064] An additional or alternative type of neural network suitable for use in the machine learning program and / or module is a Convolutional Neural Network ("CNN"). A CNN is a type of feedforward neural network that may be utilized to model data associated with input data having a grid-like topology. In some embodiments, at least one layer of a CNN may include a sparsely connected layer, in which each output of a first hidden layer does not interact with each input of the next hidden layer. For example, the output of the convolution in the first hidden layer may be an input of the next hidden layer, rather than a respective state of each node of the first layer. CNNs are typically trained for pattern recognition, such as speech processing, language processing, and visual processing. As such, CNNs may be particularly useful for implementing optical and pattern recognition programs required from the machine learning program. A CNN includes an input layer, a hidden layer, and an output layer, typical of feedforward networks, but the nodes of a CNN input layer are generally organized into a set of categories via feature detectors and based on the receptive fields of the sensor, retina, input layer, etc. Each filter may then output data from its respective nodes to correspondingnodes of a subsequent layer of the network. A CNN may be configured to apply the convolution mathematical operation to the respective nodes of each filter and communicate the same to the corresponding node of the next subsequent layer. As an example, the input to the convolution layer may be a multidimensional array of data. The convolution layer, or hidden layer, may be a multidimensional array of parameters determined while training the model.
[0065] An exemplary convolutional neural network CNN is depicted and referenced as 280 in FIG. 2B.As in the basic feedforward network 260 of FIG. 2A, the illustrated example of FIG. 2B has an input layer 282 and an output layer 286. However where a single hidden layer 264 is represented in FIG. 2A, multiple consecutive hidden layers 284A, 284B, and 284C are represented in FIG. 2B. The edge neurons represented by white-filled arrows highlight that hidden layer nodes can be connected locally, such that not all nodes of succeeding layers are connected by neurons. FIG. 2C, representing a portion of the convolutional neural network 280 of FIG. 2B, specifically portions of the input layer 282 and the first hidden layer 284A, illustrates that connections can be weighted. In the illustrated example, labels W1 and W2 refer to respective assigned weights for the referenced connections. Two hidden nodes 283 and 285 share the same set of weights W1 and W2 when connecting to two local patches.
[0066] Weight defines the impact a node in any given layer has on computations by a connected node in the next layer. FIG. 3 represents a particular node 300 in a hidden layer. The node 300 is connected to several nodes in the previous layer representing inputs to the node 300. The input nodes 301, 302, 303 and 304 are each assigned a respective weight W01, W02, W03, and W04 in the computation at the node 300, which in this example is a weighted sum.
[0067] An additional or alternative type of feedforward neural network suitable for use in the machine learning program and / or module is a Recurrent Neural Network ("RNN"). An RNN may allow for analysis of sequences of inputs rather than only considering the current input data set. RNNs typically include feedback loops / connections between layers of the topography, thus allowing parameter data to be communicated between different parts of the neural network. RNNs typically have an architecture including cycles, where past values of a parameter influence the current calculation of the parameter, e.g., at least a portion of the output data from the RNN may be used as feedback / input in calculating subsequent output data. In some embodiments, the machine learning module may include an RNN configured for language processing, e.g., an RNN configured to perform statistical languagemodeling to predict the next word in a string based on the previous words. The RNN(s) of the machine learning program may include a feedback system suitable to provide the connection(s) between subsequent and previous layers of the network.
[0068] An example for a Recurrent Neural Network RNN is referenced as 400 in FIG. 4. As in the basic feedforward network 260 of FIG. 2A, the illustrated example of FIG. 4 has an input layer 410 (with nodes 412) and an output layer 440 (with nodes 442). However, where a single hidden layer 264 is represented in FIG. 2A, multiple consecutive hidden layers 420 and 430 are represented in FIG. 4 (with nodes 422 and nodes 432, respectively). As shown, the RNN 400 includes a feedback connector 404 configured to communicate parameter data from at least one node 432 from the second hidden layer 430 to at least one node 422 of the first hidden layer 420. It should be appreciated that two or more and upto all of the nodesofa subsequent layer may provide or communicate a parameter or other data to a previous layer of the RNN 400. Moreover and in some embodiments, the RNN 400 may include multiple feedback connectors 404 (e.g., connectors 404 suitable to communicatively couple pairs of nodes and / or connector systems 404 configured to provide communication between three or more nodes). Additionally or alternatively, the feedback connector 404 may communicatively couple two or more nodes having at least one hidden layer between them, i.e., nodes of nonsequential layers of the RNN 400.
[0069] In an additional or alternative embodiment, the machine-learning program may include one or more support vector machines. A support vector machine may be configured to determine a category to which input data belongs. For example, the machine-learning program may be configured to define a margin using a combination of two or more of the input variables and / or data points as support vectors to maximize the determined margin. Such a margin may generally correspond to a distance between the closest vectors that are classified differently. The machine-learning program may be configured to utilize a plurality of support vector machines to perform a single classification. For example, the machine-learning program may determine the category to which input data belongs using a first support vector determined from first and second data points / variables, and the machine-learning program may independently categorize the input data using a second support vector determined from third and fourth data points / variables. The support vector machine(s) may be trained similarly to the training of neural networks, e.g., by providing a known input vector (including values for the input variables) and a known output classification. The support vector machine is trainedby selecting the support vectors and / or a portion of the input vectors that maximize the determined margin.
[0070] As depicted, and in some embodiments, the machine-learning program may include a neural network topography having more than one hidden layer. In such embodiments, one or more of the hidden layers may have a different number of nodes and / or the connections defined between layers. In some embodiments, each hidden layer may be configured to perform a different function. As an example, a first layer of the neural network may be configured to reduce a dimensionality of the input data, and a second layer of the neural network may be configured to perform statistical programs on the data communicated from the first layer. In various embodiments, each node of the previous layer of the network may be connected to an associated node of the subsequent layer (dense layers). Generally, the neural network(s) of the machine-learning program may include a relatively large number of layers, e.g., three or more layers, and may be referred to as deep neural networks. For example, the node of each hidden layer of a neural network may be associated with an activation function utilized by the machine-learning program to generate an output received by a corresponding node in the subsequent layer. The last hidden layer of the neural network communicates a data set (e.g., the result of data processed within the respective layer) to the output layer. Deep neural networks may require more computational time and power to train, but the additional hidden layers provide multistep pattern recognition capability and / or reduced output error relative to simple or shallow machine learning architectures (e.g., including only one or two hidden layers).
[0071] According to various implementations, deep neural networks incorporate neurons, synapses, weights, biases, and functions and can be trained to model complex non-linear relationships. Various deep learning frameworks may include, for example, TensorFlow, MxNet, PyTorch, Keras, Gluon, and the like. Training a deep neural network may include complex input / output transformations and may include, according to various embodiments, a backpropagation algorithm. According to various embodiments, deep neural networks may be configured to classify images of handwritten digits from a dataset or various other images. According to various embodiments, the datasets may include a collection of files that are unstructured and lack predefined data model schema or organization. Unlike structured data, which is usually stored in a relational database ("RDBMS") and can be mapped into designated fields, unstructured data comes in many formats that can be challenging to process and analyze.Examples of unstructured data may include, according to non-limiting examples, dates, numbers, facts, emails, text files, scientific data, satellite imagery, media files, social media data, text messages, mobile communication data, and the like.
[0072] Referring now to FIG. 5 and some embodiments, an Al program 502 may include a front-end algorithm 504 and a back-end algorithm 506. The artificial intelligence program 502 may be implemented on an Al processor 520, such as the processing device 120, the processing device 220, and / or a dedicated processing device. The instructions associated with the front-end algorithm 504 and the back-end algorithm 506 may be stored in an associated memory device and / or storage device of the system (e.g., storage device 124, memory device 122, storage device 224, and / or memory device 222) communicatively coupled to the Al processor 520, as shown. Additionally or alternatively, the system may include one or more memory devices and / or storage devices (represented by memory 524 in FIG. 5) for processing use and / or including one or more instructions necessary for operation of the Al program 502. In some embodiments, the Al program 502 may include a deep neural network (e.g., a front-end algorithm 504 configured to perform pre-processing, such as feature recognition, and a back- end algorithm 506 configured to perform an operation on the data set communicated directly or indirectly to the back-end algorithm 506 that together form the deep neural network). For instance, the front-end algorithm 504 can include at least one CNN 508 communicatively coupled to send output data to the back-end algorithm 506.
[0073] Additionally or alternatively, the front-end algorithm 504 can include one or more Al algorithms 510, 512 (e.g., statistical models or machine learning programs such as decision tree learning, associate rule learning, recurrent artificial neural networks, support vector machines, and the like). In various embodiments, the front-end algorithm 504 may be configured to include built in training and inference logic or suitable software to train the neural network prior to use (e.g., machine learning logic including, but not limited to, image recognition, mapping and localization, autonomous navigation, speech synthesis, document imaging, or language translation such as natural language processing). For example, a CNN 508 and / or Al algorithm 510 may be used for image recognition, input categorization, and / or support vector training. In some embodiments and within the front-end algorithm 504, an output from an Al algorithm 510 may be communicated to a CNN 508 or 509, which processes the data before communicating an output from the CNN 508, 509 and / or the front-end algorithm 504 to the back-end algorithm 506. In various embodiments, the back-endalgorithm 506 may be configured to implement input and / or model classification, speech recognition, translation, and the like. For instance, the back-end network 506 may include one or more CNNs (e.g., CNN 514) or dense networks (e.g., dense networks 516), as described herein.
[0074] For instance and in some embodiments of the Al program 502, the program may be configured to perform unsupervised learning, in which the machine learning program performs the training process using unlabeled data, e.g., without known output data with which to compare. During such unsupervised learning, the neural network may be configured to generate groupings of the input data and / or determine how individual input data points are related to the complete input data set (e.g., via the front-end algorithm 504). For example, unsupervised training may be used to configure a neural network to generate a self-organizing map, reduce the dimensionally of the input data set, and / or to perform outlier / anomaly determinations to identify data points in the data set that falls outside the normal pattern of the data. In some embodiments, the Al program 502 may be trained using a semi-supervised learning process in which some but not all of the output data is known, e.g., a mix of labeled and unlabeled data having the same distribution.
[0075] In some embodiments, the Al program 502 may be accelerated via a machine-learning framework 522 (e.g., hardware). The machine learning framework may include an index of basic operations, subroutines, and the like (primitives) typically implemented by Al and / or machine learning algorithms. Thus, the Al program 502 may be configured to utilize the primitives of the framework 522 to perform some or all of the calculations required by the Al program 502. Primitives suitable for inclusion in the machine learning framework 522 include operations associated with training a convolutional neural network (e.g., pools), tensor convolutions, activation functions, basic algebraic subroutines and programs (e.g., matrix operations, vector operations), numerical method subroutines and programs, and the like.
[0076] It should be appreciated that the machine-learning program may include variations, adaptations, and alternatives suitable to perform the operations necessary for the system, and the present disclosure is equally applicable to such suitably configured machine learning and / or artificial intelligence programs, modules, etc. For instance, the machine-learning program may include one or more long short-term memory ("LSTM") RNNs, convolutional deep belief networks, deep belief networks DBNs, and the like. DBNs, for instance, may be utilized to pre-train the weighted characteristics and / or parameters using an unsupervisedlearning process. Further, the machine-learning module may include one or more other machine learning tools (e.g., Logistic Regression ("LR"), Naive-Bayes, Random Forest ("RF"), matrix factorization, and support vector machines) in addition to, or as an alternative to, one or more neural networks, as described herein.
[0077] FIG. 6 is a flow chart representing a method 600, according to at least one embodiment, of model development and deployment by machine learning. The method 600 represents at least one example of a machine learning workflow in which steps are implemented in a machine-learning project.
[0078] In step 602, a user authorizes, requests, manages, or initiates the machine-learning workflow.This may represent a user such as human agent, or customer, requesting machine-learning assistance or Al functionality to simulate intelligent behavior (such as a virtual agent) or other machine-assisted or computerized tasks that may, for example, entail visual perception, speech recognition, decision-making, translation, forecasting, predictive modelling, and / or suggestions as non-limiting examples. In a first iteration from the user perspective, step 602 can represent a starting point. However, with regard to continuing or improving an ongoing machine learning workflow, step 602 can represent an opportunity for further user input or oversight via a feedback loop. Such feedback may flow through a user, or in various embodiments, the method automatically provides feedback, retrains and redeploys the retrained model.
[0079] In step 604, data is received, collected, accessed, or otherwise acquired and entered as can be termed data ingestion. In step 606, the data ingested in step 604 is pre-processed, for example, by cleaning, and / or transformation such as into a format that the following components can digest. The incoming data may be versioned to connect a data snapshot with the particularly resulting trained model. As newly trained models are tied to a set of versioned data, preprocessing steps are tied to the developed model. If new data is subsequently collected and entered, a new model will be generated. If the preprocessing step 606 is updated with newly ingested data, an updated model will be generated. Step 606 can include data validation, which focuses on confirming that the statistics of the ingested data are as expected, such as that data values are within expected numerical ranges, that data sets are within any expected or required categories, and that data comply with any needed distributions such as within those categories. Step 606 can proceed to step 608 to automatically alert the initiating user, other human or virtual agents, and / or other systems,if any anomalies are detected in the data, thereby pausing or terminating the process flow until corrective action is taken.
[0080] In step 610, training test data such as a target variable value is inserted into an iterative training and testing loop. In step 612, model training, a core step of the machine learning work flow, is implemented. A model architecture is trained in the iterative training and testing loop. For example, features in the training test data are used to train the model based on weights and iterative calculations in which the target variable may be incorrectly predicted in an early iteration as determined by comparison in step 614, where the model is tested. Subsequent iterations of the model training in step 612 are conducted with updated weights in the calculations.
[0081] During each iteration of the training and testing loop, the accuracy of the model may be evaluated. In one embodiment, the re-evaluation of the model can include comparing an output of the model with an actual target result or variable to determine the accuracy of the prediction. If the model is not satisfying a minimum threshold level of accuracy (i.e., the model is underfitted), the system may automatically determine that the threshold level of accuracy is not satisfied and may adjust the weights for a subsequent iteration of the training and testing loop.
[0082] The weights may be iteratively adjusted during each iteration of the training and testing loop based on the comparison to the threshold level of accuracy. However, there is a balance for training the model in order to avoid overfitting when the model would not perform well on predictions of new data. Rather, the model is automatically trained to be well-fitted such that it satisfies a threshold level of accuracy without learning the noise in the data to the extent that the model would not apply to new data by preventing additional iterations of the training and testing once a maximum accuracy threshold value has been obtained. Thus, with each iteration of the training and testing loop, the accuracy of the model is improved and the iterative training and testing of the model provides an improvement to the performance of a computer and computing technology because the system may automatically determine how many iterations to perform so that the model is well-fitted by surpassing the minimum threshold level of accuracy while automatically stopping the iterative training and testing of the model before the maximum accuracy threshold is obtained.
[0083] In some embodiments, the training and testing loop utilizes a backpropagation algorithm and a gradient descent algorithm. Gradient descent is an optimization algorithm used to minimizedifferentiable real-valued multivariate functions. Gradient descent is an optimization algorithm used to minimize differentiable real-valued multivariate functions. The gradient descent algorithm may be used to iteratively adjust model parameters using calculated derivatives to minimize a loss function. Backpropagation may be used to calculate the gradient of the error function with respect to the neural network's weights.
[0084] When compliance and / or success in the model testing in step 614 is achieved, process flow proceeds to step 616, where model deployment is triggered. The model may be utilized in Al functions and programming, for example to simulate intelligent behavior, to perform machine-assisted or computerized tasks, of which visual perception, speech recognition, decision-making, translation, forecasting, predictive modelling, and / or automated suggestion generation serve as non-limiting examples.
[0085] As discussed above, oversight of a deployed machine learning model may be automatically performed via a feedback loop whereby the method assesses performance of the deployed model (see step 616) and the feedback loop automatically provides feedback for further training of the machine learning model to improve its performance, and upon completion of the other method steps such as 612, the machine learning model that has been automatically retrained based on the feedback loop is then redeployed (step 614). In some embodiments, the system is continually receiving training data as new predictions are made and more data is collected. The continuous training data may be discretized to generate input data to retrain the model. Discretization methods can convert continuous data to discrete data by binning, clustering, and numerical discretization. The model may monitor incoming data sets to make predictions. When predictions are made, the system analyzes the predictions to determine whether the model needs to be retrained.
[0086] In some embodiments, the model may detect anomalies in the predictions. Anomaly detection can provide a benefit by identifying instances of the prediction that deviate from expected data or a general pattern. A difficulty in anomaly detection is that the system must define the boundary between ordinary data and anomalous data to accurately classify the data as ordinary or anomalous. The line between ordinary and anomalous may be difficult to determine with cases approaching a boundary and based on the specific application. For example, small variations may trigger an identification of an anomaly in the data while relatively larger deviations may be considered normal in less sensitive applications. The disclosed systems and methods may provide solutions to detecting anomalies in order tomore accurately and quickly determine whether a model needs to be retrained. If data would be inapplicable or would corrupt the model by reducing the quality of the input data or training process (e.g., due to missing values, outliers, inconsistent formatting, incorrect labels, noisy data, etc.) that data may be automatically dropped and the source of that data may be blocked from providing data that would be used to train the model. This reflects an improvement in the process of training and deploying a model that is accurate and specific to the type of prediction sought. In particular, this provides an improvement in the field of model training, which provides a practical application.
[0087] In other applications, the anomaly detections processes described herein may be used to provide enhanced security to the overall computing system by detecting malicious attacks on network security. For example, the system may take proactive measures to remediate danger by detecting the source address associated with potentially malicious packets and dropping potentially malicious packets. This provides an improvement in network security by dropping potentially malicious packets and blocking future traffic from the source address of the potentially malicious source address.Natural Language Processing Technology
[0088] The systems and methods disclosed herein may also be used to analyze text to form the predictions. In particular, the systems and methods described herein include a combination of elements that are utilized in a specific manner for automatically performing automated processes based on technological efficiency, which provides a specific improvement over prior art systems resulting in improved computer processing for faster automated processing functions. For example, the systems and method may apply robotic process automation for digital transformation of the data based on specific criteria to interpret text and unstructured data using text processing software techniques. The interpretation of the text may be implemented using the models described herein including unsupervised learning techniques or supervised learning techniques. The processor may track how much memory and / or processing time has been allocated to perform a function and the system may be trained to automatically detect and identify processes eligible for increased efficiencies based on existing inefficiencies in the process.
[0089] For example, the machine learning models may use unsupervised learning to identify and characterize hidden structures of unstructured and unlabeled content data, or supervisedtechniques that operate on labeled content data and include instructions informing the system which outputs are related to specific input values. In such instances, software processing can rely on iterative training techniques and training data to configure neural networks with an understanding of individual words, phrases, subjects, sentiments, and parts of speech.
[0090] Supervised learning software systems are trained using content data that is labeled or "tagged." During training, the supervised software systems learn the best mapping function between a known data input and expected known output (i.e., labeled or tagged content data). Supervised natural language processing software then uses the best approximating mapping learned during training to analyze unforeseen input data (never seen before) to accurately predict the corresponding output. Supervised learning software systems often require extensive and iterative optimization cycles to adjust the input-output mapping until they converge to an expected and well-accepted level of performance, such as an acceptable threshold error rate between a calculated probability and a desired threshold probability.
[0091] The software systems are supervised because the way of learning from training data mimics the same process of a teacher supervising the end-to-end learning process. Supervised learning software systems are typically capable of achieving excellent levels of performance, but this excellent level of performance requires labeled data to be available. Developing, scaling, deploying, and maintaining accurate supervised learning software systems can take significant time, resources, and technical expertise from a team of skilled data scientists. Moreover, precision of the systems is dependent on the availability of labeled content data for training that is comparable to the corpus of content data that the system will process in a production environment.
[0092] Supervised learning software systems implement techniques that include, without limitation, Latent Semantic Analysis ("LSA"), Probabilistic Latent Semantic Analysis ("PLSA"), Latent Dirichlet Allocation ("LDA"), and more recent Bidirectional Encoder Representations from Transformers ("BERT"). Latent Semantic Analysis software processing techniques process a corporate of content data files to ascertain statistical co-occurrences of words that appear together, which then give insights into the subjects of those words and documents.
[0093] Unsupervised learning software systems can perform training operations on unlabeled data and less requirement for time and expertise from trained data scientists. Unsupervised learning software systems can be designed with integrated intelligence and automation to1automatically discover information, structure, and patterns from content data. Unsupervised learning software systems can be implemented with clustering software techniques that include, without limitation, K-means clustering, Mean-Shift clustering, Density-based clustering, Spectral clustering, Principal Component Analysis, and Neural Topic Modeling ("NTM").
[0094] Clustering software techniques can automatically group semantically similar words together to accelerate the derivation and verification of an underneath common intent— i.e., ascertain or derive a new classification or subject, and not just classification into an existing subject or classification. Unsupervised learning software systems are also used for association rules mining to discover relationships between features from content data.
[0095] The system can incorporate an interactive content software service that utilizes one or more supervised or unsupervised software processing techniques to perform a subject classification analysis to generate subject data. Suitable software processing techniques can include, without limitation, LSA, PLSA, and LDA. Latent Semantic Analysis software processing techniques generally process a corpus of alphanumeric text files, or documents, to ascertain statistical co-occurrences of words that appear together, which then give insights into the subjects of those words and documents. The interactive content software service can utilize software processing techniques that include Non-Matrix Factorization, Correlated Topic Model ("CTM"), and K-Means or other types of clustering.
[0096] Neural networks may be trained using training set content data that comprise sample tokens, phrases, sentences, paragraphs, or documents for which desired subjects, content sources, interrogatories, or sentiment values are known. A labeling analysis may be performed on the training set content data to annotate the data with known subject labels (e.g., topics discussed during a shared experience also referred to herein as topic identifications), interrogatory labels (e.g., questions asked during a shared experience also referred to herein as interrogatory identifications), content source labels (e.g., identifications of participants to a shared experience), segment labels (e.g., opening, issue identification, or another segment), or sentiment labels, thereby generating annotated training set content data. For example, a person can utilize a labeling software application to review training set content data to identify and tag or "annotate" various parts of speech, subjects, questions, content sources, and sentiments.
[0097] The training set content data is fed to neural networks to identify subjects, content sources, or sentiments and the corresponding probabilities. For example, the analysis might identify that particular text represents a question with a 35% probability. If the annotations indicate the text is, in fact, a question, the error rate is 65% or the difference between the calculated probability and the known certainty. Then parameters to the neural network are adjusted (i.e., constants and formulas that implement the nodes and connections between node), to increase the probability from 35% to ensure the neural network produces more accurate results, thereby reducing the error rate. The process is run iteratively on different sets of training set content data to continue to increase the accuracy of the neural network.
[0098] The content data is first pre-processes using a reduction analysis to create reduced content data. The reduction analysis first performs a qualification operation that removes unqualified content data that does not meaningfully contribute to the subject classification analysis. The qualification operation removes certain content data according to criteria defined by a provider. For instance, the qualification analysis can determine whether content data files are "empty" and contain no recorded linguistic expressions between a provider agent and a user and designate such empty files as not suitable for use in a subject classification analysis. As another example, the qualification analysis can designate files below a certain size or having a shared experience duration below a given threshold (e.g., less than one minute) as also being unsuitable for use in an analysis, such as subject identification, sentiment analysis, or segmentation.
[0099] The reduction analysis can also perform a contradiction operation to remove contradictions and punctuations from the content data. Contradictions and punctuation include removing or replacing abbreviated words or phrases that can cause inaccuracies in a subject classification analysis. Examples include removing or replacing the abbreviations "min" for minute, "u" for you, and "wanna" for "want to," as well as apparent misspellings, such as "mssed" for the word missed. In some embodiments, the contradictions can be replaced according to a standard library of known abbreviations, such as replacing the acronym "brb" with the phrase "be right back." The contradiction operation can also remove or replace contractions, such as replacing "we're" with "we are."
[0100] The reduction analysis can also streamline the content data by performing one or more of the following operations, including: (i) tokenization to transform the content data into a collection of words or key phrases having punctuation and capitalization removed; (ii) stop wordremoval where short, common words or phrases such as "the" or "is" are removed; (iii) lemmatization where words are transformed into a base form, like changing third person words to first person and changing past tense words to present tense; (iv) stemming to reduce words to a root form, such as changing plural to singular; and (v) hyponymy and hypernym replacement where certain words are replaced with words having a similar meaning so as to reduce the variation of words within the content data.
[0101] Following a reduction analysis, the reduced content data is vectorized to map the alphanumeric text into a vector or matrix form— an operation that is also known as embedding. The content data files are converted to a series of machine encoded communication elements that make up the vectors or matrices (referred to herein as vectors as a shorthand even where matrices can be used). Machine encoded communication elements, also referred to as "communication elements" or sometimes "words," can be words, phrases, symbols (e.g., an emoji, logo, etc.), numbers, or other elements of that make up a written or transcribed communication.
[0102] One approach to vectorizing content data includes applying "bag-of-words" modeling. The bag-of-words approach counts the number of times a particular word appears in content data to convert the words into a numerical value. The bag-of-words model can include parameters, such as setting a threshold on the number of times a word must appear to be included in the vectors.
[0103] Another technique for vectorization includes the Word2Vec method. The Word2Vec method employs skip-grams or a continuous bag of words ("CBOW") and reconstructs the linguistic context of words by considering both the order of words in history as well as the predicted order of words. Shallow neural networks that have an input layer, an output layer, and a projection layer to iterate over a corpus of text to learn the association between the machine encoded communication elements. The method assumes that neighboring words in a text have semantic similarities with each other, and semantically similar machine encoded communication elements are mapped to geometrically close embedding vectors. Sematic similarity is measured using the cosine similarity metric. Cosine similarity is equal to the cosine of an angle where the angle is measured between the vector representation of text (e.g., sentences, phrases, or documents). If the cosine angle is one, it means that the words are overlapping. If the cosine angle is a right angle, the machine encoded communication elements hold no contextual similarity and are independent of each other.
[0104] Techniques to encode the content communication elements may, in part, determine how often communication elements appear together. Communication elements are often individual words but can also be groups of works like phrases, speech patterns, tone, cadence, or other characteristics of text. Determining the adjacent pairing of communication elements can be achieved by creating a co-occurrence matrix with the value of each member of the matrix counting how frequently one communication element coincides with another, either just before or just after it. That is, the words or communication elements form the row and column labels of a matrix, and a numeric value appears in matrix elements that correspond to a row and column label for communication elements that appear adjacent in the content data.
[0105] As an alternative to counting communication elements (e.g., words) in a corpus of content data and turning it into a co-occurrence matrix, another software processing technique may be used where a communication element in the content data corpus predicts the next communication element. Looking through a corpus, counts may be generated for adjacent communication elements, and the counts are converted from frequencies into probabilities (i.e., using n-gram predictions with Kneser-Ney smoothing) using a simple neural network. Suitable neural network architectures for such purpose include a skip-gram architecture. The neural network may be trained by feeding through a large corpus of content data, and embedded middle layers in the neural network are adjusted to best predict the next word.
[0106] The predictive processing creates weight matrices that densely carry contextual, and hence semantic, information from the selected corpus of content data. Pre-trained, contextualized content data embedding can have high dimensionality. To reduce the dimensionality, a uniform manifold approximation and projection algorithm ("UMAP") can be applied to reduce dimensionality while maintaining essential information.
[0107] Prior to conducting a subject analysis to ascertain subject identifications in the content data (i.e., topics or subjects addressed in the content data) or interaction driver identifications in the content data (i.e., reasons why the customer initiated the interaction with the provider, such as the reason underlying a support request), the system can perform a concentration analysis on the content data. The concentration analysis concentrates, or increases the density of, the content data by identifying and retaining communication elements that have significant weight in the subject analysis and discarding or ignoring communication elements that have relativity little weight.
[0108] In one embodiment, the concentration analysis includes executing a term frequency-inverse document frequency ("tf-idf") software processing technique to determine the frequency or corresponding weight quantifier for communication elements with the content data. The weight quantifiers are compared against a pre-determined weight threshold to generate concentrated content data that is made up of communication elements having weight quantifiers above the weight threshold.
[0109] Vectorization can be better understood with reference to the following simplified example. A corpus of machine encoded communication elements might include the following where each sentence is a row in a matrix: [I, forgot, my, account, password | | The, account, is, locked | | Please, reset, my, password, and, account]. Each machine encoded communication element can then be replaced by its frequency, such as: [1, 1, 2, 3, 2 1 1 1, 3, 1, 1 1 1 1, 1, 2, 2, 1, 3], Here, the highest frequency is three, so each frequency value is divided by 3 to yield: [.33, .33, .66, 1, .66 | | .33, 1, .33, .33 | | .33, .33, .66, .66, .33, 1],
[0110] In other examples, the vectorization creates a "sparse matrix" where each sentence, or row of the matrix, includes a frequency value for all distinct machine encoded communication elements within the corpus of content data. Where a communication element does not appear in a sentence, the frequency of the communication element is set to zero. Continuing with the foregoing example, the distinct communication elements include [I, forgot, my, account, password, the, is, locked, please, reset, and]. Each sentence is represented as follows: [1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0 | | 0, 0, 0, 1, 0, 1, 1, 1, 0, 0, 0 | | 1, 1, 2, 2, 1, 3 | | 0, 1, 1, 1, 0, 0, 0, 1, 1, 1],
[0111] Techniques to encode the context of words, or machine encoded communication elements, determine how often machine encoded communication elements appear together. Determining the adjacent pairing of machine encoded communication elements can be achieved by creating a co-occurrence matrix with the value of each member of the matrix counting how often one machine encoded communication element coincides with another, either just before or just after it. That is, the words or machine encoded communication elements form the row and column labels of a matrix, and a numeric value appears in matrix elements that correspond to a row and column label for communication elements that appear adjacent in the content data.
[0112] The concentrated content data is processed using a subject classification analysis to determine subject identifications (i.e., topics) addressed within the content data. The subjectclassification analysis can specifically identify one or more interaction driver identifications that are the reason why a user initiated a shared experience or support service request. An interaction driver identification can be determined by, for example, first determining the subject identifications having the highest weight quantifiers (e.g., frequencies or probabilities) and comparing such subject identifications against a database of known interaction driver identifications.
[0113] In one embodiment, the subject classification analysis is performed on the content data using a LDA analysis to identify subject data that includes one or more subject identifications (e.g., topics addressed in the underlying content data). Performing the LDA analysis on the reduced content data may include transforming the content data into an array of text data representing key words or phrases that represent a subject (e.g., a bag-of-words array) and determining the one or more subjects through analysis of the array. Each cell in the array can represent the probability that given text data relates to a subject. A subject is then represented by a specified number of words or phrases having the highest probabilities (i.e., the words with the five highest probabilities), or the subject is represented by text data having probabilities above a predetermined subject probability threshold.
[0114] Clustering software processing techniques include K-means clustering, which is an unsupervised processing technique that does not utilized labeled content data. Clusters are defined by "K" number of centroids where each centroid is a point that represents the center of a cluster. The K-means processing technique run in an iterative fashion where each centroid is initially placed randomly in the vector space of the dataset, and the centroid moves to the center of the points that is closest to the centroid. In each new iteration, the distance between each centroid and the points are recalculated, and the centroid moves again to the center of the closest points. The processing completes when the position or the groups no longer change or when the distance in which the centroids change does not surpass a pre-defined threshold.
[0115] The clustering analysis yields a group of words or communication elements associated with each cluster, which can be referred to as subject vectors. Subjects may each include one or more subject vectors where each subject vector includes one or more identified communication elements (i.e., keywords, phrases, symbols, etc.) within the content data as well as a frequency of the one or more communication elements within the content data. The interactive content software service can be configured to perform an additionalconcentration analysis following the clustering analysis that selects a pre-defined number of communication elements from each cluster to generate a descriptor set, such as the five or ten words having the highest weights in terms of frequency of appearance (or in terms of the probability that the words or phrases represent the true subject when neural networking architecture is used). In one embodiment, the descriptor sets were analyzed to determine if the reasons driving a customer support request were identified by the descriptor set subject identifications.
[0116] The software model may be evaluated according to three categories, including a "good match" where the support request reason(s) are identified by the top words in the subject vector (i.e., the words with the highest weight or frequency), a "moderate" match where the support request reason(s) are identified by the second tier of words in the subject vector (i.e., words six to ten), and a "poor" match where, for instance, the top words in a subject vector do not match or identify the reasons the support request was initiated.
[0117] Alternatively, instead of selecting a pre-determined number of communication elements, post-clustering concentration analysis can analyze the subject vectors to identify communication elements that are included in several subject vectors having a weight quantifier (e.g., a frequency) below a specified weight threshold level that are then removed from the subject vectors. In this manner, the subject vectors are refined to exclude content data less likely to be related to a given subject. To reduce an effect of spam, the subject vectors may be analyzed, such that if one subject vector is determined to include communication elements that are rarely used in other subject vectors, then the communication elements are marked as having a poor subject correlation and is removed from the subject vector.
[0118] In another embodiment, the concentration analysis is performed on unclassified content data by mapping the communication elements within the content data to integer values. The content data is thus turned into a bag-of-words that includes integer values and the number of times the integers occur in content data. The bag-of-words is turned into a unit vector, where all the occurrences are normalized to the overall length. The unit vector may be compared to other subject vectors produced from an analysis of content data by taking the dot product of the two-unit vectors. All the dot products for all vectors in a given subject are added together to provide a weighting quantifier or score for the given subject identification, which is taken as subject weighting data. A similar analysis can be performed on vectorscreated through other processing, such as K-means clustering or techniques that generate vectors where each word in the vector is replaced with a probability that the word represents a subject identification or request driver data.
[0119] To illustrate generating subject weighting data, for any given subject there may be numerous subject vectors. Assume that for most of subject vectors, the dot product will be close to zero — even if the given content data addresses the subject at issue. Since there are some subjects with numerous subject vectors, there may be numerous small dot products that are added together to provide a significant score. Put another way, the particular subject is addressed consistently throughout a document, several documents, sessions of the content data, and the recurrence of the carries significant weight.
[0120] In another embodiment, a predetermined threshold may be applied where any dot product that has a value less than the threshold is ignored and only stronger dot products above the threshold are summed for the score. In another embodiment, this threshold may be empirically verified against a training data set to provide a more accurate subject analysis.
[0121] In another example, a number of subject identifications may be substantially different, with some subjects having orders of magnitude fewer subject vectors than do other subjects. The weight scoring might significantly favor relatively unimportant subjects that occur frequently in the content data. To address this problem, a linear scaling on the dot product scoring based on the number of subject vectors may be applied. The result provides a correction to the score so that important but less common subjects are weighed more heavily.
[0122] Once all scores are calculated for all subjects, then subjects may be sorted, and the most probable subjects are returned. The resulting output provides an array of subjects and strengths. In another embodiment, hashes may be used to store the subject vectors to provide a simple lookup of text data (e.g., words and phrases) and strengths. The one or more subject vectors can be represented by hashes of words and strengths, or alternatively an ordered byte stream (e.g., an ordered byte stream of 4-byte integers, etc.) with another array of strengths (e.g., 4-byte floating-point strengths, etc.).
[0123] The interactive content software service can also use term frequency-inverse document frequency software processing techniques to vectorize the content data and generating weighting data that weight words or particular subjects. The tf-idf is represented by a statistical value that increases proportionally to the number of times a word appears in the content data. This frequency is offset by the number of separate content data instances thatcontain the word, which adjusts for the fact that some words appear more frequently in general across multiple shared experiences or content data files. The result is a weight in favor of words or terms more likely to be important within the content data, which in turn can be used to weigh some subjects more heavily in importance than others. To illustrate with a simplified example, the tf-idf might indicate that the term "password" carries significant weight within content data. To the extent any of the subjects identified by a natural language processing analysis include the term "password," that subject can be assigned more weight by the interactive content software service.
[0124] The content data can be visualized and subject to a reduction into two-dimensional data using a UMAP to generate a cluster graph visualizing a plurality of clusters. The interactive content software service feeds the two-dimensional data into a DBSCAN and identify a center of each cluster of the plurality of clusters. The process may, using the two dimensional data from the UMAP and the center of each cluster from the DBSCAN, apply a KNN to identify data points closest to the center of each cluster and shade each of the data points to graphically identify each cluster of the plurality of clusters. The processor may illustrate a graph on the display representative of the data points that are shaded following application of the KNN.
[0125] The system can also incorporate Part of Speech ("POS") tagging software code that assigns words a part of speech depending upon the neighboring words, such as tagging words as a noun, pronoun, verb, adverb, adjective, conjunction, preposition, or other relevant parts of speech. The interactive content software service can utilize the POS tagged words to help identify questions and subjects according to pre-defined rules, such as recognizing that the word "what" followed by a verb is also more likely to be a question than the word "what" followed by a preposition or pronoun (e.g., "What is this?" versus "What he wants is an answer.").
[0126] POS tagging in conjunction with Named Entity Recognition ("NER") software processing techniques can be used by the interactive content software service to identify various content sources within the content data. NER techniques are utilized to classify a given word into a category, such as a person, product, organization, or location. Using POS and NER techniques to process the content data allow the system to identify particular words and text as a noun and as representing a person participating in the discussion (i.e., a content source). Recognizing participants in content data is also called diarization and is discussed more fully below.Language Models
[0127] Natural language processing can be implemented by a LM, which is a sophisticated artificial intelligence model that is configured specifically for natural language processing tasks. Language Models are designed to understand and generate human-like text based on the patterns and structures that the LM learns from processing substantial volumes of training data.
[0128] Language Models are implemented with a deep learning architecture called a transformer.Transformers consist of multiple layers of self-attention mechanisms that allow a LM to weigh the importance of different words or tokens in a sequence and capture the relationships between them. The attention mechanism allows LM to effectively process and generate text with contextually relevant and coherent patterns. By assigning different weights to different words, LMs can effectively focus on the most relevant information, which facilitates accurate and contextually appropriate content generation.
[0129] Language Models learn by predicting the next word in a given context using an unsupervised learning process. Through repetition and exposure to training data, the LM develops a proficiency for grammar, semantics, and the knowledge contained in the training data. During a pre-training phase, training data is ingested by a LM and tokenized to break the training data down into smaller units called tokens. Tokens can be words, subwords, or characters. Tokenization allows a LM to process and understand text at a granular level. The LM learns to predict the next token in a sequence, given the preceding tokens.
[0130] After pre-training, a LM can be fine tuned and applied to a wide range of applications that entail performing specific tasks, such as sentiment analysis, diarization, a chat bot, virtual assistant, or a content generation system. Fine-tuning involves providing the LM with taskspecific labeled data, so the LM learns intricacies of a particular task.
[0131] Once the LM is trained and fine-tuned, it can be used for inference. Inference involves utilizing the LM to generate text or perform specific language-related tasks. During inference, LMs may employ a beam search technique to generate the most likely sequence of tokens. Beam search is an algorithm that explores several possible paths in the sequence generation process while keeping track of the most likely candidates based on a scoring mechanism. This approach helps generate more coherent and high-quality text outputs.
[0132] There are various types of LMs, including, without limitation: (i) autoregressive language models; (ii) encoder-decoder models; (iii) transformer-based models; (iv) trained and fine-tuned models; and (v) hybrid models. Autoregressive models generate text by predicting the next word given the preceding words in a sequence. Encoder-decoder models are commonly used for machine translation, summarization, and question-answering tasks. Encoderdecoder models consist of two main components: (i) an encoder for processing input sequences; and (ii) a decoder that generates the output sequence. A transformer LM is a type of encoder-decoder architecture.
[0133] Hybrid LMs combine the strengths of different architectures. For example, some LMs may incorporate both transformer-based architectures and recurrent neural networks. RNNs are commonly used for sequential data processing and can be integrated into a LM to capture sequential dependencies in addition to the self-attention mechanisms of transformers.
[0134] Figure 7 shows a flow-chart of an example method 100, performed by a computing device.The computing device is any of the computer devices disclosed herein, such as the computing device 10 of figure 1.
[0135] The method 100 comprises inputting 102, at a language model, LM, a first text message.
[0136] The method comprises inputting 104, at the LM, a conversational state history, CSH, vector, wherein the CSH vector comprises an array of numbers and wherein the CSH vector is representative of text messages received at the LM prior to the first text message.
[0137] As used herein, "CSH vector" should be construed as an array of numbers that represent a content of messages received and / or sent by the computing device over a period of time. Put differently, the array of numbers may denote a context-aware representation (as the messages received and / or sent by the computing device provide context) of timedependent memory.
[0138] As used herein, an "intent" should be construed as should be construed as an indicator of a type of behaviour.
[0139] The LM may be built on BERT, Bidirectional Encoder Representations from Transformers.The architecture of BERT may take an encoder part of the transformer and may be pretrained on unlabelled text. The pretrained model may then be fine-tuned with a dedicated domain specific data for classification. In deployment, the fine-tuned model may take a text input, such as the first text message, and may output probabilities in one or more classes, such as cyberbullying, self-harm, online grooming, and others.
[0140] In order to translate the content of a text into an array of numbers, the word sequence in the text may first be converted to an array of token IDs, which are indexed in a vocabulary.The BERT embedding layer with trained weights may then map the token IDs to a vector of floating point numbers.
[0141] The method comprises outputting 106, by the LM, a measurement of the likelihood the first text message comprises an intent.
[0142] In an example, the measurement may be outputted, by the LM, by means of a classify, CLS, token that may be added at the beginning of a sequence input in BERT. For classification, the final state corresponding to this token may be used as the aggregate sequence representation in BERT and may be fed into a fully connected layer to get a probability output. The CSH vector may be provided as an additional embedding to the BERT input and get updated inside the BERT.
[0143] The method may comprise updating 108 the CSH vector with the first text message.
[0144] The method may comprise retaining 110 a copy of the CSH vector in a vector database of the computing device.
[0145] The method may comprise sending 112, to an external server, the CSH vector, wherein the CSH vector is updated with the first text message and wherein the CSH vector is representative of text messages received at the LM up to the first text message.
[0146] The method may comprise receiving 114, at the LM and after the first text message, a second text message.
[0147] The method may comprise sending 116, to the external server, the second text message.
[0148] The method may comprise receiving 118, from the external server, a server measurement of the likelihood the second text message comprises an intent.
[0149] The Ul may be a Ul screen.
[0150] The invention is defined in the claims. However, below there is provided a non-exhaustive list of non-limiting aspects. Any one or more of the features of these aspects may be combined with any one or more features of another example, embodiment or disclosure described herein.Aspect 1. A method for identifying an intent in a real time conversation performed by a computing device, the method comprising:inputting, at a language model, LM, a first text message;inputting, at the language model, a conversational state history, CSH, vector, wherein the CSH vector comprises an array of numbers and wherein the CSH vector is representative of text messages received at the LM prior to the first text message ;outputting, by the language model, a measurement of the likelihood the first text message comprises an intent.Aspect 2. The method of aspect 1, further comprising:updating the CSH vector with the first text message.Aspect 3. The method of any one of aspects 1 to 2, wherein the first text message and the text messages received at the LM prior to the first text message belong to a conversation.Aspect 4. The method of any one of aspects 1 to 3, further comprising:retaining a copy of the CSH vector in a vector database of the computing device. Aspect 5. The method of any one of aspects 1 to 4, wherein the array of numbers has a predetermined length, and wherein an update of the CSH vector replaces at least a previous number of the array of numbers with an updated number of the array of numbers.Aspect 6. The method of any one of aspects 1 to 5, wherein the array of numbers is stored in a recurrent architecture, preferably a long short-term memory, LSTM or a Recurring Memory Transformer.Aspect 7. The method of aspect 6, wherein the array of numbers represents a cell state of the recurrent architecture.Aspect 8. The method of any one of aspects 1 to 7, further comprising:sending, to an external server, the CSH vector, wherein the CSH vector is updated with the first text message and wherein the CSH vector is representative of text messages received at the LM up to the first text message;receiving, at the LM and after the first text message, a second text message; sending, to the external server, the second text message; andreceiving, from the external server, a server measurement of the likelihood the second text message comprises an intent.Aspect 9. The method of aspect 8, further comprising:establishing a session between the computing device and the external server.Aspect 10. The method of any one of aspects 1 to 9, wherein the intent comprises one or more of: cyberbullying intent, self-harm intent or online grooming intent.Aspect 11. The method of any one of aspects 1 to 10, wherein the computing device is user equipment, UE.Aspect 12. The method of aspect 11, wherein the UE is a mobile UE.Aspect 13. The method of any one of aspects 1 to 12, wherein the LM is a small LM, SLM.Aspect 14. The method of any one of aspects 1 to 13, wherein the first text message is associated with a chat application installed in the computing device.Aspect 15. A computer readable programmable medium carrying a computer programme stored thereon which, when executed by a processor, implements the method according to any of aspects 1 to 14.Aspect 16. A computing device comprising:a memory; andone or more processors operatively coupled to the memory, the one or more processors configured to:input, at a language model, LM, a first text message;input, at the LM, a conversational state history, CSH, vector , wherein the CSH vector comprises an array of numbers and wherein the CSH vector is representative of text messages received at the LM prior to the first text message;output, by the language model, a measurement of the likelihood the first text message comprises a harmful intent.Aspect 17. The computing device of aspect 16, wherein the one or more processors are further configured to:update the CSH vector with the first text message.Aspect 18. The computing device of any one of aspects 16 to 17, wherein the first text message and the text messages received at the LM prior to the first text message belong to a conversation.Aspect 19. The computing device of any one of aspects 16 to 18, wherein the one or more processors are further configured to:retain a copy of the CSH vector in a vector database of the computing device.Aspect 20. The computing device of any one of aspects 16 to 19, wherein the array of numbers has a predetermined length, and wherein an update of the CSH vector replaces at least a previous number of the array of numbers with an updated number of the array of numbers.Aspect 21. The computing device of any one of aspects 16 to 20, wherein the array of numbers is stored in a recurrent architecture, preferably a long short-term memory, LSTM or a Recurring Memory Transformer.Aspect 22. The computing device of aspect 21, wherein the array of numbers represents a cell state of the recurrent architecture.Aspect 23. The computing device of any one of aspects 16 to 22, wherein the one or more processors are further configured to:send, to an external server, the CSH vector, wherein the CSH vector is updated with the first text message and wherein the CSH vector is representative of text messages received at the LM up to the first text message;receive, at the LM and after the first text message, a second text message; send, to the external server, the second text message; andreceive, from the external server, a server measurement of the likelihood the second text message comprises an intent.Aspect 24. The computing device of aspect 23, wherein the one or more processors are further configured to:establish a session between the computing device and the external server.Aspect 25. The computing device of any one of aspects 16 to 24, wherein the intent comprises one or more of: cyberbullying intent, self-harm intent or online grooming intent. Aspect 26. The computing device of any one of aspects 16 to 25, wherein the computing device is user equipment, UE.Aspect 27. The computing device of any one of aspects 16 to 26, wherein the UE is a mobile UE.Aspect 28. The computing device of any one of aspects 16 to 27, wherein the LM is a small LM, SLM.Aspect 29. The computing device of any one of aspects 16 to 28, wherein the first text message is associated with a chat application installed in the computing device.Aspect 30. A computing system comprising the computing device of any one of aspects 16 to 29 when depending on aspect 23 and an external server, wherein the external server comprises:a server memory; andone or more server processors operatively coupled to the server memory, the one or more server processors configured to:upon receiving, by the external server, the CSH vector and the second text message, output, by a server LM, the server measurement of the likelihood the second text message comprises a harmful intent.
[0151] The invention is not limited to the embodiments hereinbefore described but may be varied in both construction and detail.
[0152] The words "comprises / comprising" and the words "having / including" when used herein with reference to the present invention are used to specify the presence of stated features, integers, steps or components but does not preclude the presence or addition of one or more other features, integers, steps, components or groups thereof. It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination.
Claims
CLAIMS1. A method for identifying an intent in a real time conversation performed by a computing device, the method comprising:inputting, at a language model, LM, a first text message;inputting, at the language model, a conversational state history, CSH, vector, wherein the CSH vector comprises an array of numbers and wherein the CSH vector is representative of text messages received at the LM prior to the first text message;outputting, by the language model, a measurement of the likelihood the first text message comprises an intent.
2. The method of claim 1, further comprising:updating the CSH vector with the first text message.
3. The method of claim 1, wherein the first text message and the text messages received at the LM prior to the first text message belong to a conversation.
4. The method of claim 1, further comprising:retaining a copy of the CSH vector in a vector database of the computing device.
5. The method of claim 1, wherein the array of numbers has a predetermined length, and wherein an update of the CSH vector replaces at least a previous number of the array of numbers with an updated number of the array of numbers.
6. The method of claim 1, wherein the array of numbers represents is stored in a recurrent architecture, preferably a long short-term memory, LSTM or a Recurring Memory Transformer.
7. The method of claim 6, wherein the array of numbers represents a cell state of the recurrent architecture.
8. The method of claim 1, further comprising:sending, to an external server, the CSH vector, wherein the CSH vector is updated with the first text message and wherein the CSH vector is representative of text messages received at the LM up to the first text message;receiving, at the LM and after the first text message, a second text message; sending, to the external server, the second text message; andreceiving, from the external server, a server measurement of the likelihood the second text message comprises an intent.
9. The method of claim 8, further comprising:establishing a session between the computing device and the external server.
10. The method of claim 1, wherein the intent comprises one or more of: cyberbullying intent, self-harm intent or online grooming intent.
11. The method of claim 1, wherein the computing device is user equipment, UE.
12. The method of claim 11, wherein the UE is a mobile UE.
13. The method of claim 1, wherein the LM is a small LM, SLM.
14. The method of claim 1, wherein the first text message is associated with a chat application installed in the computing device.
15. A computer readable programmable medium carrying a computer programme stored thereon which, when executed by a processor, implements the method according to any of claims 1 to 14.
16. A computing device identifying harmful intent in a real time conversation comprising:a memory; and45one or more processors operatively coupled to the memory, the one or more processors configured to:input, at a language model, LM, a first text message;input, at the LM, a conversational state history, CSH, vector, wherein the CSH vector comprises an array of numbers and wherein the CSH vector is representative of text messages received at the LM prior to the first text message;output, by the language model, a measurement of the likelihood the first text message comprises a harmful intent.
17. The computing device of claim 16, wherein the one or more processors are further configured to:send, to an external server, the CSH vector, wherein the CSH vector is updated with the first text message and wherein the CSH vector is representative of text messages received at the LM up to the first text message;receive, at the LM and after the first text message, a second text message; send, to the external server, the second text message; andreceive, from the external server, a server measurement of the likelihood the second text message comprises an intent.
18. A computing system comprising the computing device of claim 17 and an external server, wherein the external server comprises:a server memory; andone or more server processors operatively coupled to the server memory, the one or more server processors configured to:upon receiving, by the external server, the CSH vector and the second text message, output, by a server LM, the server measurement of the likelihood the second text message comprises a harmful intent.