System and method for detecting chat robot
By generating challenge messages and twisting dialogue sequels using a generative language model, the problem of distinguishing between humans and chatbots in existing technologies is solved, achieving more accurate identity verification and reducing fraud risks.
Patent Information
- Application Number
- CN202480028731.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-19
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-25
AI Technical Summary
Existing Turing test methods have become outdated due to the development of AI and language modeling, making it difficult to effectively distinguish between humans and chatbots, leading to problems such as fraud and unsolicited communication.
Generative language models are used to generate alternative dialogue sequels and add challenge messages. The similarity between the dialogue partner's response and the alternative response is used to determine whether it is a bot. This includes generating challenge messages and twisting the dialogue sequels to trigger anomalous behavior.
It improves the accuracy of distinguishing between humans and chatbots, reduces unsolicited communication and fraud, and enhances the ability to identify conversation partners.
Smart Images

Figure CN121014191A_ABST
Abstract
Description
Background Technology
[0001] This invention relates to artificial intelligence, and more particularly, to using artificial intelligence (AI) to automatically determine whether an entity participating in electronic messaging is a robot.
[0002] Recent advancements in AI are driving radical changes in virtually every aspect of human activity. Certain categories of AI applications involve human-computer interaction and include developing language models that enable computers to communicate in natural languages such as English or Chinese. While creating enormous opportunities for business growth by automating various tasks such as customer service and information retrieval, software agents that support language models (often called chatbots) also present unprecedented technological and ethical challenges. As chatbots become increasingly sophisticated and capable of engaging in genuine conversations, distinguishing between humans and machines becomes extremely difficult. Unethical entities may exploit this confusion for purposes such as fraud, unsolicited communication (spam), identity theft, large-scale disinformation campaigns, and political manipulation.
[0003] The problem of determining whether a conversational partner is human or computer is commonly known as the Turing Test and is almost as old as information technology itself. Several methods for implementing the Turing Test have been proposed for various applications recently related to communication via the Internet. One example includes a class of methods collectively known as the Fully Automated Public Turing Test (CAPTCHA) for distinguishing between computers and humans. In various embodiments, CAPTCHA involves issuing a challenge (e.g., an image recognition problem) in response to an attempt to access an online service (e.g., a specific webpage) and selectively allowing access to the corresponding service based on the response to the challenge. Such methods implicitly rely on the fact that typical software agents cannot currently solve the specific type of problem used as the challenge. However, recent developments in AI and language modeling are rapidly rendering such conventional Turing Tests obsolete.
[0004] For the reasons outlined above, there has been renewed interest in developing robust and effective Turing tests. Summary of the Invention
[0005] According to one aspect, a computer system includes at least one hardware processor configured to employ a generative language model to generate alternative dialogue continuations and alternative responses. The alternative dialogue continuations include predicted continuations of an ongoing online conversation comprising a sequence of messages. The alternative responses include predicted responses from a conversation partner to the alternative dialogue continuations. The at least one hardware processor is further configured to, in response, distort the alternative dialogue continuations to generate a challenge message and add the challenge message to the ongoing online conversation. The at least one hardware processor is further configured to, in response to receiving a partner response from the conversation partner, the partner response including a response to the challenge message, determine whether the conversation partner includes a bot based on the similarity between the partner response and the alternative responses. Attached Figure Description
[0006] The foregoing aspects and advantages of the invention will be better understood after reading the following detailed description and referring to the accompanying drawings, wherein:
[0007] Figure 1 This invention illustrates multiple client devices participating in electronic communication and a set of server computers implementing a chatbot detector, according to some embodiments of the invention.
[0008] Figure 2 An exemplary configuration of a chatbot detector according to some embodiments of the present invention is described.
[0009] Figure 3 An exemplary sequence of steps performed by a chatbot detector according to some embodiments of the present invention is shown.
[0010] Figure 4 An exemplary sequence of steps performed by the challenge generator component of a chatbot detector according to some embodiments of the present invention is shown.
[0011] Figure 5 Exemplary dialogue contexts, alternative dialogue sequels, challenges, and alternative responses are described according to some embodiments of the present invention.
[0012] Figure 6 An exemplary process for generating alternative text according to some embodiments of the present invention is described.
[0013] Figure 7 An exemplary method for generating challenges by distorting alternative dialogue sequels is presented according to some embodiments of the invention.
[0014] Figure 8 An exemplary encoder is shown that computes an embedding vector from an input lexical sequence according to some embodiments of the present invention.
[0015] Figure 9 This describes an exemplary method for assessing the similarity between two lexical sequences according to some embodiments of the present invention.
[0016] Figure 10 Exemplary hardware configurations of computer systems programmed to perform some of the methods described herein are shown. Detailed Implementation
[0017] In the following description, it should be understood that all referenced connections between structures may be direct operational connections or indirect operational connections via intermediate structures. A set of elements comprises one or more elements. Any reference to an element should be understood to refer to at least one element. Multiple elements comprise at least two elements. Any use of "or" implies a non-exclusive "or". Unless otherwise required, any described method steps do not necessarily need to be performed in the specific order stated. A first element derived from a second element (e.g., data) encompasses a first element equal to the second element, as well as a first element generated by processing the second element and optionally other data. Making a determination or decision based on parameters encompasses making a determination or decision based on parameters and optionally other data. Unless otherwise specified, an indicator of quantity / data may be the quantity / data itself, or an indicator different from the quantity / data itself. A computer program is a sequence of processor instructions that performs a task. The computer program described in some embodiments of the invention may be a standalone software entity or a sub-entity (e.g., a subroutine, a library) of another computer program. Computer-readable media encompasses non-transitory media, such as magnetic, optical, and semiconductor storage media (e.g., hard disk drives, optical disks, flash memory, DRAM), and communication links, such as conductive cables and fiber optic links. According to some embodiments, the present invention particularly provides a computer system comprising hardware (e.g., one or more processors) programmed to perform the methods described herein, and a computer-readable medium coded with instructions to perform the methods described herein.
[0018] Embodiments of this invention relate to detecting chatbots masquerading as humans in online conversations. A chatbot, as used herein, refers to any computer program configured to automatically generate text in natural languages (e.g., English, Russian, and Chinese) and interact with messaging applications to transmit corresponding text to another computer system. Generating text involves effectively creating corresponding text fragments (e.g., by automatically concatenating words extracted from a dictionary according to algorithms and / or language models), rather than simply encoding text provided by a human operator. Online conversations, as used herein, include sequences of electronic messages exchanged among a group of partners via messaging applications and / or online platforms. The format of the messages may vary depending on the respective messaging platform, protocol, and / or application, but generally, electronic messages may include encoding of text portions and / or encoding of media files (e.g., images, movies, sound, etc.). The text portions may include text written in natural languages as well as other alphanumeric and / or special characters, such as emoticons.
[0019] For clarity, the following description will focus on processing text messages. However, those skilled in the art will recognize that the described systems and methods are adaptable to other types of messaging, such as audio / video (e.g., spoken text) or combinations thereof. For example, some embodiments may determine whether an audio file containing spoken messages was generated by a chatbot. In one exemplary embodiment, a speech-to-text translator may be used to convert an audio file into text fragments, and then the methods described herein related to text messaging may be applied.
[0020] Online messaging encompasses peer-to-peer messaging as well as messaging via public chat rooms, forums, social media websites, etc. Examples of online conversations include the exchange of Short Message Service (SMS) messages, email message sequences, and message sequences exchanged via instant messaging applications such as WhatsApp Messenger®, Telegram®, WeChat®, and Facebook® Messenger®. Other exemplary online conversations include content on Facebook® Walls, chats on online forums such as Reddit® and Discord®, and a set of comments on a blog post. Exemplary messaging applications according to embodiments of the invention include client-side examples of mobile applications such as WhatsApp®, Facebook®, Instagram®, and Snapchat®, and server-side software that performs the corresponding messaging operations. Other examples of messaging applications include examples of email clients and internet browsers.
[0021] The following description illustrates embodiments of the invention by way of example and not necessarily by way of limitation.
[0022] Figure 1 A set of exemplary client systems 10a to 10c are shown, demonstrating participation in online conversations via a communication network 15 (e.g., the Internet). Parts of network 15 may include local area networks (LANs) and telecommunications networks (e.g., mobile phones). Client systems 10a to 10c generally represent any electronic device having at least one processor and components connected to network 15. Exemplary clients 10a to 10c include personal computers, mainframe computers, mobile computing devices (laptops, smartphones, tablets, etc.), wearable computing devices (e.g., smartwatches, etc.), and home appliances (smart refrigerators, home security systems, etc.) and computerized vehicles, etc.
[0023] In typical online messaging scenarios (e.g., social media platforms and online forums), messages are centralized, routed, and / or distributed by messaging server 12, for example, using a client-server protocol. In other words, individual messages sent by multiple client systems 10a to c may accumulate at messaging server 12, which may selectively display or otherwise deliver the messages to their intended destinations. In alternative embodiments, electronic messaging utilizes a decentralized network of individual peer-to-peer connections between client systems 10a to c.
[0024] In some embodiments, a chatbot detector determines whether at least a portion of an online conversation (i.e., a set of messages) is automatically generated. For example, the chatbot detector may determine whether selected client systems 10a to c include chatbots masquerading as humans in the online conversation. Exemplary embodiments of the chatbot detector include a set of interconnected computer programs, dedicated hardware modules, or combinations thereof. In some embodiments, at least a portion of the chatbot detector may be executed on a utility server 16, which may include a set of interconnected computer systems further connected to a communications network 15. In some embodiments as described herein, the chatbot detector includes an artificial intelligence (AI) system comprising a set of pre-trained neural networks. Training the corresponding AI system may be performed on a dedicated AI training facility 14.
[0025] Figure 2 This describes exemplary components of a chatbot detector 20 according to some embodiments of the present invention. The detector 20 may be included in a client system (e.g., Figure 1The computer program is executed on at least one hardware processor of the client systems 10a to c). In an alternative embodiment, some or all of the illustrated components of the chatbot detector 20 may execute on the messaging server 12 and / or the utility server 16. Those skilled in the art will also appreciate that, in alternative embodiments, some or all of the illustrated components may be embodied as hardware modules, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
[0026] Figure 3 Further explanation is provided of an exemplary sequence of steps performed by a chatbot detector 20 according to some embodiments of the present invention. The illustrated method includes automatically generating a challenge (e.g., a new message) based on the context of the current conversation and inserting the appropriate challenge into the corresponding conversation. A response to the challenge received from a conversation partner can then be compared with an alternative response automatically generated by the detector 20, which then determines whether the corresponding conversation partner includes a chatbot based on the comparison result. Exemplary components of the detector 20 are described below.
[0027] In some embodiments, detector 20 includes a dialogue agent 32 interconnected with challenge generator 34 and response analyzer 36. Dialogue agent 32 may interface with messaging application 30, which typically represents any software configured to enable users of a given client system to exchange electronic messages with other users. Exemplary messaging application 30 includes native examples of Facebook® mobile applications and native examples of Internet browsers, as well as server-side components of messaging platforms as described above. Application 30 may display the content of each electronic message on an output device (e.g., a screen) of the given client system and may further organize messages according to sender, recipient, time, subject, or other criteria. Application 30 may further receive input from users (e.g., from a keyboard, touchscreen, dictation interface, etc.), formulate electronic messages based on the received input, and transmit electronic messages to messaging server 12 and / or directly to other client systems 10a to c. Message transmission may include, for example, adding the encoding of the corresponding message to an outbound queue of the communication interface of the given client system.
[0028] Interfacing with application 30 may involve retrieving and / or transmitting data to / from application 30. For example, agent 32 may be configured to parse an ongoing online conversation and extract a conversation sample 22 containing selected portions copied from the corresponding conversation. Extracting the conversation sample may include identifying individual messages in the conversation and determining message-specific characteristics, such as the sender and / or receiver, the time of transmission (e.g., a timestamp), the text of the corresponding message, and possibly other content data, such as an image attached to the corresponding message. Interfacing with messaging application 30 may further include transmitting a challenge 44 to application 30, the challenge 44 comprising a deliberately crafted message for insertion into the corresponding online conversation, as described in more detail below.
[0029] The dialogue agent 32 can be implemented using any method known in the art. In one exemplary embodiment, agent 32 may be incorporated into messaging application 30, for example, as an add-on or plugin (e.g., a browser extension). Such embodiments may use the functionality of application 30 to extract data from and / or insert data into an ongoing dialogue. Other exemplary embodiments of agent 32 may employ techniques known in the art of robotic process automation, such as using a local driver to automatically identify elements of the user interface of messaging application 30 and to simulate how a human user interacts with it to read and write messages. Some such embodiments may use built-in features of the local operating system to extract message content, such as an accessibility application programming interface (API)—typically software used to capture information currently displayed on the screen for the purpose of making such information accessible to people with disabilities. Agent 32 may further parse various data structures, such as a user interface tree or document object model (DOM), to identify individual messages and extract their content and other information, such as the sender's identity.
[0030] In an alternative embodiment, the messaging application 30 may be secretly modified (e.g., via hooking) to install the dialogue agent 32. For example, a code patch may be inserted into the application 30, with corresponding code configured to transparently invoke the execution of the agent 32 whenever the messaging application 30 performs a specific action (e.g., receiving or transmitting a message). Another embodiment of the agent 32 may extract message content directly from intercepted network traffic entering the messaging application 30 and / or via network adapters of the respective client devices 10a to c. Such communication interceptors may implement communication protocols such as HTTP, WebSocket, and MQTT to parse communications and extract structured message data. When instant messages are encrypted, some embodiments employ techniques such as man-in-the-middle (MITM) to decrypt traffic for message content extraction.
[0031] Some embodiments of the invention rely on the observation that modern chatbots are typically aware of the context of a conversation, i.e., they generate individual messages based on the content of other messages previously exchanged during an ongoing conversation. Therefore, some embodiments of the conversation agent 32 are further configured to organize messages into conversations. A conversation, as used herein, includes a sequence of individual messages exchanged between the same pair of conversationalists (in the case of one-on-one communication) or within the same group (e.g., in the case of a group chat). In one exemplary embodiment, the conversation agent 32 may identify the sender and / or receiver of each intercepted message and attach a tag to each message that creates an association between the corresponding message and the conversation. After such tagging, messages can be selectively retrieved based on the conversation and ordered according to a message-specific timestamp. In an alternative embodiment, the conversation agent 32 may store each conversation as a separate data object, the separate data object comprising a concatenation of messages selected according to the sender and / or receiver and arranged in the order of transmission according to their respective timestamps.
[0032] An instance of a dialogue data object may further include a set of media indicators, such as a copy of an image / video / audio file attached to a message belonging to the corresponding dialogue, or the network address / URL of the corresponding media file. Other exemplary media indicators may include indicators of media format (encoding protocol), etc. Those skilled in the art will understand that the actual data format of the dialogue object may differ in embodiments; exemplary formats include a version of Extensible Markup Language (XML) and JavaScript Object Notation (JSON), etc.
[0033] In some embodiments, in step 302 ( Figure 3In this process, conversation agent 32 may extract conversation sample 22 from messaging application 30. In text processing embodiments, sample 22 may include text fragments, such as the content of individual text messages. Step 304 may further determine conversation context 42 based on sample 22. In some embodiments, conversation context 42 includes selected portions of the conversation, such as a concatenation of a set of recent messages forming a portion of the selected conversation. In a simple instance, context 42 may consist of the entire conversation. Another exemplary conversation context 42 includes a concatenation of N recent messages arranged according to timestamps. In a typical instance, N may range from one to dozens of messages. In some embodiments, conversation context 42 is configured to exclusively contain messages specified by a specific target user or messages exchanged within a specific target group of speakers. The target user or group may be indicated by an operator, for example, via the user interface of chatbot detector 20. In one such instance, an operator may be invited to click / tap or otherwise select a portion of the conversation or a group of speakers from the user interface of messaging application 30. In response, conversation agent 32 may select messages to include in conversation context 42 based on author, timestamp, etc. The following is about Figure 5 Further details on some instances of dialogue context 42.
[0034] In response to determining the dialogue context 42, the agent 32 may transfer the context 42 to the challenge generator 34. In some embodiments, the challenge generator 34 is configured to automatically construct a challenge 44 and at least one alternative response 48 to the corresponding challenge based on the dialogue context 42. Figure 3 (Step 306 in the original text). In some embodiments, challenge 44 includes artificially generated text fragments, such as text messages designed as a potential continuation of an ongoing conversation. To generate challenge 44 and / or alternative responses 48, some embodiments of challenge generator 34 employ an artificial intelligence module implementing a generative language model, such as a set of generative pre-trained transformer (GPT) neural networks. The corresponding AI module may execute on a corresponding client system. In alternative embodiments, the AI module responsible for generating challenge 44 and / or response 48 may execute remotely, for example, on utility server 16. In such embodiments, generator 34 may transmit the encoding of conversation context 42 to server 16 and receive the encoding of challenge 44 and / or response 48 in exchange.
[0035] Figure 4 An exemplary sequence of steps, according to some embodiments of the present invention, is demonstrated by challenge generator 34 to generate challenge 44 and alternative response 48. Thus, Figure 4 The flowchart described in the text is described in detail. Figure 3 In step 306, generator 34 receives an indicator of the dialogue context from dialogue agent 32 in step 402. Figure 5 The illustration describes an exemplary dialogue context 42, which includes a sequence of messages 52a to c from an ongoing conversation. For example, messages 52a to c may form part of an instant messaging conversation (e.g., WhatsApp® communication) or may be part of a conversation posted on an online forum (e.g., Reddit® or Discord®). In some embodiments, the dialogue context 42 selectively contains a subset of messages from the corresponding conversation, the subset of messages being selected based on timestamps (e.g., the most recent 10 messages received in the conversation) and / or based on authors (e.g., the 3 most recent messages received from a particular conversation partner).
[0036] Next, in step 404, generator 34 may employ a generative language model to determine an alternative dialogue continuation 43 based on the dialogue context 42. The alternative continuation 43 is intentionally constructed to predict or simulate a continuation (e.g., a new message) of the ongoing dialogue represented by context 42. The modifier "alternative" applied to both "continuation" and "response" is used herein to indicate that the response is an artifact generated by challenge generator 34, rather than a real message exchanged via messaging application 30. Figure 5 In the example, all items within the dashed box are alternatives. Exemplary alternative sequel 43 simulates or predicts a response to message 52c. Exemplary alternative responses 48a to 48b simulate or predict responses to challenge 44 and alternative dialogue sequel 43, respectively. Conversely, exemplary partner response 46 includes another response to challenge 44, which is received from the actual dialogue partner via messaging application 30.
[0037] Generator 34 can apply any generative language model known in its domain to produce alternative dialogue sequels 43. A typical language model receives an input text fragment and produces an output text fragment that is computationally continued from the input text fragment. The model can be iteratively invoked, where at each step the input is modified based on the output determined at a previous step, for example, by concatenation. Language models are typically pre-trained on large text corpora and are language-specific.
[0038] An exemplary architecture for implementing an AI module of a generative language model includes a convolutional neural network (CNN) layer, followed by dense (i.e., fully connected) layers, which are further coupled to a rectifier (e.g., ReLU or other activation function) and / or a loss layer. Alternative embodiments may include a CNN layer fed into a recurrent neural network (RNN), followed by fully connected layers and a ReLU / loss layer. The convolutional layers effectively multiply the internal representation of each word in the sequence with a weight matrix in its domain, called a filter, to produce an embedding tensor such that each element of the corresponding tensor has contributions from the corresponding word and also from other words neighboring the selected word. Thus, the embedding tensor collectively represents the input word sequence with a coarser granularity than individual words. The filter weights are adjustable parameters that can be tuned during the training process.
[0039] Recurrent Neural Networks (RNNs) form a special category of artificial neural networks where the connections between network nodes form a directed graph. Several types of RNNs are known in this field, including Long Short-Term Memory (LSTM) networks and Graph Neural Networks (GNNs). A typical RNN consists of a set of hidden units (e.g., individual neurons), and the network topology is specifically configured such that each hidden unit receives not only the representation of the corresponding word t... j The input (e.g., an embedding vector) also receives input from a neighboring hidden unit, which in turn receives a representation of the word sequence located at word t. j Previous lexical t j-1 The input is t. As a result, the output of each hidden unit is not only affected by the corresponding word t. j Influenced by, and also by the preceding word element t j-1 The impact of this. In other words, RNN layers can process information about each word within the context of previous words. Bidirectional RNN architectures can process information about each word within the context of both previous and subsequent words in the input word sequence.
[0040] Another exemplary embodiment of the AI module for generating a specific dialogue sequel 43 includes a stack of transformer neural network layers. For example, a transformer architecture is described by A. Vaswani et al. in 'Attention is all you need' (arXiv:1706.03762), etc. For each input lexical sequence, the transformer layer can produce a contextualized sequence of lexical embedding vectors, where each lexical embedding vector encodes multiple (e.g., all) lexical units t from the input sequence. j The information. The output of the transformer layer can be fed into multiple different classifier modules (e.g., dense layers), which are referred to as prediction heads in their respective domains. The prediction heads then determine the output terms and thus construct a continuation of the input term sequence.
[0041] Figure 6 An exemplary process for generating an alternative dialogue sequel 43 is demonstrated in an embodiment that implements a bidirectional encoder representation (BERT) language model from a transformer. Figure 6 The top and bottom of each section illustrate the continuous iterations. A generative language module 50, comprising a set of pre-trained neural networks, is configured to receive dialogue context 42 (see also...). Figure 5 The input word sequence 56a to b (text fragments) is used. In various embodiments, individual words 54a to e may include individual words 54a to d, phrases, numbers, alphanumeric characters, emojis, and punctuation marks 54e to f, etc. Some embodiments further use special words (hereafter referred to as [masked]) to represent placeholders that can receive any word.
[0042] In some embodiments as illustrated, the input lexical sequence further includes a temporary sequel 143 that projects the dialogue context 42 into the future. Initially, the temporary sequel 143 may exclusively consist of placeholder lexical units, such as... Figure 6 The upper part describes this. However, in alternative embodiments, by changing the position and number of [masked] lexical units, module 50 can be coaxed to generate alternative sequels 43 with various desired features, such as predetermined length, predetermined grammar, predetermined sentence patterns (interrogative sentences, exclamatory sentences, etc.), and / or a predetermined set of fixed lexical units (e.g., certain desired keywords and / or special characters, emojis, etc.). Such parameters can be tuned during the training process to produce dialogue sequels that conform to various criteria. The illustrated generative language module 50 is configured to determine output lexical units 54g to h based on the input lexical unit sequences 56a to b, respectively. In some embodiments, in order to proceed to the next iteration, the input lexical unit sequence is modified by replacing one of the [masked] placeholders with the output lexical units generated in the current iteration. Figure 6 In the example, the input lexical sequence 56a is modified by replacing the first [masked] lexical with the output lexical 54g to produce an updated input sequence 56b, which is then modified by replacing another [masked] placeholder with lexical 54h to produce another input lexical sequence for subsequent iterations, and so on. The process can be repeated until all [masked] placeholders are replaced with output lexicals, at which point temporary sequel 143 is derived as alternative dialogue sequel 43.
[0043] In response to determining an alternative dialogue continuation 43, step 406 may distort the dialogue continuation 43 to generate a challenge 44. Some embodiments rely on the observation that deep neural networks are extremely complex systems and therefore prone to exhibiting chaotic behavior. A characteristic of chaotic dynamics is its extreme sensitivity to initial conditions, where tiny differences in the initial state of a chaotic system grow exponentially rapidly—a phenomenon sometimes referred to in popular culture as the "butterfly effect." The consequence of this sensitivity is that a chaotic system can be sent on drastically different trajectories by the slightest push. This behavior translates into the realm of generative language models, where various computer experiments reveal that the same model can produce drastically different outputs when fed slightly different input text. Based on this observation, some embodiments distort the alternative dialogue continuation 43 in an intentional attempt to construct challenge 44, which causes the corresponding chatbot to deviate along a forked trajectory, thus achieving chatbot detection.
[0044] In some embodiments, the challenge generator 34 includes, for example, Figure 7 The sequence modifier 60 described herein is configured to distort the alternative dialogue sequel 43 by selectively applying a set of transformations 64. The exemplary transformation T2 described converts affirmative sentences into interrogative sentences. Another exemplary transformation can rewrite the sequel 43b to change its mood, for example, from neutral to angry, aggressive, sad, happy, excited, etc. Some transformations T... i The selected lexical unit / word / phrase in sequel 43 can be replaced with an alternative that has a predetermined semantic relationship with the replaced item (e.g., antonym, synonym, etc.). Other exemplary transformations T i The selected morphological changes are applied to the selected lexical units, thus altering various grammatical attributes such as tense, mood, person, number, case, and gender. Other exemplary transformations change the spelling of the selected lexical units, for example, by capitalizing them, introducing invisible lexical units, and / or replacing some characters with homographs (characters that look the same but belong to different alphabets or character sets). Yet another exemplary category of distortions includes introducing special characters, punctuation marks, symbols, emoticons, and / or abbreviations with specific meanings in online conversation (e.g., #, @, LOL, :P, etc.).
[0045] exist Figure 7 In the example described above, by selectively replacing some of the lexical units in sequel 43 with [masked] placeholders and iteratively calling generative language module 50 to fill in the masked lexical units (as described above regarding...). Figure 6(As described) to distort the sequel 43. The replaced lexical units may be selected based on a set of grammatical / syntactic rules and / or the results of experiments. For example, experiments may show that certain lexical units / words, characters, and / or emojis are more likely to trigger aberrant chatbot behavior; therefore, some embodiments may intentionally distort the dialogue sequel 43 by including such triggering lexical units.
[0046] In response to the determination of challenge 44, in step 408, challenge generator 34 may determine a set of alternative responses 48a to b that respectively simulate the responses to challenge 44 and sequel 43 (see Figure 5 Step 408 can be performed in a manner similar to step 404 described above, wherein challenge 44 replaces alternative dialogue sequel 43. In some embodiments, the sequence of steps 410 to 412 may then determine whether challenge 44 satisfies a certain predetermined number of conditions, and if not, generator 34 may return to step 406 to generate an alternative challenge, for example by applying different types of distortion to sequel 43.
[0047] Evaluating challenge 44 may include determining the likelihood that the chatbot's response to challenge 44 may differ significantly from a human response. In some embodiments, step 410 includes evaluating the similarity between two lexical sequences (e.g., between alternative dialogue sequel 43 and challenge 44) and comparing the corresponding similarity to a predetermined threshold. Challenge 44 that has not been sufficiently removed from sequel 43 may then be rejected as unsatisfactory.
[0048] In an alternative embodiment, when the difference Dc between challenge 44 and sequel 43 is less than a predetermined upper limit ∆ U Challenge 44 is considered satisfactory based on the observation that successful challenges only derail the robot, while overly strange or out-of-context challenges may trigger non-standard responses from both humans and the robot. Other exemplary embodiments may evaluate the difference Dr between alternative responses 48a and 48b, and determine when the difference exceeds a predetermined lower threshold ∆. L It was determined that Challenge 44 was satisfactory. ∆ U and ∆ L Both can be determined through computer experiments and may be specific to a particular type of chatbot.
[0049] Such criteria can also be combined. In one such instance, when Dc < ∆ U And Dr>∆ L In this case, challenge 44 is considered satisfactory. In another instance, challenge generator 34 determines the composite distance: [1]
[0050] Furthermore, when D is below a predetermined threshold, challenge 44 is deemed satisfactory.
[0051] Challenge generator 34 can quantify the similarity between text sequences using any method known in the relevant domain. Exemplary similarity measures include various variants of edit distance (Levenshtein distance) and distances evaluated in the embedding space, which are described in further detail below. Other similarity measures known in the relevant domain determine a measure of the emotion conveyed by the target text fragment, enabling some embodiments to detect changes in emotion (e.g., from neutral to angry, etc.). Still other exemplary similarity measures can be derived from measures known in the field of machine translation, such as the Bilingual Evaluation Substitute (BLEU) score or the Recall-Oriented Summary Evaluation Substitute (ROUGE) score. When challenge 44 is deemed satisfactory, in step 414 ( Figure 4 Challenge 44 will be output together with the corresponding alternative responses 48a to b.
[0052] In step 308 ( Figure 3 In the process, challenge generator 34 transmits challenge 44 to dialogue agent 32, which in turn instructs messaging application 30 to insert it into the corresponding ongoing dialogue. Further steps 310 to 312 monitor the partner's response to challenge 44, i.e., the response received from the dialogue partner via messaging application 30. In response to receiving partner response 46, in step 314, the response analyzer 36 component of chatbot detector 20 determines a bot decision 26 indicating whether the author of partner response 46 is a chatbot.
[0053] In some embodiments, the response analyzer 36 determines the bot decision 26 based on a measure of similarity between the partner response 46 determined in step 306 and at least one of the alternative responses 48a to b. For example, the partner response 46 may be considered to originate from the chatbot when the difference between the partner response 46 and the alternative response 48b exceeds a predetermined threshold (which may be determined experimentally). Alternative embodiments may further determine the difference between the partner response 46 and the alternative response 48a, and determine that response 46 was created by the chatbot when items 46 and 48a are sufficiently similar. Still other embodiments may make a decision based on the similarity between item 46 and item 48b and based on the similarity between item 46 and item 48a. For example, if response 46 is more similar to alternative response 48a than alternative response 48b, then an exemplary embodiment may determine that the partner response 46 was generated by the chatbot.
[0054] The exemplary similarity metric used in step 314 can be determined based on the distance separating two lexical sequences in an abstract hyperspace (sometimes called the embedding space). In such embodiments, each lexical sequence can be represented as a multidimensional embedding vector comprising multiple numerical coordinates that collectively indicate the position of the corresponding lexical sequence in the embedding space. The individual coordinates of the embedding vector are determined by a component of the chatbot detector 20, typically referred to as the encoder. Figure 8 This demonstrates an exemplary encoder 66 that transforms an input word sequence into an embedding vector 72. Each word t in the input sequence... i It can be represented by a multidimensional representation vector, such as a one-hot encoded vector determined according to a lexical dictionary. Encoder 66 may include an AI module, such as a set of pre-trained neural networks. In some embodiments, encoder 66 forms part of generative language module 50 and is trained co-trained with other components of module 50.
[0055] Figure 8 Further explanation is provided of an exemplary embedding space 70 according to some embodiments of the invention, and a pair of embedding vectors 72a to b representing exemplary alternative responses 48b and partner responses 46, respectively. Those skilled in the art will appreciate that while the illustrated embedding space has only two dimensions, a typical embedding space may have hundreds or thousands of dimensions. Figure 9 Further, an exemplary distance d from the separating vectors 72a to b in the embedding space 70 is presented, where distance d can be used as a measure of the similarity between vectors 72a and b. Various methods for evaluating such similarity measures are known in the field.
[0056] In an alternative embodiment, encoder 66 may compute an embedding vector for each word in the sequence, and the similarity metric may then be computed as an aggregate distance combining the distances between multiple words. Instance-specific text word embeddings may be computed using a version of Word2Vec or GloVe algorithms, etc. To generate such word embedding vectors, encoder 66 may be pre-trained on a text corpus, for example, using bag-of-words and / or skip grammar algorithms.
[0057] The bot decision 26 may include a label (e.g., human, bot, etc.) or a Boolean value indicating whether the author of the partner response 46 is a chatbot. Alternatively, decision 26 may include a number indicating the probability that response 46 was generated by a bot (e.g., a probability scaled between 0 and 1, where 1 indicates determinism). In response to determining decision 26, chatbot detector 20 may display an indicator of decision 26 to the user of the corresponding client system. For example, some embodiments may employ a dialogue agent 32 to accordingly tag the corresponding dialogue partner in the user interface of messaging application 30. In one such instance, unique labels, colors, icons, etc., may be used to highlight chatbot dialogue partners.
[0058] As indicated above, various components of the chatbot detector 20 (e.g., generative language module 50 and encoder 66, etc.) may include pre-trained neural networks. Training as used herein refers to the following process: presenting a set of training samples (e.g., a corpus of texts in natural languages such as English, Chinese, Russian, etc.) to the corresponding neural network; using the network to determine the output based on the training samples; and, in response, adjusting a set of parameters (e.g., synaptic weights, etc.) of the corresponding neural network based on the output. Several training strategies are known in the art, such as supervised and unsupervised training. Training generative language models is typically expensive in terms of computational resources (processing power and memory); therefore, in a typical embodiment, training is conducted at a dedicated AI training facility 14 ( Figure 1 The facility, which executes on the client systems 10a to 10c and / or the utility server 16, includes dedicated hardware such as a graphics processing unit (GPU) array. Facility 14 can further manage a set of training corpora 18 tailored to various applications, natural language processing, and / or chatbot detector 20 components. Training results may include a set of optimal detector parameter values 24 (e.g., neural network synaptic weights, etc.), which can be transferred to a runtime example of the chatbot detector 20 executing on client systems 10a to 10c and / or the utility server 16.
[0059] Some embodiments may completely bypass language model training and use publicly available language models and / or chatbots (such as the implementation from OpenAI's ChatGPT) to embody some of the functionality of chatbot detector 20. For example, challenge generator 34 may invoke a remote chatbot to generate alternative dialogue sequels 43 and / or alternative responses 48a to b. In alternative embodiments, for example, generative language module 50 ( Figures 6 to 7 The components can be executed remotely on the server computer and / or can be outsourced as a service.
[0060] Figure 10 Exemplary hardware configurations of a computer system 80, programmed to perform some of the methods described herein, are shown. Computer system 80 may generally represent... Figure 1 The computer system described herein is any one of client systems 10a to c, message server 12, utility server 16, and AI training facility 14. The computer system described is a personal computer; other devices (e.g., servers, mobile phones, tablets, and wearable devices) may have slightly different configurations. Processor 82 includes physical devices (e.g., microprocessors, multi-core integrated circuits formed on a semiconductor substrate) configured to perform computational and / or logical operations using a set of signals and / or data. Such signals or data may be encoded in the form of processor instructions (e.g., machine code) and delivered to processor 82.
[0061] Processor 82 is typically characterized by an instruction set architecture (ISA), which specifies a set of corresponding processor instructions (e.g., x86 family vs. ARM® family) and register sizes (e.g., 32-bit vs. 64-bit processors). The architecture of processor 82 can be further varied depending on its intended primary use. While the central processing unit (CPU) is a general-purpose processor, the graphics processing unit (GPU) can be optimized for image / video processing and some forms of parallel computing. Processor 82 may further include application-specific integrated circuits (ASICs), such as the Tensor Processing Unit (TPU) from Google®, Inc., and Neural Processing Units (NPUs) from various manufacturers. TPUs and NPUs may be particularly suitable for AI and machine learning applications as described herein.
[0062] Storage unit 84 may include volatile computer-readable media (e.g., dynamic random access memory - DRAM) that stores data / signal / instruction codes accessed or generated by processor 82 during operation. Input device 86 may include a computer keyboard, mouse, microphone, etc., and includes corresponding hardware interfaces and / or adapters that allow users to introduce data and / or instructions into computer system 80. Output device 88 may include a display device (e.g., monitor, speaker, etc.) and hardware interfaces / adapters (e.g., graphics card) that enable the corresponding computing facility to transmit data to the user. In some embodiments, input and output devices 86 to 88 share common hardware (e.g., touchscreen). Storage device 92 includes non-volatile computer-readable media that implements software instructions and / or data storage, retrieval, and writing. Exemplary storage devices include disk and optical disk and flash memory devices, and removable media, such as CD and / or DVD optical discs and drives. Network adapter 94 enables computer system 80 to connect to an electronic communication network (e.g., Figure 1 Network 15) and / or to other devices / computer systems.
[0063] Controller hub 90 typically represents multiple system, peripheral, and / or chipset buses, and / or all other circuitry that enables communication between processor 82 and the remaining hardware components of system 80. For example, controller hub 90 may include memory controllers, input / output (I / O) controllers, and interrupt controllers. Depending on the hardware manufacturer, some of these controllers may be incorporated into a single integrated circuit and / or integrated with processor 82. In another example, controller hub 90 may include a northbridge connecting processor 82 to memory 84, and / or a southbridge connecting processor 82 to devices 86, 88, 92, and 94.
[0064] The exemplary systems and methods described above enable the execution of Turing tests to determine whether online conversation partners include chatbots. Chatbot technology benefits from recent advancements in natural language processing, and specifically from the success of large language models such as Generative Pre-trained Transformers (GPTs). Currently, determining whether a conversation partner is a bot is crucial for many applications, including fraud prevention and combating online misinformation.
[0065] Conventional bot detection methods encompass various types of CAPTCHA, including inviting users to solve specific puzzles (e.g., displaying multiple images and asking users to indicate which one represents a specific item, such as a car or a bicycle), and determining whether a user is a bot based on their response to the puzzle. Conventional CAPTCHA explicitly relies on the limitations of current AI systems in handling certain problems (e.g., image recognition). Recent advancements in AI are rapidly rendering such conventional Turing tests obsolete. In contrast to conventional CAPTCHA tests, some embodiments do not assume any capabilities or limitations of the target chatbot (aside from the ability to generate plausible text), but instead leverage deeper, intrinsic features of generative language models, such as their inherent chaotic nature. Therefore, the systems and methods described herein can be demonstrated to be more reliable than conventional Turing tests based on CAPTCHA techniques in detecting AI.
[0066] The chatbot detector described herein can directly interface with messaging applications to effectively participate in ongoing online conversations. Exemplary conversation agent components can be added to the corresponding messaging application as extensions or plugins. These components determine the context of an ongoing conversation, assign individual images to their respective users, and instruct the messaging application to submit continuations to the corresponding conversation. The chatbot detector can formulate challenges in the form of at least one conversational message, listen for responses to the challenges, and determine whether the corresponding conversation partner is a chatbot based on the responses.
[0067] Some embodiments rely on the observation that generative language models adopted by chatbots may exhibit chaotic behavior due to their complexity. A characteristic of chaotic systems is their extreme sensitivity to initial conditions; that is, subtle differences in initial conditions can lead to drastically different futures. The chaotic nature of language models can explain, for example, why chatbots sometimes output surprising, out-of-context statements—a property known in the field as illusion. Some embodiments of the invention explicitly use such trajectory bifurcation for chatbot detection by carefully and deliberately constructing challenges that will cause the chatbot to deviate from its expected behavior.
[0068] To take advantage of sensitivity to initial conditions, some embodiments construct alternative dialogue sequels 43 (e.g., see...). Figure 5 The process involves a reasonable continuation of the ongoing dialogue. This dialogue continuation is then slightly distorted to attempt to elicit a response from the bot in a way that is not expected in the current context of the dialogue. The resulting challenge 44 can then be added as a new message to the ongoing dialogue. The chatbot detector assesses the similarity between a response 46 received from the corresponding dialogue partner to the corresponding challenge and a manually generated, reasonable response 48b to the undistorted dialogue continuation. Significant differences between the partner's response 46 and the alternative response 48b can indicate a chaotic trajectory bifurcation and thus reveal that the corresponding dialogue partner is a chatbot.
[0069] The described system and method further rely on the observation that a successful challenge 44 should deviate slightly from a reasonable dialogue continuation 43 to avoid causing the human dialogue partner to react in an unexpected way, thus leading to false positive detections. Therefore, some embodiments run a set of quality tests on candidate challenges before actually submitting the corresponding challenge to the messaging application. In one instance, a challenge is considered satisfactory when the difference between the corresponding challenge 44 and the undistorted dialogue continuation 43 is within predetermined upper and lower limits. In another instance, the chatbot detector can generate a reasonable response 48a to the corresponding challenge. Then, the corresponding challenge 44 can be considered satisfactory if it remains far from the predetermined upper limit of the undistorted dialogue continuation 43, while an alternative response 48a to the corresponding challenge is farther from the alternative response 48b to the undistorted dialogue continuation than the predetermined lower limit.
[0070] Instances of distortions applied to the construction challenge include modifying the generated dialogue continuation by introducing specific keywords, special characters, and / or emojis; changing the mood of the dialogue continuation (e.g., from neutral to angry); and changing sentence structure (e.g., from affirmative to negative or interrogative). Specific types of distortions used by the chatbot detector may be updated from time to time to keep pace with advancements in chatbot technology. The selection of distortions can be informed through direct computer experiments, which include various attempts to induce hallucinations or trajectory divergences in real online chatbots.
[0071] It will be apparent to those skilled in the art that the above embodiments can be modified in various ways without departing from the scope of the invention. Therefore, the scope of the invention should be determined by the appended claims and their legal equivalents.
Claims
1. A computer system comprising at least one hardware processor, said at least one hardware processor being configured to: Generative language models are applied to generate alternative dialogue sequels and alternative responses, wherein: The alternative dialogue sequel includes a predicted continuation of an ongoing online conversation, which includes a sequence of electronic messages. The alternative response includes a predicted response from the dialogue partner to the alternative dialogue sequel; In response, the alternative dialogue sequel is distorted to generate a challenge message; Add the challenge message to the ongoing online conversation; and In response to receiving a partner response from the dialogue partner, the partner response including a response to the challenge message, the similarity between the partner response and the alternative response determines whether the dialogue partner includes a bot.
2. The computer system of claim 1, wherein the at least one hardware processor is configured to determine whether the dialogue partner includes the robot based on the result of comparing a similarity metric with a predetermined threshold, wherein the similarity metric quantifies the similarity between the partner's response and an alternative response.
3. The computer system of claim 1, wherein the at least one hardware processor is configured to further determine whether the dialogue partner includes the robot based on the similarity between the partner's response and another alternative response including a predicted response to the challenge message from the dialogue partner.
4. The computer system of claim 3, wherein the at least one hardware processor is configured to determine whether the dialogue partner includes the robot based on the result of comparing a first similarity metric with a second similarity metric, wherein the first similarity metric quantifies the similarity between the partner's response and the alternative response, and the second similarity metric quantifies the similarity between the partner's response and the other alternative response.
5. The computer system of claim 1, wherein distorting the alternative dialogue sequel comprises selecting items from a set consisting of: replacing selected words in the alternative dialogue sequel with alternative words, adding words to the alternative dialogue sequel, rewriting the alternative dialogue sequel as a question, and rewriting the alternative dialogue sequel to change the mood of the alternative dialogue sequel.
6. The computer system of claim 1, wherein the at least one hardware processor is further configured to: In response to the generation of the challenge message, it is determined whether the challenge message meets the quality conditions based on the similarity between the challenge message and the alternative dialogue sequel, and In response, the challenge message is added to the ongoing online conversation only if the challenge message meets the quality conditions.
7. The computer system of claim 6, wherein the at least one hardware processor is configured to determine whether the challenge message satisfies the quality condition based on the result of comparing a similarity metric with a predetermined threshold, wherein the similarity metric quantifies the similarity between the challenge message and the alternative dialogue sequel.
8. The computer system of claim 6, wherein the at least one hardware processor is configured to further determine whether the challenge message satisfies the quality condition based on the similarity between the alternative response and another alternative response, including another predicted response to the challenge message from the dialogue partner.
9. The computer system of claim 1, wherein the ongoing online conversation comprises items selected from a group consisting of: the exchange of messages performed via an instant messaging application running on the computer system, a sequence of messages posted to an online forum, and a sequence of messages posted to a social media page.
10. The computer system of claim 1, wherein applying the generative language model includes transmitting an encoded fragment of the ongoing online conversation to a remote chatbot, and in response, receiving the alternative conversation sequel or the alternative response from the remote chatbot.
11. A chatbot detection method, comprising employing at least one hardware processor of a computer system to: Generative language models are applied to generate alternative dialogue sequels and alternative responses, wherein: The alternative dialogue sequel includes a predicted continuation of an ongoing online conversation, which comprises a message sequence, and The alternative response includes a predicted response from the dialogue partner to the alternative dialogue sequel; In response, the alternative dialogue sequel is distorted to generate a challenge message; Add the challenge message to the ongoing online conversation; and In response to receiving a partner response from the dialogue partner, the partner response including a response to the challenge message, the similarity between the partner response and the alternative response determines whether the dialogue partner includes a bot.
12. The method of claim 11, further comprising determining whether the dialogue partner includes the robot based on the result of comparing a similarity metric with a predetermined threshold, wherein the similarity metric quantifies the similarity between the partner's response and an alternative response.
13. The method of claim 11, further comprising determining whether the dialogue partner includes the robot based on the similarity between the partner's response and another alternative response including a predicted response to the challenge message from the dialogue partner.
14. The method of claim 13, further comprising determining whether the dialogue partner includes the robot based on the result of comparing a first similarity metric with a second similarity metric, wherein the first similarity metric quantifies the similarity between the partner's response and the alternative response, and the second similarity metric quantifies the similarity between the partner's response and the other alternative response.
15. The method of claim 11, wherein distorting the alternative dialogue sequel comprises selecting items from a set consisting of: replacing selected words in the alternative dialogue sequel with alternative words, adding words to the alternative dialogue sequel, rewriting the alternative dialogue sequel as a question, and rewriting the alternative dialogue sequel to change the mood of the alternative dialogue sequel.
16. The method of claim 11, further comprising employing the at least one hardware processor to: In response to the generation of the challenge message, it is determined whether the challenge message meets the quality conditions based on the similarity between the challenge message and the alternative dialogue sequel, and In response, the challenge message is added to the ongoing online conversation only if the challenge message meets the quality conditions.
17. The method of claim 16, further comprising determining whether the challenge message satisfies the quality condition based on the result of comparing a similarity metric with a predetermined threshold, wherein the similarity metric quantifies the similarity between the challenge message and the alternative dialogue sequel.
18. The method of claim 16, further comprising determining whether the challenge message satisfies the quality condition based on the similarity between the alternative response and another alternative response, including another predicted response to the challenge message from the dialogue partner.
19. The method of claim 11, wherein the ongoing online conversation comprises items selected from a group consisting of: the exchange of messages performed via an instant messaging application running on the computer system, a sequence of messages posted to an online forum, and a sequence of messages posted to a social media page.
20. The method of claim 11, wherein applying the generative language model includes employing the at least one hardware processor to transmit encoded segments of the ongoing online conversation to a remote chatbot, and in response, receiving the alternative conversation sequel or the alternative response from the remote chatbot.
21. A non-transitory computer-readable medium storing instructions that, when executed by at least one hardware processor of a computer system, cause the computer system to: Generative language models are applied to generate alternative dialogue sequels and alternative responses, wherein: The alternative dialogue sequel includes a predicted continuation of an ongoing online conversation, which comprises a message sequence, and The alternative response includes a predicted response from the dialogue partner to the alternative dialogue sequel; In response, the alternative dialogue sequel is distorted to generate a challenge message; Add the challenge message to the ongoing online conversation; and In response to receiving a partner response from the dialogue partner, the partner response including a response to the challenge message, the similarity between the partner response and the alternative response determines whether the dialogue partner includes a bot.