Virtual assistant for game discovery
By training a neural network to replicate a character's voice and behavior, and combining this with a natural language processing module and database queries, a personalized virtual assistant is generated. This solves the problem of users lacking human-like assistance in applications, and improves user experience and interaction.
Patent Information
- Application Number
- CN202480050270.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-31
- Filing Date
- 2024-05-23
- Publication Date
- 2026-03-06
AI Technical Summary
When users need help operating the application, existing technologies typically rely on reading manuals or searching online, lacking a friendly and user-friendly interaction. Furthermore, in a sales environment, there is a lack of direct communication with sales personnel, which affects user experience and willingness to purchase.
By training a neural network with machine learning algorithms to replicate a character's voice and behavior, and combining this with a natural language processing module and database queries, a personalized virtual assistant is generated, providing customized responses and application operation assistance.
It enables user-friendly interaction within the application, providing personalized help and guidance, enhancing the user experience, especially in gaming and sales environments to increase user engagement and purchase intent.
Smart Images

Figure CN121620797A_ABST
Abstract
Description
Technical Field
[0001] Various aspects of this disclosure relate to providing virtual assistance; more specifically, various aspects of this disclosure relate to providing customized virtual assistance agents for applications. Background Technology
[0002] Literature, film, video games, and other media have created numerous memorable characters. These characters typically possess unique personalities, habits, and voice patterns. Unfortunately, people can usually only experience these memorable characters within the mediums that originally created them. Places like theme parks allow people to experience characters using performers trained to mimic and dress as them. It would be even more fun and exciting if people could interact with their favorite characters without having to go to a theme park.
[0003] When operating an application, users often encounter problems or need assistance. The thought of having to read a manual or search for help online can be daunting. Users may prefer answers to their questions delivered through a friendly explanation. Additionally, in a sales environment, people are generally more likely to make a purchase when they can communicate with a salesperson.
[0004] It is against this backdrop that the various aspects of this disclosure have emerged. Attached Figure Description
[0005] The teachings of this disclosure can be readily understood by considering the following detailed description in conjunction with the accompanying drawings, wherein:
[0006] Figure 1 These are screenshots of an example implementation of a virtual assistant based on aspects of this disclosure.
[0007] Figure 2 These are screenshots of an example implementation of a virtual assistant with a game-playing module according to aspects of this disclosure.
[0008] Figure 3 This is a block diagram illustrating a device for implementing a virtual agent according to aspects of this disclosure.
[0009] Figure 4 This is a flowchart depicting the operation method of a role with optional application operation capabilities, based on aspects of this disclosure.
[0010] Figure 6 This is a flowchart depicting the operation method of replicating a role according to aspects of this disclosure.
[0011] Figure 7A It is a simplified node graph of a recurrent neural network according to aspects of this disclosure.
[0012] Figure 7BIt is a simplified node graph of an expanded recurrent neural network according to aspects of this disclosure.
[0013] Figure 7C This is a simplified diagram of a convolutional neural network according to aspects of this disclosure.
[0014] Figure 7D This is a flowchart of a method for training a neural network according to aspects of this disclosure.
[0015] Figure 8 This is a block diagram of a system for implementing a virtual assistant according to aspects of this disclosure. Detailed Implementation
[0016] Although the following detailed description contains many specific details for illustrative purposes, those skilled in the art will understand that many variations and modifications to these details are within the scope of this disclosure. Therefore, examples of embodiments of the present disclosure described below are set forth without losing the generality of the claimed embodiments and without imposing limitations on the claimed embodiments.
[0017] Movies, books, radio dramas, short stories, and video games offer unique characters with distinct personalities and traits. The sources of these unique characters can provide a wealth of data, which can be used to train generative systems to replicate those characters. Generative machine learning systems allow for the creation of unique responses to questions. Furthermore, generative machine learning systems can connect to databases to query answers to questions, or be trained on corpora containing answers to questions that users of the generative machine learning system might ask. In this way, help systems can be created that use the familiar personalities of characters to provide useful information to users in unique and engaging ways.
[0018] Figure 1 These are simulated screenshots of an example implementation of a virtual assistant according to aspects of this disclosure. Here, screenshot 100 shows a video game application displayed on the system screen. A dialog box 101 is displayed on the screen, where the user can enter questions or comments about the virtual assistant. The virtual assistant 103 is displayed as an animated avatar of a character near a corner of the screen. The virtual assistant 103 is a Viking-type character from a different application, with a rough and older style of English speech. The user asks the virtual assistant, for example via text or voice-to-text input, how to defeat enemies in the scene. The virtual assistant 103 responds with instructions 102 in the character's voice regarding how to defeat the enemies. Furthermore, the virtual assistant is integrated with the game, allowing boxes 104 to be displayed around the enemies on the screen, showing the enemies' weak points.
[0019] The creation of this virtual assistant can utilize the integration of several different components within the system to generate representations of a character that also provides assistance to the user. First, the virtual assistant may include a neural network trained with machine learning algorithms to replicate the character. As will be discussed later, the neural network can be trained on a corpus containing one or more game, movie, web, and literary data containing the character to be replicated. Replicating the character may include, for example, but not limited to, simulating the character's style of phrasing, vocabulary, slang, and semiotics. In some implementations, the neural network may use contextual information to construct responses. By leveraging a machine learning-trained virtual assistant and a diverse corpus of data including the character to be replicated, unique stylistic aspects of the character's behavior can be simulated using the virtual assistant. Here, character replication allows for new expressions of the character's manner of speaking in text that are not found in the corpus. A natural language processing module can be used to convert textual questions from the user into a machine-readable form and generate vectors representing questions that can be provided to the virtual assistant. The virtual assistant can be further trained on data corresponding to answers to questions about the application or other data sources.
[0020] Virtual assistants can be trained on a wide variety of questions and data. These can include application-specific questions. For example, if the virtual assistant is designed for a specific video game, it can be trained on questions related to game mechanics, story elements, character backstories, optimal strategies, controls, and so on. The agent can also be trained to handle general game queries, such as "Based on my provided preferences, what game genres might I like?" or "What are some of the top games in this genre?" Additionally, virtual assistants can be trained to answer questions related to trivia and general knowledge. For games with complex backstories and universes, the agent can answer questions about in-game lore, character histories, game world geography, and more. For example, in a game like World of Warcraft, the agent can answer detailed questions about complex aspects of the game.
[0021] In some implementations, the virtual assistant can be further trained to provide assistance during in-game activities. As an example, a player might ask how to use available in-game items, for instance, during a specific activity. Alternatively, a player might ask for information about a particular character in the game and their relationship to the player character.
[0022] In some implementations, machine-readable text can be provided to a search module connected to a database. The search module can search the database for responses to the text, and if found, can provide the response from the database to the virtual assistant. The virtual assistant can then use the response to replicate a character. A character module can place a visual representation onto the response generated by the virtual assistant. The character module can include animations and / or images of the character. For example, but not limited to, the character module can have a database of animated speech of the character. In some implementations, the character module can include a neural network trained with machine learning algorithms to create customized images of the character. The character neural network can be trained using images of the character from a corpus of training data. Finally, a text-to-speech module can be used to convert the virtual assistant's text output into sound data that can be played through a speaker. The text-to-speech module can include a database of audio samples of the character's voice. In some implementations, machine learning techniques such as Hidden Markov Models or deep learning neural networks can be developed or trained to model and synthesize speech using audio samples. The text-to-speech module can be operable to replicate the character's speech style. The character's speech style can include one or more of a group of intonation, speed, pronunciation, and prosody.
[0023] Figure 2 These are simulated screenshots of an example implementation of a virtual assistant with a game-playing module according to aspects of this disclosure. In the illustrated exemplary implementation, a dialogue 201 with the user is evaluated by a natural language processing module, and the content triggers the game-playing module. The game-playing module may be operable to operate the application 200. Here, a video module is operable to work in conjunction with the application to display the virtual assistant 203 as a character's customized car. The virtual assistant responds with dialogue 201, which replicates how a character might respond to dialogue. While the game-playing module is operating, the virtual assistant may monitor the game status and periodically comment on the game's progress. For example, but not limited to, if the virtual assistant 203's car is passed by the player 204, the virtual assistant may respond with a short statement like "Good job!" or "Back here!" The character module of the application may display a label 202 on the character's visual representation. In this way, the virtual assistant may be able to operate the application with the user or show the user how to perform actions within the application. The application may be, for example, but not limited to, a game, an educational application, a word processor, a spreadsheet program, a storefront application, or a multimedia player. In a store-based implementation, the virtual assistant can receive information about the products in the store or be trained on product data to provide responses about products in a style that mimics the user's personality.
[0024] Figure 3This is a block diagram illustrating an apparatus for implementing a virtual agent according to various aspects of the present disclosure. The apparatus shown may include one or more processors 313 and memory 302. The one or more processors may include one or more central processing units and / or one or more graphics processors. The memory may be implemented as random access memory (RAM), read-only memory (ROM), or a computer-readable medium. The memory may include instructions that cause the processors to implement a virtual assistant 314. Additionally, the memory may include instructions for implementing an application 304 using the processors and / or a database 305 including information about the application.
[0025] The virtual assistant 314 may include one or more neural networks 306 trained with machine learning algorithms to replicate a character. In some implementations, the neural network may be trained to replicate two or more different characters, and the user may be given the option to choose between different characters. Additionally, the virtual assistant may utilize dedicated modules such as a text-to-speech module 307, natural language processing (NLP) 308, a game-playing module 309, a character module 310, a video module 311, and / or a negation representation module 312.
[0026] The virtual assistant 314 can receive data from the application 304 and / or the database 305. For example, the virtual assistant can receive application state information, which can be used to determine the context surrounding a question raised by the user. The application state information can also provide periodic updates for player actions within the application, allowing the virtual assistant to periodically generate comments based on the application state. The database may include search algorithms and can receive information about queries from the NLP module and return response data to the virtual assistant.
[0027] The text-to-speech module 307 can be programmed sufficiently to synthesize speech data from the output of the virtual assistant neural network 306. The text-to-speech module enables the processor to perform a speech encoding step, which maps words provided by the virtual assistant neural network to phonemes using speech rules and a speech dictionary. The phonemes can then be modeled using speech modeling, which captures the stress patterns, velocity, rhythm, and intonation of the speech. This can be performed using a speech neural network trained with machine learning algorithms to determine the stress patterns, velocity, rhythm, and intonation of the speech. Alternatively, the neural network can be trained on a corpus of character data to replicate the character's speech style. In some implementations, the speech neural network can be trained on speech recordings using machine learning algorithms.
[0028] Once speech modeling is complete, speech can be synthesized using, for example, but not limited to, concatenation synthesis or formant synthesis. Concatenation synthesis uses pre-recorded sounds from a database, which can be obtained from a corpus of character data. In other implementations, pre-recorded speech of the character is not available, and a general recording of speech can be used. Alternatively, formant synthesis can be used, which uses manipulation of audio waveforms to create phonemes with characteristics selected to simulate the character's speech or a speech prototype associated with the character's prototype. For example, a character may have a warrior prototype, and certain speech characteristics corresponding to the warrior prototype can be used to simulate the character's speech. The audio waveforms can be further post-processed to smooth between phoneme waveforms or recorded speech segments, add intonation, normalize samples, or adjust the amplitude of sound between phoneme waveforms, etc.
[0029] NLP module 308 converts user-input text into a machine-readable form. The NLP module can tokenize the user-input string and identify keywords and / or phrases. Tokenization of the user-input string can include tasks such as sentence boundary detection, identification of individual tokens (words, punctuation marks, symbols, etc.), and parsing sentences into phrases. Keyword identification can include identifying interrogative words, locations, exclamation marks, etc. The NLP module can extract words and phrases related to one or more applications 304 from database 305. Database 305 can contain a dictionary of words and / or phrases related to one or more applications 304, which can be used as a data source for the NLP module. Additionally, the NLP module can include one or more machine learning techniques for keyword identification. For example, the NLP module can utilize neural networks, support vector machines, or hidden Markov models trained with machine learning algorithms. The NLP module can provide machine-readable text to virtual assistant neural network 306.
[0030] The game execution module 309 can be programmed sufficiently to provide input to one or more applications 304, thereby allowing the virtual assistant to operate the applications. In some implementations, the application 304 is operable to allow the game execution module 309 to collaborate with the user. In implementations where the application 304 is a collaborative game, the game execution module 309 can be operable to collaboratively play the game with the user.
[0031] In some implementations, the game execution module 309 can generate scripts or commands to be executed in response to changes in the game state. The system implementing the virtual assistant can be designed to interpret the game state. As an example, and not by limitation, the game state can include information about the user's actions, in-game events or environmental changes, AI behavior, etc. The virtual assistant can then use this input to inform its own actions. This continuous process of receiving game state input, interpreting it, and then generating scripts to interact with the game application allows the game module to act as an effective agent in the game, working collaboratively with the user or independently. This concept also realizes the potential of collaborative games, where the game module and the user continue to interact and collaborate to achieve common game goals.
[0032] As an example, consider a scenario where the game execution module 309 is provided with a game wiki, which it uses to generate answers to player questions. The game wiki comprises a corpus detailing activities, plots, characters, and other game-related information. If a player asks the virtual assistant how to solve a specific problem, the game execution module 309 can retrieve the correct knowledge excerpt relevant to that problem from the wiki. This excerpt can then be submitted to a large language model, prompting the user to use it to answer the player's question. The large language model responds and presents the response to the player.
[0033] In some implementations, the game-playing module can implement a simple AI character built into the application. The game-playing module is operable to operate in conjunction with the application, causing the application to label the simple AI character as a virtual assistant. In alternative implementations, the game-playing module can include a neural network trained with machine learning algorithms to input input into the application and operate the application based on application state updates. The game-playing module can also monitor application state updates provided by one or more applications 304. For example, but not limited to, application state data can indicate to the game-playing module that the application is a racing game. In this case, the game-playing module can activate a neural network trained to play a racing game application. The game-playing neural network can typically be trained on a given type of application (e.g., racing games, shooting games, puzzle games, learning games, etc.), or a separate neural network can exist for each application, where each separate neural network is specifically trained to operate the application based on application state updates.
[0034] The character module 310 can be programmed to display a visual representation of a character. The character module may include a database of images of one or more characters that can be copied by a virtual assistant. The images may include characters in different orientations, positions, and with different facial expressions. Additionally, in some embodiments, the character module may also have character animations. For example, but not limited to, the character module may include animations of characters speaking and moving their heads. The character module may send visual representation data of the character to a processor, where the visual representation data can be rendered, and the rendered image can be sent to a display. In some embodiments, the character module may provide overlay information to the processor, instructing the processor to overlay the visual representation of the character onto image frames from the application. The overlay information can provide the position and size of the visual representation of the character in the overlay. The overlay map can be customized by the user, allowing the user to select the size of the visual representation and / or its position on the display. For example, the character module 310 may be operable to allow the user to customize the character's appearance, voice, and other attributes, such as clothing. Furthermore, the character module 310 is operable to dynamically depict the character's emotions based on the dialogue context. Emotions may be displayed as facial expressions, body language, gestures, etc. In addition, the character module can be operated to provide options for hidden subtitles, sign language, etc.
[0035] The character module can also synchronize with the output of the virtual assistant's neural network, so that the character's visual representation and the virtual assistant's neural network's response appear on the screen simultaneously. The character module can also be responsible for generating the display of responses from the virtual assistant's neural network. For example, the character module can generate visual data for displaying responses generated by the virtual assistant's neural network. Furthermore, in implementations with animated visual representations, the animation of the visual representation can be timed to match the display of text responses from the virtual assistant's neural network or sound from the text-to-speech module.
[0036] The video module 311 can be programmed sufficiently to provide video based on responses from the virtual assistant neural network 306. The video module may include a video database organized by keywords and / or phrases. When keywords and / or phrases are detected in the responses from the virtual assistant neural network, the video module can extract video data from the video database. The video data can then be sent to a display. Additionally, the video module can instruct the processor to overlay video and / or text and / or graphical content onto image frames of the application. In some alternative implementations, the video module can search a remote video database for keywords and / or phrases detected in the responses from the virtual assistant neural network 306.
[0037] The negation module 312 can be programmed sufficiently to review responses from the virtual assistant neural network 306 that are inappropriate, disruptive to the application, or make illegal statements. The negation module may include a negation neural network trained with machine learning algorithms to detect inappropriate responses from those generated by the virtual assistant neural network 306. Additionally, the negation module neural network may be trained to identify legally problematic statements. Finally, in some implementations, the negation module includes spoiler prevention. Spoiler prevention may receive application state data and use it to track the user's progress within the application. The negation module may also include a spoiler table of keywords and / or phrases indexed by the application state data. When a keyword or phrase in a response from the virtual assistant neural network is detected to be inconsistent with the current application state, that keyword or phrase may be removed. Alternative phrases indicating that a question or response is inappropriate, spoiler-like, or legally problematic can be selected from a table of alternative phrases. For example, but not limited to, responses from the virtual assistant's neural network can reference the name of an item from the application, and a spoiler table can indicate that the application state in which the item appears has not yet been satisfied (e.g., the item entry found in the table does not match the current application state). Therefore, the negation module can replace the response generated by the virtual assistant's neural network with a replacement statement (e.g., "No spoilers" or "I can't tell you yet"). In some implementations, a key statement entered by the user into the virtual assistant 314 (e.g., "Spoilers for me" or "Tell me quickly") can pause the operation of the negation module of the previous statement and provide the response generated by the virtual assistant's neural network to the user. If the response is inappropriate or has legal issues, the key statement may not pause the operation of the negation module. In alternative implementations, instead of generating a replacement response, the virtual assistant's neural network can be forced to generate different responses to user prompts.
[0038] In some implementations, the virtual assistant neural network 306 may be a generative language model trained with a corpus of character information to produce text that replicates the character. In some implementations, the virtual assistant neural network 306 may include one or more neural networks trained to replicate multiple different characters. In some alternative implementations, multiple virtual assistant neural networks may exist, each customized to replicate a different character. One or more virtual assistant neural networks 306 may initially be based on a pre-trained model, which can then be customized via transfer learning to produce text that replicates the character. Examples of pre-trained models that can be used for transfer learning include Generative Pre-trained Transformer (GPT), Bidirectional Encoder Representation from Transformer (BERT), or Efficient Encoder for Accurate Classification of Label Replacements (ELECTRA). These pre-trained models can then be refined using transfer learning. Transfer learning can refine the pre-trained model using a corpus of character information to respond as if the character might respond to a question. Transfer learning can further endow the model with application-specific knowledge. In unsupervised training, the task of the virtual assistant neural network may be to generate the next token in a token sequence (e.g., to generate the next word or letter in a sentence). For more information on transfer learning, see “A Comprehensive Survey on Transfer Learning” by Fuzhen Zhuang et al., Proceedings of the IEEE, Vol. 109, No. 1, pp. 43-76, Jan. 2021, which is incorporated herein by reference. For more information on generative models, see “Improving Language Understanding by Generative Pre-training” by Radford et al., Open AI, (2018) and “ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators” by Clark et al., ArXivabs / 2003.10555, (2020), both of which are incorporated herein by reference. In some alternative implementations, one or more neural networks can be trained from scratch using semi-supervised techniques and can have a transformer-type architecture. The trained neural network can then be further customized with a corpus of character data to model the response styles of one or more characters.
[0039] Figure 4This is a flowchart depicting a method of replicating the operation of a character with optional application operational capabilities according to aspects of this disclosure. Initially, the user can input a prompt 401 into the system. The prompt can be, for example, but not limited to, a text string entered into a text box or some other field in the system. In an alternative implementation, speech recognition can be used to convert the user-spoken prompt into a text prompt. The prompt can be, for example, but not limited to, a question or statement that the user intends to elicit a response from the system. NLP 407 then evaluates the prompt. NLP tokenizes the prompt and performs some keyword recognition tasks on the prompt. Keywords from the prompt can be detected by the game execution module 403, and a game execution can be initiated with the application 408. Simultaneously, the virtual assistant neural network acquires the tokenized prompt and keywords and predicts a response 402 that the replicated character might produce. The prediction of the response can use application state data 404 from the application 408 to create an appropriate and / or customized response. Once the response is generated, it can be provided to the character module, which can display the generated response on the display 406. The character module can also use the response subsequently displayed 406 to generate a visual representation 405 of the character. Additionally, when the game module 403 is active, the application's operations can be customized by the character module to present operations performed using the game module as if performed by a character. For example, but not limited to, the character module can provide labels in the overlay of the application's image frames that represent AI entities as characters, such as... Figure 2 As shown.
[0040] Figure 5 This is a flowchart depicting an operational method for searching a replica role using a database according to aspects of this disclosure. Initially, the user can input a prompt 501 into the system. The prompt can be, for example, but not limited to, a text string entered into a text box or some other field in the system. In an alternative implementation, speech recognition can be used to convert the user-spoken prompt into a text prompt. The prompt can be, for example, but not limited to, a question or statement intended by the user to elicit a response from the system. NLP 507 then evaluates the prompt. NLP tokenizes the prompt and performs some keyword recognition tasks on the prompt. The keywords detected by NLP can be used to query the database 502. The database 503 can then return a basic response based on the query 504. The basic response is then passed to a virtual assistant neural network, in which a response for the replica role is predicted based on the basic response from the query 505. The predicted response is then provided to the user 506.
[0041] Figure 6This is a flowchart depicting the operation of a replicated character according to aspects of this disclosure. In some embodiments, the application may be a storefront application, such as the Google Play Store, PlayStation Store, Xbox Store, etc. In these embodiments, the character's statements may be particularly problematic because they may contain incorrect statements or promises that could have legal implications. Therefore, the legal representation module can prevent the virtual assistant from generating legally problematic, including spoiler-filled, or inappropriate responses.
[0042] Initially, the user can input a prompt 601 into the system. The prompt can be, for example, but not limited to, a text string entered into a text box or some other field in the system. In an alternative implementation, speech recognition can be used to convert the user-spoken prompt into a text prompt. The prompt can be, for example, but not limited to, a question or statement that the user intends to elicit a response from the system. The NLP 602 then evaluates the prompt. The NLP tokenizes the prompt and performs some keyword recognition task on it. The virtual assistant neural network can then use the tokenized prompt to predict a response 602 that replicates the response from the character. For example, the NLP logic can retrieve relevant information 603 from a knowledge base and use game context 605 to refine the information, such as what equipment the player possesses, the player's location in the game world, the player's active tasks, etc. The resulting refined information 607 can be sent to one or more AI models (e.g., LLM or GPT models) to generate a response, as indicated at 608. The response can then be validated against a policy at 610 before being sent to the player at 612. In some implementations, the validation of the response can include evaluating a negative representation 611 by a negative representation module. Negative representations can be, for example, but not limited to, responses that are erroneous, misleading, inappropriate, and / or contain spoilers. The negative representation module may include a neural network trained with machine learning algorithms to identify statements that constitute negative representations. Additionally, the negative representation module may receive application state information 613, which can be used to determine whether a response is a spoiler, since the response includes one or more keywords or phrases that have not yet been revealed to the player in its current application state. In some embodiments, the negative representation module may also receive user profile information 615, which can be used to determine whether the user has reached a location further than the current application state or has settings that allow application information to be corrupted. In some embodiments, if a response is found to contain a negative representation, the virtual assistant neural network may regenerate the response at 614 based on a prompt. In some alternative embodiments, the response may be replaced by an audio response instead of a generated response. The audio response may be associated with a negative representation and / or may be selected from a list of audio responses. For example, but not limited to, if the response is a spoiler, the default response might be "no spoilers," and if the response is inappropriate, the default response might be "sorry, but I know nothing about it." Once the negative representation of the response has been checked and no negative representation is found in the response, it can be passed to the user at 612.
[0043] Generalized Neural Network Training
[0044] The neural network (NN) in the virtual assistant discussed above can include one or more of several different types of neural networks and can have many different layers. As an example and not a limitation, the neural network can include one or more convolutional neural networks (CNN), recurrent neural networks (RNN), dynamic neural networks (DNN), and / or generative pre-trained transformers (GPT). Many of the neural networks described herein can be trained using the general training methods disclosed herein. For more information on generative pre-trained transformers and their training, see Ashis Vaswani et al., “Attention Is All You Need,” arXiv:1706.03762 (December 5, 2017), which is incorporated herein by reference.
[0045] As an example, not a limitation, Figure 7A This describes a basic form of an RNN that can be used, for example, in training a model. In the example shown, the RNN has one layer of 720 nodes, each characterized by an activation function S, an input weight U, recurrent hidden node transition weights W, and an output transition weight V. The activation function S can be any nonlinear function known in the art and is not limited to the hyperbolic tangent (tanh) function. For example, the activation function S can be a sigmoid or ReLU function. Unlike other types of neural networks, an RNN has a set of activation functions and weights used throughout the layer. Figure 7B As shown, an RNN can be considered as a series of nodes 720 that move through times T and T+1 with the same activation function. Therefore, an RNN maintains historical information by feeding the results from the previous time T to the current time T+1.
[0046] In some implementations, convolutional RNNs can be used. Another type of RNN that can be used is a Long Short-Term Memory (LSTM) neural network, which adds memory blocks with input gate activation functions, output gate activation functions, and forget gate activation functions to the RNN nodes to create gated memory, allowing the network to retain certain information for a longer period of time, as described in Hochreiter & Schmidhuber, “Long Short-Term Memory”, Neural Computation 9(8): 1735-1780 (1997), which is incorporated herein by reference.
[0047] Figure 7CAn example layout of a convolutional neural network (such as a convolutional recurrent neural network (CRNN)) is depicted, which can be used, for example, in a trained model according to various aspects of this disclosure. In this depiction, the convolutional neural network is generated for an input 732, which has a size of 4 units in height and 4 units in width, for a total area of 16 units. The depicted convolutional neural network has a filter 733 with a height of 2 units and a width of 2 units, a jump value of 1, and a channel size of 9 for 736. For clarity, in... Figure 7C Only the connection 734 between the first column channel and its filter window is depicted. However, aspects of this disclosure are not limited to such an implementation. According to aspects of this disclosure, the convolutional neural network may have any number of additional neural network node layers 731, and may include layer types of any size such as additional convolutional layers, fully connected layers, pooling layers, max pooling layers, local contrast normalization layers, etc.
[0048] like Figure 7D As shown, training the neural network begins with initializing the weights at position 741. Typically, the initial weights should be randomly distributed. For example, a neural network with a tanh activation function should have weights distributed as follows: and The random value between n and n, where n is the number of inputs to the node.
[0049] After initialization, the activation function and optimizer are defined. Then, at 742, the feature vectors or input dataset are provided to the NN. Each of the different feature vectors generated using a unimodal NN can be provided with an input having a known label. Similarly, a multimodal NN can be provided with feature vectors corresponding to inputs with known labels or classifications. Then, at 743, the NN predicts the label or classification of the features or input. At 744, the predicted label or class is compared to the known label or class (also called the ground truth), and the loss function measures the total error between the predictions and the ground truth for all training samples. As an example, and not by limitation, the loss function can be the cross-entropy loss function, quadratic cost, triple contrastive function, exponential cost, etc. Multiple different loss functions can be used depending on the purpose. As an example, and not by limitation, the cross-entropy loss function can be used to train a classifier, while the triple contrastive function can be used to learn pre-trained embeddings. The results of the loss function are then used to optimize and train the NN using known neural network training methods (e.g., backpropagation with adaptive gradient descent, etc.), as shown at 745. During each training epoch, the optimizer attempts to select model parameters (i.e., weights) that minimize the training loss function (i.e., the total error). The data is divided into training, validation, and test samples.
[0050] During training, the optimizer minimizes the loss function on the training samples. After each training epoch, the model is evaluated on the validation samples by calculating the validation loss and accuracy. If there is no significant change, training can be stopped, and the resulting trained model can be used to predict the labels for the test data.
[0051] Therefore, neural networks can be trained on inputs with known labels or classifications to recognize and classify those inputs. Similarly, the described methods can be used to train NNs to generate feature vectors from inputs with known labels or classifications. While the above discussion relates to RNNs and CRNNs, it can be applied to NNs that do not include recurrent or hidden layers.
[0052] system
[0053] Figure 8 A system according to aspects of this disclosure is described. The system may include a computing device 800 coupled to a user peripheral device 802. The peripheral device 802 may be a controller, a touchscreen, a microphone, or other device allowing the user to input voice data into the system. Additionally, the peripheral device 802 may also include one or more inertial measurement units (IMUs).
[0054] The computing device 800 may include one or more processor units 803, which may include one or more central processing units (CPUs) and / or one or more graphics processing units, which may be configured according to well-known architectures, such as single-core, dual-core, quad-core, multi-core, processor-coprocessor, unit processor, etc. The computing device may also include one or more memory units 804 (e.g., random access memory (RAM), dynamic random access memory (DRAM), read-only memory (ROM), etc.).
[0055] One or more processor units 803 can execute one or more programs, portions of which can be stored in memory 804, and processor units 803 can be operatively coupled to memory, for example, by accessing memory via data bus 805. The programs can be configured to train one or more other NN portions of the Virtual Assistant Neural Network (VA NN) 808 and / or other software modules, such as text-to-speech module 821, NLP module 821, game-playing module 825, and / or character module 822. Furthermore, memory 804 may contain one or more applications 826 that can generate application state data 827 that can be used with the virtual assistant. Memory 804 may also contain software modules constituting the virtual assistant, such as the virtual assistant neural network module 808, negation representation module 810, text-to-speech module 821, NLP module 821, character module 822, video module 824, and game-playing module 825. The VNN module 808 and other modules are components of the virtual assistant, such as... Figure 1 and Figure 2 The virtual assistant component is depicted in the diagram. Memory 804 may also include one or more databases 823, which may include information about one or more applications 826 and can be searched using keywords detected from prompts in user input by an NLP module. The overall structure and probabilities of the neural network may also be stored as data 818 in mass storage 815. Processor unit 803 is also configured to execute one or more programs 817 stored in mass storage 815 or memory 804, which cause the processor to perform methods for training the NN from a corpus of character information and / or application information. The system can generate neural networks as part of the NN training process. These neural networks may be stored in memory 804 as part of a VNN module 808 or other software modules. The completed NN may be stored in memory 804 or as data 818 in mass storage 815.
[0056] The computing device 800 may also include known support circuitry, such as input / output (I / O) 807, circuitry, power supply (P / S) 811, clock (CLK) 812, and cache 813, which may communicate with other components of the system, for example, via a data bus 805. The computing device may include a network interface 814. The processor unit 803 and the network interface 814 may be configured to implement a local area network (LAN) or personal area network (PAN) via a suitable network protocol for a PAN (e.g., Bluetooth). The computing device may optionally include a mass storage device 815, such as a disk drive, CD-ROM drive, tape drive, flash memory, etc., and the mass storage device may store programs and / or data. The computing device may also include a user interface 816 to facilitate interaction between the system and the user. The user interface may include a keyboard, mouse, light pen, gamepad, touch interface, or other devices.
[0057] Computing device 800 may include a network interface 814 to facilitate communication via an electronic communication network 820. The network interface 814 may be configured to enable wired or wireless communication via a local area network (LAN) and a wide area network (WAN) such as the Internet. Device 800 may send and receive data and / or requests for files via one or more message packets through network 820. Message packets sent via network 820 may be temporarily stored in a buffer in memory 804.
[0058] The aspects disclosed herein allow for the creation of virtual assistants that adopt a role that can be identified by the user. Furthermore, the virtual assistant can be integrated with applications, allowing the virtual assistant to emulate a role that responds to questions about the application and even interacts with the user. This can provide users with an engaging experience while also assisting them in operating the application or responding to prompts.
[0059] While the foregoing is a complete description of preferred embodiments of the present disclosure, various alternatives, modifications, and equivalents may be used. Therefore, the scope of this disclosure should not be determined by reference to the foregoing description, but rather by reference to the full scope of the appended claims and their equivalents. Any feature described herein (whether preferred or not) may be combined with any other feature described herein (whether preferred or not). In the appended claims, the indefinite article “a” or “an” refers to the number of one or more items following the article, unless otherwise expressly stated. The appended claims should not be construed as including a component with a functional limitation, unless such limitation is expressly stated in a given claim using the phrase “component for…”.
Claims
1. A device for providing a dedicated agent, comprising: an application; and a virtual assistant, wherein the virtual assistant comprises a neural network trained with a machine learning algorithm to mimic a communication style of a character, and wherein the virtual assistant is further trained on application metadata to predict responses to questions about the application.
2. The device of claim 1, further comprising a natural language processing module operable to convert language inputs into a computer readable form.
3. The device of claim 1, further comprising a text-to-speech module operable to convert text responses from the virtual assistant into speech.
4. The apparatus of claim 3, wherein, the text-to-speech module comprises a speech neural network trained with a machine learning algorithm to replicate a speech style of a character.
5. The apparatus of claim 4, wherein, the speech neural network is trained to replicate the speech style by replicating one or more of a group of intonation, pace, pronunciation, and rhythm.
6. The apparatus of claim 3, wherein, the text-to-speech module comprises a speech neural network trained with a machine learning algorithm on speech recordings.
7. The apparatus of claim 1, wherein, the virtual assistant further comprises a game play module, wherein the game play module is operable to operate the application.
8. The apparatus of claim 7, wherein, the application is operable to allow the game play module to operate in coordination with a user.
9. The apparatus of claim 7, wherein, the application is a cooperative game and the game play module is operable to play the game in coordination with a user.
10. The apparatus of claim 7, wherein, the game play module comprises a neural network trained with machine learning to operate the application.
11. The device of claim 1, further comprising a negative representation module operable to detect and remove false, misleading, or inappropriate predicted responses.
12. The apparatus of claim 11, wherein, the negative representation module comprises a representation neural network trained with a machine learning algorithm to detect false, misleading, or inappropriate predicted responses.
13. The apparatus of claim 11, wherein, the negative representation module is further operable to detect and remove predicted responses that are spoilers for a user for a portion of the application that the user has not yet seen.
14. The apparatus of claim 1, wherein, the virtual assistant further comprises a character module, wherein the character module is operable to coordinate visual displays of the character with predicted responses to questions.
15. The device of claim 1, further comprising a video module, wherein the video module is operable to retrieve videos based on predicted responses.
16. The device of claim 1, further comprising a video module, wherein the video module is operable to overlay video and / or text and / or graphical content over image frames of the application.
17. A system for providing a dedicated agent, comprising: a processor; a memory coupled to the processor; non-transitory processor-executable instructions embodied in the memory that, when executed, cause the processor to perform a method, the method comprising: executing an application; and providing a response to a question of a user about the application, wherein the response is generated by a virtual assistant, wherein the virtual assistant includes a neural network trained with a machine learning algorithm to mimic a communication style of a character, wherein the virtual assistant is further trained on application metadata to predict responses to questions about the application.
18. The system of claim 17, wherein, The non-transitory processor-executable instructions further include one or more instructions that, when executed by the processor, cause the processor to convert the question from a non-computer readable format to a computer readable format using a natural language processing neural network, wherein the natural language processing neural network is communicatively coupled with the virtual assistant.
19. The system of claim 17, wherein, The non-transitory processor-executable instructions further include one or more instructions that, when executed by the processor, cause the processor to convert the response of the user to the question into speech using a text-to-speech module.
20. The system of claim 19, wherein, The text-to-speech module includes a speech neural network trained with a machine learning algorithm to replicate a speech style of a character.
21. The system of claim 20, wherein, The speech neural network is trained to replicate the speech style of the character by replicating one or more of a group of inflection, pace, pronunciation, and prosody.
22. The system of claim 20, wherein, The speech neural network is trained with a machine learning algorithm using speech recordings.
23. The system of claim 17, wherein, The non-transitory processor-executable instructions further include one or more instructions that, when executed by the processor, cause the processor to operate the application to provide the response to the question from the user using a game play module.
24. The system of claim 23, wherein, The application is operable to allow the game play module to operate in coordination with the user.