Virtual game interaction method and system based on large language model
By loading virtual scenes into VR devices and using large language models to generate logically rich NPC responses, the problems of insufficient immersion and high cost of offline script-killing games are solved, and a low-cost and highly immersive virtual script-killing experience is achieved.
Patent Information
- Application Number
- CN202511080010.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-04
AI Technical Summary
Offline script-killing games lack immersion, are socially awkward, and are costly, relying on external factors and a large number of venues and staff.
Virtual scenes are loaded through VR devices, and natural conversations with virtual NPCs are achieved by combining them with large language models, reducing the demand for venues and staff. VR devices are used to collect player voice information and generate logically rich responses through large language models.
It improves the player's immersion in the script-killing game, reduces the implementation cost, and enables natural dialogue between virtual NPCs and players.
Smart Images

Figure CN120571230B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer software technology, and in particular to a virtual game interaction method and system based on a large language model. Background Art
[0002] Script Killing is an immersive role-playing reasoning game in which players play specific roles, collect clues, discuss reasoning within the story set in the script, and ultimately solve the case (or complete the plot mission).
[0003] In related technologies, script-killing games are mainly realized through offline interactions (i.e., offline script-killing). However, the immersion of offline script-killing often depends on external factors. For example, novice NPCs (Non-Player Characters) may be unskilled or prone to laughing, and some players may feel socially awkward when interacting offline. These external factors will affect the player's immersion. In addition, offline script-killing has a large demand for venues and staff. For example, it usually requires multiple physical rooms (such as discussion rooms, evidence search rooms, etc.) and multiple staff (such as real NPCs, hosts, set designers, etc.). Therefore, offline script-killing has the problems of insufficient immersion and high cost. Summary of the Invention
[0004] In order to solve or partially solve the problems existing in the related art, the present application provides a virtual game interaction method and system based on a large language model, which can enhance the player's sense of immersion in the script-killing game and reduce the implementation cost of the script-killing game.
[0005] In a first aspect, the present application provides a virtual game interaction method based on a large language model, the method comprising:
[0006] After detecting that a player is wearing a VR device, a virtual scene of a script-killing game is loaded for the player through the VR device; the virtual scene includes a virtual NPC;
[0007] collecting first voice information of the player through the VR device; the first voice information includes questions asked by the player to the virtual NPC;
[0008] Performing speech recognition on the first speech information to obtain first text information, and inputting the first text information into a large language model to generate, through the large language model, second text information that matches the first text information; the second text information includes a response made by the virtual NPC to the question;
[0009] The second text information is formatted to obtain second voice information, and the second voice information is output to the player through the VR device.
[0010] In one embodiment, inputting the first text information into a large language model to generate second text information matching the first text information through the large language model includes:
[0011] Querying auxiliary data information related to the problem from a pre-built knowledge base; wherein the auxiliary data information includes game setting information of the virtual scene and character setting information of the virtual NPC;
[0012] The first text information, the game setting information, and the character setting information are input into a large language model, so as to generate second text information matching the first text information based on the game setting information and the character setting information through the large language model.
[0013] In one embodiment, the inputting the first text information, the game setting information, and the character setting information into a large language model to generate, through the large language model, second text information that matches the first text information based on the game setting information and the character setting information includes:
[0014] Determining whether the player has seen through the lies of the virtual NPC;
[0015] If the player sees through the virtual NPC's lie, the virtual NPC's behavior mode is switched from a misleading mode to a frank mode;
[0016] Obtaining a first prompt word corresponding to the confession mode; the first prompt word is used to guide the large language model to output truth;
[0017] The first text information, the game setting information, the character setting information, and the first prompt word are input into the large language model, so as to generate, through the large language model, a second text information that matches the first text information and does not contain any logical loopholes based on the game setting information, the character setting information, and the first prompt word.
[0018] In one embodiment, determining whether the player has seen through the lies of the virtual NPC includes:
[0019] Collecting the player's spatial motion information through the VR device;
[0020] determining whether the player has found a key clue based on the spatial action information, and determining whether the player has raised questions regarding the virtual NPC's response based on the first text information;
[0021] If the player finds a key clue and questions the virtual NPC's reply, it is determined that the player has seen through the virtual NPC's lie.
[0022] In one embodiment, after determining whether the player has seen through the lies of the virtual NPC, the method further includes:
[0023] If the player fails to see through the lie of the virtual NPC, the behavior mode of the virtual NPC is maintained in the misleading mode;
[0024] Obtaining a second prompt word corresponding to the misleading mode; the second prompt word is used to guide the large language model to output falsehood;
[0025] The first text information, the game setting information, the character setting information, and the second prompt word are input into the large language model, so as to generate, through the large language model, a second text information that matches the first text information and contains a logical loophole based on the game setting information, the character setting information, and the second prompt word.
[0026] In one embodiment, the outputting the second voice information to the player through the VR device includes:
[0027] Determining acoustic features of the virtual NPC based on the character setting information; the acoustic features including at least one of timbre, pitch, speech speed, and emotion;
[0028] The second voice information is output to the player through the VR device according to the acoustic characteristics.
[0029] In one embodiment, the virtual scene further includes game props, and the method further includes:
[0030] The visual picture presented to the player by the virtual scene is updated according to the spatial action information; the visual picture includes a perspective synchronization picture and / or an object interaction picture, the perspective synchronization picture is used to represent the synchronization of the visual picture with the perspective corresponding to the spatial action information, and the object interaction picture is used to represent the state change of the object when the player interacts with the object, and the object includes at least one of the virtual NPC and the game props.
[0031] A second aspect of the present application provides a virtual game interaction system based on a large language model, the system comprising:
[0032] A virtual scene loading module is used to load a virtual scene of the script-killing game for the player through the VR device after detecting that the player is wearing a VR device; the virtual scene includes a virtual NPC;
[0033] A voice collection module, configured to collect first voice information of the player through the VR device; the first voice information includes questions raised by the player to the virtual NPC;
[0034] a natural language processing module configured to perform speech recognition on the first speech information to obtain first text information, and input the first text information into a large language model to generate, through the large language model, second text information that matches the first text information; the second text information including a response from the virtual NPC to the question;
[0035] The voice playback module is used to convert the format of the second text information to obtain second voice information, and output the second voice information to the player through the VR device.
[0036] A third aspect of the present application provides an electronic device, including:
[0037] processor; and
[0038] The memory stores executable codes thereon, and when the executable codes are executed by the processor, the processor is caused to execute the method described above.
[0039] A fourth aspect of the present application provides a computer-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method described above.
[0040] The technical solution provided by this application may include the following beneficial results:
[0041] The technical solution of the present application is to load a virtual scene of a script-killing game for the player through the VR device after detecting that the player is wearing a VR device; the virtual scene includes a virtual NPC; the player's first voice information is collected through the VR device; the first voice information includes questions asked by the player to the virtual NPC; voice recognition is performed on the first voice information to obtain first text information, and the first text information is input into a large language model to generate a second text information matching the first text information through the large language model; the second text information includes the virtual NPC's reply to the question; the second text information is formatted to obtain a second voice information, and the second voice information is output to the player through the VR device. The present application realizes online script-killing through VR devices and large language models, wherein the VR device can provide players with a virtual scene of the script-killing game (including a virtual NPC), thereby reducing the demand for venues and staff, and thereby reducing the implementation cost of the script-killing game, and the large language model can drive the language logic of the virtual NPC according to the player's real-time input, thereby realizing natural dialogue between the virtual NPC and the player, thereby improving the player's immersion in the script-killing game.
[0042] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The above and other objects, features and advantages of the present application will become more apparent by describing in more detail exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.
[0044] Figure 1 1 is a flow chart of a virtual game interaction method based on a large language model according to an embodiment of the present application;
[0045] Figure 2 is another flowchart of a virtual game interaction method based on a large language model according to an embodiment of the present application;
[0046] Figure 3 This is a diagram showing the overall architecture of a virtual game interaction system based on a large language model according to an embodiment of the present application;
[0047] Figure 4 This is a diagram of a spatial action-natural language dual-modal interaction framework shown in an embodiment of the present application;
[0048] Figure 5 This is a diagram of the voice interaction between a virtual NPC and a player shown in an embodiment of the present application;
[0049] Figure 6 Schematic diagram of the structure of a virtual game interaction system based on a large language model shown in an embodiment of the present application;
[0050] Figure 7 It is a structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although the accompanying drawings illustrate embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0052] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0053] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0054] In related technologies, script-killing games are realized through offline interactions (i.e., offline script-killing). However, the immersion of offline script-killing often depends on external factors. For example, novice NPCs may be unskilled or prone to laughing, and some players may be socially awkward when interacting offline. These external factors will affect the player's immersion. In addition, offline script-killing has a large demand for venues and staff. For example, it usually requires multiple physical rooms (such as discussion rooms, evidence search rooms, etc.) and multiple staff (such as real NPCs, hosts, set designers, etc.). Therefore, offline script-killing has the problems of insufficient immersion and high cost.
[0055] In response to the above problems, an embodiment of the present application provides a virtual game interaction method based on a large language model, which realizes online script-killing through VR equipment and a large language model. Among them, the VR equipment can provide players with virtual scenes of script-killing games (including virtual NPCs), thereby reducing the demand for venues and staff, and thus reducing the implementation cost of script-killing games. The large language model can drive the language logic of the virtual NPC according to the player's real-time input, thereby realizing natural dialogue between the virtual NPC and the player, thereby improving the player's immersion in the script-killing game.
[0056] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0057] Figure 1 This is a flow chart of a virtual game interaction method based on a large language model shown in an embodiment of the present application.
[0058] See also Figure 1 The virtual game interaction method based on the large language model of the present application may include:
[0059] S110, after detecting that the player is wearing the VR device, a virtual scene of the script-killing game is loaded for the player through the VR device; the virtual scene includes a virtual NPC.
[0060] In an embodiment of the present application, a virtual game interaction system based on a large language model (hereinafter referred to as the "interaction system") can be applied. When a player wants to start an online script-killing game, he or she can enter the virtual world by wearing a VR device (such as VIVE HTC Pro). VR (Virtual Reality) is a hardware system that simulates a three-dimensional virtual environment through computer technology to allow players to obtain an immersive sensory experience. VIVE HTC Pro is a head-mounted VR device, which includes a head-mounted display device, a handle, and an audio device. The audio device includes an audio acquisition device and an audio playback device. After detecting that the player is wearing the VR device, the interaction system can present the player with multiple theme options of the script-killing game through the head-mounted display device. Different themes correspond to different virtual scenes. The interaction system can collect the player's selection instruction through the handle or audio acquisition device. The selection instruction carries the target theme selected by the player. In response to the selection instruction, the interaction system loads the virtual scene corresponding to the target theme, and then presents the visual image of the virtual scene to the player through the head-mounted display device.
[0061] Among them, virtual scenes include virtual NPCs. Virtual NPCs are different from real-life NPCs. Virtual NPCs are virtual characters driven by large language models and the Unity engine layer, while real-life NPCs are played by real humans. Therefore, virtual NPCs can replace real-life NPCs and achieve natural interaction with players.
[0062] S120, collecting first voice information of the player through the VR device; the first voice information includes questions asked by the player to the virtual NPC.
[0063] Players can use audio capture devices to interact verbally with virtual NPCs in the virtual scene, and can use controllers to interact with virtual NPCs or other objects in the virtual scene through gestures. Specifically, the interactive system can use the audio capture device to capture the player's first voice information. The first voice information can include questions posed by the player to the virtual NPC, conversations between the player and the virtual NPC, or instructions issued by the player. These first voice information become the starting data for the player's interaction with the interactive system.
[0064] The processing procedures for players' questions, communications, and instructions are basically the same, so the embodiment of the present application uses the first voice information as an example of a question asked by a player to a virtual NPC.
[0065] S130, performing speech recognition on the first speech information to obtain first text information, and inputting the first text information into the large language model to generate second text information matching the first text information through the large language model; the second text information includes the virtual NPC's response to the question.
[0066] The interactive system can convert the first voice information into an audio stream format, and then call the ASR (Automatic Speech Recognition) interface of SenseVoice (speech recognition model). The SenseVoice model has powerful speech recognition capabilities. It will analyze and process the input audio stream and convert the voice content into a text form that can be understood by the computer. Therefore, the interactive system can input the audio stream of the first voice information into the SenseVoice model based on the ASR interface, so that the audio stream of the first voice information can be converted into first text information that can be understood by the computer through the SenseVoice model.
[0067] It should be noted that the content of the first text message and the first voice message is the same, both including questions asked by the player to the virtual NPC, but the data formats are different, for example, the first text message is in text format, and the first voice message is in voice format.
[0068] The first text message is key data for subsequent interactive logic processing. The interactive system uses this data to understand the player's intent and respond accordingly. In a specific implementation, the interactive system can invoke a large language model (such as Ultra 4.0). The large language model (LLM) has powerful natural language processing capabilities. Therefore, the interactive system can input the first text message into the LLM. The LLM then leverages its powerful natural language processing capabilities and knowledge base to synthesize multiple aspects of information and generate a second text message that matches the first text message. The second text message is the virtual NPC's response to the player's voice input. For example, if the first voice message is a question posed by the player to the virtual NPC, the second text message is the virtual NPC's reply to the question; if the first voice message is a conversation between the player and the virtual NPC, the second text message is the conversation between the virtual NPC and the player; if the first voice message is a command issued by the player, the second text message is the virtual NPC's response to the command.
[0069] In addition, in addition to interacting with players through language, virtual NPCs also interact with other virtual NPCs in the same virtual scene. Specifically, the large language model can trigger the generation of third text information according to the game setting information of the virtual scene. The third text information is a dialogue between different virtual NPCs to promote the development of the plot, guide the player's reasoning direction, and enhance the player's immersion.
[0070] It can be seen that whether it is the virtual NPC's response to the player's question, the communication between the virtual NPC and the player, the virtual NPC's response to the instruction, or the conversation between different virtual NPCs, they are all generated by the large language model. Therefore, the large language model can provide rich and logical language support for the interaction in the script-killing game.
[0071] S140: Convert the format of the second text information to obtain second voice information, and output the second voice information to the player through the VR device.
[0072] After obtaining the second text information output by the large language model, the interactive system can call the TTS (Text-to-Speech) interface of ChatTTS (ChatText-to-Speech, speech synthesis model). The ChatTTS model has powerful speech synthesis capabilities. It will use speech synthesis technology to generate the corresponding voice stream based on the input text content. Therefore, the interactive system can input the second text information into the ChatTTS model so that the second text information can be converted into second voice information through the ChatTTS model.
[0073] It should be noted that the content of the second voice message and the second text message is the same, both including the virtual NPC's response to the question, but the data format is different, for example, the second voice message is in voice format, and the second text message is in text format.
[0074] After obtaining the second voice information output by the ChatTTS model, the interactive system can play the second voice information to the player through the audio playback device built into the VR device, thereby achieving a natural and smooth voice interaction closed loop.
[0075] It should be noted that the embodiments of the present application can be applied to scenarios such as virtual reality (VR) game entertainment (such as script-killing games), intelligent agent interaction (such as virtual tour guides, educational auxiliary roles), etc.
[0076] As can be seen from this example, the solution provided by this application, after detecting that the player is wearing a VR device, loads the virtual scene of the script-killing game for the player through the VR device; the virtual scene includes a virtual NPC; the player's first voice information is collected through the VR device; the first voice information includes questions asked by the player to the virtual NPC; voice recognition is performed on the first voice information to obtain a first text information, and the first text information is input into the large language model to generate a second text information matching the first text information through the large language model; the second text information includes the virtual NPC's reply to the question; the second text information is formatted to obtain a second voice information, and the second voice information is output to the player through the VR device. This application realizes online script-killing through VR devices and large language models, wherein the VR device can provide players with a virtual scene (including a virtual NPC) of the script-killing game, thereby reducing the demand for venues and staff, thereby reducing the implementation cost of the script-killing game, and the large language model can drive the language logic of the virtual NPC according to the player's real-time input, thereby realizing natural dialogue between the virtual NPC and the player, thereby improving the player's immersion in the script-killing game.
[0077] Figure 2 This is another flowchart of the virtual game interaction method based on the large language model shown in this application.
[0078] See also Figure 2 The virtual game interaction method based on the large language model of the present application may include:
[0079] S210, after detecting that the player is wearing a VR device, a virtual scene of the script-killing game is loaded for the player through the VR device; the virtual scene includes a virtual NPC.
[0080] like Figure 3As shown, the interactive system can include user layer, VR interaction layer, voice interaction layer, scene management layer, RAGFlow layer, large language model, and Unity engine layer. Each layer works closely together to realize the core functions of the system. Among them, the user layer, as the starting point and end point of the entire interactive system, is the direct interface for players to interact with the interactive system; the VR interaction layer is responsible for processing the interaction data between players and VR devices; the voice interaction layer is mainly responsible for voice recognition and voice synthesis; the scene management layer is developed based on the XR Interaction Toolkit (XR interaction toolkit) of Unity (a cross-platform real-time 3D / 2D content development engine), and is mainly responsible for managing and rendering the virtual scenes of the script-killing game; the RAGFlow layer is built based on the open source RAGFlow (Retrieval-Augmented Generation Flow, retrieval enhancement generation framework), and plays a key bridging role; the Unity engine layer, as the basic platform for the development of the entire script-killing game, provides rich functions and tools for creating and optimizing the virtual scenes of the script-killing game, managing game resources, and implementing various game logics. It provides underlying support for the VR interaction layer and scene management layer, ensuring that the entire script-killing game can run efficiently and stably, achieving high-quality graphics rendering and a smooth interactive experience. The large language model layer is connected to the Ultra4.0 version of the large language model, providing rich and logical language support for interactions in script-killing games; the knowledge base layer stores rich information, including game setting information for each virtual scene and character setting information for each virtual NPC. This information has been sorted and structured so that the RAGFlow layer can quickly and accurately retrieve and utilize it, providing knowledge support for the dialogues and behaviors of virtual NPCs, making the performance of virtual NPCs more realistic and credible.
[0081] Specifically, the user layer includes VR devices. When players want to start an online script-killing game, they can enter the virtual world by wearing VR devices (such as VIVE HTC Pro). VR devices include head-mounted displays, handles, and audio devices (such as audio acquisition devices and audio playback devices). The head-mounted displays are used to provide players with picture display functions in virtual scenes, the handles are used to provide players with action interaction functions in virtual scenes, and the audio devices are used to provide players with language interaction functions in virtual scenes. Therefore, the player's input (such as first voice information, spatial action information) is collected by the user layer to the interaction system, and the output of the interaction system (such as second voice information, virtual scene changes) is presented to the player by the user layer.
[0082] In actual applications, after detecting that the player is wearing a VR device, the interactive system can present the player with multiple theme options of script-killing games through the head-mounted display device. Different themes correspond to different virtual scenes. The interactive system can collect the player's selection instructions through the handle or audio acquisition device. The selection instruction carries the target theme selected by the player. The interactive system responds to the selection instruction and loads the virtual scene corresponding to the target theme through the scene management layer, so that the scene management layer controls the generation, layout and physical properties of objects in the virtual scene. Among them, the objects may include at least one of virtual NPCs and game props, and the physical properties may include at least one of position information, size information and rotation information. After the scene management layer renders the visual picture of the virtual scene, the interactive system can present the visual picture of the virtual scene to the player through the head-mounted display device.
[0083] S220, collecting first voice information of the player through the VR device; the first voice information includes questions asked by the player to the virtual NPC.
[0084] The interactive system can collect the player's first voice information through the audio collection device. The first voice information may include questions asked by the player to the virtual NPC, or communication between the player and the virtual NPC, or instructions issued by the player, etc. These first voice information become the starting data for the player to interact with the interactive system.
[0085] The processing procedures for players' questions, communications, and instructions are basically the same, so the embodiment of the present application uses the first voice information as an example of a question asked by a player to a virtual NPC.
[0086] S230, performing voice recognition on the first voice information to obtain first text information, and searching for auxiliary data information related to the question from a pre-built knowledge base; wherein the auxiliary data information includes game setting information of the virtual scene and character setting information of the virtual NPC.
[0087] The interactive system can transmit the first voice information to the voice interaction layer through the VR interaction layer. The voice interaction layer includes a SenseVoice model. After receiving the first voice information, the SenseVoice model can use its powerful voice recognition capabilities to convert the player's first voice information into first text information that can be understood by the computer, wherein the first text information may include questions asked by the player to the virtual NPC.
[0088] Afterwards, the interactive system can pass the first text information to the RAGFlow layer, and the RAGFlow layer can perform query operations in the knowledge base layer based on the built-in retrieval algorithm. Specifically, the knowledge base layer includes a pre-built knowledge base, which stores rich information related to script-killing games, including game setting information for each virtual scene and character setting information for each virtual NPC. Therefore, the RAGFlow layer queries the knowledge base for auxiliary data information related to the question. The auxiliary data information includes the game setting information of the current virtual scene and the character setting information of each virtual NPC in the current virtual scene. The game setting information can include plot background (such as school, hospital, theater, etc.) and game clues (such as physical evidence clues, text clues, environmental clues, etc.). The character setting information can include the character background of the virtual NPC (such as life information, relationship with other virtual NPCs, etc.) and personality traits (such as selfishness / kindness, calmness / irritability, honesty / cunningness, etc.).
[0089] S240, inputting the first text information, the game setting information and the character setting information into the large language model, so as to generate a second text information matching the first text information based on the game setting information and the character setting information through the large language model; the second text information includes the virtual NPC's response to the question.
[0090] Auxiliary data information is used to assist the large language model in outputting more accurate responses. In a specific implementation, the RAGFlow layer inputs the auxiliary data information together with the first text information into the large language model layer. The large language model layer includes the Ultra4.0 version of the large language model. The large language model can use its powerful natural language processing capabilities to generate a second text information that matches the first text information based on the auxiliary data information provided by the RAGFlow layer. The second text information includes the virtual NPC's response to the question. It can be seen that the second text information can not only answer the player's question, but also conform to the game setting information of the virtual scene and the character setting information of the virtual NPC.
[0091] In one example, assuming the virtual NPC is Li Bai, players can ask questions to the virtual NPC, and the virtual NPC will reply to the player as Li Bai. For example, if the player asks Li Bai how to write poetry, the RAGFlow layer will query the auxiliary information related to the question from the knowledge base. The auxiliary information includes the historical and cultural background knowledge files of the Tang Dynasty and Li Bai's personal life and background files. The large language model combines Li Bai's personal life and background files with "At the age of fifteen, he was good at swordsmanship and served all the princes" to give an example of how Li Bai wrote poetry.
[0092] In addition, if the first voice message is the communication between the player and the virtual NPC, the second text message is the communication between the virtual NPC and the player; if the first voice message is an instruction issued by the player, the second text message is the response of the virtual NPC to the instruction.
[0093] In addition, in addition to interacting with players through language, virtual NPCs also interact with other virtual NPCs in the same virtual scene. Specifically, the large language model can trigger the generation of third text information according to the game setting information of the virtual scene. The third text information is a dialogue between different virtual NPCs to promote the development of the plot, guide the player's reasoning direction, and enhance the player's immersion.
[0094] It can be seen that whether it is the virtual NPC's response to the player's question, the communication between the virtual NPC and the player, the virtual NPC's response to the instruction, or the conversation between different virtual NPCs, they are all generated by the large language model. Therefore, the large language model can provide rich and logical language support for the interaction in the script-killing game.
[0095] In one embodiment, inputting the first text information, the game setting information, and the character setting information into a large language model to generate, through the large language model, second text information that matches the first text information based on the game setting information and the character setting information may include:
[0096] Determine whether the player sees through the lies of the virtual NPC; if the player sees through the lies of the virtual NPC, switch the behavior mode of the virtual NPC from the misleading mode to the confessing mode; obtain a first prompt word corresponding to the confessing mode; the first prompt word is used to guide the large language model to output the truth; input the first text information, game setting information, character setting information and the first prompt word into the large language model, so as to generate a second text information that matches the first text information and does not contain any logical loopholes through the large language model based on the game setting information, character setting information and the first prompt word.
[0097] The core reasoning logic of script-killing relies on players to discover contradictions and reveal the truth through multi-dimensional interactions (such as exploring scenes and talking to NPCs). However, script-killing in related technologies is often a single-mode interaction (such as pure text dialogue), which makes it difficult to support complex reasoning requirements. Therefore, the embodiment of this application proposes a dual-mode interaction framework of spatial action-natural language, such as Figure 4 As shown in the figure, by integrating VR scene actions and LLM semantic analysis capabilities, it is dynamically judged whether the player has seen through the lies of the virtual NPC. When the player sees through the lies of the virtual NPC, the behavior mode of the virtual NPC is switched. Figure 4The goal of the framework shown is to serve the "two-faced" design (lying → honesty) of virtual NPCs in virtual scenes, and to determine whether the player meets the mode switching conditions through dual-modal collaboration.
[0098] In a specific implementation, any response from the virtual NPC may be a lie to deceive the player (that is, any second text information generated by the large language model may be a lie). Appropriate lies by the virtual NPC can increase the complexity of reasoning, hide key truths, guide plot branches, etc., thereby promoting plot development, increasing the dimension of reasoning, and enhancing the player's immersion. If the virtual NPC's response contains pre-set contradictions that conflict with game clues (such as physical evidence and testimony), then the virtual NPC's response is a lie. After obtaining the player's first voice message, the interactive system can first determine whether the player has seen through the virtual NPC's lie. If the player has seen through the virtual NPC's lie, it is determined that the player meets the mode switching conditions. The plot setting of the script-killing game is that the virtual NPC misleads in the early stage and confesses in the later stage. Therefore, the interactive system can switch the virtual NPC's behavior mode from misleading mode to confessing mode. The misleading mode can also be called lying mode, and the confessing mode can also be called honest mode. Different behavior modes correspond to different prompt words. For example, the confession mode (or honest mode) corresponds to the first prompt word p triggered , the misleading mode (or lying mode) corresponds to the second prompt word p normal , where the first prompt word p triggered Used to guide the large language model to output the truth, the second prompt word p normal Used to guide large language models to output false statements.
[0099] When the virtual NPC is in confession mode, the interactive system can obtain the first prompt word p corresponding to the confession mode. triggered , then the first text information, game setting information, character setting information and the first prompt word p triggered Input into the large language model so that the large language model can use its powerful natural language processing capabilities to calculate the game setting information, character setting information and the first prompt word p triggered , generating a second text message that matches the first text message, since the first prompt word p triggered It is used to guide the large language model to output the truth, so that the second text information does not contain logical loopholes. Logical loopholes refer to pre-set contradictions that conflict with game clues (such as physical evidence and testimony). That is, the second text information output by the large language model is the truth, which changes the virtual NPC's response from a lie to the truth.
[0100] In one embodiment, determining whether the player has seen through the lies of the virtual NPC may include:
[0101] The player's spatial motion information is collected through the VR device; whether the player has found the key clue is determined based on the spatial motion information, and whether the player has doubts about the virtual NPC's reply is determined based on the first text information; if the player has found the key clue and the player has doubts about the virtual NPC's reply, it is determined that the player has seen through the virtual NPC's lie.
[0102] Players can use the controller to interact with objects in the virtual scene (such as virtual NPCs and game props). For example, players can use the controller to wave, shake hands, or pass objects to virtual NPCs. Another example is that players can use the controller to grab a fan, open a treasure chest, or flip through files. In specific implementations, the VR interaction layer can receive controller operation data from the user layer. Controller operation data is the player's spatial movement information collected by the VR device. The VR interaction layer can convert this spatial movement information (i.e., controller operation data) into actual movements in the virtual scene, ensuring that players can interact with objects in the virtual scene naturally and smoothly.
[0103] exist Figure 4 In the VR game, the motion capture module is the controller of the VR device. The player operates the controller, enabling the controller to capture the player's spatial motion information (i.e., controller operation data). The user layer transmits this spatial motion information to the VR interaction layer, which then semantically converts it into action, generating spatial evidence. Specifically, the VR interaction layer converts the player's spatial motion information into actual actions in the virtual scene. Actual actions are the physical interactions of clue exploration. Its function is to explore game clues through physical interactions and verify whether the player has discovered key evidence (i.e., whether the key clue has been found). The VR interaction layer uses preset logic to identify key clues and trigger actions. When the player finds a key clue, the action for that key clue is marked as a "key exploration behavior." The VR interaction layer records this action and generates "spatial evidence," which is used to indicate that the player has discovered the key clue to reveal the lie. This allows the player to be judged as one step closer to revealing the lie. The VR interaction layer then transmits this "spatial evidence" to the interactive system's dual-modal engine.
[0104] The audio capture device can be built into a VR device or into a PC device (Personal Computer). When the audio capture device is built into a VR device, the user layer includes the VR device, and the VR device has both voice capture and voice playback functions. When the audio capture device is built into a PC device, the user layer includes the VR device and the PC device, and the PC device has the voice capture function, and the VR device has the voice playback function.
[0105] exist Figure 4In this example, a player uses an audio capture device (such as a microphone) connected to a PC to input voice information, allowing the PC's audio capture device to capture the player's first voice message. The speech recognition module is the SenseVoice model in the speech interaction layer. The SenseVoice model converts the first voice message into first text information. The speech interaction layer then passes the first text information to the RAGFlow layer. The RAGFlow layer queries the knowledge base for auxiliary data and obtains a third prompt word. The third prompt word guides the large language model to output a binary judgment ("yes" / "no") to indicate whether the player's response to the virtual NPC is questioned based on the first text information. The RAGFlow layer inputs the first text information, auxiliary data, and third prompt word into the large language model. The large language model then performs semantic and sentiment analysis on the first text information based on the auxiliary data and the third prompt word, generating "linguistic evidence" that indicates the player's questioning of the virtual NPC's response. This indicates that the player is one step closer to uncovering the lie. The RAGFlow layer then passes the "linguistic evidence" to the interactive system's bimodal engine.
[0106] It should be noted that the third prompt word and the auxiliary data information are similar in that they both aim to make the large language model's response more accurate. The difference is that the third prompt word is pre-set, so the third prompt word is fixed, while the auxiliary data information is the RAGFlow layer that queries the knowledge base based on the player's input (i.e., the first voice information) to obtain a description related to the question, so the auxiliary data information is associated with the player's input.
[0107] In one example, assuming the virtual NPC is Zhang San, the third prompt word used to determine whether Zhang San is lying can be: "You are Zhang San's semantic understanding assistant. Please summarize the contents of the knowledge base to answer the question.
[0108] Please note that you now need to judge whether the player has noticed that Zhang San is lying based on the player's input. If the player's input mentions "I found a box of gold in your room" and asks Zhang San "Didn't you say that gold was used to repair the city walls and be used?", then it is considered that the player has noticed that Zhang San is lying, otherwise he has not noticed it.
[0109] Please note that the answers can only be "yes" or "no".
[0110] Here is the knowledge base:
[0111] {knowledge}.”
[0112] exist Figure 4In the bimodal engine, after receiving "spatial evidence" and "linguistic evidence", it can make a state judgment according to the bimodal fusion decision. Its function is to integrate "spatial evidence" and "linguistic evidence" to trigger the switching of the virtual NPC's behavior mode. Among them, the bimodal fusion decision adopts the "logical and" rule to determine whether the player has seen through the lies of the virtual NPC. Specifically, if "spatial evidence" represents that the player has explored key clues (such as grabbing the fan and triggering the mechanism), and "linguistic evidence" represents that the player has questioned the virtual NPC's reply, then the bimodal engine determines that the player has seen through the lies of the virtual NPC. In other words, only when both are true at the same time, the virtual NPC's behavior mode changes from misleading mode (i.e. Figure 4 Misleading template in the switch to candid mode (i.e. Figure 4 Confession template in ).
[0113] For a two-sided intelligent agent, at the mathematical modeling level, the state space function of the intelligent agent is defined as:
[0114]
[0115] Among them, s normal represents the normal state (i.e., misleading mode) when the cue is not triggered, s clue-triggered Indicates the exposure state after the trigger cue (i.e., confession mode).
[0116] It should be noted that if the SenseVoice model encounters speech recognition errors, it will cause the large language model to misunderstand the speech, which will in turn affect the response of the virtual NPC. The core formula for error propagation is:
[0117]
[0118] Among them, Z err represents the probability of an incorrect response from the virtual NPC, ε ASR It represents the speech recognition error of the SenseVoice model, that is, its error rate in the speech recognition process. δ represents the semantic amplification coefficient, which reflects the sensitivity of the large language model to errors. γ represents the error correction capability of the ChatTTS model.
[0119] It can be seen that through the above core formula of error propagation, the impact of speech recognition errors on the semantic understanding of large language models can be quantified, revealing the relationship between speech recognition errors, semantic processing sensitivity, text-to-speech error correction ability and the probability of incorrect responses of virtual NPCs.
[0120] In one embodiment, after determining whether the player has seen through the lies of the virtual NPC, the method may further include:
[0121] If the player fails to see through the virtual NPC's lie, the virtual NPC's behavior pattern is maintained in a misleading mode; a second prompt word corresponding to the misleading mode is obtained; the second prompt word is used to guide the large language model to output a lie; the first text information, game setting information, character setting information and the second prompt word are input into the large language model, so as to generate a second text information that matches the first text information and contains a logical loophole based on the game setting information, character setting information and the second prompt word through the large language model.
[0122] If the player does not see through the lies of the virtual NPC, it is determined that the player does not meet the mode switching condition, so the interactive system can maintain the behavior mode of the virtual NPC in the misleading mode.
[0123] When the virtual NPC is in the misleading mode, the interactive system can obtain the second prompt word p corresponding to the misleading mode. normal , then the first text information, game setting information, character setting information and the second prompt word p normal Input into the large language model, so that the large language model can use its powerful natural language processing capabilities to calculate the game setting information, character setting information and the second prompt word p normal , generating a second text message that matches the first text message, because the second prompt word p normal It is used to guide the large language model to output lies, so the second text information contains logical loopholes. Logical loopholes refer to pre-set contradictions that conflict with game clues (such as physical evidence and testimony). That is, the second text information output by the large language model is a lie, causing the virtual NPC's response to continue to tell lies on key issues.
[0124] The response function of the agent can be expressed as:
[0125]
[0126] Among them, q represents the natural semantic understanding result, K represents the spatial action clue exploration result, and p normal The second prompt word corresponding to the misleading mode, p triggered Indicates the first prompt word corresponding to the confession mode, f mislead and f truth Both are generative functions based on large language models.
[0127] It can be seen that when the virtual NPC is in misleading mode (s normal ), the second text message is false: ; When the virtual NPC is in confession mode (s clue-triggered ), the second text message is true: .
[0128] Figure 4The framework shown realizes the dynamic and controllable switching of the lying state of the virtual NPC in the script-killing game through the dual-modal collaboration of "spatial action exploration of key clues-natural language questioning and verification".
[0129] S250: Convert the format of the second text information to obtain second voice information, and output the second voice information to the player through the VR device.
[0130] After obtaining the second text information output by the large language model, the interactive system can call the TTS interface of ChatTTS, and then input the second text information into the ChatTTS model so that the ChatTTS model can use its powerful speech synthesis capabilities to convert the second text information into a second voice information, wherein the second voice information may include the virtual NPC's response to the question.
[0131] After obtaining the second voice information output by the ChatTTS model, the interactive system can play the second voice information to the player through the audio playback device built into the VR device, thereby achieving a natural and smooth voice interaction closed loop.
[0132] Since the second voice message is the virtual NPC's reply to the player's question, while the second voice message is played to the player, the virtual NPC can be controlled to be in a speaking state through the Unity engine layer, thereby making the virtual NPC more vivid.
[0133] In one embodiment, outputting the second voice information to the player through the VR device may include:
[0134] The acoustic characteristics of the virtual NPC are determined based on the character setting information; the acoustic characteristics include at least one of timbre, pitch, speaking speed and emotion; and the second voice information is output to the player through the VR device according to the acoustic characteristics.
[0135] The ChatTTS model can determine the acoustic characteristics of the virtual NPC based on the character setting information of the virtual NPC (such as character background, personality traits, etc.). The acoustic characteristics may include at least one of timbre, pitch, speaking speed and emotion. Then the interactive system can output the second voice information to the player according to the acoustic characteristics of the virtual NPC through the built-in audio playback device of the VR device, so that the player can get a more natural and realistic interactive experience.
[0136] In one example, suppose the virtual NPC is William, whose character background is a 60-year-old retired ocean freighter captain who tells maritime legends in a tavern. His personality traits are calm and kind but a little vicissitudes of life, he habitually pauses for thought, and speaks slowly. The acoustic characteristics of this virtual NPC can be: the timbre is a low and hoarse baritone, the pitch is a low base frequency (around 110Hz), the end of the sentence is slightly depressed (reflecting a sense of authority), the speaking speed is 90 words per minute (20% slower than normal), a 0.3 second pause is added before the key noun, and the emotions are divided into three stages: the regular narrative stage, the storm plot stage, and the mention of During the shipwreck incident, in the regular narration stage, the virtual NPC's emotion is calm with a touch of nostalgia (heavy breathing); during the storm plot stage, the virtual NPC's emotion is a sudden 5% increase in pitch and a faster speaking speed; during the shipwreck incident stage, the virtual NPC's emotion is a slight vibrato and sigh, so the output of the second voice message can be: (deep and hoarse) "The waves that night... (0.4 second pause) were higher than the mast... (exhalation) We retracted the three sails... (cough) But the second mate... (vibrato) failed to grab the rope after all..."
[0137] In one embodiment, the virtual scene may further include game props, and the method may further include:
[0138] The virtual scene is updated based on the spatial action information to present a visual picture to the player; the visual picture includes a perspective synchronization picture and / or an object interaction picture, the perspective synchronization picture is used to represent the perspective synchronization between the visual picture and the spatial action information, and the object interaction picture is used to represent the state change of the object when the player interacts with the object, and the object includes at least one of a virtual NPC and a game prop.
[0139] The VR interaction layer receives handle operation data and head tracking data from the user layer. Both handle operation data and head tracking data are spatial motion information of the player collected by the VR device. Among them, handle operation data is collected by the handle, and head tracking data is collected by the head-mounted display device. The VR interaction layer can convert spatial motion information (such as handle operation data and head tracking data) into actual actions in the virtual scene.
[0140] The scene management layer works in conjunction with the VR interaction layer. Specifically, the VR interaction layer transmits the player's actual actions in the virtual scene to the scene management layer, enabling it to update the visual image presented to the player in real time based on the player's actual actions in the virtual scene. The visual image can include at least one of a perspective synchronization image and an object interaction image. The perspective synchronization image represents the synchronization of the visual image with the spatial motion information (i.e., head tracking data). For example, when the player moves in the virtual scene, the visual image of the virtual scene changes in sync with the player's perspective. For another example, when the player rotates in the virtual scene, the visual image of the virtual scene changes in sync with the player's perspective. The object interaction image represents the state changes of objects (such as virtual NPCs and game props) when the player interacts with them. For example, if a player shakes hands with a virtual NPC in the virtual scene, the virtual NPC will appear friendly. Or, if a player triggers a mechanism in the virtual scene, a cabinet will open.
[0141] It can be seen that updating the visual images presented to the player by the virtual scene according to the player's spatial action information can ensure the authenticity and interactivity of the virtual scene.
[0142] In order to enable those skilled in the art to better understand the embodiments of the present application, the embodiments of the present application are described below with the help of the following examples.
[0143] Figure 5 It shows the voice interaction process between players and virtual NPCs in a virtual game interaction system based on a large language model, clearly presenting the entire process from player voice input to the voice output of the interaction system. Each link is closely connected to ensure the smoothness and intelligence of virtual game interaction.
[0144] See also Figure 5 ,The virtual game interaction process based on the large language model is as follows:
[0145] S1. Player voice input to the PC: When playing an online script-killing game, the player uses an audio capture device (such as a microphone) connected to the PC to provide voice input. The player's questions to the virtual NPC, their interactions with the virtual NPC, and their commands (i.e., the first voice message) become the starting data for the player's interaction with the interactive system.
[0146] S2, the PC device transmits the audio stream to the Unity engine layer: After receiving the player's voice input (i.e., the first voice message), the PC device converts it into an audio stream format and transmits it to the game development environment running in the Unity engine layer on the PC device. During this process, the PC device collects, initially processes, and transmits the player's voice input (i.e., the first voice message).
[0147] S3: The Unity engine layer calls the ASR interface of the SenseVoice model. After receiving the audio stream, the Unity engine layer calls the ASR interface of the SenseVoice model. The SenseVoice model has powerful speech recognition capabilities. It analyzes and processes the input audio stream, converting the speech content into text that computers can understand.
[0148] S4: The SenseVoice model returns text to the Unity engine layer. After completing the speech recognition task, the SenseVoice model returns the recognized text (i.e., the first text message) to the Unity engine layer. This text (i.e., the first text message) is key data for subsequent interaction logic processing. The interactive system uses this data to understand the player's intent and respond accordingly.
[0149] S5: The Unity engine layer initiates a knowledge base query to the RAGFlow layer. After receiving the text (i.e., the first text message) returned by the SenseVoice model, the Unity engine layer sends it to the RAGFlow layer. Based on the received text (i.e., the first text message), the RAGFlow layer performs a query within a pre-built knowledge base. This knowledge base contains a wealth of information about script-killing games, including but not limited to game settings for each virtual scene (such as plot background and game clues) and character settings for each virtual NPC (such as character background and personality traits). This information supports the large language model in generating accurate responses.
[0150] In step S6, the RAGFlow layer calls the large language model to generate a response and returns it to the Unity engine layer. The RAGFlow layer combines the knowledge base query results (i.e., auxiliary data, including game settings for the current virtual scene and character settings for virtual NPCs within the scene) with the large language model (e.g., Ultra 4.0) to generate the response text (i.e., the second text message). The large language model leverages its powerful natural language processing capabilities and knowledge base to synthesize multiple aspects of information (e.g., auxiliary data, the first prompt word / second prompt word) to output the corresponding response content (i.e., the second text message). The RAGFlow layer then returns the generated response text (i.e., the second text message) to the Unity engine layer.
[0151] S7, the Unity engine layer calls the ChatTTS model's TTS interface: After receiving the reply text (i.e., the second text message) from the RAGFlow layer, the Unity engine layer calls the ChatTTS model's TTS interface. The ChatTTS model is responsible for converting the text reply (i.e., the second text message) into a voice reply (i.e., the second voice message), allowing the player to receive the virtual NPC's response in voice form.
[0152] S8: The ChatTTS model returns the audio stream to the Unity engine layer. Based on the input text (i.e., the second text message), the ChatTTS model uses speech synthesis technology to generate a corresponding audio stream (i.e., the second audio message) and returns it to the Unity engine layer. The audio stream (i.e., the second audio message) generated by the ChatTTS model simulates human acoustic characteristics (such as timbre, pitch, speaking speed, and emotion), providing players with a more natural and authentic interactive experience.
[0153] S9, Unity engine layer outputs voice to VRDevice: The Unity engine layer sends the received voice stream (i.e., the second voice message) to the VR device (such as VIVE HTC Pro). As the game development environment, the Unity engine layer coordinates and manages each layer throughout the interaction process, ensuring that the voice stream (i.e., the second voice message) is accurately output to the VR device.
[0154] S10, VRDevice plays the voice to the player (Player): After the VR device receives the voice stream (i.e., the second voice information) from the Unity engine layer, it plays the voice to the player through its built-in audio playback device.
[0155] At this point, the player has completed a voice interaction with the interactive system, forming a complete closed loop from voice input to final voice output, and realizing natural language interaction between the player and the virtual NPC.
[0156] As can be seen from this example, the solution provided in this application realizes online script-killing through VR equipment and a large language model. Among them, the VR equipment can provide players with virtual scenes of script-killing games (including virtual NPCs), thereby reducing the demand for venues and staff, and thus reducing the implementation cost of script-killing games. The large language model can drive the language logic of the virtual NPC according to the player's real-time input, thereby realizing natural dialogue between the virtual NPC and the player, thereby enhancing the player's immersion in the script-killing game.
[0157] Furthermore, the solution provided in this application discloses the technology stack corresponding to the interactive system framework. Specifically, Unity+XR Interaction Toolkit is used to create the virtual scene of the script-killing game, SenseVoice+ChatTTS is used to realize the voice interaction between players and virtual NPCs, and RAGFlow+knowledge base+large language model is used to drive the language logic of virtual NPCs. The technical implementation details are quantified through various formulas (such as the core formula of error propagation, the state space function of the intelligent agent, and the response function of the intelligent agent) and architecture diagrams (such as the voice interaction process between virtual NPCs and players, and the dual-modal interaction framework), which are reproducible.
[0158] Furthermore, the solution provided by this application introduces a large language model (LLM) and a retrieval-augmented generation framework (RAGFlow). This dual-modal architecture, which identifies whether the player has seen through a virtual NPC's lies, enables intelligent dialogue and mode switching for virtual NPCs. Specifically, the RAGFlow layer queries the knowledge base for auxiliary information related to the question, which can assist the large language model in outputting more accurate responses. Furthermore, the dual-modal interaction framework (spatial action + natural language) enhances the depth of reasoning. By combining the player's spatial action information with the large speech model to analyze the semantics of the player's input, it can determine whether the player has seen through the virtual NPC's lies and dynamically trigger the virtual NPC's behavior mode to switch from misleading mode to honest mode.
[0159] Corresponding to the aforementioned application function implementation method embodiment, the present application also provides a virtual game interaction system, electronic device and corresponding embodiments based on a large language model.
[0160] Figure 6 It is a structural diagram of a virtual game interaction system based on a large language model shown in an embodiment of the present application.
[0161] See also Figure 6 The present application provides a virtual game interaction system based on a large language model, which may include:
[0162] The virtual scene loading module 610 is used to load the virtual scene of the script-killing game for the player through the VR device after detecting that the player is wearing the VR device; the virtual scene includes a virtual NPC;
[0163] The voice collection module 620 is used to collect the first voice information of the player through the VR device; the first voice information includes questions raised by the player to the virtual NPC;
[0164] The natural language processing module 630 is configured to perform speech recognition on the first speech information to obtain a first text message, and input the first text message into a large language model to generate a second text message that matches the first text message through the large language model; the second text message includes a response from the virtual NPC to the question;
[0165] The voice playing module 640 is used to convert the format of the second text information to obtain second voice information, and output the second voice information to the player through the VR device.
[0166] In one embodiment, the natural language processing module 630 may include:
[0167] The data query submodule is used to query auxiliary data information related to the problem from a pre-built knowledge base; the auxiliary data information includes game setting information of the virtual scene and character setting information of the virtual NPC;
[0168] The natural language processing submodule is used to input the first text information, game setting information and character setting information into the large language model, so as to generate second text information matching the first text information based on the game setting information and the character setting information through the large language model.
[0169] In one embodiment, the natural language processing submodule may include:
[0170] A conditional judgment unit, used to determine whether the player has seen through the lies of the virtual NPC;
[0171] A mode switching unit, configured to switch the behavior mode of the virtual NPC from a misleading mode to a candid mode if the player sees through the virtual NPC's lies;
[0172] A first prompt word acquisition unit is used to acquire a first prompt word corresponding to the confession mode; the first prompt word is used to guide the large language model to output the truth;
[0173] The truth output unit is used to input the first text information, game setting information, character setting information and the first prompt word into the large language model, so as to generate a second text information that matches the first text information and does not contain any logical loopholes based on the game setting information, character setting information and the first prompt word through the large language model.
[0174] In one embodiment, the condition judgment unit may include:
[0175] The motion collection subunit is used to collect the player's spatial motion information through the VR device;
[0176] a conditional judgment subunit, configured to judge whether the player has found the key clue based on the spatial action information, and to judge whether the player has raised doubts about the virtual NPC's response based on the first text information;
[0177] The lie detection judgment subunit is used to judge whether the player has detected the lie of the virtual NPC if the player has found a key clue and has raised doubts about the virtual NPC's reply.
[0178] In one embodiment, after determining whether the player has seen through the lies of the virtual NPC, the natural language processing submodule may further include:
[0179] A mode maintaining unit, configured to maintain the behavior mode of the virtual NPC in a misleading mode if the player fails to see through the lies of the virtual NPC;
[0180] A second prompt word acquisition unit is used to acquire a second prompt word corresponding to the misleading mode; the second prompt word is used to guide the large language model to output false words;
[0181] The lie output unit is used to input the first text information, game setting information, character setting information and the second prompt word into the large language model, so as to generate the second text information that matches the first text information and contains a logical loophole based on the game setting information, character setting information and the second prompt word through the large language model.
[0182] In one embodiment, the voice playback module 640 may include:
[0183] an acoustic feature determination submodule for determining the acoustic features of the virtual NPC based on the character setting information; the acoustic features include at least one of timbre, pitch, speech rate, and emotion;
[0184] The voice output submodule is used to output the second voice information to the player through the VR device according to the acoustic characteristics.
[0185] In one embodiment, the virtual scene may further include game props, and the system may further include:
[0186] The screen update module is used to update the visual screen presented to the player by the virtual scene based on the spatial action information; the visual screen includes a perspective synchronization screen and / or an object interaction screen. The perspective synchronization screen is used to represent the perspective synchronization between the visual screen and the spatial action information. The object interaction screen is used to represent the state change of the object when the player interacts with the object. The object includes at least one of a virtual NPC and a game prop.
[0187] As can be seen from this example, the solution provided by this application, after detecting that the player is wearing a VR device, loads the virtual scene of the script-killing game for the player through the VR device; the virtual scene includes a virtual NPC; the player's first voice information is collected through the VR device; the first voice information includes questions asked by the player to the virtual NPC; voice recognition is performed on the first voice information to obtain a first text information, and the first text information is input into the large language model to generate a second text information matching the first text information through the large language model; the second text information includes the virtual NPC's reply to the question; the second text information is formatted to obtain a second voice information, and the second voice information is output to the player through the VR device. This application realizes online script-killing through VR devices and large language models, wherein the VR device can provide players with a virtual scene (including a virtual NPC) of the script-killing game, thereby reducing the demand for venues and staff, thereby reducing the implementation cost of the script-killing game, and the large language model can drive the language logic of the virtual NPC according to the player's real-time input, thereby realizing natural dialogue between the virtual NPC and the player, thereby improving the player's immersion in the script-killing game.
[0188] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.
[0189] Figure 7 It is a structural diagram of an electronic device shown in an embodiment of the present application.
[0190] See also Figure 7 , the electronic device 700 includes a memory 710 and a processor 720.
[0191] The processor 720 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0192] Memory 710 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage. ROM may store static data or instructions required by processor 720 or other computer modules. Permanent storage may be a readable and writable storage device. Permanent storage may be a non-volatile storage device that maintains stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device utilizes a mass storage device (e.g., a magnetic or optical disk, flash memory). In other embodiments, the permanent storage device may be a removable storage device (e.g., a floppy disk, optical drive). System memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory (DRAM). System memory may store some or all instructions and data required by the processor during operation. Furthermore, memory 710 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), as well as magnetic disks and / or optical disks. In some embodiments, the memory 710 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0193] The memory 710 stores executable codes. When the executable codes are processed by the processor 720 , the processor 720 may execute part or all of the above-mentioned methods.
[0194] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.
[0195] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium), which stores executable code (or computer program or computer instruction code) and, when executed by a processor of an electronic device (or server, etc.), enables the processor to perform part or all of the steps of the above-mentioned method according to the present application.
[0196] The present application also provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.
[0197] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A virtual game interaction method based on a large language model, characterized in that: The method comprises: After detecting that a player is wearing a VR device, a virtual scene of a script-killing game is loaded for the player through the VR device; the virtual scene includes a virtual NPC; collecting first voice information of the player through the VR device; the first voice information includes questions asked by the player to the virtual NPC; Performing speech recognition on the first speech information to obtain first text information, and inputting the first text information into a large language model to generate, through the large language model, second text information that matches the first text information; the second text information includes a response made by the virtual NPC to the question; Converting the second text information into a second voice message, and outputting the second voice message to the player through the VR device; Inputting the first text information into a large language model to generate second text information matching the first text information through the large language model includes: Querying auxiliary data information related to the problem from a pre-built knowledge base; wherein the auxiliary data information includes game setting information of the virtual scene and character setting information of the virtual NPC; Inputting the first text information, the game setting information, and the character setting information into a large language model, so as to generate second text information matching the first text information based on the game setting information and the character setting information through the large language model; The step of inputting the first text information, the game setting information, and the character setting information into a large language model, so as to generate, through the large language model, second text information that matches the first text information based on the game setting information and the character setting information, includes: Determining whether the player has seen through the lies of the virtual NPC; If the player sees through the virtual NPC's lie, the virtual NPC's behavior mode is switched from a misleading mode to a frank mode; Obtaining a first prompt word corresponding to the confession mode; the first prompt word is used to guide the large language model to output truth; Inputting the first text information, the game setting information, the character setting information, and the first prompt word into the large language model, so as to generate, through the large language model and based on the game setting information, the character setting information, and the first prompt word, a second text information that matches the first text information and does not contain any logical loopholes; The determining whether the player has seen through the lies of the virtual NPC includes: Collecting the player's spatial motion information through the VR device; determining whether the player has found a key clue based on the spatial action information, and determining whether the player has raised questions regarding the virtual NPC's response based on the first text information; If the player finds a key clue and questions the virtual NPC's reply, it is determined that the player has seen through the virtual NPC's lie.
2. The method according to claim 1, characterized in that After determining whether the player has seen through the lies of the virtual NPC, the method further includes: If the player fails to see through the lie of the virtual NPC, the behavior mode of the virtual NPC is maintained in the misleading mode; Obtaining a second prompt word corresponding to the misleading mode; the second prompt word is used to guide the large language model to output falsehood; The first text information, the game setting information, the character setting information, and the second prompt word are input into the large language model, so as to generate, through the large language model, a second text information that matches the first text information and contains a logical loophole based on the game setting information, the character setting information, and the second prompt word.
3. The method according to claim 1, characterized in that The outputting the second voice information to the player through the VR device includes: Determining acoustic features of the virtual NPC based on the character setting information; the acoustic features including at least one of timbre, pitch, speech speed, and emotion; The second voice information is output to the player through the VR device according to the acoustic characteristics.
4. The method according to claim 1, wherein The virtual scene also includes game props, and the method further includes: The visual picture presented to the player by the virtual scene is updated according to the spatial action information; the visual picture includes a perspective synchronization picture and / or an object interaction picture, the perspective synchronization picture is used to represent the synchronization of the visual picture with the perspective corresponding to the spatial action information, and the object interaction picture is used to represent the state change of the object when the player interacts with the object, and the object includes at least one of the virtual NPC and the game props.
5. A virtual game interaction system based on a large language model, characterized in that: The system comprises: A virtual scene loading module is used to load a virtual scene of the script-killing game for the player through the VR device after detecting that the player is wearing a VR device; the virtual scene includes a virtual NPC; A voice collection module, configured to collect first voice information of the player through the VR device; the first voice information includes questions raised by the player to the virtual NPC; a natural language processing module configured to perform speech recognition on the first speech information to obtain first text information, and input the first text information into a large language model to generate, through the large language model, second text information that matches the first text information; the second text information including a response from the virtual NPC to the question; a voice playing module, configured to convert the format of the second text information to obtain second voice information, and output the second voice information to the player via the VR device; The natural language processing module includes: A data query submodule is used to query auxiliary data information related to the problem from a pre-built knowledge base; wherein the auxiliary data information includes game setting information of the virtual scene and character setting information of the virtual NPC; a natural language processing submodule, configured to input the first text information, the game setting information, and the character setting information into a large language model, so as to generate, through the large language model, second text information that matches the first text information based on the game setting information and the character setting information; The natural language processing submodule includes: A condition judgment unit, used to judge whether the player has seen through the lie of the virtual NPC; a mode switching unit, configured to switch the behavior mode of the virtual NPC from a misleading mode to a confessing mode if the player sees through the lie of the virtual NPC; A first prompt word acquisition unit is configured to acquire a first prompt word corresponding to the confession mode; the first prompt word is used to guide the large language model to output truthful speech; a truth output unit, configured to input the first text information, the game setting information, the character setting information, and the first prompt word into the large language model, so as to generate, through the large language model and based on the game setting information, the character setting information, and the first prompt word, a second text information that matches the first text information and does not contain any logical loopholes; The condition judgment unit includes: A motion collection subunit, configured to collect spatial motion information of the player through the VR device; a conditional judgment subunit, configured to judge whether the player has found a key clue based on the spatial action information, and to judge whether the player has raised questions about the reply to the virtual NPC based on the first text information; The lie detection judgment subunit is used to judge whether the player has detected the lie of the virtual NPC if the player has found a key clue and has raised doubts about the reply of the virtual NPC.
6. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that An executable code is stored thereon, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Game process guiding system and method based on VR
CN113617018A
Virtual space interaction method, system, device and medium for offline script killing
CN116382466A