Method and device for processing interaction of virtual scene, electronic device, computer readable storage medium and computer program product

CN122643700APending Publication Date: 2026-08-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610825974.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0003]本申请实施例提供一种虚拟场景的互动处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品,能够解决多名非玩家角色同时回复同一条消息造成的频道刷屏问题,使群聊对话清晰有序,接近真人队友的交流体验

Benefits of technology

本申请实施例在虚拟对局运行阶段的同队交互场景中,针对玩家角色发送的面向多个非玩家角色的广播消息,基于各非玩家角色的角色特征与广播消息之间的匹配度,从多个非玩家角色中确定匹配度最高的第一非玩家角色输出回复消息,并控制其余非玩家角色保持静默,如此,一方面可以避免多个非玩家角色同时回复同一广播消息而导致的频道混乱和信息冗余,提升队伍群聊交互的清晰性和有序性,另一方面,可以使所输出的回复信息与对应非玩家角色的角色特征更加匹配,增强回复内容的针对性、角色一致性和交互沉浸感。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122643700A_ABST
    Figure CN122643700A_ABST
Patent Text Reader

Abstract

The application provides an interactive processing method and device of a virtual scene, an electronic device, a computer readable storage medium and a computer program product; the method comprises: displaying a first virtual scene, wherein the first virtual scene comprises a player character in the same team and a plurality of non-player characters, and different non-player characters have different character characteristics; in response to a message sending trigger operation, displaying a first message sent by the player character in the first virtual scene, wherein the first message is a broadcast message facing the plurality of non-player characters; controlling a first non-player character to output a second message for replying to the first message, and controlling other non-player characters to be in a silent state, wherein the first non-player character is a non-player character with the highest matching degree of character characteristics and the first message among the plurality of non-player characters. Through the application, the problem of channel screen flooding caused by multiple non-player characters replying to the same message at the same time can be solved, and the group chat conversation is clear and orderly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for interactive processing of virtual scenes. Background Technology

[0002] With the development of tactical competitive and multiplayer cooperative virtual games, the interaction method of configuring player characters and multiple non-player characters in the same team is gradually increasing. Especially during the virtual game, players often need to communicate with their teammates in real time to complete operations such as information sharing, tactical coordination, target confirmation, and status feedback. To enhance the sense of companionship and collaborative experience, non-player characters are usually equipped with automatic dialogue capabilities, enabling them to respond to messages sent by players. In multi-character scenarios during virtual games, messages sent by players are often not directed at a single non-player character, but are broadcast messages to multiple non-player characters within the team. In related technologies, when multiple non-player characters receive the same broadcast message, each non-player character usually generates a reply independently. However, this method is prone to multiple non-player characters responding to the same message simultaneously, causing problems such as message channel flooding, overlapping reply content, unclear dialogue targets, and confusing information expression, which in turn affects the player's identification of key information and interferes with the operation during the virtual game. Summary of the Invention

[0003] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for interactive processing in virtual scenes, which can solve the problem of channel flooding caused by multiple non-player characters replying to the same message at the same time, making group chat dialogue clear and orderly, and providing a communication experience close to that of real teammates.

[0004] The technical solution of this application embodiment is implemented as follows: This application provides an interactive processing method for a virtual scene, including: The first virtual scene is displayed, wherein the first virtual scene is a virtual scene in the virtual game operation stage, and the first virtual scene includes player characters in the same team and multiple non-player characters, and the different non-player characters have different character characteristics. In response to a message sending trigger operation, a first message sent by the player character is displayed in the first virtual scene, wherein the first message is a broadcast message to the plurality of non-player characters; The system controls a first non-player character to output a second message in response to the first message, and controls other non-player characters to remain silent. The first non-player character is the non-player character whose character characteristics match the first message the most among the plurality of non-player characters, and the other non-player characters are the non-player characters other than the first non-player character among the plurality of non-player characters.

[0005] This application provides an interactive processing device for a virtual scene, comprising: The display module is used to display a first virtual scene, wherein the first virtual scene is a virtual scene during the virtual game operation phase, and the first virtual scene includes player characters in the same team and multiple non-player characters, and the different non-player characters have different character characteristics. The display module is further configured to respond to a message sending trigger operation and display a first message sent by the player character in the first virtual scene, wherein the first message is a broadcast message to the plurality of non-player characters; The control module is used to control the first non-player character to output a second message in response to the first message, and to control other non-player characters to remain silent. The first non-player character is the non-player character whose character characteristics match the first message the most among the plurality of non-player characters, and the other non-player characters are the non-player characters other than the first non-player character among the plurality of non-player characters.

[0006] This application provides an electronic device, including: Memory is used to store executable instructions for a computer; The processor, when executing computer-executable instructions stored in the memory, implements the interactive processing method for virtual scenes provided in the embodiments of this application.

[0007] This application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the interactive processing method for a virtual scene provided in this application.

[0008] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the interactive processing method for virtual scenes provided in this application.

[0009] The embodiments of this application have the following beneficial effects: In the team interaction scenario during the virtual game operation phase, this application embodiment addresses broadcast messages sent by player characters to multiple non-player characters. Based on the matching degree between the character characteristics of each non-player character and the broadcast message, the first non-player character with the highest matching degree is determined to output a reply message, while the remaining non-player characters remain silent. This avoids channel confusion and information redundancy caused by multiple non-player characters replying to the same broadcast message simultaneously, improving the clarity and orderliness of team group chat interaction. Furthermore, it makes the output reply information more closely match the character characteristics of the corresponding non-player character, enhancing the relevance of the reply content, character consistency, and interactive immersion. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the architecture of the virtual scene interactive processing system 100 provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the electronic device 500 provided in the embodiments of this application; Figure 3 This is a first flowchart illustrating the interactive processing method for virtual scenes provided in this application embodiment; Figure 4 This is a schematic diagram of the second process of the interactive processing method for virtual scenes provided in the embodiments of this application; Figure 5 This is a schematic diagram of the third process of the interactive processing method for virtual scenes provided in the embodiments of this application; Figure 6 This is a schematic diagram of the first process of the interactive processing method for virtual scenes provided in the embodiments of this application; Figure 7 This is a schematic diagram of the second process of the virtual scene interaction processing method provided in the embodiments of this application; Figure 8 This is a schematic diagram of the third process of the virtual scene interaction processing method provided in the embodiments of this application; Figure 9 This is a schematic diagram of the fourth process of the virtual scene interaction processing method provided in the embodiments of this application; Figure 10 This is a schematic diagram of the fifth process of the virtual scene interaction processing method provided in the embodiments of this application. Figure 11 This is a schematic diagram of an emotion coordinate axis composed of pleasure and activation provided in an embodiment of this application; Figure 12 This is a schematic diagram of the emotional PA distribution provided in the embodiments of this application. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0013] It is understood that in the embodiments of this application, data such as user information are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0014] In the following description, the terms “first, second, ...” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first, second, ...” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0016] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0017] 1) Responding to: used to indicate the conditions or states on which the operation is performed depends. When the conditions or states on which it depends are met, one or more operations can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0018] 2) Virtual Scene: This refers to the scene displayed (or provided) by the application when it runs on the terminal device. This scene can be a simulation of the real world, a semi-simulated / semi-fictional virtual environment, or a purely fictional virtual environment. A virtual scene can be any of a two-dimensional, 2.5-dimensional, or three-dimensional virtual scene; this application does not limit the dimension of the virtual scene. For example, a virtual scene may include the sky, land, ocean, etc., and the land may include environmental elements such as deserts and cities. Players can control their character to move within this virtual scene.

[0019] 3) Player Character: A player character is a virtual character in a virtual environment that corresponds to a real user and is controlled or at least partially controlled by the real user. A player character can be a character in the game, a combat unit, a control device, or other interactive virtual entity. Player characters are typically used to represent the behavior, state, and interactive intentions of a real user in a virtual game, such as moving, attacking, sending messages, and executing commands.

[0020] 4) Non-player characters: These are virtual characters that exist in a virtual environment but are not directly controlled by the current real user. Non-player characters can be driven by game logic, script rules, artificial intelligence (AI) models, decision-making systems, or a combination thereof. Non-player characters can be teammates on the same team as player characters, or they can possess abilities such as cooperative combat, information feedback, dialogue responses, and task coordination. In the embodiments of this application, multiple non-player characters typically coexist with player characters in the same virtual game environment.

[0021] 5) Character Traits: These refer to one or more attribute information used to characterize the differences between non-player characters. Character traits may include at least one of the following: identity setting, job positioning, combat responsibilities, skill characteristics, voice style, expression habits, personality tendencies, relationship with player characters, and historical interaction characteristics. Different non-player characters have different character traits, which allows them to have different adaptability and response tendencies when faced with the same message.

[0022] 6) Silent State: This refers to the state in which a non-player character does not output a reply message to a broadcast message sent by a player in the current message response round. A silent state does not necessarily mean that the non-player character has completely stopped operating, nor does it necessarily mean that it has stopped performing other game actions. That is, a non-player character in a silent state can still perform non-dialogue actions such as moving, fighting, following, searching, and rescuing; they simply will not output any speech in the reply processing corresponding to the current broadcast message.

[0023] 7) Human-Computer Interaction Interface: This refers to an interface used to provide human-computer interaction functions or to display virtual scenes. For example, a human-computer interaction interface can be a graphical user interface (GUI), an augmented reality (AR) interface, a virtual reality (VR) interface, a voice user interface (VUI), an interactive projection interface (using projection technology to display information on a flat surface), an eye-tracking interface (an interface controlled by detecting the user's gaze), a holographic interface (a three-dimensional hologram formed by projecting images using holographic projection technology, allowing the user to see stereoscopic images without special glasses), a multimodal interface (an interface combining multiple interaction methods, such as an interface combining touch, vision, and hearing), and a brain-machine interface (BMI), etc.

[0024] 8) Cloud Gaming: Also known as Gaming on Demand, this involves deploying a game program on a server and running an instance of the game program (referred to as a game instance). The game instance sends the game data output during its operation to the user's browser page. The page uses the browser's media components to decode the game data and renders the real-time game screen based on the decoding results. When the page detects user actions in the game screen, it reports this to the game instance running on the server. Upon receiving the game data generated by the game instance in response to the action, the page repeats the decoding and rendering process, thus displaying the changes in the game screen based on the user's actions.

[0025] In other words, cloud gaming is an online gaming technology based on cloud computing. Cloud gaming technology enables thin clients with relatively limited graphics processing and data processing capabilities to run high-quality games. In a cloud gaming scenario, the game does not run on the user's terminal (e.g., the player's gaming device), but rather on a cloud server. The cloud server renders the game scene as an audio and video stream, which is then transmitted to the user's terminal via the network. Therefore, the user's terminal does not need powerful graphics processing and data processing capabilities; it only needs basic streaming media playback capabilities and the ability to receive player input commands and send them to the cloud server.

[0026] 9) Large Language Model (LLM): This refers to a deep learning model trained on a large amount of text data, enabling it to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on a wide range of topics by being trained on massive datasets. The core idea is to learn the patterns and structures of natural language through large-scale unsupervised training, thus simulating the human language cognition and generation process to some extent.

[0027] In implementing the embodiments of this application, the applicant discovered the following problems with the solutions provided by related technologies: 1) Dialogue and team system disconnected: In related technologies, AI teammate dialogues mostly only complete question and answer generation, without forming a closed loop with multi-player team formation, command distribution, and combat status perception in the game; 2) Chaotic conversations when multiple AI teammates are online at the same time: The relevant technology is not designed with relevance judgment and silence mechanism, which leads to multiple AI teammates replying to the same message at the same time, causing the channel to be flooded; 3) Lack of specialized matching in command distribution: When players issue commands to multiple AI teammates, the relevant technology lacks a differentiated matching mechanism, resulting in similar response styles; 4) AI teammates' dialogue interferes with player operations during combat: The relevant technology does not have a combat event-driven dialogue interruption mechanism, which causes AI teammates to send idle messages during intense combat, interfering with player operations; 5) AI teammates have no emotional connection with players during the spawn island phase: In related technologies, AI teammates only play fixed lines during the spawn island phase, without personalized dialogue based on historical memory; 6) Differences in AI teammate personas: In related technologies, the differences in AI teammate response styles are mainly reflected in the content of questions and answers, without adaptive adjustment based on the player's emotional state.

[0028] In view of this, embodiments of this application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for interactive processing in virtual scenes. This solution addresses the channel flooding problem caused by multiple non-player characters simultaneously replying to the same message, making group chat conversations clear and orderly, and providing a communication experience close to that of real teammates. The electronic device provided in this application embodiment is described below. The electronic device provided in this application embodiment can be implemented as a terminal device (corresponding to a standalone game application), or implemented collaboratively by a terminal device and a server (corresponding to an online game application). The following description uses the interactive processing method for virtual scenes provided in this application embodiment, implemented collaboratively by a server and a terminal device, as an example.

[0029] Before introducing the architecture of the virtual scene interactive processing system provided in this application embodiment, the game modes involved in this application embodiment are first introduced. For the scheme implemented collaboratively by terminal devices and servers, two main game modes are involved: local game mode and cloud game mode. In local game mode, the terminal device and server collaboratively run the game processing logic. The operation commands input by the player on the terminal device are partly processed by the terminal device's game logic and partly by the server's game logic. Furthermore, the game logic processed by the server is often more complex and requires more computing power. In cloud game mode, the game logic is entirely processed by the server (e.g., a cloud server), and the cloud server renders the game scene data into audio and video streams, which are then transmitted to the terminal device for display via the network. In other words, the terminal device only needs basic streaming media playback capabilities and the ability to obtain player operation commands and send them to the server.

[0030] The architecture of the virtual scene interactive processing system provided in the embodiments of this application will be described below.

[0031] In some embodiments, see Figure 1 , Figure 1 This is a schematic diagram of the architecture of the virtual scene interactive processing system 100 provided in the embodiments of this application, as shown below. Figure 1 As shown, the virtual scene interactive processing system 100 provided in this application embodiment includes: a server 200 (e.g., a game backend server), a network 300, and a terminal device 400. The network 300 can be a local area network (LAN) or a wide area network (WAN), or a combination of both. The terminal device 400 is a terminal device associated with the player. A client 410 runs on the terminal device 400. The client 410 can be various types of game clients, such as shooting game clients, action game clients, open-world game clients, role-playing game clients, and browsers.

[0032] For example, taking client 410 as a shooting game client, the human-computer interaction interface of client 410 can display a first virtual scene. This first virtual scene can be a virtual scene during a virtual game session. Within this first virtual scene, player characters in the same team (e.g., game character A controlled by the current player) and multiple non-player characters (e.g., multiple game characters controlled by an AI model, referred to as AI teammates) can be displayed. Different non-player characters can have different character characteristics (e.g., different AI teammates have different personas). Then, in response to a message sending trigger, client 410 can display a first message sent by a player character in the first virtual scene. This first message can be a broadcast message to multiple non-player characters in the same team. After receiving the first message from the player, server 200 can, based on the character characteristics of each non-player character, select the first non-player character that best matches the first message and call a pre-trained language model to generate a second message to reply to the first message. Afterwards, server 200 can send the character identifier of the first non-player character and the second message to client 410, so that client 410 can control the first non-player character to output the second message in response to the first message. In this way, the channel flooding problem caused by multiple non-player characters replying to the same message at the same time can be solved, making the group chat clear and orderly, and providing a communication experience close to that of real teammates. It should be noted that the virtual scene in the interactive processing method of the virtual scene provided in this application embodiment can be entirely based on the output of the terminal device, or based on the collaborative output of the terminal device and the server. For example, it can rely entirely on... Figure 1 The terminal device 400 shown in the diagram utilizes its graphics processing hardware computing capabilities to perform data calculations and outputs related to the virtual scene. The graphics processing hardware includes both a central processing unit (CPU) and a graphics processing unit (GPU). For example, when visual perception of a virtual scene is formed, the terminal device 400 uses its graphics computing hardware to calculate the data required for display, and completes the loading, parsing, and rendering of the display data. The graphics output hardware outputs video frames capable of forming visual perceptions of the virtual scene; for example, displaying two-dimensional video frames on a smartphone screen, or projecting video frames to achieve a three-dimensional display effect onto the lenses of augmented reality / virtual reality glasses. Furthermore, to enrich the perceptual effects, the terminal device 400 can also utilize different hardware to form one or more of the following: auditory perception, tactile perception, motion perception, and gustatory perception.

[0033] Of course, the computing power of server 200 can also be used to complete the virtual scene calculation and output the virtual scene to terminal device 400. For example, taking the visual perception of forming a virtual scene as an example, server 200 calculates the display data (such as scene data) related to the virtual scene and sends it to terminal device 400 through network 300. Terminal device 400 relies on graphics computing hardware to complete the loading, parsing and rendering of the calculated display data, and relies on graphics output hardware to output the virtual scene to form a visual perception. For example, two-dimensional video frames can be presented on the display screen of a smartphone, or video frames that achieve a three-dimensional display effect can be projected onto the lenses of augmented reality / virtual reality glasses. As for the perception of the form of the virtual scene, it can be understood that the corresponding hardware output of terminal device 400 can be used, such as using a microphone to form auditory perception, using a vibrator to form tactile perception, and so on.

[0034] In addition, it should be noted that, Figure 1 The server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal device 400 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart TV, in-vehicle terminal, etc., but is not limited to these. The terminal device 400 and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.

[0035] In other embodiments, the terminal device can also implement the interactive processing method for virtual scenes provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as shooting game APPs; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0036] The structure of the electronic device provided in the embodiments of this application will be further described below. Taking the electronic device as a terminal device as an example, see... Figure 2, Figure 2 This is a schematic diagram of the structure of the electronic device 500 provided in the embodiments of this application. Figure 2 The illustrated electronic device 500 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 540.

[0037] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0038] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0039] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.

[0040] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.

[0041] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0042] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 553 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with user interface 530. The input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532. In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 An interactive processing device 555 for a virtual scene stored in memory 550 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a display module 5551, a control module 5552, a playback module 5553, a feature extraction module 5554, a feature encoding module 5555, a determination module 5556, a generation module 5557, and an input module 5558. These modules are logically connected and can therefore be arbitrarily combined or further separated according to the functions implemented. It should be noted that... Figure 2 For ease of explanation, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the interactive processing device 555 in the virtual scene, which may only include the display module 5551 and the control module 5552. The functions of each module will be described below.

[0043] The interactive processing method for virtual scenes provided in this application will be specifically described below with reference to exemplary applications and implementations of the terminal devices provided in the embodiments of this application.

[0044] For example, see Figure 3 , Figure 3 This is a first flowchart illustrating the interactive processing method for virtual scenes provided in this application embodiment, which will be combined with... Figure 3 The steps shown are explained.

[0045] It should be noted that, Figure 3The method illustrated can be executed by various forms of computer programs running on the terminal device, and is not limited to a client. For example, it can also be the operating system, software module, script, and applet mentioned above. Therefore, the client-side examples used below should not be considered as limiting the embodiments of this application. Furthermore, for ease of description, no specific distinction will be made between the terminal device and the client running on the terminal device below.

[0046] In step 101, the first virtual scene is displayed.

[0047] Here, the first virtual scenario refers to the virtual scenario during the virtual match's operational phase. This first virtual scenario can include player characters on the same team and multiple non-player characters, where different non-player characters can have different character characteristics. Furthermore, the virtual match's operational phase can refer to the stage where the virtual match has begun and is in a continuous interactive process. This phase is typically distinct from the login phase, matchmaking phase, loading phase, or settlement phase. In this phase, player characters and multiple non-player characters have entered the same match environment and can move, fight, cooperate, communicate, and exchange messages. The same team can refer to a set of characters in the virtual match who share a common faction, common goal, cooperative relationship, or friendly relationship. Player characters and non-player characters on the same team typically share consistency in mission objectives, tactical coordination, and information sharing; therefore, messages sent by a player character can be directed to multiple non-player characters within that team.

[0048] In some embodiments, the terminal device can display a first virtual scene after the virtual game enters the running phase. The virtual game running phase refers to a stage where the virtual game has completed matchmaking, loading, or character deployment, and has entered a phase where movement, combat, cooperative interaction, and message communication are possible. The first virtual scene can be the game scene corresponding to the current virtual game, such as a tactical competitive scene, a team battle scene, a mission cooperation scene, or other multiplayer cooperative scene. The terminal device can render and display the first virtual scene based on scene data, character data, and game status data sent by the server, allowing players to observe the current game environment and the status information of each character in the team through the display interface. During the display of the first virtual scene, the first virtual scene can include a player character and multiple non-player characters. The player character can be a game character corresponding to the current player's account and controlled by the current player; the multiple non-player characters can be virtual teammates existing in the first virtual scene, which are not directly controlled by the current player but are driven by game logic, scripts, decision-making models, artificial intelligence models, or a combination thereof. Player characters and multiple non-player characters can belong to the same team, for example, by sharing the same faction identifier, team number, friendly identification, or common mission objective, thus forming a cooperative relationship in virtual matches. Accordingly, the first virtual scene can also display team-related interface elements, such as teammate names, health status, location markers, team numbers, voice or message status markers, etc., to help players identify teammates. Furthermore, different non-player characters can have different character characteristics. These characteristics can be attribute information used to characterize the differences between non-player characters, such as at least one of the following: identity setting, class type, combat role, skill abilities, weapon preferences, interaction style, language style, personality traits, response habits, relationship information with player characters, or historical cooperation information. For example, non-player character A can be set as a reconnaissance character, tending to respond to enemy situation, path, and vision-related content; non-player character B can be set as a support character, tending to respond to rescue, supply, and protection-related content; and non-player character C can be set as an assault character, tending to respond to attack, charge, and suppression-related content. By configuring different character traits for different non-player characters, multiple non-player characters can form differentiated character settings and interaction bases within the same virtual environment. Furthermore, character traits can be pre-configured static traits or dynamically adjusted traits based on the current game state. Static traits include character class, basic personality, default language style, etc.; dynamic traits include current equipment status, current task assignment, current position, remaining health, real-time distance from the player character, most recently performed action, or current tactical role, etc.Terminal devices or servers can load character feature data corresponding to each non-player character when entering the first virtual scene, and update some character features according to the real-time status during the game to make subsequent message processing based on character features more in line with the current game scene.

[0049] For example, when players enter a four-player cooperative match, the terminal device can display a first virtual scene containing two player characters and two non-player characters, where the two player characters and two non-player characters belong to the same team. The two non-player characters (e.g., two AI teammates) can correspond to two different role characteristics: medic support and melee assault. When displaying the first virtual scene, the terminal device can not only show the position, actions, and status of each character on the map, but also load appearance, equipment, action style, or interaction style corresponding to their respective character characteristics, thus providing a basis for differentiated message responses among multiple non-player characters later.

[0050] In some other embodiments, before performing step 101 above, the following processing may also be performed: displaying a second virtual scene, wherein the second virtual scene is a virtual scene of the accurate stage of a virtual game; controlling a second non-player character to output an opening message in the second virtual scene, wherein the opening message may be generated based on the player character's historical behavior data, and the second non-player character may be any one of the aforementioned multiple non-player characters; in response to a message sending trigger operation, displaying a third message sent by the player character to reply to the opening message in the chat area of ​​the second virtual scene, and controlling the second non-player character to continue outputting a fourth message to reply to the third message.

[0051] For example, before displaying the first virtual scene, the terminal device can also display a second virtual scene. This second virtual scene can be a virtual scene for the virtual match preparation phase, such as the team waiting scene after character matching is complete, before officially entering the match map, the spawn island scene, the preparatory scene after loading, or other virtual scenes used for team pre-interaction. The second virtual scene can display the player character and at least some of the non-player characters, and can display interface elements such as a chat area, character identification area, preparation status area, and voice or text interaction entry points to establish interaction relationships between the player character and non-player characters before the official match begins. The terminal device can control the second non-player character to output an opening message in the second virtual scene. This second non-player character can be any one of multiple non-player characters, pre-specified by the system, or determined from multiple non-player characters according to a preset strategy. For example, based on character activity, character settings, historical interaction frequency with the player character, or current team configuration, a suitable second non-player character to initiate pre-interaction can be selected from multiple non-player characters. The opening message can be the first greeting, small talk, status confirmation, tactical probe, or relationship-building message output by the second non-player character to the player character in the second virtual scenario. The opening message can be generated based on the player character's historical behavioral data. Historical behavioral data can include at least one of the following: the player character's operating style, commonly used weapons, commonly used tactics, commonly used landing spots, rescue tendencies, communication frequency, historical interaction preferences, win / loss performance, teaming habits, or historical interaction records with various non-player characters in past virtual matches. The terminal device or server can extract features from the historical behavioral data to obtain a behavioral profile of the player character, and based on this profile, determine opening topics, forms of address, semantic content, or expression style more relevant to the player character, thereby generating a personalized opening message. For example, if a player character's historical behavior data indicates a high frequency of support and strong teamwork in past matches, the opening message from the second non-player character could be, "Shall we stick to your usual playstyle this game? I'll cover for you." If the player character's historical behavior data indicates a preference for fast attacks and proactive enemy scouting, the opening message from the second non-player character could be, "Are you going to push the point first? I can scout ahead for you." Thus, the opening message is no longer a uniform, fixed phrase, but rather personalized interactive content that reflects the current player character.

[0052] For example, continuing from the above, players can also reply to the opening message sent by a second non-player character. For instance, in response to a message sending trigger, the terminal device can display a third message sent by the player character in the chat area of ​​the second virtual scene as a reply to the opening message. This message sending trigger can include the player clicking the send control, confirming a quick reply, completing voice input, clicking a recommended reply option, or performing other input operations that trigger message sending. The third message can be the reply content entered or selected by the player in response to the opening message, such as confirming, denying, asking follow-up questions, providing supplementary explanations, or extending the communication. The chat area can be the text chat display area, chat bubble display area, team message display area, or message display area in the comprehensive interaction panel within the second virtual scene. After displaying the third message, the terminal device can also control the second non-player character to continue outputting a fourth message in response to the third message. This fourth message can be a series of reply messages generated by the second non-player character based on the semantic content of the third message, the context of the opening message, and the character characteristics of the second non-player character. In other words, in the second virtual scenario, the second non-player character can engage in at least two rounds of dialogue with the player character around the opening message, establishing an initial communication link before officially entering the first virtual scenario. The fourth message can be a direct response to the third message, or it can be further tactical coordination, emotional feedback, personal expression, or relationship reinforcement. For example, after the second non-player character outputs the opening message "Are you still planning to go to the high-resource area first this round?", the player can send the third message "Check the flight path first, if it's not suitable, play it safe" through message sending trigger operation; subsequently, the second non-player character can output the fourth message "Understood, then I'll prioritize keeping an eye on the surrounding landing points for you, and I'll report the location if someone tries to take the point." As another example, after the second non-player character outputs the opening message "You rescued someone very timely last round, I'll go with you this round," the player sends the third message "Okay, make sure to cover my position"; correspondingly, the second non-player character can continue to output the fourth message "Received, I'll prioritize watching your flanks and rear." Through the above methods, a round of pre-interaction with contextual continuity can be completed before officially entering the first virtual scenario.

[0053] It should be noted that the opening, third, and fourth messages in the second virtual scene can be displayed in the same chat area and arranged chronologically to form a continuous dialogue record. This continuous dialogue record can also be retained, collapsed, or used as contextual input to continue generating subsequent messages after switching back to the first virtual scene. Therefore, the pre-interaction content in the second virtual scene can form a contextual connection with the formal interaction content in the first virtual scene, making the responses of non-player characters more coherent in subsequent stages of the game. Furthermore, the second non-player character can belong to the same character set as multiple non-player characters who subsequently participate in selecting broadcast message responses in the first virtual scene. In other words, the second non-player character initiating the opening interaction in the second virtual scene can be any of the subsequent non-player characters. This design allows non-player characters to have a certain degree of personality exposure and relationship setup before the game begins, thus providing supplementary context for selecting a more suitable target response character based on the broadcast message content in the first virtual scene.

[0054] As can be seen, this application embodiment, on the one hand, by adding a pre-interaction process in the second virtual scene before displaying the first virtual scene, allows the player character to establish an initial communication relationship with the non-player character before officially entering the game, avoiding the non-player character only passively responding during the official game, thereby improving the naturalness of the interaction. Furthermore, by controlling the second non-player character to output an opening message generated based on the player character's historical behavior data, the opening content can be matched with the player character's past habits, communication style, and combat preferences. Compared to a fixed template greeting, this is more likely to reflect personalization and familiarity, enhancing the player's perception that the non-player character "understands them." On the other hand, by forming at least two rounds of continuous dialogue—"opening message - third message - fourth message"—in the second virtual scene, context establishment can be completed before the official game begins, giving the non-player character's subsequent replies in the first virtual scene better contextual continuity and reducing the abruptness of the replies. Simultaneously, since the second non-player character can be any one of multiple non-player characters, differentiated character displays can be made for different non-player characters before the game, which helps players perceive the expression style and character characteristics of each non-player character in advance, thereby enhancing subsequent collaborative interaction among multiple non-player characters.

[0055] In step 102, in response to the message sending trigger operation, the first message sent by the player character is displayed in the first virtual scene.

[0056] Here, the first message can be a broadcast message aimed at multiple non-player characters. A broadcast message can be a message sent by a player character and intended for multiple non-player characters. Broadcast messages are typically not targeted at a single non-player character, but rather are visible, perceptible, or actionable to multiple candidate responding characters simultaneously. Broadcast messages can be text messages, voice messages, shortcut messages, preset phrase messages, text messages obtained through speech recognition, or other information carriers that express the player's interactive intent. Broadcast messages can be natural language content, structured commands, or semi-structured prompts; for example, they can be messages sent by player characters in public channels, team channels, or group chats. In other words, the player does not specify a particular non-player character in the team when sending the first message.

[0057] In some embodiments, the terminal device can respond to a message sending trigger operation by displaying a first message sent by a player character in a first virtual scene. The message sending trigger operation may include, but is not limited to, at least one of the following operations: the player enters text in a chat input box and triggers the send control; the player selects a preset message in a quick message panel; the player presses a voice input control and completes voice acquisition; the terminal device recognizes the voice content and generates corresponding text; or the player confirms message sending through other interactive controls. After detecting a message sending trigger operation, the terminal device can obtain the message content corresponding to the operation and generate a first message sent by the player character. The first message may be a broadcast message directed to multiple non-player characters. A broadcast message means that the reception scope of the first message is not limited to a single non-player character, but is sent to multiple candidate non-player characters in the first virtual scene. In other words, multiple non-player characters can receive, perceive, parse, or participate in the subsequent response determination to the first message. The broadcast message may be a natural language message, or a structured message, a semi-structured message, a preset phrase message, a tactical signal message, or a text message converted from speech. Furthermore, broadcast messages can include inquiries, tactical instructions, status confirmation requests, resource requests, support requests, location hints, action reminders, or other information suitable for sharing within the team. For example, the first message could be something like "Is anyone ahead?", "Advance with me", "Check the right side for me", "Don't fire yet", "Does anyone have a medkit?", or "Preparing to move to the safe zone". Because these messages are typically shared among multiple characters in the same team, they are more suitable as broadcast messages to trigger subsequent response choices among multiple non-player characters.

[0058] It should be noted that the terminal device can display the first message in the chat area of ​​the first virtual scene. This chat area can be the team chat bar, message list area, floating dialog panel, or other interface area used to display game messages. Furthermore, the terminal device can also display a message bubble, voice icon, or avatar prompt corresponding to the first message near the player character, making the first message visible not only in the interface chat area but also reflecting the player character's message sending action at the scene level. In addition, while displaying the first message, the terminal device or server can also determine the message type, reception range, or broadcast identifier corresponding to the first message. For example, a team broadcast identifier can be attached to the first message to indicate that the message is intended for multiple non-player characters within the same team; message topic tags, semantic intent tags, or context identifiers can also be attached so that subsequent non-player characters can make appropriate judgments about the first message based on their respective character characteristics. Thus, the first message not only exists as displayed content but can also serve as input information for the subsequent non-player character reply generation process. Moreover, the first message can be initiated by the player character or selected by the player character from system-recommended messages. For example, when the system identifies an enemy target, resource shortage, or relocation need near a player character based on the current game status, it can recommend preset broadcast messages related to the current scenario in the quick message panel. When the player selects one of the recommended messages and confirms sending, the terminal device can display the selected message as the first message in the first virtual scenario. This reduces the input cost for players in highly real-time games. For example, in one scenario, two player characters and two non-player characters are on the same team and are advancing in the first virtual scenario. At this time, the current player can trigger a message sending operation by clicking the quick message control, and the terminal device will then display the first message sent by the player character controlled by the current player in the chat area, such as "There may be enemies to the right front, pay attention." Since this first message is open to multiple non-player characters in the same team, both non-player characters can perceive and process it as a broadcast message. As another example, the current player can input "Who will defend this place with me?" via voice. After completing voice recognition, the terminal device can display the recognition result as the first message in the chat area of ​​the first virtual scenario and use this message as a broadcast message to multiple non-player characters in the subsequent reply process. Furthermore, the display of the first message in the first virtual scene can be chronologically ordered to create a continuous dialogue context with subsequent replies from non-player characters. Additionally, the first message can be associated with the scene state at the time of sending, such as the player character's current location, orientation, surrounding enemy situation, current mission objective, or team status, thus providing more complete contextual support for subsequent non-player characters to understand the true intent of the first message.

[0059] In step 103, the first non-player character is controlled to output a second message in response to the first message.

[0060] Here, the first non-player character can be the non-player character whose character characteristics match the first message the most among the aforementioned non-player characters; that is, the non-player character determined to execute a response output for the first message. The matching degree can be a quantitative value or ranking result used to characterize the fit between the non-player character's character characteristics and the first message. The matching degree can be determined based on the semantic content, topic, intent, instruction type, and contextual information of the first message, as well as the non-player character's identity setting, role positioning, expression style, and historical interaction characteristics. A higher matching degree generally indicates that the corresponding non-player character is more suitable to respond to the first message. The second message can be the response message output by the first non-player character in response to the first message. The second message can be in text form, voice form, or a combination of text and voice, and can also include character actions, facial expressions, speech bubble patterns, or other expressive information. The content of the second message is usually related to the first message and is constrained by the character characteristics of the corresponding non-player character.

[0061] In some embodiments, step 103 can be implemented as follows: in the chat area of ​​the first virtual scene, a second message sent by the first non-player character to reply to the first message is displayed, wherein the second message can be displayed in the form of a message bubble, and the message bubble can display the character identifier of the first non-player character (such as including the character avatar or character name).

[0062] For example, after displaying the first message sent by a player character in the first virtual scene, the terminal device can also display a second message sent by a first non-player character in the chat area of ​​the first virtual scene as a reply to the first message. The first non-player character can be a target non-player character selected from multiple non-player characters to reply to the first message. The second message can be a reply message generated based on the semantic content of the first message, the current game state, and the character characteristics corresponding to the first non-player character. The chat area can be a team chat display area, a dialogue information display area, a side message bar, or a floating message interaction area in the first virtual scene. The terminal device can display the first and second messages in the chat area according to the chronological order of their generation, creating a continuous dialogue between the first message sent by the player character and the second message sent by the first non-player character. This allows players to intuitively view the current interaction context during the virtual game. Furthermore, the second message can be displayed in the form of a message bubble. The message bubble can be a visual area surrounding the text content of the second message, such as a rounded rectangular bubble, a chat bubble with a pointer, or other graphical interface elements suitable for display in the chat area. The terminal device can display the text content, speech-to-text content, preset phrase content, or semantically generated content of the second message within a message bubble. It can also style the message bubble using preset colors, border styles, transparency, font size, or background textures to distinguish it from message bubbles sent by player characters. Additionally, the message bubble can display the character identifier of the first non-player character. This character identifier identifies the sender of the second message, allowing players to recognize the non-player character corresponding to the current reply. The character identifier can include at least one of the following: character avatar, character name, character abbreviation, character number, class icon, role label, faction mark, team rank identifier, or a personalized symbol corresponding to the character's characteristics. For example, when the first non-player character is a reconnaissance character, the character identifier can include a telescope icon or reconnaissance label corresponding to the reconnaissance role; when the first non-player character is a support character, the character identifier can include a medical cross icon or support label. The character identifier can be displayed in a preset position within the message bubble, such as the upper left corner, left side area, top title area, or a position displayed alongside the message text. Furthermore, the terminal device can associate the character identifier with the visual style of the message bubble. For example, message bubbles for different non-player characters can use different border colors, background patterns, or font colors for their names to further enhance differentiation when multiple non-player characters speak consecutively. For instance, when multiple non-player characters output reply messages in the first virtual scene, the second message from each non-player character can be displayed as a message bubble with their respective character's identifier, thus forming a distinguishable multi-character dialogue interface.

[0063] It should be noted that, in addition to being displayed in the chat area, the second message can also be associated with the first non-player character's behavior in the first virtual scene. For example, while the second message is displayed in the chat area, the terminal device can also display a simplified bubble, speech status marker, or voice prompt marker corresponding to the second message near the first non-player character, allowing players to simultaneously perceive that the first non-player character has performed a response from both the chat area and the scene character level. This enhances the consistency between message display and character behavior within the scene. Furthermore, the character identifier in the message bubble can also be used to participate in subsequent interactions. For example, when a player clicks on the character identifier or the corresponding message bubble, the terminal device can display brief information about the first non-player character, their current status, role, or a quick interaction entry. In this way, the character identifier not only displays the message source but also serves as an interaction entry point for players to further identify and operate the corresponding non-player character.

[0064] As can be seen, by displaying the second message sent by the first non-player character in the chat area of ​​the first virtual scene, the embodiments of this application provide clear and visual feedback for broadcast messages sent by player characters, making it easier for users to understand the response results within the team in a timely manner. Furthermore, by displaying the second message in the form of a message bubble, the reply content has clear boundaries and high readability on the interface, allowing players to quickly capture key information during the game and reducing the reading burden. Simultaneously, by displaying the character identifier of the first non-player character in the message bubble, the sender of the second message can be clearly distinguished, avoiding confusion about the message source in scenarios where multiple non-player characters exist simultaneously, thereby improving the recognizability of multi-character interactions.

[0065] In other embodiments, following the above examples, the above-described display of the second message sent by the first non-player character in response to the first message can be achieved by: displaying the second message sent by the first non-player character in response to the first message using a different display style than the first message, wherein the second message and the first message are displayed adjacent to each other in the chat area to represent the reply relationship of the second message to the first message, or, in the chat area, the reply relationship of the second message to the first message is indicated by a connection mark.

[0066] For example, when displaying a second message sent by a first non-player character in the chat area of ​​a first virtual scene as a reply to a first message, the terminal device can use a different display style to display the second message, allowing players to intuitively distinguish between the first message sent by the player character and the second message sent by the first non-player character in the chat area. The different display styles can include at least one of the following differentiated display methods: different message bubble colors, different message bubble shapes, different border styles, different transparency, different font colors, different font styles, different message arrangement positions, different character identifier display methods, different message background textures, or different message prefix tags. For instance, the first message sent by the player character can use a message bubble of the first color and be displayed on the first side of the chat area, while the second message sent by the first non-player character can use a message bubble of the second color and be displayed on the second side of the chat area; or, the first message can be displayed using a regular bubble style, while the second message can be displayed using a bubble style with a character identifier, a highlighted border, or a role label. Through these methods, players can quickly identify that the second message is a reply from a non-player character to the first message without having to read the entire text. Furthermore, the second message and the first message can be displayed adjacently in the chat area to represent the second message's response to the first message. Adjacent display can mean that the second message and the first message are displayed immediately next to each other in the message arrangement order of the chat area, without any irrelevant messages inserted in between; or, the first message and the second message form the same group of message units in the visual layout, such as being arranged vertically adjacent, horizontally adjacent, indented, or partially grouped. The terminal device can insert the second message into the adjacent position of the first message after detecting that a non-player character has generated a response to the first message, instead of simply arranging them indiscriminately according to the global timeline, thus highlighting the response correspondence between the two at the display level. Adjacent display can also include indented display, offset display, or following display of the second message. For example, the first message can be displayed in the regular team chat format, and the second message can be displayed below the first message with an indentation of a preset distance to indicate that the second message is a response to the first message; or, the second message can be displayed next to the first message to form a visual "question-answer pair" structure. For example, the terminal device can apply the same local grouping background, group bounding box, or group time tag to the first and second messages to further strengthen their correlation. Alternatively, the terminal device can indicate the reply relationship of the second message to the first message in the chat area through a connection marker. The connection marker can be a graphical identifier used to establish a visual connection between two messages, and its form can include at least one of the following: a connecting line, arrow, zigzag guide line, bracket marker, quotation bar, reply anchor, association icon, thread identifier, or highlighted indicator area.The connection marker can be displayed between the first and second messages, or at a preset position in the second message, pointing to the display area corresponding to the first message, to clearly indicate that the second message is a reply to the first message. For example, the terminal device can display a reference marker to the left of the second message, linking it to the first message; or, a short connecting line can be displayed between the first and second messages to indicate that they belong to the same round of interaction; or, a reply prompt message "Reply to message X" can be displayed at the top of the second message, associating this prompt message with the first message. For scenarios with a large number of messages in the chat area and a high message refresh frequency, using connection markers can avoid the identification difficulties caused by relying solely on time sequence to determine the reply object.

[0067] It should be noted that in practical applications, adjacent display and connection marker indication can be used individually or in combination. That is, the terminal device can display the second message adjacent to the first message, and can also further indicate the reply relationship between the two through connection markers. For scenarios with a small number of messages or sufficient interface space, adjacent display can be used first; for scenarios with a large number of messages, continuous scrolling of dialogue content, or multiple non-player characters replying concurrently, connection markers can be added on top of adjacent display to improve the accuracy of reply relationship identification.

[0068] As can be seen, by displaying the second message in a different display style than the first message, the embodiments of this application can clearly distinguish between player character messages and non-player character messages at the interface level, improve the efficiency of message source identification, and reduce user understanding costs. In addition, by displaying the second message and the first message adjacent to each other in the chat area, users can quickly determine that the two belong to the same round of interaction based on spatial proximity, thereby more intuitively perceiving the reply chain. Alternatively, by setting connection markers in the chat area to indicate the reply relationship between the second message and the first message, even when messages refresh quickly, there are many speakers, or the dialogue content is complex, the reply object can be accurately identified, reducing misreading and ambiguity.

[0069] In some embodiments, following the above examples, the above-mentioned display of the second message sent by the first non-player character in response to the first message can also be achieved in the following way: within a preset time period after receiving the first message, the second message sent by the first non-player character in response to the first message is displayed in a dynamically expanded manner, wherein the dynamically expanded manner may include at least one of word-by-word display, fade-in display, and bubble pop-up display.

[0070] For example, after receiving a first message from a player character in a first virtual scene, the terminal device can dynamically display a second message sent by a first non-player character as a reply to the first message within a preset duration. The preset duration can be the time interval from receiving the first message to the start of displaying the second message, or the time interval from receiving the first message to the completion of displaying the second message. The preset duration can be a fixed, pre-set duration or a dynamic duration adaptively determined based on message length, urgency, the personality traits of the first non-player character, the current game pace, scene event density, or the terminal display status. For example, the preset duration can be set relatively short for short and urgent replies; and relatively long for replies containing more text or requiring the character's "thinking process." The dynamic expansion method can include at least one of word-by-word display, fade-in display, and bubble pop-up display. The terminal device can display the second message in the chat area using a single dynamic expansion method, or it can combine two or more dynamic expansion methods to enhance the perceptibility and expressiveness of the reply message. For example, when using a character-by-character display method, the terminal device can display the text content of the second message character by character, phrase by phrase, or semantic fragment by semantic fragment in the message bubble according to a preset character output order. The character-by-character display speed can be fixed or dynamically adjusted based on the text length of the second message, the speaking style of the first non-player character, or the rhythm of the current scene. For example, if the first non-player character is a calm character, the character-by-character display speed can be relatively stable; if the first non-player character is an active character, the character-by-character display speed can be relatively fast. Furthermore, before or during the character-by-character display, the terminal device can first display an input prompt marker, a waiting marker, or an ellipsis animation to indicate that the first non-player character is generating or outputting a response. When using a gradual display method, the terminal device can gradually transition the message bubble, message text, or character identifier corresponding to the second message from a low transparency to a target transparency, achieving a visual presentation effect of the second message from weak to strong. The gradual display can apply only to the text content or simultaneously to the message bubble and the character identifier. For example, the terminal device can also control the second message to change brightness, border highlight, or local shadow during the gradual appearance process to improve visual recognition when the message appears. When using a bubble pop-up display method, the terminal device can control the message bubble carrying the second message to be displayed in the chat area in the form of a pop-up animation. For example, the message bubble can gradually enlarge from a shrunk state to the target size, pop up from near the character icon of the first non-player character to the target display position, or perform one or more slight bounces after reaching the target position to create a strong message arrival notification effect.The pop-up display method can be used alone to highlight the appearance of the second message, or it can be used in conjunction with word-by-word display or fade-in display.

[0071] It should be noted that the dynamic expansion method can also be associated with the character characteristics of the first non-player character. That is, different types of non-player characters can adopt different dynamic display strategies for the second message. For example, the second message of a reconnaissance-type non-player character can use a faster bubble pop-up followed by rapid word-by-word display to reflect quick reaction; the second message of a support-type non-player character can use a smooth fade-in method to reflect stability and reliability; characters with distinct personalities can also be configured with exclusive pop-up rhythms, text appearance speeds, or bubble animation styles to enhance the expressive differences between different non-player characters. In addition, the terminal device can control the second message to be displayed dynamically near the first message based on the reply relationship between the first and second messages. For example, the second message can pop up and appear adjacent to the first message, and then the text of the second message can be displayed word-by-word within the message bubble; or, the second message can first appear in the chat area as a bubble pop-up, and then the text content can be fully displayed through a fade-in method. In this way, not only can the appearance process of the second message be reflected, but the reply attribute of the second message to the first message can also be further strengthened.

[0072] As can be seen, by displaying the second message within a preset time after receiving the first message, the embodiments of this application can make the replies of non-player characters exhibit a reasonable response delay, avoiding the mechanical appearance of instantaneous replies and thus improving the naturalness of the interaction. Furthermore, by employing at least one dynamic expansion method among word-by-word display, fade-in display, and bubble pop-up display, the visual cue effect when the second message appears can be enhanced, making it easier for users to notice the reply content and improving the information perceptibility of the chat area.

[0073] In other embodiments, following the above examples, the above-mentioned display of the second message sent by the first non-player character in response to the first message can also be achieved in the following way: within a preset time period after receiving the first message, the second message sent by the first non-player character in response to the first message is displayed in the form of a reference reply, wherein the reference reply form may include displaying at least part of the content of the first message, and at least part of the content may be displayed above the second message or in the reference area inside the second message.

[0074] For example, after receiving a first message sent by a player character in a first virtual scene, the terminal device can display a second message sent by a first non-player character as a reply to the first message within a preset time period, in the form of a reference reply. The reference reply format can include displaying at least a portion of the content of the first message. For example, the at least a portion of the content can be the entire content of the first message, or it can be a portion of text extracted from the first message, keyword content, semantic summary content, intent fragment content, or core fragment content related to the generation result of the second message. When the length of the first message does not exceed a preset threshold, the terminal device can display the entire content of the first message in the reference area; when the length of the first message exceeds the preset threshold, the terminal device can display only the first N characters, keywords related to the reply topic, a simplified expression obtained through semantic compression, or display the truncated content of the first message using ellipses, in order to control the space occupied by the interface while ensuring that the reply relationship is recognizable. Furthermore, at least a portion of the content of the first message can be displayed above the second message. Specifically, the terminal device can configure a reference header area for the second message in the chat area and display at least a portion of the content of the first message within the reference header area. The header area of ​​the quote can adopt a different display style from the body area of ​​the second message, such as a lighter background color, thinner font, quote prefix symbol, vertical guide bar, border separator, or indentation style, so that players can distinguish between the "quoted content" and the "reply content". Based on this, the terminal device can display the body content of the second message below the header area, thus forming a display structure of "quoted content on top, reply content below". Alternatively, at least part of the first message can also be displayed in the quote area within the second message. The quote area can be a partial display area set within the message bubble of the second message, such as an embedded sub-area located at the top, left, right, or front of the body of the message bubble. The terminal device can display at least part of the first message in the quote area and display the second message in the body area within the same message bubble. By integrating the quoted content and reply content into the same message bubble, the information organization compactness can be improved in a limited chat area, making it easier for players to read quickly. Furthermore, the quote area can be distinguished in style through at least one of the following methods: background color distinction, border style distinction, font size distinction, transparency distinction, text color distinction, prefix icon distinction, quotation mark style distinction, or indentation distinction. For example, the terminal device can display a vertical quote line on the left side of the quote area, showing at least part of the first message in a smaller font size, while displaying the second message in a regular font size in the body text area; alternatively, a light-colored quote block can be set inside the second message bubble, displaying quoted content such as "Is anyone in the right warehouse?" within the quote block, and a reply message "I just checked, no one is here right now" can be displayed below the quote block. This can further improve the efficiency of players in recognizing quote relationships.

[0075] It should be noted that in practical applications, the terminal device can also determine the display method of at least part of the content of the first message based on the semantic correspondence between the first and second messages. For example, when the second message replies to the entire content of the first message, it can display the complete content of the first message or the main content at the beginning; when the second message replies only to a specific issue, request, or keyword in the first message, it can only display the corresponding issue fragment, request fragment, or keyword fragment in the reference area. For example, if the first message is "Who can check the high point on the left and report the situation downstairs?", and the second message only replies with "high point on the left", then the reference area can only display "Who can check the high point on the left?" to improve the matching degree between the referenced content and the reply content. In addition, the reference reply format can also be displayed in conjunction with the role identifier of the first non-player character. That is, when displaying the reference area and the body area of ​​the second message, the terminal device can also display the character avatar, character name, duty icon, or other role identifier of the first non-player character in the message bubble corresponding to the second message to represent the sender of the second message. Among them, the reference reply style of different non-player characters can be associated with their role characteristics. For example, reconnaissance non-player characters can use a simpler quote bar style, while support non-player characters can use a quote header style with a role label to enhance differentiation in multi-role chat scenarios.

[0076] As can be seen, the embodiments of this application, on the one hand, by displaying the second message in the form of a reference reply within a preset time after receiving the first message, can make the reply of the non-player character have both a reasonable time delay and a clear contextual reference, thereby improving the naturalness of the interaction; on the other hand, by displaying at least part of the content of the first message in the second message, the original message corresponding to the reply can be directly presented to the player, avoiding the ambiguity caused by relying solely on the time sequence to judge the reply object, and improving the accuracy of reply relationship identification.

[0077] In some embodiments, following the above example, the above-mentioned display of the second message sent by the first non-player character in response to the first message can also be achieved in the following way: within a preset time after receiving the first message, the second message sent by the first non-player character in response to the first message is displayed in an expression style that is adapted to the emotional state of the player controlling the player character, wherein different emotional states can correspond to different message display styles.

[0078] For example, after receiving a first message from a player character in a first virtual scene, the terminal device can, within a preset time period, display a second message from a first non-player character, used to reply to the first message, employing an expression style adapted to the emotional state of the player controlling the player character. The second message can be a reply generated based on the message content of the first message, the current scene state, the character characteristics of the first non-player character, and the player's emotional state. By adapting the expression style of the second message to the player's current emotional state, the relevance and naturalness of the non-player character's reply can be enhanced. The preset time period can be adaptively determined based on the urgency of the first message, the player's current emotional state, the response priority of the first non-player character, the message density in the chat area, or the intensity of events in the current virtual scene. For instance, when the player is in a tense, urgent, or anxious state, the preset time period can be relatively shorter to allow the second message to display faster; when the player is in a calm or non-urgent communication state, the preset time period can be a normal duration to maintain a natural interaction rhythm. Specifically, the terminal device can determine the emotional state of the player controlling the player character. The emotional state can include at least one of the following: calm, tense, anxious, urgent, excited, frustrated, angry, or other distinguishable emotional states. The determination of the emotional state can include at least one of the following: analysis based on the semantic content of the first message text; recognition based on the player's voice tone, speed, and volume; inference based on the player's frequency of actions, perspective switching, movement trajectory, attack rhythm, or interactive behavior; or comprehensive judgment based on contextual information such as combat intensity, resource status, hit status, and teammate status in the first virtual scene. For example, the terminal device can determine the player's emotional state based on the fusion of multiple information sources to improve the accuracy of emotion recognition results. Furthermore, the expression style adapted to the player's emotional state can include at least one of the following: text expression style and message display style. The text expression style can include at least one of the following: word choice style, sentence structure style, tone intensity, reply length, level of reassurance, level of encouragement, clarity of instructions, rhythm, or emotional tone; the message display style can include at least one of the following: message bubble color, border style, font style, font size, transparency, animation, emphasis markers, character tags, icon elements, or background decoration. In other words, the terminal device can not only adjust what the second message "says" and "says," but also how it "is displayed," so that the first non-player character's reply matches the player's emotional state in both content and presentation. Furthermore, different emotional states can correspond to different message display styles.Specifically, the terminal device can establish a preset mapping relationship between emotional states and message display styles. Based on the determined player's emotional state, it selects the target message display style corresponding to that emotional state from multiple candidate message display styles to display the second message. For example, when the player is calm, the second message can use message bubbles of regular colors, smooth transition display animations, and neutral font styles; when the player is tense or anxious, the second message can use highly visible borders, strong contrast colors, a faster appearance rhythm, or message styles with prompts to highlight the reply information; when the player is anxious or frustrated, the second message can use softer bubble colors, gentler display animations, and soothing visual elements to reduce the oppressive feeling of the interface; when the player is excited, the second message can use a lighter display rhythm or a more vibrant visual style to enhance the sense of communication and feedback.

[0079] It's worth noting that the style of expression can also be adjusted according to the player's emotional state. For example, when the player is tense or anxious, the terminal device can control the first non-player character to output shorter, clearer, and more command-like responses, such as "I'm going to the left," "I'll provide support immediately," or "Enemy ahead." When the player is anxious or frustrated, the terminal device can control the first non-player character to output responses with a reassuring, confirmatory, or encouraging tone, such as "Don't worry, I'll provide support," "I have medicine here, I'll give it to you right away," or "Stay calm, I'll keep an eye on the right side." When the player is calm, the second message can use more natural and conventional dialogue. When the player is excited, the second message can appropriately enhance interactivity and responsiveness. In this way, the responses of the first non-player character can better match the player's current psychological expectations. Furthermore, the message display style corresponding to different emotional states can also be associated with the character characteristics of the first non-player character. Even under the same player's emotional state, different types of non-player characters can use different adaptive expression methods. For example, when a player is anxious, the second message from a supportive non-player character can use gentle colors and reassuring wording, the second message from a reconnaissance non-player character can use a concise and efficient prompt style, and the second message from a commanding non-player character can use a more explicit guiding reply. This allows for the preservation of individual differences and responsibilities among different non-player characters while maintaining emotional adaptation. Simultaneously, the terminal device can dynamically adjust the expression style of the second message based on changes in the player's emotional state. For instance, after receiving the first message, if the terminal device detects a change in the player's emotional state from calm to urgent, it can switch to a higher-priority expression style and message display style before or after displaying the second message; or, in subsequent replies from other non-player characters in the chat area, the display strategy adapted to the current player's emotional state can continue to be used. In this way, the responses of non-player characters can more closely resemble the situational changes in real-world interactions.

[0080] As can be seen, by displaying the second message in a style that matches the emotional state of the player controlling the player character within a preset time after receiving the first message, the responses of non-player characters can better match the player's current psychological state and scene expectations, thereby improving the naturalness of the interaction.

[0081] In other embodiments, following the foregoing, the aforementioned emotional state may include at least one of an aggressive state, a conservative state, and an anxious state. This can be achieved by displaying a second message sent by a first non-player character in response to the first message, using an expression style adapted to the emotional state of the player controlling the player character: when the player controlling the player character is in an aggressive state, a second message sent by the first non-player character including a cooperative offensive response is displayed; when the player controlling the player character is in a conservative state, a second message sent by the first non-player character including a defensive suggestion response is displayed; and when the player controlling the player character is in an anxious state, a second message sent by the first non-player character including a reassuring response is displayed.

[0082] For example, the terminal device can pre-establish a mapping relationship between emotional states and response content types. For instance, when a player controlling a character is in an aggressive state, the mapping relationship can determine a coordinated offensive response corresponding to that aggressive state. This coordinated offensive response can express the first non-player character's response to actions such as proactive attack, synchronized suppression, coordinated assault, flanking maneuver, fire support, or rapid follow-up. The coordinated offensive response can include at least one of the following: attack confirmation, tactical coordination, cover support, or target suppression. For example, the first non-player character could send a second message such as "I'll push with you," "I'll smoke, you go first," "I'll flank from the side," "You take the lead, I'll finish you off," or "I'll scout first, you continue the advance." Furthermore, when the first non-player character's role type is assault, reconnaissance, or fire support, the terminal device can further adjust the coordinated offensive response based on their role characteristics, making the second message more consistent with the first non-player character's role. For example, a reconnaissance-type non-player character might send a message like "I'll scout the front, you follow," while a firepower-suppression-type non-player character might send a message like "I'll hold the window, you advance." When the player controlling the player character is in a conservative state, a defensive suggestion response can be determined based on a mapping relationship. This defensive suggestion response can express the first non-player character's advice on actions such as holding the position, reducing exposure, waiting for support, observing the enemy situation, controlling key points, or delaying the advance. The defensive suggestion response can include at least one of the following: point defense, risk avoidance, observation reminders, retreat suggestions, or regrouping. For example, the first non-player character could send a second message like "Don't push yet, hold the cover," "I'll watch the right side, you hold the door," "Wait for teammates to arrive before advancing," "Retreat to a safe point first, don't get surrounded," or "I'm watching the high ground, you watch the close range." This method allows the second message to reflect the conservative player's state at the content level, thereby reducing unnecessary high-risk actions. When a player controlling a player character is in an anxious state, a reassuring response can be determined based on a mapping relationship. This reassuring response can express reassurance, confirmation, support commitment, or clear guidance from a non-player character regarding the player's current state of tension, helplessness, or resource shortage. Reassuring responses can include at least one of the following: emotional reassurance, support confirmation, resource replenishment, evacuation guidance, or risk mitigation. For example, the non-player character could send a second message such as "Don't worry, I'm coming right away," "Stay calm, I'll cover you," "I have meds here, move closer to me," or "Retreat behind the wall, I'll cover you." This type of response can buffer and soothe the player's anxiety while conveying tactical information, thus improving the user-friendliness of the interaction.

[0083] It should be noted that in practical applications, the aforementioned collaborative offensive, defensive suggestion, and reassuring responses can differ not only in text content but also in message display style. That is, after generating the response text for the second message, the terminal device can configure a corresponding display style based on the determined emotional state. For example, for a collaborative offensive response corresponding to an aggressive state, the second message can use a faster-paced, higher-contrast display style to enhance the sense of responsiveness; for a defensive suggestion response corresponding to a conservative state, the second message can use a relatively stable, clear, and low-interference display style to reflect stability; and for a reassuring response corresponding to an anxious state, the second message can use softer colors, gentler animations, or less stimulating visual elements to enhance the reassuring effect. Thus, the second message can be adapted to the player's emotional state at both the "response content" and "presentation format" levels. Furthermore, the terminal device can further refine the generation of the second message by combining the specific semantics of the first message and the current context. In other words, even if players are in the same emotional state, different first message content can trigger different second message expressions.

[0084] In some embodiments, step 103 described above can also be implemented as follows: In the first virtual scene, a voice content corresponding to the second message used to reply to the first message is played using a voice timbre pre-configured for the first non-player character, wherein the voice timbre configured for different non-player characters is different. That is, after receiving the first message sent by the player character, after the terminal device determines that the first non-player character is the target non-player character to reply to the first message and generates the second message, it can not only display the text content of the second message, but also further output the voice content corresponding to the second message based on the preset voice timbre corresponding to the first non-player character, so as to form a voice reply matching the identity of the first non-player character.

[0085] For example, the pre-configured voice for the first non-player character can be understood as a set of voice features associated with that character. This set of voice features can include at least one of the following: fundamental frequency features, formant features, speech rate features, intonation features, volume features, pause features, accent features, breathy voice features, age-related features, personality-specific tone features, or voice identification parameters. The terminal device can pre-create voice configuration files for different non-player characters, or create basic voice templates for different character types, and then overlay character personality parameters onto these templates to obtain the target voice for the corresponding non-player character. Therefore, even if different non-player characters reply with similar message content, players can quickly identify the replying subject through differences in voice voice. Furthermore, the different voice configurations for different non-player characters can be based on at least one of the following attributes of the non-player character: character identity, profession type, role division, faction attribute, age setting, personality setting, appearance setting, or historical storyline setting. For example, a reconnaissance-type non-player character can be configured with a fast-paced, soft-spoken, and concise voice; a support-type non-player character can be configured with a steady and friendly voice; a command-type non-player character can be configured with a highly distinctive, clear, and directive voice; and a defensive-type non-player character can be configured with a calm and composed voice. This method ensures that the voice configuration aligns with the functional role of the non-player character. Furthermore, after generating the text content of the second message, the terminal device can use the second message as input for voice generation and combine it with the preset voice corresponding to the first non-player character to generate the corresponding voice content for the second message. The voice content generation method can include at least one of the following: speech synthesis based on a text-to-speech (TTS) model, splicing pre-recorded speech segments, waveform generation based on a parametric speech engine, or joint generation of the target speech based on a speech model and a voice conversion model. For example, the terminal device can directly call the voice resource package corresponding to the first non-player character to generate the voice output corresponding to the text content of the second message. Furthermore, the playback of the voice content corresponding to the second message used to reply to the first message can occur before, simultaneously with, or after the display of the second message text. The terminal device can determine the timing of voice playback based on the interaction rhythm, combat intensity, message priority, or terminal audio playback status within the current first virtual scene. For example, for second messages involving enemy situation reports, immediate support, or tactical coordination, output can be prioritized using synchronization with text or a voice-first-text approach to improve information transmission speed; for general communication-type second messages, text can be displayed first, followed by voice playback, to maintain the natural rhythm of scene communication. Simultaneously, the playback method of the voice content can also be correlated with the positional relationship of the first non-player character within the first virtual scene.If the first non-player character is near the player character, the terminal device can play the voice content corresponding to the second message using in-scene spatial audio, matching the playback direction, distance attenuation effect, or left and right channel distribution to the first non-player character's location within the scene. If the first non-player character is not near the player character, or if the second message is a team communication message, the terminal device can output the voice content using a team voice channel playback method. This ensures that the voice response format is consistent with the interactive context in the first virtual scene. Furthermore, the terminal device can optimize the spoken expression of the second message based on the first non-player character's voice timbre configuration. That is, different first non-player characters can present different language styles and voice performances when outputting the same semantic response. By combining voice timbre with spoken expression, the individual differences between different non-player characters can be further enhanced.

[0086] In other embodiments, following the examples above, the playback of the voice content corresponding to the second message used to reply to the first message can be achieved in the following way: within a preset time period after receiving the first message, the voice content corresponding to the second message used to reply to the first message is played, wherein the voice content can be played at least one of the following: speech rate, tone, and pauses adapted to the emotional state of the player controlling the player character. In this way, the voice reply of the first non-player character to the first message can not only reflect the semantics of the reply, but also match the player's current state at the auditory expression level.

[0087] For example, the terminal device can establish a mapping relationship between the player's emotional state and voice playback parameters. These voice playback parameters can include at least one of speech rate, tone, and pause patterns. Speech rate can characterize the speed at which voice content is played per unit of time; tone can characterize the pitch variation trend, amplitude, stress distribution, or intensity of the voice content; pause patterns can characterize the pause location, duration, frequency, or rhythm of sentence transitions. After determining the player's emotional state, the terminal device can obtain the target voice playback parameters corresponding to that emotional state from a preset mapping relationship and play the voice content corresponding to the second message based on these parameters. For example, when the player controlling the character is in an aggressive state, the terminal device can play the voice content corresponding to the second message expressed in a coordinated offensive style. In this case, the voice content can be played with a relatively fast speech rate, a clear stress distribution, and short pause patterns to enhance the sense of responsiveness and coordination. If the second message involves immediate coordinated actions such as flanking, follow-up shots, or suppression, the terminal device can further shorten the pause duration to improve the efficiency of tactical information transmission. When the player controlling the character is in a conservative stance, the terminal device can play a second message with voice prompts in a defensive style. In this case, the voice prompts can be played at a relatively steady pace, with gentle intonation and clear pauses to enhance a sense of stability and comprehensibility. When the player controlling the character is in an anxious state, the terminal device can play a second message with voice prompts in a reassuring style. In this case, the voice prompts can be played in a relatively soft tone, at a moderate pace, and with reassuring pauses to reduce tension and enhance a sense of support.

[0088] It should be noted that the audio content corresponding to the second message can be generated through text-to-speech or by splicing pre-recorded audio segments and adjusting prosodic parameters. If text-to-speech is used, the terminal device can input the text content of the second message, along with target speech rate, target intonation, and target pause parameters corresponding to the player's emotional state, into the speech synthesis model to generate the adapted target audio content. If pre-recorded audio segment splicing is used, the terminal device can extract audio segments matching the second message from the audio resources associated with the first non-player character, and dynamically adjust the playback level of the audio segments through methods such as duration stretching, pitch adjustment, pause insertion, or pause compression. Alternatively, the terminal device can also perform emotional adaptation on the audio content corresponding to the second message while maintaining the basic timbre of the first non-player character. That is, different non-player characters can still retain their pre-configured character timbre, while the speech rate, intonation, and pause patterns are dynamically adjusted according to the player's emotional state. This allows for both the recognizability of the non-player character's identity and the emotional adaptability of the response expression. For example, when saying "I'm coming right away," a support-type non-player character in an anxious state can deliver the message with a gentle, slightly slower start and clear pauses, while an assault-type non-player character in an aggressive state can deliver it with a faster pace and a more compact rhythm. Furthermore, the terminal device can further adjust the voice playback parameters based on the content of the second message. For instance, when the player is anxious, if the second message is an enemy alert, the terminal device can appropriately increase the emphasis on keywords while maintaining a soothing tone to balance soothing and warning effects; when the player is aggressive, if the second message is a confirmation of coordinated attack, the speech speed can be increased and pauses reduced; when the player is conservative, if the second message is a location alert, sentence pauses and keyword clarity can be emphasized. This approach avoids using a completely fixed voice template for the same emotional state, improving the relevance of voice responses.

[0089] In some embodiments, following the above example, when playing the voice content corresponding to the second message used to reply to the first message, the following processing can also be performed: displaying a speech prompt message corresponding to the first non-player character in the first virtual scene, wherein the speech prompt message may include at least one of a speaking icon, a voice ripple icon, and a text prompt. In this way, in addition to receiving the reply content of the first non-player character through hearing, the player can also quickly identify the current speaker and its speaking status through visual cues.

[0090] For example, a voice prompt can be understood as a scene prompt element indicating that the first non-player character is currently outputting a voice reply. The voice prompt can indicate at least one of the following: the current speaking character's identity, the start of a voice message, the duration of a voice message, the end of a voice message, changes in voice volume, or a summary of the reply content. The terminal device can trigger the display of the voice prompt when it detects that the voice content corresponding to the second message has started playing, and stop displaying or switch the voice prompt when the voice playback ends. The speaking icon can be a voice status indicator displayed near the first non-player character. For example, the speaking icon can be set in at least one of the following areas: above the first non-player character's head, in their portrait area, in their name identifier area, in the team list area, or in the edge indicator area of ​​the screen. For example, when the first non-player character is within the player's current field of view, the terminal device can display the speaking icon above their name; when the first non-player character is not within the player's field of view or is far away, the terminal device can highlight the corresponding character's portrait in the team information bar and overlay the speaking icon. The speaking icon can be a microphone style, a speech bubble style, a megaphone style, or other preset speaking indicator styles. Voice ripple icons can be used to represent the dynamic state of the current voice output of the first non-player character. Voice ripple icons can include dynamic ripple effects spreading around the first non-player character's portrait, head icon, character model, or speaking icon. The dynamic ripple effect can change synchronously according to the volume, speech rate, stress distribution, or pause rhythm of the voice content corresponding to the second message. For example, when the volume of the voice content is relatively high or the tone is strong, the amplitude of the voice ripple icon can increase relatively; when there is a pause in the voice content, the voice ripple icon can decrease synchronously or briefly stop spreading. This enhances the consistency between visual cues and actual voice playback. Text prompts can be used to provide players with text information or summary information corresponding to the second message. Text prompts can be the complete text of the second message, or keywords, phrase prompts, or intent summaries extracted from the second message. For example, when the second message is "Don't worry, I'm coming right away, you retreat behind cover first," the terminal device can display the complete text, or only display a summary text prompt such as "Support immediately" or "Retreat first." Furthermore, text prompts can be displayed in the chat area, subtitle area, above the first non-player character's head, team information panel, or side prompt area of ​​the screen.

[0091] It should be noted that the display format of the voice prompts can be adjusted based on the positional relationship between the first non-player character and the player character. When the first non-player character is near the player character and within the current field of view, the terminal device can prioritize displaying a speaking icon and voice ripple indicator above the first non-player character's head or near the character model to enhance the sense of immersion in the scene. When the first non-player character is not within the current field of view, the terminal device can display voice prompts pointing to the location of the first non-player character at the edge of the screen, or highlight the corresponding character in the team list to help players quickly locate the source of the voice message. When the second message is a reply for long-distance team communication, the terminal device can also display text prompts in the chat area or subtitle area to ensure that players accurately understand the reply content. In addition, the voice prompts can also be adapted to the emotional state of the player controlling the player character. That is, when playing the voice content corresponding to the second message, in addition to the speech speed, tone, and pauses being adapted to the player's emotional state, the display style of the voice prompts can also be dynamically adjusted according to the player's emotional state. For example, when a player is aggressive, the speaking icon corresponding to a coordinated offensive response can flash faster and provide more direct feedback, while the voice ripple icon can have a stronger diffusion effect to enhance the sense of coordinated action. When a player is conservative, the speaking prompts can use a more stable and clear display style to reduce visual interference. When a player is anxious, the speaking prompts can use softer colors, a gentler dynamic rhythm, and clearer text prompts to enhance the calming effect and comprehensibility. Furthermore, the display duration of the speaking prompts can be correlated with the playback duration of the corresponding voice content in the second message. Specifically, the terminal device can display the speaking prompts when the voice playback begins and remove them when the voice playback ends; or, it can continue to display the text prompts for a preset duration after the voice playback ends, allowing the player to review them later. For example, for short confirmation-type second messages, the speaking prompts can disappear quickly after playback; for second messages containing location reminders, tactical suggestions, or support guidance, the text prompts can remain for a short duration after the voice playback ends to improve information retention.

[0092] In other embodiments, the second message may include a dialogue response message and a confirmation response message. Step 103 can then be implemented as follows: when the first message is a casual chat message, the first non-player character is controlled to output a dialogue response message semantically related to the casual chat message; when the first message is a command message, the first non-player character is controlled to output a confirmation response message corresponding to the command message, and based on the command message, the first non-player character is controlled to execute corresponding cooperative actions. These cooperative actions may include at least one of target switching, path movement, attack cooperation, defense cooperation, healing support, and retreat avoidance. This approach allows the first non-player character to adopt different response strategies for different types of player input, thus balancing a natural dialogue experience with tactical cooperative execution effectiveness.

[0093] For example, casual messages can refer to communication messages that are not directly used to drive the first non-player character to perform tactical actions. Casual messages can include at least one of the following types: greeting messages, emotional expression messages, scene commentary messages, relationship interaction messages, status inquiry messages, or mild suggestion messages. For example, casual messages could be "You reacted quickly just now," "This side doesn't look very safe," "Are you still keeping up?" or "Let's play it safe this round." For this type of message, the terminal device can select a dialogue response message that is semantically related to it from a preset response corpus based on semantic understanding results, or call a dialogue generation model to generate response content that conforms to the identity settings of the first non-player character. Thus, the first non-player character can form a natural communication with contextual relevance around the player's current input. Dialogue response messages that are semantically related to casual messages can be understood as response content that has a semantic response relationship with the player's input but does not require triggering explicit action execution. Dialogue response messages can reflect agreement, supplementation, reassurance, evaluation, inquiry, or extension of communication. For example, when the first message is "The pressure here is a bit high," the first non-player character can output "Indeed, the enemy's firepower is quite concentrated"; when the first message is "You did a good job saving me just now," the first non-player character can output "You reported your position very promptly"; when the first message is "I feel there might be someone on the right," the first non-player character can output "I had that feeling too, I'll take a closer look at the right side later." Furthermore, the terminal device can also stylize the dialogue response messages based on the first non-player character's character attributes, historical interaction records, or current scene state. Command messages can refer to messages used to instruct the first non-player character to perform targeted cooperative actions. Command messages can include at least one of the following types: target switching commands, path movement commands, attack cooperation commands, defense cooperation commands, healing support commands, or retreat / evasion commands. For example, command messages could be "Attack the enemy on the left first," "Go around to the right," "Help me push the front," "Hold this doorway," "Heal me first," "Retreat to cover," etc. After identifying the first message as an instruction message, the terminal device can further parse its instruction intent, target object, target location, priority, or timing relationship to generate a corresponding confirmation reply message for the first non-player character and trigger corresponding collaborative behavior. The confirmation reply message can be a response indicating that the first non-player character has received and understood the instruction message. The confirmation reply message can include at least one semantic element among confirming execution, confirming understanding, confirming support, or confirming adjustment. For example, for the instruction message "Follow me to push the right flank," the first non-player character can output "Received, I'll push the right flank with you"; for the instruction message "Don't rush in yet, hold the entrance," the first non-player character can output "Understood, I'll hold the entrance first"; for the instruction message "Save me first," the first non-player character can output "Received, I'll come to revive you right away."By first outputting a confirmation message, the player is clearly informed that the first non-player character has received and responded to the command message, thereby improving the certainty of feedback during human-computer collaboration. Furthermore, the terminal device can establish a mapping relationship between command messages and collaborative behaviors. After determining that the first message is a command message, the terminal device can control the first non-player character to execute the corresponding collaborative behavior according to the mapping relationship. Collaborative behaviors can include at least one of target switching, path movement, attack collaboration, defense collaboration, healing support, and retreat / evasion. For example, when the command message is a target switching command, the terminal device can control the first non-player character to execute a target switching collaborative behavior. Target switching collaborative behaviors can include switching the current target of attention from the first target to the second target and updating the attack priority, aiming direction, target acquisition range, or skill release target. When the command message is a path movement command, the terminal device can control the first non-player character to execute a path movement collaborative behavior. Path movement collaborative behaviors can include moving to a target path point, following the player character, flanking to a designated flank, moving to a designated cover position, or occupying a designated high ground position. When the command message is an attack collaboration command, the terminal device can control the first non-player character to execute an attack collaborative behavior. Attack coordination actions can include synchronized firing, crossfire suppression, follow-up fire, skill-linked attacks, or cover advances. This improves the team's overall output rhythm and tactical coordination efficiency. When the command message is a defensive coordination command, the terminal device can control the first non-player character to perform defensive coordination actions. Defensive coordination actions can include guarding a designated direction, holding a designated area, blocking entrances, constructing crossfire defenses, or providing early warning of specific passages. When the command message is a healing support command, the terminal device can control the first non-player character to perform healing support coordination actions. Healing support coordination actions can include approaching a player character, using healing skills, providing recovery items, performing revival, or establishing temporary protection. When the command message is a retreat / evasion command, the terminal device can control the first non-player character to perform retreat / evasion coordination actions. Retreat / evasion coordination actions can include retreating, moving to cover, avoiding enemy fire lines, ceasing advances, or regrouping.

[0094] In some embodiments, before performing step 103 described above, the following processing may also be performed: controlling the first non-player character to enter an output prompt state, wherein the output prompt state may include at least one of a replying prompt, a thinking prompt, and a speaking prompt; in response to the completion of the generation of a second message for replying to the first message, proceeding to step 103. In this way, the player can be continuously fed back to the response progress of the first non-player character before the second message is officially output, thereby avoiding the player's perception of response interruption or lack of system feedback while waiting for a reply.

[0095] For example, the output prompt status can be understood as a transitional status prompt during the process of the first non-player character generating and preparing to output the second message. The output prompt status can at least represent one of the following information: the first non-player character has received the first message, the first non-player character is processing the first message, the first non-player character is generating reply content, the first non-player character is about to output the second message, or the first non-player character has entered the voice playback preparation stage or the text display preparation stage. The terminal device can control the first non-player character to display different output prompt statuses at different stages according to the message processing progress. Among them, the "Replying" prompt can be used to indicate that the first non-player character has responded to the first message and entered the reply process. After receiving the first message and confirming the first non-player character, the terminal device can prioritize controlling the first non-player character to enter the "Replying" prompt state. For example, the terminal device can display "Replying," "Replying," or other prompts indicating that a response has begun in the area above the first non-player character's head, avatar area, team list area, subtitle area, or chat area; it can also display dynamic ellipses, flashing chat bubble icons, pre-lit talk icons, or lightweight status effects. In this way, players can perceive that the first non-player character has "received and begun processing" the first message even before the second message has been fully generated. The "Thinking" prompt can indicate that the first non-player character is semantically understanding, identifying intent, generating response content, or making collaborative behavioral decisions regarding the first message. The "Thinking" prompt can be triggered when the first message is semantically complex, has strong contextual dependencies, the current virtual scene state changes rapidly, or multiple candidate responses need to be combined for filtering. For example, when the first message is a casual chat message containing complex intents, or a complex instruction message containing a target object, execution path, and priority, the terminal device can control the first non-player character to enter the "Thinking" prompt state. The "Thinking" prompt can be displayed using a rotating icon next to the avatar, an ellipsis bubble above the head, weak dynamic voice ripples, a slightly highlighted outline, a "Thinking" text prompt in the subtitle area, or a gradual breathing effect on the corresponding character card. This allows players to more intuitively understand that the first non-player character is not unresponsive, but rather generating a response or making a strategy judgment. The "Speaking" prompt can indicate that the first non-player character has completed the generation of the second message and entered the output preparation stage or begun the output stage. The "Speaking" notification can be triggered during a brief transition period after the second message is generated and before it is actually displayed or played, or it can be triggered synchronously with the output action when the second message begins to be output. For example, the terminal device can display a speaking icon above the head of the first non-player character, display a voice ripple icon around the avatar, display the character name in the subtitle area with a highlighted status, or reserve a corresponding message space in the chat area and display a "Speaking Soon" dynamic effect.This method allows for a smooth transition between the completion status of the second message generation and subsequent output status.

[0096] It should be noted that the terminal device can control the switching of output prompt states according to a preset state transition logic. Specifically, after receiving the first message, the terminal device can first control the first non-player character to enter the "Replying" prompt state; if the second message requires further semantic reasoning, context completion, or action decision-making, it will switch to the "Thinking" prompt state; when it detects that the second message used to reply to the first message has been generated, it will switch from the "Replying" and / or "Thinking" prompt states to the "Speaking" prompt state, and execute the step of controlling the first non-player character to output the second message. Alternatively, if the second message is generated quickly, the terminal device can directly switch from the "Replying" prompt state to the "Speaking" prompt state, or only display one of the output prompt states to reduce unnecessary interface jumps. Furthermore, the step of responding to the completion of the second message generated in response to the first message, and then proceeding to control the first non-player character to output the second message, may include at least one of the following processes: controlling the first non-player character to display the text content corresponding to the second message; controlling the first non-player character to play the voice content corresponding to the second message; controlling the first non-player character to synchronously execute character actions related to the semantics of the second message; controlling the first non-player character to display speaking prompts related to speaking; and controlling the first non-player character to execute collaborative behaviors corresponding to the command message. That is, the output prompt status can not only serve as transitional feedback during the message generation stage but also as a state transition node before the formal output of the second message. Additionally, the display style of the output prompt status can be associated with the message type of the first message. When the first message is a casual chat message, the terminal device can adjust the output prompt status of the first non-player character to reflect a natural conversation style, such as using a softer dynamic ellipsis, a slight avatar highlight, or a weaker thinking prompt effect to reflect the naturalness of the dialogue. When the first message is a command message, the terminal device can adjust the output prompt status of the first non-player character to reflect a tactical response style, such as using a clearer "received" prompt, a faster-paced replying status, or a more prominent speaking pre-prompt to reflect quick confirmation of the command and its imminent execution. Simultaneously, the display duration of the output prompt status can adaptively adjust based on the generation time of the second message. If the second message takes a short time to generate, the terminal device can only display a short "replying" prompt, then quickly switch to a "speaking" prompt; if the second message takes a long time to generate, the terminal device can extend the display duration of the "thinking" prompt and, upon reaching a preset duration threshold, periodically change the prompt style, such as switching from "replying" to "thinking," or enhancing the dynamic effects to continuously provide feedback to the player on the processing progress. This method avoids inconsistencies between the prompt status and the actual processing progress caused by differences in waiting time.

[0097] In other embodiments, when performing step 103 above, at least one of the following processes may also be performed: controlling the first non-player character to perform a speaking action adapted to the second message, wherein the speaking action may include at least one of a nodding action, a waving action, and a body orientation adjustment action; controlling the first non-player character to display a speaking expression adapted to the second message, wherein the speaking expression may include at least one of a smiling expression, a serious expression, and a prompting expression; highlighting the first non-player character, wherein the highlighting may include at least one of outline highlighting, partial illumination display, and character identifier highlighting. In this way, when the first non-player character outputs the second message, it can not only convey the reply content through text or voice, but also represent the reply intention through actions, expressions, and visual reinforcement prompts, thereby improving the perceptibility of the reply process and the scene expressiveness.

[0098] For example, adaptation to the second message can be understood as the terminal device determining the appropriate speaking action, facial expression, and highlighting method based on at least one of the following: message type, semantic content, emotional tone, urgency, tactical attributes, reply recipient, and the current state of the first virtual scene. For instance, when the second message is a conversational reply in a casual chat scenario, it can be matched with natural communicative actions and a softer facial expression; when the second message is a confirmation reply in a command scenario, it can be matched with a more explicit confirmation action, a clearer prompting facial expression, and a more prominent highlighting method. Speaking actions can be triggered when the second message begins to be output, or pre-triggered within a preset time before the second message is output, and are synchronized with the corresponding voice or text content of the second message. For example, when the second message is a confirmation reply, the terminal device can control the first non-player character to perform a nodding action. The nodding action can be used to represent the first non-player character's confirmation, understanding, or acceptance of the command message sent by the player character. For example, when the first message is "Guard the right entrance," the first non-player character can simultaneously perform a nodding action when outputting the second message "Understood, I guard the right entrance." The nodding gesture further visualizes the "received and executed" feedback, allowing players to quickly confirm the response status of the first non-player character. When the second message is a greeting, acknowledgment, or reassuring reply, the terminal device can control the first non-player character to perform a waving gesture. This waving gesture can be a full wave or a simplified action like a slight raise or a brief signal. For example, when the first message is "Are you alright?", the first non-player character can simultaneously perform a slight wave while replying "No problem, keep going"; similarly, when the first message is "You played well just now," the first non-player character can simultaneously perform an interactive wave or signal while replying "You cooperated very well." This enhances the natural feel of communication in casual conversation scenarios. When the second message involves target hints, directional alerts, coordinated attacks, defensive assignments, or retreat / evasion, the terminal device can control the first non-player character to adjust their body orientation. This can include facing the player character, facing a target enemy, facing a designated waypoint, facing a warning direction, or facing cover. For example, when the second message is "Someone is approaching from the right," the first non-player character can adjust their body orientation to the right of the warning direction while simultaneously sending the second message; when the second message is "I'm going to surround from the left," the first non-player character can first move to the left to adjust their orientation. This method establishes a clearer correspondence between the spoken content and the character's posture. Expressions can be implemented through switching facial expression resources, changing eyebrow and eye states, controlling lip movements, or using partial facial animation, and can be synchronized with the tone of voice and speaking actions corresponding to the second message.For example, when the second message is a positive response, friendly interaction, lighthearted chat, or encouraging reassurance, the terminal device can control the first non-player character to display a smiling emoticon. For instance, when the second message is "Don't worry, we can handle it" or "That was good teamwork," the first non-player character can use a smiling emoticon to enhance the positive interaction and increase the character's approachability. When the second message is a tactical confirmation, enemy situation alert, risk warning, retreat order response, or other tense response, the terminal device can control the first non-player character to display a serious emoticon. For instance, when the second message is "Retreat behind cover" or "The firepower ahead is too strong, don't push any further," the first non-player character can use a serious emoticon. A serious emoticon strengthens the warning and execution guidance of the second message, helping players quickly identify high-priority information. When the second message contains directional hints, target hints, coordination suggestions, or key reminders, the terminal device can control the first non-player character to display a prompt emoticon. The prompt emoticon can be expressed as raised eyebrows, focused gaze, slightly open mouth, or other facial expressions used to indicate reminders or attention. For example, when the second message is "There is an enemy on the high platform" or "Watch the left passage first," the first non-player character can display a prompt emoticon. This makes it easier for players to visually recognize such second messages as key information with a reminder nature. In addition, the above-mentioned highlight display can be triggered when the second message begins to be output, or it can remain displayed for a preset period of time before and after the second message is output, and then gradually disappear after the second message is output. Among them, outline highlighting can be used to improve the overall recognizability of the first non-player character in the first virtual scene. The terminal device can overlay outline lines, highlight strokes, or edge lighting effects around the first non-player character model to represent the current speaker. For example, when multiple non-player characters are moving or fighting simultaneously in the scene, outline highlighting can help players quickly identify the first non-player character currently outputting the second message. Local illumination display can be used to enhance local areas related to the speech. Local areas can include the head area, mouth area, chest communication device area, gesture area, or weapon mounting area. For example, when the first non-player character outputs a second message, a short-term glowing effect can be overlaid near their head or avatar area to highlight their speaking status; when performing a waving gesture, the hand area can also be locally illuminated to enhance motion recognition. Character identifier highlighting can be used to reinforce the name, avatar, team, or screen edge indicators corresponding to the first non-player character. For example, when the first non-player character is not in the player's current field of view, the terminal device can highlight the corresponding character's avatar in the team list or highlight the directional indicator at the edge of the screen, so that the player can clearly identify the subject of the current second message output even if they do not directly see the first non-player character.

[0099] It should be noted that in practical applications, the terminal device can control the combination of the above processing methods according to the importance of the second message. For example, for a casual chat-type second message, the first non-player character can be controlled to only display a smiling expression or perform a slight waving gesture; for a normal confirmation-type second message, the first non-player character can be controlled to perform a nodding gesture and have their character icon highlighted; for a high-priority tactical reminder-type second message, the first non-player character can be controlled to simultaneously adjust their body orientation, display a serious expression, and have their outline highlighted or partially illuminated. Thus, differentiated feedback can be achieved according to different interaction scenarios. Furthermore, the terminal device can synchronize the speaking actions, speaking expressions, and highlighting with the output rhythm of the second message. For example, outline highlighting can be triggered when the second message begins playing voice content, prompt expression enhancement can be triggered when keywords are output, and the highlighting effect can be turned off after a preset time delay after the voice ends; or, a nodding gesture can be performed at the beginning of a confirmation reply message, and the highlighting can be gradually canceled as subsequent collaborative actions begin. This method allows for a more natural temporal correspondence between visual feedback and the actual message output process.

[0100] In some embodiments, when performing step 103 above, the following processing can also be performed: in response to a combat event occurring in the first virtual scene, the first non-player character is controlled to pause outputting a second message, wherein the second message may be a casual chat message; in response to the end of the combat event, the first non-player character is controlled to continue outputting the second message. In this way, the first non-player character can dynamically switch between non-combat communication and combat behavior, thereby prioritizing real-time responsiveness in combat scenarios while also ensuring the continuity of casual chat replies.

[0101] For example, a combat event can be an event that signifies that the first virtual scene has entered a combat interaction state. A combat event can include at least one of the following: an enemy target enters the warning range, the player character or the first non-player character is attacked, enemy firing is detected, skill release is detected, damage calculation is detected, a focus fire warning is detected, a tactical warning is detected, or the first non-player character enters an attack state, evasion state, or support state. The terminal device can identify combat events based on scene state data, character behavior state data, enemy-ally distance data, damage event data, or combat marker data. The second message is a casual message, which can be understood as a response to the casual first message and does not directly undertake the function of confirming high-priority tactical commands. For example, the second message can be an evaluation reply, a greeting reply, a reassuring reply, a supplementary explanation reply, or a mild judgment reply. Since this type of second message usually does not have strong time-sensitive tactical execution constraints, when a combat event occurs, the terminal device can prioritize interrupting the output process of this type of second message so that the first non-player character can promptly transition to combat-related behaviors. Pausing the output of the second message can include at least one of the following: pausing the voice playback corresponding to the second message, pausing the word-by-word display of the text corresponding to the second message, pausing the scrolling display of the subtitles corresponding to the second message, pausing the lip movements associated with the second message, pausing the speaking actions associated with the second message, pausing the speaking emoticons associated with the second message, or ending the highlighting associated with the second message. Furthermore, the terminal device can record the current output progress of the second message so that output can resume from the interrupted position after the combat event ends. For example, after detecting a combat event, the terminal device can first determine the type of the currently output second message. If the second message is a casual message, the first non-player character is controlled to pause the output of the second message and switch to the combat response flow; if the second message is a confirmation reply message, a warning reply message, or other high-priority tactical message, the current output of the second message can be maintained, or the output can be quickly completed using a compressed output method. This method can distinguish the priority of different types of second messages, avoiding unnecessary interruptions to key tactical feedback. Simultaneously, switching to the combat response flow can include controlling the first non-player character to perform at least one of the following: attack coordination, defense coordination, healing support, retreat evasion, target switching, or path movement. For example, when the first non-player character is outputting the casual message "It looks relatively safe here for now," if the terminal device detects an enemy target suddenly entering its attack range, it can immediately pause the voice and text output of the second message and control the first non-player character to turn towards the enemy target, enter an alert or attack state. This improves the plausibility of the first non-player character's behavior in sudden combat situations. Simultaneously, to ensure continuity after interruption, the terminal device can cache the output context related to the second message when it pauses outputting the second message.The output context can include the second message text content, output segments, unoutput segments, voice playback progress, subtitle display progress, corresponding speaking action status, corresponding speaking expression status, and semantic tags associated with the generation of the second message. In response to the end of the battle event, the terminal device can control the first non-player character to continue outputting the second message based on the output context. Controlling the first non-player character to continue outputting the second message in response to the end of the battle event can include at least one of the following methods: continuing to output the remaining content from the point where the second message was paused, re-outputting the unfinished portion of the second message, simplifying the second message before outputting it, or supplementing the second message with the current scene state after the battle ends before outputting it. For example, if the battle event lasts for a short time, the terminal device can prioritize using the breakpoint resume method to continue outputting; if the battle event lasts for a long time, the terminal device can reorganize the second message according to the current context before outputting it to avoid the resumed content being out of sync with the current scene. Furthermore, the terminal device can determine that the battle event has ended based on preset conditions. The preset conditions can include at least one of the following: no new damage events are detected within a preset time period, no enemy attack behavior is detected within a preset time period, the first non-player character exits combat, the player character exits combat, the enemy target leaves the alert range, the current combat target is defeated, or the combat marker in the first virtual scene is cleared. After the preset conditions are met, the terminal device can trigger the processing to continue outputting the second message. The continued output after the combat event ends can also be adapted to the current state of the first non-player character. For example, if the first non-player character is still moving after the combat ends, the remaining voice content can continue playing during movement; if the first non-player character remains outside the player's field of vision after the combat ends, the remaining second message can continue to be presented through the team list avatar, subtitle area, or chat area; if the first non-player character's position changes significantly after the combat ends, a transitional short sentence, such as "I was just interrupted," can be output first, followed by the remaining content of the original second message, thereby enhancing the naturalness of the dialogue transition.

[0102] It's worth noting that the terminal device can also assess the timeliness of the second message before resuming output. If the casual conversational semantics of the second message still have communicative value after the battle, the first non-player character can continue outputting the second message. If the second message is clearly mismatched with the current scene state, a more coherent recovery reply can be generated based on the original second message. For example, if the original second message was "It's quite quiet here for now," but the scene situation has clearly changed after the battle, the terminal device can adjust the output to "I was interrupted just now, but it's safe here for now." This improves the consistency between the recovery reply and the current scene. Furthermore, the handling of pausing and resuming output can be linked to the output prompt status. When a battle event is detected, the terminal device can cancel the "speaking" prompt for the first non-player character and switch to a battle-related prompt status. After the battle event ends, the first non-player character is brought back into the "speaking" prompt status and the second message can continue to be output. This allows players to clearly perceive that the second message was not abnormally lost, but rather temporarily interrupted due to the battle event.

[0103] In other embodiments, following the examples above, when controlling the first non-player character to pause outputting the second message, the following processing can also be performed: displaying a pending reply prompt in the chat area of ​​the first virtual scene, wherein the pending reply prompt may include at least one of a suspension indicator, a combat interruption indicator, and a reply hold indicator; when controlling the first non-player character to continue outputting the second message, the following processing can also be performed: displaying a topic rewind prompt in the chat area of ​​the first virtual scene, wherein the topic rewind prompt may include at least a portion of the content of the first message, keywords of the preceding topic, and at least one of a reply prompt indicator. In this way, players can clearly understand that the second message was not canceled, but temporarily suspended due to a combat event, and can quickly recall the topic context before the interruption when outputting is resumed.

[0104] For example, the terminal device can display the pending reply prompt and the topic rewind prompt in the message position corresponding to the first non-player character, or in a preset prompt position in the chat area, so that players can quickly identify which first non-player character, which first message, and which second message the prompt corresponds to. The pending reply prompt can indicate that the second message is currently paused but can be resumed later. It can include at least one of the following: a suspension indicator, a combat interruption indicator, and a reply retention indicator. The suspension indicator can indicate that the current reply process has not ended, for example, displaying text prompts such as "Pending reply," "Suspended," or "Continue later," or displaying a pause icon, a collapsed bubble icon, or a grayed-out ellipsis. The combat interruption indicator can indicate that the reason for the paused output of the second message is related to the current combat event, for example, displaying text prompts such as "Combat interrupted," "Reply paused during combat," or "Temporarily suspended due to combat," or displaying a combat sign or a warning icon. The reply retention indicator can indicate that the content of the second message has not been discarded and will be resumed later, for example, displaying prompts such as "Reply retained," "Continue this reply later," or "Conversation retained." For example, while controlling the first non-player character to pause outputting a second message, the terminal device can insert a status message corresponding to the second message as a pending reply prompt in the chat area. The status message can be associated with the first non-player character's name, avatar, message thread, or time identifier. For instance, if a combat event occurs halfway through the first non-player character's second message output, the terminal device can collapse the incomplete message in the chat area and append a pending reply prompt such as "[Combat Interrupted, Reply Reserved]". Thus, players can directly perceive the interruption status of the current conversation in the chat area without relying on other interfaces. Furthermore, the pending reply prompt can also be displayed in association with the pause progress of the second message. For example, when the second message is a text message displayed word by word, the terminal device can retain the displayed text fragment and append a pause marker after it; when the second message is a voice message, the terminal device can display the corresponding voice entry in the chat area and append a combat interruption marker or a reply reserved marker. If the second message has not yet begun to be output and is only in an upcoming state, the terminal device can also only display the pending reply prompt without displaying the main text of the second message.

[0105] For example, in response to the end of a combat event and when the first non-player character is controlled to continue outputting a second message, the terminal device can first display a topic rewind prompt in the chat area, and then control the first non-player character to continue outputting the second message. The topic rewind prompt can include at least one of the following: at least part of the first message content, preceding topic keywords, and a reply prompt indicator. At least part of the first message content can be the complete text, summary text, the first half of the text, key phrases, or a truncated core segment of the first message; the preceding topic keywords can be one or more keywords used to represent the topic of the conversation before the interruption, such as characters, locations, enemy situations, routes, equipment, status evaluations, etc.; the reply prompt indicator can be used to indicate that the second message will be resumed, such as displaying prompts like "Continue the previous topic," "Continue from the last reply," or "Reply in progress." The topic rewind prompt can be displayed as a separate prompt bar, a quoted message box, a collapsed topic card, or a message prefix mark. For example, the terminal device can first display a prompt in the chat area such as "Continuing the previous topic: 'Is someone coming around to the right?'" before outputting the second message after the first non-player character has resumed; or it can display a prompt with keywords related to the preceding topic, such as "Topic Recap: Enemy Situation on the Right," before resuming outputting the second message. In this way, even if the combat event lasts for a long time, players can quickly re-establish the dialogue context. Furthermore, the terminal device can determine the richness of the topic recap prompt based on the duration of the combat interruption. If the interval between pausing and resuming output is short, only a simplified reply prompt, such as "Continue Reply," can be displayed; if the interval is long, at least part of the first message and the preceding topic keywords can be displayed simultaneously to enhance the recap effect; if other messages are inserted in the chat area during the pause, the terminal device can prioritize referencing the content of the first message to avoid confusing the player with the correct message.

[0106] In some embodiments, see Figure 4 , Figure 4 This is a second flowchart illustrating the interactive processing method for virtual scenes provided in this application embodiment, as shown below. Figure 4 As shown, during execution Figure 3 Before step 103 shown, the following steps can also be performed: Figure 4 Steps 105 to 109 shown will combine Figure 4 The steps shown are explained.

[0107] In step 105, feature extraction is performed on the first message to obtain the message feature vector corresponding to the first message.

[0108] In some embodiments, the first message can be a text message, a voice message, or a text message converted from speech recognition. The terminal device can perform at least one of the following on the first message: word segmentation, semantic parsing, intent recognition, keyword extraction, emotion recognition, and target object recognition, to obtain semantic information related to the first message. The semantic information can include at least one of message type, task intent, target object, location information, urgency level, emotional tendency, role title, and topic keywords. The terminal device further generates a message feature vector based on the semantic information. For example, when the first message is "Medical soldier, please help me recover my status," the terminal device can extract features such as "medical support" intent, "medical soldier" target title, and "high urgency level," and map them to corresponding message feature vectors. The message feature vector can be generated using a pre-trained semantic coding model, a neural network coding model, a word vector mapping model, or a multi-feature concatenation mapping model. For example, the terminal device can concatenate keyword features, intent label features, emotion features, and context features, and map them through a fully connected network to a fixed-dimensional message feature vector for subsequent similarity calculation with role feature vectors.

[0109] In step 106, feature encoding is performed on the multiple character features corresponding to the multiple non-player characters to obtain the character feature vector corresponding to each non-player character.

[0110] In some embodiments, character features may include static character features and dynamic character features. Static character features may include at least one of the following: character class, character identity, weapon type, skill type, personality tags, preferred topics, and team role; dynamic character features may include at least one of the following: current position, current orientation, distance from the player character, current health, current ammo status, whether in combat, currently performing a mission, visible enemy information, whether idle, whether currently speaking, and recent interaction status. The terminal device can normalize, embed map, and fuse the character features corresponding to each non-player character to obtain their respective character feature vectors. For example, for categorical character features, the terminal device can use embedding encoding to generate feature representations; for numerical character features, the terminal device can use normalization mapping to generate feature representations; and for status-based character features, the terminal device can use status label encoding to generate feature representations. Then, the terminal device can concatenate multiple types of character features and obtain the character feature vectors corresponding to each non-player character through a feature encoding network. Thus, character information from different sources and of different types can be uniformly mapped to the same vector space.

[0111] In step 107, the initial matching score for each non-player character for the first message is determined based on the similarity between the message feature vector and the feature vectors of each character.

[0112] In some embodiments, the terminal device can determine the initial matching score for each non-player character regarding the first message based on the similarity between the message feature vector and the feature vectors of each character. The similarity can be calculated using cosine similarity, dot product similarity, bilinear matching similarity, or other vector matching methods. The initial matching score can be used to characterize the basic semantic matching degree between a non-player character and the first message. For example, when the first message contains semantics such as "high platform" or "observe the right side," a non-player character with high-level observation features can obtain a higher initial matching score.

[0113] In step 108, based on the initial matching score, and combined with at least one of the role function matching score, scene state matching score, and historical interaction matching score, the comprehensive matching score corresponding to each non-player role is determined.

[0114] In some embodiments, the initial matching score, based solely on message semantics and role vector similarity, primarily reflects the basic semantic compatibility between the first message and each non-player role. To further improve the accuracy of role selection, the terminal device can also determine a comprehensive matching score for each non-player role by combining at least one of the following: role function matching score, scene state matching score, and historical interaction matching score, based on the initial matching score. The role function matching score can be used to characterize the correspondence between the task requirements of the first message and the functions of the non-player roles. Role functions can include medical support functions, firepower suppression functions, early warning functions, defensive coordination functions, and follow-up support functions. For example, when the first message is "Who should keep an eye on the entrance on the left?", a non-player role with defensive coordination or early warning functions can obtain a higher role function matching score; when the first message is "Help me recover," a non-player role with healing or assistance capabilities can obtain a higher role function matching score. The scene state matching score can be used to characterize whether a non-player role is suitable to immediately undertake a response or perform related tasks in the current first virtual scene. Scene state matching scores can be determined based on at least one of the following factors: distance between the non-player character and the player character, path accessibility, orientation matching, whether they are in intense combat, whether they have relevant field of vision information, whether they are performing high-priority actions, whether resources are sufficient, and whether the current target area is within their control. For example, when the first message is "Is anyone in the right window?", a non-player character currently facing the right window, with field of vision coverage and not being suppressed, can obtain a higher scene state matching score. Historical interaction matching scores can be used to characterize the contextual continuity between a non-player character and the current first message. Historical interaction matching scores can be determined based on at least one of the following information: whether the non-player character was the speaker of the previous message, whether the player character has recently initiated dialogue with the non-player character multiple times, the historical success rate of the non-player character's responses to similar messages, the frequency of interaction between the non-player character and the player character, and whether the current topic is a continuation of the conversation by the non-player character. For example, if the first message is a follow-up question to the previous dialogue, a non-player character who participated in the previous dialogue can obtain a higher historical interaction matching score, thus maintaining topic coherence.

[0115] For example, terminal devices can use a weighted fusion method to determine the overall matching score. For instance, the overall matching score for each non-player character can be expressed as: Overall Matching Score = a × Initial Matching Score + b × Character Role Matching Score + c × Scene State Matching Score + d × Historical Interaction Matching Score, where a, b, c, and d can be preset weights, weights obtained from model training, or weights dynamically adjusted according to the current scene. For example, when the first message is a tactical command message, the weights of the character role matching score and the scene state matching score can be increased; when the first message is a casual chat message, the weight of the historical interaction matching score can be increased.

[0116] In step 109, the non-player character with the highest overall matching score among multiple non-player characters is selected as the first non-player character.

[0117] In some embodiments, after obtaining the comprehensive matching scores for multiple non-player characters, the terminal device can select the non-player character with the highest comprehensive matching score as the first non-player character. Furthermore, if multiple non-player characters have the same or similar comprehensive matching scores, the terminal device can also determine the first non-player character based on secondary filtering rules. These secondary filtering rules may include at least one of the following: shorter response latency, closer proximity to a player character, currently not in a busy state, or stronger continuity of previous dialogue.

[0118] As can be seen, this embodiment of the application extracts features from the first message and generates a message feature vector, which can structurally express the semantic intent, target orientation, and emotional attributes in the first message, thereby providing a unified computational basis for subsequent role matching. Furthermore, by encoding the role features corresponding to multiple non-player roles, multi-dimensional information such as role profession, ability, location, status, and interaction background can be incorporated simultaneously, thus avoiding the mismatch problem caused by selecting a reply role based solely on a single tag. Simultaneously, by determining the initial matching score based on the similarity between the message feature vector and the role feature vector, multiple non-player roles can be quickly filtered from a semantic perspective, improving the basic relevance between candidate roles and the content of the first message. Moreover, by further combining at least one of the role function matching score, scene state matching score, and historical interaction matching score on top of the initial matching score to determine the comprehensive matching score, it can simultaneously consider "who understands this sentence best," "who is currently the most suitable to respond," and "who is most suitable to continue the topic," thereby improving the accuracy and rationality of the selection of the first non-player role.

[0119] In other embodiments, see Figure 5 , Figure 5 This is a schematic diagram of the third process of the virtual scene interaction processing method provided in the embodiments of this application, as shown below. Figure 5As shown, during execution Figure 3 Before step 103 shown, the following steps can also be performed: Figure 5 Steps 110 to 112 shown will combine Figure 5 The steps shown are explained.

[0120] In step 110, the target emotion tag to be used by the first non-player character when replying to the first message is determined.

[0121] In some embodiments, step 110 can be implemented as follows: Obtain the current emotional state coordinates of the first non-player character, and determine the emotional transfer amount based on the difference between the current emotional state coordinates and the player's emotional state coordinates. The player's emotional state coordinates are the coordinates corresponding to the player's emotional state in a preset emotional space, which includes a pleasure dimension and an activation dimension. The emotional transfer amount includes a first transfer amount in the pleasure dimension and a second transfer amount in the activation dimension. When both the first and second transfer amounts are less than or equal to a preset first transfer amount threshold, the emotional tag corresponding to the current emotional state coordinates is used by the first non-player character when replying to the first message. Target emotion tag; when both the first and second migration amounts are greater than the preset second migration amount threshold, the emotion tag corresponding to the current emotion state coordinates is used as the target emotion tag for the first non-player character to reply to the first message; when both the first and second migration amounts are greater than the first migration amount threshold and less than the second migration amount threshold, the emotion migration amount is modulated in combination with the character characteristics of the first non-player character, and the current emotion state coordinates are updated based on the modulated emotion migration amount to obtain the updated emotion state coordinates, and the emotion coordinates corresponding to the updated emotion state coordinates in the preset emotion tag set are used as the target emotion tag for the first non-player character to reply to the first message.

[0122] For example, when determining the target emotion tag to be used by the first non-player character when replying to the first message, the terminal device can adopt a segmented determination method based on the difference in emotion state coordinates. Specifically, the terminal device can obtain the current emotion state coordinates of the first non-player character and determine the emotion transfer amount based on the difference between the current emotion state coordinates and the player's emotion state coordinates. Here, the player's emotion state coordinates are the coordinates corresponding to the player's emotion state in a preset emotion space. The preset emotion space may include a pleasure dimension and an activation dimension. The emotion transfer amount may include a first transfer amount in the pleasure dimension and a second transfer amount in the activation dimension. The terminal device can further determine the target emotion tag to be used by the first non-player character when replying to the first message based on the relationship between the first and second transfer amounts. Here, the preset emotion space can be a two-dimensional emotion space. The pleasure dimension can be used to represent the positive or negative tendency of the emotion, such as from negative to positive; the activation dimension can be used to represent the activity level of the emotion, such as from calm to excited. The terminal device can map different emotion tags to different coordinate positions in the preset emotion space. For example, "calm" can correspond to a low-activation, neutral-pleasure region; "happy" can correspond to a high-pleasure, medium-high-activation region; "tense" can correspond to a low-pleasure, high-activation region; and "calm" can correspond to a medium-low-activation, slightly high-pleasure region. Thus, discrete emotion labels can be converted into calculable continuous coordinate representations. The current emotion state coordinates of the first non-player character can be determined based on the character's historical emotion states, current dialogue context, current behavioral state, combat state, intimacy state, or scene event state. The player's emotion state coordinates can be determined based on at least one of the player's text emotion, voice emotion, facial expression input, operation intensity, behavioral rhythm, or historical interaction emotion trajectory in the first message. For example, the terminal device can analyze the first message using an emotion recognition model to obtain the player's current emotional position in the pleasure and activation dimensions, and use this as the player's emotion state coordinates. For example, the terminal device can label the current emotion state coordinates of the first non-player character as (P0, A0) and the player's emotion state coordinates as (P*, A*). Where P0 and P* represent coordinate values ​​on the pleasure dimension, and A0 and A* represent coordinate values ​​on the activation dimension. The terminal device can determine the amount of emotion transfer based on the difference between the two. For example, the first transfer amount can be |P*-P0|, and the second transfer amount can be |A*-A0|. In some embodiments, the direction information of the difference can also be retained so that when updating the emotional state coordinates later, it can be determined whether to transfer to a higher pleasure dimension, a lower pleasure dimension, a higher activation dimension, or a lower activation dimension. The terminal device can preset a first transfer amount threshold and a second transfer amount threshold, and the second transfer amount threshold is greater than the first transfer amount threshold.The first migration threshold characterizes the boundary for determining "small emotional differences," while the second migration threshold characterizes the boundary for determining "large emotional differences." By setting these two thresholds, the emotional differences between the first non-player character and the player can be divided into small difference intervals, large difference intervals, and intermediate modulation intervals, allowing for different target emotion label determination strategies for different levels of difference. When both the first and second migration thresholds are less than or equal to the preset first migration threshold, the terminal device can use the emotion label corresponding to the current emotional state coordinates as the target emotion label used by the first non-player character when replying to the first message. In other words, when the current emotion of the first non-player character is already quite close to that of the player, the terminal device can maintain the current emotional expression of the first non-player character without additional emotion migration. This avoids unnecessary emotional fluctuations when emotions are already similar, ensuring the stability and naturalness of the reply emotion. When both the first and second migration thresholds are greater than the preset second migration threshold, the terminal device can also use the emotion label corresponding to the current emotional state coordinates as the target emotion label used by the first non-player character when replying to the first message. In other words, when the difference between the player's current emotion and the first non-player character's current emotion is too large, the terminal device can temporarily refrain from migrating directly towards the player's emotion, instead maintaining the first non-player character's original emotion tag. This avoids drastic, abrupt, or distorted emotional jumps in the first non-player character due to excessive fluctuations in the player's emotions, thus ensuring consistency in character personality and boundaries of emotional expression. When both the first and second migration amounts are greater than the first migration threshold and less than the second migration threshold, the terminal device can consider the current emotional difference to be within a moduloable migration range. At this point, the terminal device can modulate the emotional migration amount based on the first non-player character's character characteristics and update the current emotional state coordinates based on the modulated emotional migration amount, obtaining updated emotional state coordinates. Then, the emotional tag corresponding to the updated emotional state coordinates from the preset emotional tag set is used as the target emotional tag for the first non-player character when replying to the first message. In this way, the first non-player character can appropriately perceive the player's emotions while still retaining an emotional response style consistent with its own settings.

[0123] It should be noted that after determining the target emotion tag, the terminal device can apply the target emotion tag to the generation process of the first non-player character's response to the first message. Specifically, the terminal device can adjust at least one of the following in the second message based on the target emotion tag: word choice, tone of voice, sentence length, pause rhythm, speech rhythm, volume variation, facial expressions, lip movements, and body movements. For example, when the target emotion tag is "soothing," the terminal device can control the first non-player character to respond with a gentler tone, lower activation level, and more stable facial expressions; when the target emotion tag is "alert," it can control the first non-player character to respond with shorter sentences, more focused gaze, and a more compact speech rhythm.

[0124] In other embodiments, step 110 can also be implemented as follows: when a preset event is detected in the first virtual scene, the event emotion tag corresponding to the preset event is used as the target emotion tag for the first non-player character to use when replying to the first message. This ensures that the first non-player character's reply to the first message is not only related to the message content but also consistent with the event atmosphere in the current first virtual scene.

[0125] For example, preset events can be scene events that trigger character emotional expression. Preset events can include at least one of the following: combat-triggered events, enemy surprise attacks, key target appearance events, player being attacked events, player being in danger events, teammate being knocked down events, mission success events, mission failure events, treasure chest discovery events, important plot progression events, entering safe zones events, leaving combat events, receiving support events, and completing rescue events. Different preset events can correspond to different event emotion tags, so that the terminal device can quickly determine the target emotion tag suitable for the current scene context after detecting the corresponding event. The terminal device can pre-establish a correspondence table between preset events and event emotion tags. This correspondence table can be generated by development configuration, rule engine, empirical statistical model, or training model. For example, combat-triggered events can correspond to "alert" or "serious" event emotion tags, enemy raids can correspond to "tense" or "urgent" event emotion tags, mission success events can correspond to "happy" or "exhilarated" event emotion tags, player near-death events can correspond to "concerned" or "anxious" event emotion tags, and entering a safe zone events can correspond to "relaxed" or "calm" event emotion tags. This allows for a stable and controllable mapping relationship between the emotional responses to different events. For instance, the terminal device can analyze operational data in the first virtual scene to determine whether a preset event has occurred. Operational data can include at least one of the following: player character status data, non-player character status data, mission system status data, combat system status data, scene object interaction data, and story flow data. For example, the terminal device can identify whether a corresponding preset event has occurred based on information such as changes in player health, enemy unit spawns, mission status changes, entry into specific areas, and activation of key objects. The event emotion tag can be directly used as the target emotion tag for the first non-player character when responding to the first message. In other words, after detecting a preset event, the terminal device can no longer rely on the current emotional state coordinates of the first non-player character or the player's emotional state coordinates. Instead, it can prioritize using the event emotion tag corresponding to the preset event to drive the first non-player character to generate a second message. This ensures that the first non-player character's response emotion prioritizes the contextual needs of the current scene event. The event emotion tag can be used to adjust at least one of the following when the first non-player character replies to the first message: text content, word style, tone strength, speech rhythm, speech speed, facial expressions, body language, and gaze direction.

[0126] In step 111, prompt words are generated based on the target emotion tag, the first message, the character characteristics of the player character, and the character characteristics of the first non-player character.

[0127] Here, the prompt words can be the input text used to constrain and guide the pre-trained language model in generating response content. The prompt words can include at least one of the following types of information: message content information corresponding to the first message, emotion control information corresponding to the target emotion tag, character feature information of the player character, character feature information of the first non-player character, scene context information, and response constraint information. The terminal device can concatenate the above information according to a preset template, or use a dynamic arrangement method to generate prompt words suitable for the current response task.

[0128] In some embodiments, the message content information corresponding to the first message may include at least one of the following: the original text of the first message, the intent label obtained through semantic parsing, keywords, urgency level, topic type, target, and contextual turn information. For example, when the first message is "Medic, please help me recover my status," the terminal device can explicitly write the original content of the first message into the prompt and add semantic identifiers such as "help needed," "treatment need," and "immediate response" so that the pre-trained language model can accurately understand the core intent of the first message. The emotion control information corresponding to the target emotion label can be used to constrain the expression style of the second message. The emotion control information may include at least one of the following: target emotion label name, emotion intensity, suggested tone, suggested sentence structure, suggested vocabulary style, speech rate tendency, and prohibited styles. For example, when the target emotion label is "soothing," the terminal device can add control information to the prompt that "the reply tone should be calm, gentle, and soothing, avoiding overly intense or imperative language"; when the target emotion label is "alert," it can add control information that "the reply should be concise and direct, highlighting risks and maintaining a sense of urgency." Thus, the second message output by the pre-trained language model can better match the expected emotional expression. The player character's characteristics can include at least one of the following: class, faction, current location, current task, current state, historical behavioral preferences, relationship with the first non-player character, and preferred titles. The terminal device can incorporate these characteristics into prompt words, allowing the pre-trained language model to adjust its response based on the player character's identity when generating the second message. For example, for a high-level commanding player character, the first non-player character's response can be more concise and emphasize action feedback; for an injured player character, the second message can highlight concern and a commitment to support. Scene context information can include at least one of the following: scene type of the first virtual scene, current combat state, task stage, distribution of nearby enemy and friendly units, state of interactive objects, recent events, and preceding dialogue content. By incorporating scene context information into prompt words, the pre-trained language model can consider the actual state of the first virtual scene when generating the second message. For example, during intense combat, the second message can tend towards short sentences and high information density; during safe zone communication, the second message can appropriately increase emotional expression and interactive content. Response constraint information can be used to limit the boundaries of second message generation. Reply constraints can include at least one of the following: reply length limits, whether action suggestions are allowed, whether titles are required, whether conversational language is required, whether disclosing system information is prohibited, whether responses are prohibited beyond the capabilities of the primary non-player character, and output format constraints. For example, terminal devices can set constraints in prompts such as "replies should not exceed twenty characters," "prioritize direct responses to player requests," and "only output the reply content of the primary non-player character, without outputting explanatory information" to improve the usability and stability of the second message.

[0129] For example, the terminal device can generate prompts using a template-based approach. These prompts can include fields for character setting, emotion control, message understanding, and response requirements. An example structure could be: "First non-player character identity is XX, personality is XX, current target emotion is XX; player character identity is XX, current state is XX; the player's first message is XX, intent is XX; please generate a second message in the voice of the first non-player character to reply to the first message, satisfying constraint XX." This method reduces the complexity of prompt generation and facilitates reuse in different scenarios. Alternatively, the terminal device can generate prompts using a dynamic enhancement approach. Specifically, the terminal device can adaptively adjust the information weights and field order in the prompts based on the first message type, target emotion tag category, and first non-player character type. For example, when the first message is a tactical command message, the terminal device can increase the proportion of task status and action suggestion information in the prompts; when the first message is a casual conversation message, it can increase the proportion of character personality and emotion expression information. This allows the pre-trained language model to focus more on the key factors of the current response task.

[0130] In step 112, the prompt word is input into the pre-trained language model so that the pre-trained language model generates a second message to respond to the first message.

[0131] In some embodiments, after generating the prompt words, the terminal device can input the prompt words into a pre-trained language model, enabling the pre-trained language model to generate a second message to reply to the first message. The pre-trained language model can be a neural network model with natural language understanding and generation capabilities, such as a large language model built on the Transformer architecture, a dialogue generation model, or a language model fine-tuned from dialogue corpora. By inputting the prompt words into the pre-trained language model, the model can generate response content that conforms to the current interaction scenario based on the semantic and constraint information contained in the prompt words. For example, the terminal device can first perform input adaptation processing on the prompt words, and then input the adapted prompt words into the pre-trained language model. Input adaptation processing can include at least one of text normalization, special tag addition, field order adjustment, length truncation, sensitive content filtering, and encoding format conversion. For example, the terminal device can add identification fields to the prompt words to distinguish between system instructions, role settings, player messages, and response requests, so that the pre-trained language model can more accurately distinguish different types of input content. This improves the model's ability to recognize the structure of the prompt words and enhances the stability of the second message generation. This method allows the second message generated by the first non-player character to simultaneously consider message semantics, character persona, and emotional expression, thereby improving the naturalness and consistency of the reply content.

[0132] In some embodiments, after performing step 112 above, the following processing can also be performed: after a preset time period, the emotional state of the first non-player character is gradually restored from a first emotional state to a second emotional state based on a time decay mechanism. Here, the first emotional state is the emotional state corresponding to the target emotional tag, and the second emotional state is the default emotional state of the first non-player character without external events. This method allows the emotional changes of the first non-player character to have a continuous transition characteristic.

[0133] For example, the preset duration can be the duration for which the first non-player character maintains the first emotional state. The start time of the preset duration can be the start time of the first non-player character outputting the second message, the end time of outputting the second message, the end time of the event corresponding to the target emotional tag, or the moment when the first non-player character completes the current interactive action. The terminal device can initiate the emotional recovery process after the preset duration is reached. This avoids the first non-player character's emotional state immediately dropping after completing the response, making the emotional expression more consistent with the actual interaction logic. Specifically, the terminal device can use emotional state parameters to represent the first and second emotional states. These emotional state parameters can be at least one of emotional tags, emotional intensity values, emotional coordinate values, or emotional vectors. For example, the terminal device can use two-dimensional emotional coordinates, three-dimensional emotional coordinates, or an emotional vector containing multiple emotional components to represent the current emotional state of the first non-player character, where the first emotional state corresponds to a higher intensity event-driven emotional parameter, and the second emotional state corresponds to the first non-player character's basic emotional parameter. By using a parameterized approach to describe the emotional state, the terminal device can easily perform subsequent decay calculations and state updates. The time decay mechanism can be used to control the gradual approach of the first emotional state to the second emotional state. Specifically, the terminal device can gradually reduce the intensity of emotions affected by events in the first emotional state based on time changes according to a preset decay function, and increase the proportion of the default emotional state in the current emotional state, until the current emotional state recovers to the second emotional state or recovers to the threshold range corresponding to the second emotional state. The decay function can be a linear decay function, an exponential decay function, a piecewise decay function, a nonlinear smoothing function, or other preset functions. For example, when using a linear decay mechanism, the terminal device can update the current emotional state as follows: within the recovery period, as time increases, the difference between the current emotional state and the first emotional state gradually increases, and the difference between the current emotional state and the second emotional state gradually decreases. When using an exponential decay mechanism, the terminal device can make the first emotional state decrease rapidly in the early stage of recovery and tend to level off in the later stage of recovery, so that the emotional recovery process is more in line with the natural emotional change pattern. For example, after the first non-player character enters a high-alert state due to a combat event, the terminal device can quickly reduce its tension level in a short period of time after the battle ends, and slowly recover to the default emotional state in the subsequent time. This method is suitable for scenarios where the event stimulus is strong, but the character recovery has a buffer process.

[0134] It should be noted that the parameters of the time decay mechanism can be dynamically determined based on various factors. These factors may include at least one of the following: target emotion tag type, event intensity, event duration, the character characteristics of the first non-player character, the relationship between the first non-player character and the player character, the urgency of the current task, and the scene type of the first virtual scene. For example, for a first non-player character with a personality set as "calm" and "restrained," the terminal device can set a shorter recovery period or a faster decay rate; for a first non-player character with a personality set as "sensitive" and "extroverted," a longer recovery period or a slower decay rate can be set. This ensures that the emotion recovery process aligns with the character's personality. Furthermore, the terminal device can simultaneously adjust the multimodal performance content of the first non-player character during the emotion recovery process. Multimodal performance content may include at least one of the following: response word style, tone of voice, vocal rhythm, speech rate, facial expressions, body language, posture, and eye contact. For example, when the first emotional state is "alert," the first non-player character may exhibit a faster speaking speed, a tense expression, and more focused movements. As the time decay mechanism takes effect, the terminal device can gradually reduce the urgency of the speech, alleviate facial tension, and restore a natural posture until it reaches the default behavior corresponding to the second emotional state. In this way, emotional recovery can be reflected not only in internal state parameters but also in perceptible external manifestations.

[0135] In step 104, other non-player characters are kept silent.

[0136] Here, "other non-player roles" refers to all non-player roles other than the first non-player role mentioned above. In other words, while other non-player roles can also receive or perceive the first message, they do not output a response to that first message during the current round of message response.

[0137] In some embodiments, once the terminal device determines that a first non-player character has responded to the first message, it can control other non-player characters to remain silent. This avoids multiple non-player characters responding to the first message simultaneously, reducing response conflicts, semantic overlap, or voice overlap. The silent state can be a control state used to restrict the output of interactive information by non-player characters. Other non-player characters in a silent state may not output reply text, voice broadcasts, subtitle bubbles, proactive interruptions, or corresponding interactive prompts for the first message. That is, while the first non-player character is outputting the second message, other non-player characters can be restricted from participating in the current round of dialogue response, thus ensuring that the current responder remains unique. Specifically, after determining the first non-player character, the terminal device can set a silence flag, silence permission parameters, or dialogue suppression parameters for other non-player characters. For example, the terminal device can switch the current dialogue state of other non-player characters from "responsive" to "silent" or "response prohibited," and prioritize blocking their response requests to the first message during subsequent dialogue scheduling. Thus, other non-player characters can be uniformly constrained from a role control perspective. The silence state can be triggered at at least one of the following times: after the first non-player character is determined, before the second message is generated, during the generation of the second message, or before the second message is output. For example, after selecting the first non-player character as the current reply character, the terminal device can immediately control the remaining candidate non-player characters to enter a silence state to avoid other non-player characters competing for the right to reply again during the generation of the second message. Furthermore, the silence state can be a complete silence state or a partial silence state. In a complete silence state, other non-player characters neither output voice nor text or speech bubble information; in a partial silence state, other non-player characters can be restricted from outputting voice replies, but can still retain light actions unrelated to the current message, or only retain necessary environmental behaviors. Therefore, the degree of silence can be flexibly controlled according to the needs of the scenario.

[0138] It should be noted that the terminal device can only silently control the "current message response behavior" of other non-player characters, without restricting their basic behavioral logic. For example, other non-player characters in a silent state can still perform actions such as moving, patrolling, fighting, defending, following, or other preset behaviors, but will not provide verbal feedback to the first message. This avoids the overall rigidity of the behavior of other non-player characters due to the silent state, thus maintaining the continuity and realism of the first virtual scene. Furthermore, the conditions for ending the silent state can include at least one of the following: completion of the second message output, end of the current session round, expiration of the preset silence duration, the player character initiating a new first message, the first non-player character leaving the current interaction range, or the detection of a high-priority event. When the conditions for ending the silent state are met, the terminal device can restore the dialogue state of other non-player characters from a silent state to a responsive state, allowing them to participate in subsequent interactions.

[0139] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.

[0140] This application provides an interactive processing method for virtual scenes to solve a series of problems in tactical competitive / battle royale games, such as AI teammate dialogue relying on fixed scripts, chaotic dialogue when multiple AI teammates are online simultaneously, inaccurate assignment of player commands to corresponding AI teammates, AI teammate dialogue interfering with player operations during combat, and lack of emotional connection between AI teammates and players on the spawn island. The core of this application is to replace the traditional AI teammate dialogue scripts handwritten by game designers with a systematic solution driven by LLM (Local Management Model) that integrates group chat orchestration, command coordination, combat awareness, and spawn island memory. Specifically, it includes the following components: 1) Relevance judgment and silence mechanism in group chat: When multiple AI teammates are online at the same time, each AI teammate independently judges whether the current message is related to its own persona. Only the relevant AI teammate generates a reply, while the other AI teammates remain silent to avoid multiple AI teammates interrupting each other at the same time. 2) Command distribution and expertise matching from a single player to multiple AI teammates: After a player issues a command, the command can be assigned to the most suitable AI teammate to execute and confirm the reply based on the differences in the personalities of each AI teammate (such as reply style, tone, emotional reaction, etc.), achieving a clear interaction of one command and one AI teammate reply; 3) Combat event-driven dialogue interruption and topic rewind: If a combat event (such as gunshots, being attacked, or entering combat) is detected during casual conversation between the AI ​​teammate and the player, the conversation is immediately interrupted and combat mode is activated. After the combat ends, the AI ​​teammate proactively rewinds and resumes the previous conversation topics to maintain dialogue continuity. 4) Personalized memory dialogue on the spawn island: When a player enters the spawn island, the AI ​​teammate reads the player's historical behavioral memory (such as the content of the last conversation, player preferences, commonly used weapons, etc.) to generate differentiated opening lines, so as to achieve a personalized experience with new topics every time they meet. 5) Differentiated AI teammate personalities and adaptive responses: Different AI teammates have different personalities (such as calm, enthusiastic, cautious, etc.), and adjust their response style according to the player's current emotional state (such as aggressive, conservative, anxious, etc.), so that the dialogue experience of each AI teammate is distinctive.

[0141] In addition, this application also provides an engineering implementation path of "message routing → LLM parallel inference → relevance filtering → response generation → TTS synthesis → client presentation", which can be directly implemented in existing tactical competitive gameplay.

[0142] In other words, this application's embodiments, through relevance judgment and a silence mechanism, ensure that when multiple AI teammates are online simultaneously, only the relevant AI teammates respond, while the others remain silent, preventing channel spamming. Furthermore, through a character-differentiated command distribution and matching mechanism, each player's command is confirmed and responded to by the most suitable AI teammate, achieving clear interaction. Additionally, through a battle event-driven dialogue interruption and topic rewind mechanism, AI teammates remain silent during battle and actively resume casual conversation afterward, maintaining dialogue continuity. Simultaneously, through a personalized memory dialogue mechanism on the spawn island, AI teammates can generate new topics based on historical memories each time they meet the player, thereby establishing emotional connections. Moreover, through a character-differentiated emotion-adaptive response mechanism, the dialogue experience of each AI teammate is distinctive, thereby enhancing player immersion.

[0143] The interactive processing method for virtual scenes provided in the embodiments of this application will be described in detail below.

[0144] In some embodiments, the technical solution provided in this application can be applied to a mixed team mode of 2 real players + 2 AI teammates in a single game of a combat-based competitive game. During the game, the player and the two AI teammates are in the same group chat channel, and all messages are visible to all team members. The player can @ a specific AI teammate to send a private message, or send a broadcast message to all AI teammates. Furthermore, during the spawn island phase, the player can engage in personalized casual conversation with the AI ​​teammates, who will generate an opening line based on their historical memory of the player. During the game, the player can issue combat commands to the AI ​​teammates, who will confirm and execute them according to their own unique characteristics. Additionally, when the player enters combat, the AI ​​teammates automatically stop casual conversation and focus on combat support. After the battle ends, the AI ​​teammates proactively resume the previous casual conversation topics.

[0145] The group chat conversation orchestration module involved in the embodiments of this application will be described below.

[0146] For example, see Figure 6 , Figure 6 This is a schematic diagram of the first process of the interactive processing method for virtual scenes provided in the embodiments of this application, as shown below. Figure 6 As shown, the group chat orchestration module is responsible for receiving all messages sent by players and routing them to each AI teammate. Its core logic is as follows: if a message contains an @mention of a specified AI teammate, it will only be received and a reply generated by the @mention AI teammate; if the message is a broadcast message (i.e., without @mentioning any AI teammate), all AI teammates will receive it in parallel, and each AI teammate will independently call the LLM to determine whether it should reply. The inputs of the LLM include: the character names, identities (real person / AI), personality descriptions, and binding relationships of the four positions in the team; the historical message records of the current session; the content and speaker information of the latest player's message; the character name and position of the current AI teammate; and the judgment rules shown in Table 1 (in descending order of priority).

[0147] Table 1. Illustrated Rules for Judgment Condition Result The speaker is the current AI teammate itself Do not reply The name or position number of the current AI teammate is directly mentioned in the speech Reply The speech is asking the current AI teammate a question Reply The speech content continues what the current AI teammate said in the previous round (for example, including follow-up questions, responses, continuation, etc.) Reply The speech uses collective addresses such as "you / everyone", or an open question starting with "you", and the current AI teammate is the speaker's bound partner Reply The current topic is suitable for intervention in accordance with the character setting of the current AI teammate Reply The speech is explicitly addressed to other people and does not involve the current AI teammate Do not reply Pure modal particles, battle situation broadcasting, private chat between real players Do not reply For example, continuing from the above, the inputs and rules can be combined into a command text, which can then be fed into a large language model for semantic understanding and reasoning, so that the large language model returns results in a fixed format, such as: {"should_speak": Yes / No, "reason": Brief reason} Wherein, `should_speak` indicates whether the current AI teammate should reply, and `reason` is a brief explanation of the judgment criteria. Judgment criteria may include: whether the message content is related to the AI's own character profile, whether it is an instruction directed at the AI, and whether the AI ​​is currently in a combat silence state. In this embodiment, all the information required for judgment can be organized into two parts and fed into the large language model using a preset instruction template. The first part is a fixed instruction, including: the name, identity, personality description, and partner relationship of each character in the team; the eight judgment rules in Table 1; and the output format requirements. The second part needs to be dynamically populated each time, including: historical dialogue records (e.g., a message list arranged chronologically); and the latest speech (including speaker ID, speaker name, and speech content). The process of determining whether a message is relevant to one's own persona can be as follows: by writing the personality description of the AI ​​teammate into a fixed instruction section, the large language model automatically judges the degree of matching between the topic and the AI ​​teammate's personality during inference. The process of determining whether an instruction is directed at oneself can be done through rules such as name matching, position matching, binding relationship matching, and semantic continuity. Furthermore, whether the user is currently in a silent combat state can be used as a prerequisite, determined by the system logic before calling the large language model. If the user is in a silent state, the process is skipped, and the above judgment process is not entered. Only when one or more AI teammates determine that a reply should be given will the LLM be called to generate a reply and send it to the channel; if no AI teammate determines that a reply should be given, the channel remains silent.

[0148] The instruction distribution and expertise matching module involved in the embodiments of this application will be described below.

[0149] For example, see Figure 7 , Figure 7 This is a schematic diagram of the second process of the virtual scene interaction processing method provided in the embodiments of this application, as shown below. Figure 7As shown, the command distribution and expertise matching module is responsible for assigning player commands to the most suitable AI teammates for confirmation and response. The core logic is as follows: After a player issues a command, the intent recognition module first parses the command type (e.g., offensive, defensive, resource, navigation, etc.). The intent recognition module can use a two-level sequential judgment method to parse player commands. The first level is coarse classification, which determines whether the player's message is a command or casual chat. The second level is fine parsing, which further parses the specific command type and accompanying parameters only for messages judged as commands. In other words, the purpose of coarse classification is to quickly determine which category the message sent by the player belongs to, such as whether it is a command, casual chat, or meaningless content. The specific judgment process is as follows: 1) Text preprocessing: Remove punctuation marks from the message, convert them to lowercase, and remove distracting words such as the AI ​​teammate's role name (e.g., titles like "X-Li" or "Number Two"), retaining the pure semantic content; 2) Rule-based matching: Maintain a pre-set table of known command phrases (e.g., including high-frequency command phrases like "Don't shoot," "Let's go," and "Give me"). If the player's message matches a command in the table after removing distracting words, it is directly judged as a command without further calculation; 3) Model inference: If the rule is not matched, the original text is fed into a pre-trained text classification model for inference. The model outputs a set of values ​​corresponding to the confidence scores of the three categories: "casual chat," "command," and "meaningless." If the confidence score of the "command" category exceeds a preset threshold (e.g., 0.55), it is judged as a "command"; if the confidence score of the "meaningless" category exceeds a high threshold (e.g., 0.95), it is judged as "meaningless content" (at which point a fallback response can be given); otherwise, it is judged as "casual chat," and the conversation proceeds into casual chat mode. The purpose of detailed analysis is to identify the specific instruction type and related parameters (such as material name, location name, etc.) of messages that are initially identified as "instructions". The specific identification process is as follows: 1) Constructing instruction templates: All supported instruction types and their output formats can be predefined, and can be divided into several major categories according to function, as shown in Table 2: Table 2. Instruction Type Diagram Category Included specific instruction types Supplies category Requesting supplies, marking supplies, storing specified supplies Offensive category Shooting and concentrating fire, marking targets, deploying smoke, throwing projectiles, looting airdrops Defensive category Retreating, finding cover, holding fire, giving up kills, requesting healing Navigation / movement category Moving to a specified area, searching a specified area, finding a vehicle, moving the vehicle to a specified area, running from the poison zone, approaching a player Vehicle control category Accelerating, decelerating, stopping the vehicle, getting in the vehicle, changing the driver, refueling the vehicle Parachuting category Parachute with me, follow me to parachute, accelerate parachuting, decelerate parachuting, direct flight, diagonal flight Others Using firearms, prohibiting firearms, map marking, low health, crouching, prone, etc. 2) Assemble Input Information: The following can be combined into a structured text and fed into the large language model: definitions and output format descriptions for all command types; examples corresponding to each command type (e.g., input → output comparison examples); the message content currently sent by the player. 3) Large Language Model Inference: Based on the command definitions and examples, the large language model performs semantic understanding of the player's message and outputs a fixed-format parsing result: ["command type", "parameter"], for example: "Give me an M4" → ["Request supplies command", "M4"]; "Let's go to point A" → ["Move to designated area", "point A"]; "There are enemies in the house, let's rush together" → ["Fire and focus fire", ""]; 4) Result Verification: If the large language model outputs an empty list, it means that although the message was initially identified as a command, it cannot match any defined command type. In this case, it can be downgraded to casual conversation. Next, the character tags of each AI teammate (e.g., including reply style, tone characteristics, emotional reaction patterns, etc.) can be read, and based on the matching degree of "command type + AI character," the AI ​​teammate with the highest matching degree is selected to generate a confirmation reply. In other words, when the multi-user command system receives a command event, it doesn't simply generate a uniform response based on the command type. Instead, it further combines the AI ​​teammate's persona information to determine the matching relationship between the command event and the AI ​​persona, thereby identifying the most suitable confirmation response for the current command scenario. Specifically, multiple sets of confirmation response texts can be pre-configured, each set associated with one command event condition and at least one AI persona condition. The command event condition can include command type, command execution reason, command execution result, scene mode, vehicle type, item type, etc.; the AI ​​persona condition can include AI character identifier, response style, tone characteristics, emotional reaction pattern, character personality tags, etc. Upon receiving a command event, the current command event information and the candidate AI teammate's persona information can be combined to form a joint condition, and a set of response texts matching this joint condition can be retrieved from a pre-set response database.

[0150] It's important to note that the matching degree can be reflected through configuration hit rate. That is, if a set of candidate responses simultaneously meets both the current command event conditions and the current AI persona conditions, then the set of responses can be considered to have a high matching degree with the current scenario; if only some conditions are met, the matching degree is relatively low; if the key conditions are not met, then it is not considered a candidate response. Next, the set of responses that meets the most complete and specific conditions can be selected as the candidate response set with the highest matching degree. For example, assuming the current command event is "failed to find a vehicle," and the current AI persona is a character with a "positive tone, direct response, and encouraging response style," then the preset set of response texts that are simultaneously bound to the command type "failed to find a vehicle" and the AI ​​persona tag can be searched first. The response texts in this set are all pre-written according to the AI's persona style, so selecting any response from this set ensures that the response content conforms to both the semantics of the command event and the AI's character expression style.

[0151] In addition, it should be noted that the input for matching degree judgment can include the following types of information: 1) Command event information: used to describe the currently occurring command and its execution status, including but not limited to: command type, such as pathfinding, vehicle search, vehicle boarding, following, attacking, picking up items, etc.; command reason or execution reason, such as success, failure, target unreachable, resource not available, distance too far, etc.; command execution result, such as execution success, execution failure, delayed execution, instant feedback, etc.; scene parameters, such as map mode, combat status, vehicle type, item type, target object, etc.; 2) AI character information, used to describe the role characteristics of different AI teammates, including but not limited to: AI character identifier; response style. Examples include: lively, calm, humorous, direct, and encouraging; emotional response patterns, such as comforting players when they fail, actively confirming when they succeed, and nervously reminding them when they are in danger; character personality tags, such as steady, commanding, supportive, and teasing; 3) preset reply library information, used to provide candidate reply texts with character constraints, including: reply texts bound to the command type; reply texts bound to the command reason; reply texts bound to the AI ​​character tags; multiple equivalent candidate replies under the same matching conditions; 4) historical dialogue information, used to avoid duplicate replies, including: confirmation replies already given by AI teammates in the current game; dialogue content from the most recent rounds; and reply texts of the same type that have been used.

[0152] The following section will explain the calculation process for the matching degree.

[0153] In some embodiments, the matching degree can be determined in the following ways: 1) Extracting command event features: First, extract key fields for matching from the command event, such as: command type, command reason, execution result, vehicle type, item type, scene mode, etc. Specifically, the current event can be abstracted as: command type = finding a vehicle; command reason = no suitable vehicle found; vehicle type = vehicle; execution result = failure. These fields are used to describe what happened; 2) Extracting AI character features: Then, the character information of the candidate AI teammates can be read, such as: AI role = role A; response style = positive encouragement; tone characteristics. =Concise and direct; Emotional response mode = When failing, comfort the player and give a tendency to continue acting. These fields can be used to describe which AI teammate is more suitable to respond in what way; 3) Construct joint matching conditions: Then, the command event features can be combined with AI character features to form joint matching conditions. For example: Matching condition = Command type + Command reason + Scene parameters + AI role / character tag, or it can also be expressed as: Matching key = {Command type, Command reason, Scene mode, AI role identifier, AI response style tag}. Through this joint matching condition, the system does not judge what kind of command it is, but judges under the current command event. Is it appropriate for the AI ​​teammate of this persona to reply? 4) Search for candidate reply sets in the preset reply library: Then, according to the joint matching conditions, the corresponding candidate reply sets can be found in the preset reply library. If a reply set satisfies the following conditions at the same time: instruction type matching; instruction reason matching; scene parameter matching; AI persona matching, then the set can be considered to have the highest matching degree. If there is no set that is completely matched, it can be degraded step by step according to the preset priority. For example: instruction type + instruction reason + AI persona → instruction type + number of AIs → instruction type + general persona → general fallback reply. In this process, the more complete the matching conditions, the higher the matching degree. 5) Determine the set of responses with the highest matching degree: The matching degree can be understood as the degree of fit between the candidate response set and the current joint matching conditions. In a rule-based implementation, the matching degree can be defined as: Matching degree = Instruction type matching score + Instruction reason matching score + Scenario parameter matching score + AI persona matching score. For example: Matching degree (let's call it M) = W1 * Instruction type matching score + W2 * Instruction reason matching score + W3 * Scenario parameter matching score + W4 * AI persona matching score. Where W1, W2, W3, and W4 are weight coefficients that can be set according to business needs. If a certain condition is matched, the matching score of that condition is 1; otherwise, it is 0.In another engineering implementation, instead of explicitly calculating numerical scores, the highest matching degree can be reflected through the precise hit of the configured index. For example, when the index conditions of a group of reply texts simultaneously include the current instruction type, reason number, and AI persona identifier, this group of texts is considered the current highest matching degree reply set; 6) Select a confirmation reply from the highest matching set: After determining the highest matching candidate reply set, a confirmation reply can be selected from this set. To avoid the AI ​​teammate repeating the same expression multiple times, deduplication can be performed first by combining with historical dialogues. For example, candidate reply set - historically used replies = optional reply set. If there are still optional replies after deduplication, one can be randomly selected from the optional reply set; if it is empty after deduplication, one can be randomly selected from the original candidate reply set. Since this random selection process occurs within the "highest matching candidate set", it will not violate the principle of the highest matching degree. The purpose of random selection is to improve the diversity of expression, rather than to change the matching results of the AI ​​persona or instruction event.

[0154] It should be noted that the output of the matching process can include the following: Target AI or AI persona identifier: indicating the AI ​​teammate or AI persona that the current confirmation reply matches; Highest matching reply set: representing the set of candidate reply texts that simultaneously meet the current command event and AI persona conditions; Confirmation reply text: the final reply content selected from the highest matching reply set; Reply type: such as command confirmation reply, active command reply, passive command reply, etc.; Matching criteria: such as the type of command hit, reason number, AI persona tag, scene parameters, etc.; Optional matching degree information, such as matching score, matching level, or matching priority. Ultimately, the output confirmation reply not only reflects the execution result of the current command event but also maintains a consistent language style with the corresponding AI teammate's persona. Meanwhile, other AI teammates, although they also receive the command (for contextual coherence), do not generate confirmation replies to avoid confusion caused by multiple AI teammates simultaneously confirming.

[0155] The following description continues the explanation of the combat interruption and topic rewind module involved in the embodiments of this application.

[0156] For example, see Figure 8 , Figure 8 This is a schematic diagram of the third process of the virtual scene interaction processing method provided in the embodiments of this application, as shown below. Figure 8As shown, the combat interruption and topic rewind module is responsible for perceiving the combat state and managing the dialogue context. Its core logic is as follows: continuously detect combat events (such as gunshot detection, attack events, player health changes, and entry into combat status markers). Once a combat event is detected, a "combat mode" marker is immediately set, and the AI ​​teammate stops sending casual chat messages. Simultaneously, the current state of casual chat topics (such as topic content, progress, and context) can be written into long-term memory (long-term memory refers to a summary of important interactive information across game sessions). After the game ends, a large language model can be used to compress the historical dialogue between the player and AI teammate over several rounds into an objective third-person summary, which is then used as a rolling memory prompt for the next round, guiding the large language model to generate appropriate responses and improving the continuity of the model's responses. After the battle ends, the "combat mode" marker is cleared, and it can check for any unfinished casual chat topics. If any exist, the AI ​​teammate proactively generates a rewind message (such as "We were just talking about the gun you used last time") to resume the dialogue.

[0157] The following description continues the explanation of the personalized memory dialogue module for the birth island involved in the embodiments of this application.

[0158] For example, see Figure 9 , Figure 9 This is a schematic diagram of the fourth process of the virtual scene interaction processing method provided in the embodiments of this application, as shown below. Figure 9 As shown, the personalized memory dialogue module on the spawn island is responsible for generating personalized dialogues based on historical memories during the spawn island phase. The core logic is as follows: When a player enters the spawn island, the system first reads the player's historical behavior data with each AI teammate from the server. This historical behavior data can include: a summary of the last dialogue, the player's preferred weapons / equipment, the player's performance characteristics in previous matches, and the player's preferred tactical style. Then, the historical behavior data can be used as input context for a large language model to generate opening lines with memory associations, such as "Last time you got a lot of eliminations with XX, what are you planning to use this time?" After each dialogue, a summary of the current dialogue is updated and written to the server for use the next time the player meets the AI ​​team.

[0159] The following description continues with the description of the emotion-adaptive response module involved in the embodiments of this application.

[0160] For example, see Figure 10 , Figure 10 This is a schematic diagram of the fifth process of the virtual scene interaction processing method provided in the embodiments of this application, as shown below. Figure 10As shown, the emotion-adaptive response module is responsible for adjusting the AI ​​teammates' response style based on the player's emotional state. The core logic is as follows: it detects the player's current emotional state through player actions (such as firing frequency, movement speed, and the urgency of item usage). Emotional state can be categorized using two dimensions: pleasure (denoted as P) and activation (denoted as A). Pleasure measures an individual's positive or negative evaluation of the current state, i.e., the degree of liking or disliking, including high pleasure (denoted as P+): such as happiness, satisfaction, and contentment; and low pleasure (denoted as P-): such as pain, displeasure, and disgust. Activation reflects an individual's physiological or psychological alertness level, and the energy activation state of the body and mind, including high activation (denoted as A+): such as excitement, tension, fear, and alertness; and low activation (denoted as A-): such as calmness, drowsiness, boredom, and relaxation. For example, based on the PA two-dimensional model, emotions can be categorized as follows: Figure 11 The four quadrants are shown. Furthermore, the emotions supported by the large language model can include: happy-high, happy-low, peaceful, curious-high, curious-low, surprised-high, surprised-low, angry-combat-high, angry-combat-low, sad-high, sad-low, angry-high, and angry-low. Additionally, the following can be used: Figure 12 The diagram illustrating the emotional PA distribution visualizes the distribution of emotions, showing the distribution of 13 emotions in a two-dimensional coordinate system of pleasure and activation, and demonstrating the precise location of emotions in the four quadrants. The definitions and numerical ranges of the emotion quadrants are explained below.

[0161] For example, the P+A+ quadrant (i.e., high pleasure and high activation) represents excitement and positivity, with a corresponding numerical range of P: 0.6-1.0, A: 0.6-1.0, including emotions such as excitement, happiness, joy, and ecstasy. Typical emotions are: high joy - high happiness, and high surprise - high astonishment. The P+A- quadrant (i.e., high pleasure and low activation) represents calmness and satisfaction, with a corresponding numerical range of P: 0.6-1.0, A: 0.0-0.4, including emotions such as satisfaction, calmness, relaxation, and contentment. Typical emotions are: calmness - calmness, and low pleasure - low happiness. The P-A+ quadrant (i.e., low pleasure and high activation) represents tension and negativity, with a corresponding numerical range of P: 0.0-0.4, A: 0.6-1.0, including emotions such as anger, anxiety, fear, and tension. Typical emotions are: high anger - high fighting urgency, and high anger - high resentment. The PA quadrant (i.e., low pleasure and low activation) represents low mood and negativity, with corresponding numerical ranges of P: 0.0-0.4 and A: 0.0-0.4. It includes emotions such as sadness, boredom, frustration, and depression, with typical emotions being: high sadness-loss and low sadness-loss. A detailed emotion name mapping table is shown in Table 3.

[0162] Table 3 Emotion Name Mapping Table Chinese Emotion English Identifier P-value A-value Quadrant Joy-High Happiness happy-high 0.6-1.0 0.6-1.0 P+A+ Joy-Low Happiness happy-low 0.6-1.0 0.0-0.4 P+A- Calm-Peacefulness peace 0.6-1.0 0.0-0.4 P+A- Confusion-High Curiosity curious-high 0.0-0.4 0.6-1.0 P-A+ Confusion-Low Curiosity curious-low 0.0-0.4 0.0-0.4 P-A- Surprise-High Surprise surprised-high 0.6-1.0 0.6-1.0 P+A+ Surprise-Low Surprise surprised-low 0.0-0.4 0.0-0.4 P-A- Anger-High Combat Urgency combat-high 0.0-0.4 0.6-1.0 P-A+ Anger-Low Combat Urgency combat-low 0.0-0.4 0.0-0.4 P-A- Sorrow-High Depression sad-high 0.0-0.4 0.0-0.4 P-A- Sorrow-Low Depression sad-low 0.0-0.4 0.0-0.4 P-A- Anger-High Anger angry-high 0.0-0.4 0.6-1.0 P-A+ Anger-Low Anger angry-low 0.0-0.4 0.0-0.4 P-A- It's important to note that emotion detection is not simply a text classifier, but a composite state machine combining LLM emotion recognition, PA two-dimensional smooth transition, role-personality modulation, special event jump, and time decay recovery. Text classification is only one step in the target emotion recognition process. The emotion label ultimately sent to the large language model is obtained after a smooth transition between the role-personality parameters and the current emotional state. This will be explained in more detail below.

[0163] Step 1: Emotion Recognition In some embodiments, zero-shot classification based on a large language model can be used. The input is the history of the last 5 rounds of dialogue, concatenated into a text segment as "Player: xxx / {Character Name}: xxx". The prompt can be: You are a dialogue emotion classification assistant. Based on the given dialogue content, please determine the emotion that the character is most likely to present in the next response. Optional emotion labels include: Joy - High Happiness / Joy - Low Happiness / Calm - Calm / Confused - High Curiosity / Confused - Low Curiosity / Surprised - High Surprise / Surprised - Low Surprise / Anger - High Fighting Urgency / Anger - Low Fighting Urgency / Sadness - High Disappointment / Sadness - Low Disappointment / Anger - High Anger / Anger - Low Anger. Only one emotion label is output, without explanation or other content. The output is one of the above 13 emotion labels. The classification goal is not the player's current emotion, but the emotion that the AI ​​character should present in the next response. This is the key difference from traditional emotion classification and a natural connection to dialogue generation.

[0164] Step Two: Mapping 13 Emotional Labels to a Two-Dimensional Continuous Spatial PA In some embodiments, such as Figure 12 As shown, each emotion label has a predefined (P, A) ∈ [0, 1] in the emotion PA distribution map. 2 The coordinates, as shown in Table 4, illustrate the distribution of some emotion labels in the emotion PA distribution map. This ensures that emotions are not isolated categories, but rather fall within a continuously interpolated space, which forms the mathematical basis for subsequent smooth transitions.

[0165] Table 4. Distribution of Some Emotion Labels in the Emotion PA Distribution Map

[0166] Step 3: State transition (instead of directly using the classification results, a soft update is performed) In some embodiments, assuming the AI ​​teammate's current emotional state is (P0, A0), the identified target emotional state is (P*, A*), and the AI ​​character's personalized parameters are (v, s) + s - ), where v represents volatility, indicating the overall sensitivity or magnitude of change in emotional state to external events. A higher value indicates that emotions are more easily affected by events and the changes are more pronounced; a lower value indicates that emotions are more stable and less prone to fluctuation; s + This represents positive sensitivity, indicating the strength of the response to positive events. A higher value indicates a greater likelihood of positive emotional shifts when faced with positive stimuli such as encouragement, victory, rewards, and successful cooperation; a lower value indicates a weaker response to positive stimuli. - This represents negative sensitivity, indicating the intensity of the response to negative events. A higher value indicates a greater likelihood of negative emotional shifts when faced with negative stimuli such as failure, setbacks, teammate errors, or task obstacles. A lower value indicates stronger resistance to negative stimuli. For example, suppose an AI teammate has a volatility of 0.14, a positive sensitivity of 1.4, and a negative sensitivity of 0.65, meaning they are more sensitive to positive events than negative ones, thus exhibiting a cheerful personality. Next, the displacement (i.e., the amount of emotional transfer) can be calculated. ΔP=P*-P0 ΔA=A*-A0 Here, ΔP represents the original change in pleasure level (i.e., the first transfer factor), and ΔA represents the original change in activation level (i.e., the second transfer factor). If |ΔP|≤θ and |ΔA|≤θ, the original state is maintained to avoid fluctuations. Here, θ is the change threshold, which is a preset threshold used to determine whether a change in emotional state meets the valid update condition. For example, the value of θ can be 0.01. If |ΔP|>0.5 and |ΔA|>0.5, the transfer is rejected, and the original state is maintained to avoid extreme jumps, such as preventing abrupt switching from "sadness-low" to "joy-high". In addition, if ΔP>0, then s + Use this as the currently selected sensitivity coefficient (let's call it s); otherwise, use s. - The sensitivity coefficient currently selected is: s=s + if ΔP>0 else s - ΔP' = ΔP × v × s ΔA' = ΔA × v × s Where ΔP' represents the change in pleasure level after processing, and ΔA' represents the change in activation level after processing. Next, we can proceed to state update and resubmission, i.e.: P1 = clamp (P0 + ΔP', 0, 1) A1 = clamp (A0 + ΔA', 0, 1) Label1=argmin_l[(P1 - P_l)² + (A1 - A_l)²] Here, `clamp` is an interval constraint function used to restrict variables within a preset upper and lower bound range. `Label1` represents the updated emotion label, indicating the result obtained after remapping the smoothed updated emotion coordinates back to the preset emotion set. `argmin_l` represents the emotion label corresponding to the minimum value, indicating that among all candidate emotion labels, the emotion label closest to the current coordinate is selected. `l` is the index variable of the candidate emotion label, used to traverse all preset emotion labels. `P1` is the updated pleasure value, indicating the current pleasure coordinate after emotion transfer and smoothing. `A1` is the updated activation value, indicating the current activation coordinate after emotion transfer and smoothing. `P_l` represents the pleasure value corresponding to the l-th preset emotion label, and `A_l` represents the activation value corresponding to the l-th preset emotion label.

[0167] Step 4: Time decay recovery (allowing emotions to cool naturally back to the personality baseline) In some embodiments, after a period of inactivity, the AI ​​teammate's emotions can decay exponentially to the character's baseline state (P_base, A_base): α = 1 - exp( -(Δt×recovery_speed) / τ) P_rec = P0 + (P_base - P0)×α A_rec = A0 + (A_base - A0)×α Here, α is the recovery coefficient, representing how close the current emotional state is to the baseline state within this time period. Its value is typically between 0 and 1; the closer to 0, the less recovery, and the closer to 1, the more complete the recovery. Δt is the time interval, representing how long has passed since the last update, and can be in seconds, minutes, or simulation steps. recovery_speed is the recovery speed parameter, used to control how quickly the emotion reverts to the baseline; a larger value indicates faster recovery, and a smaller value indicates slower recovery. τ is the recovery scale parameter (e.g., 1200 seconds), which controls the time scale of exponential recovery; a larger value indicates slower recovery, and a smaller value indicates faster recovery. τ and recovery_speed together determine the recovery speed. exp is the exponential function, used here to describe a smooth, gradual recovery process, conforming to the natural decay phenomenon. P_rec represents the post-recovery happiness level, indicating the new happiness value after the recovery process; P_base represents the baseline happiness level, indicating the stable value the system naturally wants to revert to, which can be understood as the default happiness level in a calm state; A_rec represents the post-recovery activation level, indicating the new activation level after recovery, and A_base represents the baseline activation level, indicating the default activation level the system naturally reverts to. In other words, the AI ​​teammate won't remain in an angry state indefinitely, but will gradually revert to the character's inherent personality over time—a physics analogy simulating emotional decay.

[0168] Step 5: Special game event direct jump (bypass, bypassing LLM) In some embodiments, when a special game event is encountered, the PA state can be forcibly overridden, skipping the classification and personalization. These special game events are shown in Table 5. Table 5. Illustration of Special Game Events Event Direct jump to target Remark Continuous kill Joy-High Happiness direct_jump=true All teammates killed Sorrow-High Depression direct_jump=true First death Varies depending on the character, for example, Character A → Anger-High Anger, Character B → Confusion-High Curiosity character_specific Victory Joy-High Happiness direct_jump=true Step Six: Inject into Dialogue Generation In some embodiments, the final emotion label influences the LLM response style (e.g., wording, tone, and performance) through the injection of prompts into the chat LLM. Furthermore, during response generation, the current emotion label (13 categories, including "Happy - High" and "Angry - High Combat Urgency") output by the emotion state transition submodule is converted into downstream unified identifiers (e.g., happy-high, combat-high, sad-high) through a predefined mapping table. This is then appended to the end of the system prompt, assembled from character persona, player profile, and dialogue summary, using a fixed template stating "You need to reply with the emotion <emotion identifier>". This prompt, along with the player's current statement, is fed into the large dialogue model, which generates an emotion-adapted response text. This response text undergoes post-processing such as sensitive word filtering, emoji parsing, and TTS emotion label trimming before being sent to the client to drive the AI ​​teammate's voice response. For example, as shown in Table 6, the end-to-end data sample is as follows: Player Input → Emotion Detection → Prompt Injection → AI Teammate Response.

[0169] Table 6. Example of end-to-end data Field Actual value Trigger input The last utterance in 5 rounds of historical conversation: "I went to XX two days ago, it was a bit cold since it was my first time going there" AI character X Li (ai_character=1) Emotion detection path PersonalityEmotion → LLM zero-shot classification → Current state (P0, A0) = (0.75, 0.35) smooth transition Emotion tag output Confusion - low curiosity Mood Mapping curious-low PA smoothed coordinates (0.55, 0.32) Statements injected into prompt words You need to use <curious-low>Respond with emotions AI actually generated the reply Oh my, it's really cold in XX this winter! Where did you go? Did you get frostbite climbing XX? Control group (without emotional input) XX is indeed cold in winter.

[0170] It should be noted that in practical applications, in addition to using the same large language model, a multi-model combination strategy can also be adopted. For example, the message routing and instruction distribution module can use a model with stronger reasoning capabilities, the emotion detection and response generation module can use a fine-tuned model with a tone closer to the character, and the spawn island memory module can use a long context model. Each module can independently switch models according to the actual effect, but it must be based on the LLM and cannot replace the core dialogue generation function of the LLM with a rule engine. In addition, besides directly generating natural language response text, structured fields such as emotion tags and action tags can be added to the output, so that the client can trigger the AI ​​teammate's actions and expressions (such as nodding while speaking, standing up when emotionally agitated) simultaneously while playing the response voice, thus enhancing expressiveness. The LLM is still responsible for generating response text and structured tags and cannot be replaced by a fixed action library. In addition, the embodiments of this application can store the player's historical behavior data of the AI ​​teammate on the server or locally on the client to reduce the server load, but the historical memory cannot be read when logging in across devices. Regardless of the storage location, personalized dialogue generation must be based on the LLM's understanding and reorganization of historical memory and cannot replace the dynamic generation of the LLM with template filling. In addition to detecting emotions through player actions, this embodiment can also incorporate voice tone analysis (requiring the player to enable microphone access) or combat data input (such as knockdown count, number of injuries, and remaining teammates) as auxiliary input for emotion judgment. The emotion detection results must still be input into the LLM (Local Management Module), which generates the final emotion-adapted response; a fixed response library cannot be used. Furthermore, this embodiment can trigger topic rewind immediately after combat, or it can be set to wait for the player to actively engage in dialogue with the AI ​​before rewinding, avoiding interference while the player is still operating after combat. Regardless of the triggering timing, the topic content must be dynamically generated by the LLM based on previously saved context and cannot be replaced with fixed phrases.

[0171] In summary, the technical solutions provided in this application have the following beneficial effects: 1) Clear and orderly dialogue when multiple AI teammates are online at the same time: This application embodiment solves the channel flooding problem caused by multiple AI teammates replying to the same message at the same time through relevance judgment and silence mechanism. Only the relevant AI teammates generate a reply, while the other AI teammates remain silent, making the group chat dialogue clear and orderly, and close to the communication experience of real teammates.

[0172] 2) Precise command distribution and differentiated response styles: This application embodiment uses a command distribution matching mechanism with differentiated character settings to ensure that each command of the player is confirmed and responded to by the most suitable AI teammate, avoiding the confusion caused by multiple AI teammates confirming at the same time, and making the response style of each AI teammate recognizable.

[0173] 3) Uninterrupted immersion during combat: This application embodiment uses a combat event-driven dialogue interruption mechanism to make AI teammates automatically remain silent during intense player combat, without interfering with player operations. At the same time, through a topic rewind mechanism, the dialogue naturally resumes after the battle ends, maintaining immersion and continuity.

[0174] 4) Establishing emotional connection during the spawn island stage: This application embodiment uses a personalized memory dialogue mechanism on the spawn island to enable AI teammates to generate new topics based on historical memories each time they meet with the player, avoiding the repetition of fixed lines, establishing an emotional connection between the AI ​​teammates and the player, and enhancing the player's sense of companionship with the AI ​​teammates.

[0175] 5) Differentiated AI teammates bring an immersive experience: This application embodiment uses an emotion-adaptive response mechanism to enable different AI teammates to give different responses when facing the same player in the same emotional state. Each AI teammate has a distinct personality, which significantly enhances the sense of immersion.

[0176] 6) Reduce the production cost of AI teammate dialogue content: The embodiments of this application use LLM to dynamically generate most of the dialogue content, eliminating the need for planners to manually write response scripts for each scene. Each new AI teammate only requires supplementing the character description, which can save more than 70% of the copywriting production workload.

[0177] The following description continues to illustrate the exemplary structure of the virtual scene interaction processing device 555 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2 As shown, the software modules in the interactive processing device 555 of the virtual scene stored in the memory 550 may include: a display module 5551 and a control module 5552.

[0178] Display module 5551 is used to display a first virtual scene, which is a virtual scene during the virtual game operation phase. The first virtual scene includes player characters in the same team and multiple non-player characters, and the different non-player characters have different character characteristics. Display module 5551 is also used to respond to a message sending trigger operation and display a first message sent by a player character in the first virtual scene, where the first message is a broadcast message to multiple non-player characters. Control module 5552 is used to control the first non-player character to output a second message to reply to the first message, and to control other non-player characters to be in a silent state, where the first non-player character is the non-player character whose character characteristics match the first message the most among the multiple non-player characters, and the other non-player characters are the non-player characters other than the first non-player character among the multiple non-player characters.

[0179] In the above scheme, the display module 5551 is also used to display a second message sent by the first non-player character in the chat area of ​​the first virtual scene as a reply to the first message. The second message is displayed in the form of a message bubble, and the character identifier of the first non-player character is displayed in the message bubble.

[0180] In the above scheme, the display module 5551 is also used to display a second message sent by the first non-player character in response to the first message, using a display style different from the first message. The second message and the first message are displayed adjacent to each other in the chat area to represent the reply relationship of the second message to the first message. Alternatively, in the chat area, the reply relationship of the second message to the first message is indicated by a connection mark.

[0181] In the above scheme, the display module 5551 is also used to display, within a preset time after receiving the first message, a second message sent by the first non-player character in response to the first message in a dynamic expansion manner, wherein the dynamic expansion manner includes at least one of word-by-word display, gradual display, and bubble pop-up display.

[0182] In the above scheme, the display module 5551 is also used to display, within a preset time after receiving the first message, a second message sent by the first non-player character in reply to the first message in the form of a reference reply, wherein the reference reply form includes displaying at least part of the content of the first message, and at least part of the content is displayed above the second message or in the reference area inside the second message.

[0183] In the above scheme, the display module 5551 is also used to display a second message sent by the first non-player character in response to the first message within a preset time after receiving the first message, using an expression style that is adapted to the emotional state of the player controlling the player character. Different emotional states correspond to different message display styles.

[0184] In the above scheme, the emotional state includes at least one of an aggressive state, a conservative state, and an anxious state; the display module 5551 is also used to perform the following processing: when the player controlling the player character is in an aggressive state, display a second message sent by a first non-player character including a cooperative offensive response; when the player controlling the player character is in a conservative state, display a second message sent by a first non-player character including a defensive suggestion response; when the player controlling the player character is in an anxious state, display a second message sent by a first non-player character including a reassuring response.

[0185] In the above scheme, the virtual scene interaction processing device 555 also includes a playback module 5553, which is used in the first virtual scene to play the voice content corresponding to the second message used to reply to the first message, using a voice timbre pre-configured for the first non-player character. The voice timbre is configured differently for different non-player characters.

[0186] In the above scheme, the playback module 5553 is also used to play the voice content corresponding to the second message used to reply to the first message within a preset time after receiving the first message. The voice content is played at least one of the following: speech rate, tone, and pause mode, which is adapted to the emotional state of the player controlling the player character.

[0187] In the above scheme, when the playback module 5553 plays the voice content corresponding to the second message used to reply to the first message, the display module 5551 is also used to display the speech prompt information corresponding to the first non-player character in the first virtual scene, wherein the speech prompt information includes at least one of a speaking icon, a voice ripple icon, and a text prompt.

[0188] In the above scheme, the second message includes a dialogue reply message and a confirmation reply message; the control module 5552 is also used to control the first non-player character to output a dialogue reply message semantically related to the chat message when the first message is a chat message; and to control the first non-player character to output a confirmation reply message corresponding to the command message when the first message is a command message, and to control the first non-player character to perform corresponding cooperative behavior based on the command message, wherein the cooperative behavior includes at least one of target switching, path movement, attack cooperation, defense cooperation, healing support, and retreat evasion.

[0189] In the above scheme, before controlling the first non-player character to output a second message to reply to the first message, the control module 5552 is also used to control the first non-player character to enter an output prompt state, wherein the output prompt state includes at least one of replying prompt, thinking prompt, and speaking prompt; in response to the completion of the generation of the second message to reply to the first message, the step of controlling the first non-player character to output the second message to reply to the first message is executed.

[0190] In the above scheme, when the control module 5552 controls the first non-player character to output a second message to reply to the first message, it is also used to perform at least one of the following processes: controlling the first non-player character to perform a speaking action adapted to the second message, wherein the speaking action includes at least one of a nodding action, a waving action, and a body orientation adjustment action; controlling the first non-player character to display a speaking expression adapted to the second message, wherein the speaking expression includes at least one of a smiling expression, a serious expression, and a prompting expression; and highlighting the first non-player character, wherein the highlighting includes at least one of outline highlighting, partial illumination display, and character identifier highlighting.

[0191] In the above scheme, when the control module 5552 controls the first non-player character to output a second message in response to the first message, it is also used to control the first non-player character to pause outputting the second message in response to a battle event occurring in the first virtual scene, wherein the second message is a chat message; and to control the first non-player character to continue outputting the second message in response to the end of the battle event.

[0192] In the above scheme, when the control module 5552 controls the first non-player character to pause outputting the second message, the display module 5551 is also used to display a pending reply prompt in the chat area of ​​the first virtual scene, wherein the pending reply prompt includes at least one of a suspension indicator, a combat interruption indicator, and a reply retention indicator; when the control module 5552 controls the first non-player character to continue outputting the second message, the display module 5551 is also used to display a topic rewind prompt in the chat area of ​​the first virtual scene, wherein the topic rewind prompt includes at least a portion of the content of the first message, the preceding topic keywords, and a reply prompt indicator.

[0193] In the above scheme, before displaying the first virtual scene, the display module 5551 is also used to display the second virtual scene, wherein the second virtual scene is the virtual scene of the virtual game preparation stage; the control module 5552 is also used to control the second non-player character to output an opening message in the second virtual scene, wherein the opening message is generated based on the player character's historical behavior data, and the second non-player character is any one of multiple non-player characters; the display module 5551 is also used to respond to the message sending trigger operation and display the third message sent by the player character to reply to the opening message in the chat area of ​​the second virtual scene; the control module 5552 is also used to control the second non-player character to continue to output a fourth message to reply to the third message.

[0194] In the above scheme, the virtual scene interaction processing device 555 further includes a feature extraction module 5554, a feature encoding module 5555, and a determination module 5556. Before the control module 5552 controls the first non-player character to output a second message to reply to the first message, the feature extraction module 5554 is used to extract features from the first message to obtain the message feature vector corresponding to the first message; the feature encoding module 5555 is used to encode the features of multiple roles corresponding to multiple non-player characters to obtain the role feature vector corresponding to each non-player character; the determination module 5556 is used to determine the initial matching score of each non-player character for the first message based on the similarity between the message feature vector and the feature vector of each role; based on the initial matching score, and combined with at least one of the role function matching score, scene state matching score, and historical interaction matching score, the comprehensive matching score corresponding to each non-player character is determined; and the non-player character with the highest comprehensive matching score among the multiple non-player characters is selected as the first non-player character.

[0195] In the above scheme, before the control module 5552 controls the first non-player character to output a second message to reply to the first message, the determining module 5556 is also used to determine the target emotion tag that the first non-player character needs to use when replying to the first message; the virtual scene interaction processing device 555 also includes a generation module 5557 and an input module 5558, wherein the generation module 5557 is used to generate prompt words based on the target emotion tag, the first message, the character characteristics of the player character, and the character characteristics of the first non-player character; the input module 5558 is used to input the prompt words into a pre-trained language model so that the pre-trained language model generates a second message to reply to the first message.

[0196] In the above scheme, the determining module 5556 is further used to obtain the current emotional state coordinates of the first non-player character, and to determine the emotional transfer amount based on the difference between the current emotional state coordinates and the player's emotional state coordinates. The player's emotional state coordinates are the coordinates corresponding to the player's emotional state in a preset emotional space, which includes a pleasure dimension and an activation dimension. The emotional transfer amount includes a first transfer amount in the pleasure dimension and a second transfer amount in the activation dimension. When both the first and second transfer amounts are less than or equal to a preset first transfer amount threshold, the emotional tag corresponding to the current emotional state coordinates is used as the target emotion required by the first non-player character when replying to the first message. Emotional tags; when both the first and second migration amounts are greater than the preset second migration amount threshold, the emotion tag corresponding to the current emotional state coordinates is used as the target emotion tag for the first non-player character to reply to the first message; when both the first and second migration amounts are greater than the first migration amount threshold and less than the second migration amount threshold, the emotion migration amount is modulated in combination with the character characteristics of the first non-player character, and the current emotional state coordinates are updated based on the modulated emotion migration amount to obtain the updated emotional state coordinates, and the emotion coordinates corresponding to the updated emotional state coordinates in the preset emotion tag set are used as the target emotion tag for the first non-player character to reply to the first message.

[0197] In the above scheme, the determining module 5556 is also used to, when a preset event is detected in the first virtual scene, use the event emotion tag corresponding to the preset event as the target emotion tag to be used by the first non-player character when replying to the first message.

[0198] In the above scheme, after the input module 5558 inputs the prompt word into the pre-trained language model so that the pre-trained language model outputs a second message to reply to the first message, the control module 5552 is also used to control the emotional state of the first non-player character to gradually recover from the first emotional state to the second emotional state based on the time decay mechanism after a preset time. Here, the first emotional state is the emotional state corresponding to the target emotional label, and the second emotional state is the default emotional state of the first non-player character without the influence of external events.

[0199] It should be noted that the description of the apparatus in this application embodiment is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment, therefore, it will not be repeated. For any technical details not covered in the virtual scene interactive processing apparatus provided in this application embodiment, please refer to... Figure 3 , Figure 4 ,or Figure 5 The meaning is understood in accordance with the description of any of the accompanying drawings.

[0200] This application provides a computer program product, which includes a computer program or computer-executable instructions. The processor of an electronic device reads the computer-executable instructions from a computer-readable storage medium, and executes the computer-executable instructions, causing the electronic device to perform the virtual scene interaction processing method described above in this application.

[0201] This application provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by a processor, they cause the processor to execute the interactive processing method for a virtual scene provided in this application. For example, ... Figure 3 , Figure 4 ,or Figure 5 The interactive processing method for the virtual scene is shown.

[0202] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0203] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0204] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be part of a file of other programs or data, such as in one or more scripts stored in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files storing one or more modules, subroutines, or code sections).

[0205] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0206] The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for interactive processing of virtual scenes, characterized in that, The method includes: The first virtual scene is displayed, wherein the first virtual scene is a virtual scene in the virtual game operation stage, and the first virtual scene includes player characters in the same team and multiple non-player characters, and the different non-player characters have different character characteristics. In response to a message sending trigger operation, a first message sent by the player character is displayed in the first virtual scene, wherein the first message is a broadcast message to the plurality of non-player characters; The system controls a first non-player character to output a second message in response to the first message, and controls other non-player characters to remain silent. The first non-player character is the non-player character whose character characteristics match the first message the most among the plurality of non-player characters, and the other non-player characters are the non-player characters other than the first non-player character among the plurality of non-player characters.

2. The method according to claim 1, characterized in that, The control of the first non-player character to output a second message in response to the first message includes: In the chat area of ​​the first virtual scene, a second message sent by the first non-player character in response to the first message is displayed. The second message is displayed in the form of a message bubble, and the character identifier of the first non-player character is displayed in the message bubble.

3. The method according to claim 2, characterized in that, The display of the second message sent by the first non-player character in response to the first message includes: A second message sent by a first non-player character in response to the first message is displayed using a different display style than the first message. The second message and the first message are displayed adjacent to each other in the chat area to represent the reply relationship of the second message to the first message. Alternatively, the reply relationship of the second message to the first message is indicated by a connection mark in the chat area.

4. The method according to claim 2, characterized in that, The display of the second message sent by the first non-player character in response to the first message includes: Within a preset time period after receiving the first message, a second message sent by the first non-player character in response to the first message is displayed in a dynamically expanded manner, wherein the dynamically expanded manner includes at least one of word-by-word display, gradual display, and bubble pop-up display.

5. The method according to claim 2, characterized in that, The display of the second message sent by the first non-player character in response to the first message includes: Within a preset time period after receiving the first message, a second message sent by the first non-player character in response to the first message is displayed in the form of a reference reply. The reference reply form includes displaying at least a portion of the content of the first message, which is displayed above or in a reference area within the second message.

6. The method according to claim 2, characterized in that, The display of the second message sent by the first non-player character in response to the first message includes: Within a preset time period after receiving the first message, a second message sent by a first non-player character to reply to the first message is displayed, using an expression style adapted to the emotional state of the player controlling the player character. Different emotional states correspond to different message display styles.

7. The method according to claim 6, characterized in that, The emotional state includes at least one of an aggressive state, a conservative state, and an anxious state; The second message, sent by the first non-player character in response to the first message, is displayed using an expression style adapted to the emotional state of the player controlling the player character, including: When the player controlling the player character is in the aggressive state, a second message containing a coordinated offensive response sent by the first non-player character is displayed; When the player controlling the player character is in the conservative state, a second message containing defensive advice is displayed, sent by the first non-player character. When the player controlling the player character is in the anxious state, a second message containing reassuring replies is displayed, sent by a first non-player character.

8. The method according to claim 1, characterized in that, The control of the first non-player character to output a second message in response to the first message includes: In the first virtual scene, a voice content corresponding to the second message used to reply to the first message is played using a voice timbre pre-configured for the first non-player character. The voice timbre is different for different non-player characters.

9. The method according to claim 8, characterized in that, The playback of the voice content corresponding to the second message used to reply to the first message includes: Within a preset time period after receiving the first message, the voice content corresponding to the second message used to reply to the first message is played, wherein the voice content is played at least one of the following: speech rate, tone, and pauses that are adapted to the emotional state of the player controlling the player character.

10. The method according to claim 8, characterized in that, When playing the voice content corresponding to the second message used to reply to the first message, the method further includes: Display speech prompts corresponding to the first non-player character in the first virtual scene, wherein the speech prompts include at least one of a speech icon, a voice ripple icon, and a text prompt.

11. The method according to claim 1, characterized in that, The second message includes a dialogue reply message and a confirmation reply message; The control of the first non-player character to output a second message in response to the first message includes: When the first message is a casual chat message, control the first non-player character to output the dialogue reply message that is semantically related to the casual chat message; When the first message is a command message, the first non-player character is controlled to output the confirmation reply message corresponding to the command message, and the first non-player character is controlled to perform the corresponding cooperative behavior based on the command message. The cooperative behavior includes at least one of target switching, path movement, attack cooperation, defense cooperation, healing support, and retreat evasion.

12. The method according to claim 1, characterized in that, Before controlling the first non-player character to output a second message in response to the first message, the method further includes: Control the first non-player character to enter an output prompt state, wherein the output prompt state includes at least one of the following: replying prompt, thinking prompt, and speaking prompt; In response to the completion of the generation of the second message used to reply to the first message, the process proceeds to the step of controlling the first non-player character to output the second message used to reply to the first message.

13. The method according to claim 1, characterized in that, When the first non-player character outputs a second message in response to the first message, the method further includes: Perform at least one of the following processes: Control the first non-player character to perform a speaking action adapted to the second message, wherein the speaking action includes at least one of a nodding action, a waving action, and a body orientation adjustment action; Control the first non-player character to display a speech emoticon that is adapted to the second message, wherein the speech emoticon includes at least one of a smiling emoticon, a serious emoticon, and a prompting emoticon; The first non-player character is highlighted, wherein the highlighting includes at least one of outline highlighting, local illumination display, and character identifier highlighting.

14. The method according to claim 1, characterized in that, When the first non-player character outputs a second message in response to the first message, the method further includes: In response to a combat event occurring in the first virtual scene, the first non-player character is controlled to pause outputting the second message, wherein the second message is a chat message; In response to the end of the combat event, control the first non-player character to continue outputting the second message.

15. The method according to claim 14, characterized in that, When controlling the first non-player character to pause outputting the second message, the method further includes: Display a pending reply message in the chat area of ​​the first virtual scene, wherein the pending reply message includes at least one of a suspension indicator, a battle interruption indicator, and a reply hold indicator; When controlling the first non-player character to continue outputting the second message, the method further includes: In the chat area of ​​the first virtual scene, a topic retracing prompt is displayed, wherein the topic retracing prompt includes at least a portion of the content of the first message, keywords of the preceding topic, and at least one of the reply prompt identifier.

16. The method according to claim 1, characterized in that, Before displaying the first virtual scene, the method further includes: Display a second virtual scene, which is a virtual scene during the virtual game preparation phase; In the second virtual scene, a second non-player character is controlled to output an opening message, wherein the opening message is generated based on the player character's historical behavior data, and the second non-player character is any one of the plurality of non-player characters; In response to the message sending trigger operation, a third message sent by the player character to reply to the opening message is displayed in the chat area of ​​the second virtual scene, and the second non-player character is controlled to continue to output a fourth message to reply to the third message.

17. The method according to any one of claims 1 to 16, characterized in that, Before controlling the first non-player character to output a second message in response to the first message, the method further includes: Feature extraction is performed on the first message to obtain the message feature vector corresponding to the first message; The feature encoding is performed on the multiple character features corresponding to the multiple non-player characters to obtain the character feature vector corresponding to each of the multiple non-player characters; Based on the similarity between the message feature vector and the feature vectors of each character, an initial matching score is determined for each non-player character for the first message; Based on the initial matching score, and combined with at least one of the role function matching score, scene state matching score, and historical interaction matching score, the comprehensive matching score corresponding to each of the non-player roles is determined; Among the multiple non-player characters, the non-player character with the highest overall matching score is selected as the first non-player character.

18. The method according to any one of claims 1 to 16, characterized in that, Before controlling the first non-player character to output a second message in response to the first message, the method further includes: Determine the target emotion tag that the first non-player character should use when replying to the first message; Based on the target emotion tag, the first message, the character characteristics of the player character, and the character characteristics of the first non-player character, a prompt word is generated; The prompt word is input into a pre-trained language model so that the pre-trained language model generates a second message to reply to the first message.

19. The method according to claim 18, characterized in that, The step of determining the target emotion tag to be used by the first non-player character when replying to the first message includes: Obtain the current emotional state coordinates of the first non-player character, and determine the emotional transfer amount based on the difference between the current emotional state coordinates and the player's emotional state coordinates. The player's emotional state coordinates are the coordinates of the player's emotional state in a preset emotional space, which includes a pleasure dimension and an activation dimension. The emotional transfer amount includes a first transfer amount in the pleasure dimension and a second transfer amount in the activation dimension. When both the first migration amount and the second migration amount are less than or equal to the preset first migration amount threshold, the emotion tag corresponding to the current emotion state coordinates will be used as the target emotion tag for the first non-player character to use when replying to the first message. When both the first migration amount and the second migration amount are greater than the preset second migration amount threshold, the emotion tag corresponding to the current emotion state coordinates will be used as the target emotion tag for the first non-player character to reply to the first message. When both the first migration amount and the second migration amount are greater than the first migration amount threshold and less than the second migration amount threshold, the emotion migration amount is modulated in combination with the character characteristics of the first non-player character, and the current emotion state coordinates are updated based on the modulated emotion migration amount to obtain the updated emotion state coordinates. The emotion coordinates corresponding to the updated emotion state coordinates in the preset emotion tag set are used as the target emotion tags that the first non-player character needs to use when replying to the first message.

20. The method according to claim 18, characterized in that, The step of determining the target emotion tag to be used by the first non-player character when replying to the first message includes: When a preset event is detected in the first virtual scene, the event emotion tag corresponding to the preset event is used as the target emotion tag that the first non-player character needs to use when replying to the first message.

21. The method according to claim 18, characterized in that, After inputting the prompt word into a pre-trained language model to output a second message in response to the first message, the method further includes: After a preset time period, the emotional state of the first non-player character is gradually restored from the first emotional state to the second emotional state based on the time decay mechanism. The first emotional state is the emotional state corresponding to the target emotional tag, and the second emotional state is the default emotional state of the first non-player character without the influence of external events.

22. An interactive processing device for a virtual scene, characterized in that, The device includes: The display module is used to display a first virtual scene, wherein the first virtual scene is a virtual scene during the virtual game operation phase, and the first virtual scene includes player characters in the same team and multiple non-player characters, and the different non-player characters have different character characteristics. The display module is further configured to respond to a message sending trigger operation and display a first message sent by the player character in the first virtual scene, wherein the first message is a broadcast message to the plurality of non-player characters; The control module is used to control the first non-player character to output a second message in response to the first message, and to control other non-player characters to remain silent. The first non-player character is the non-player character whose character characteristics match the first message the most among the plurality of non-player characters, and the other non-player characters are the non-player characters other than the first non-player character among the plurality of non-player characters.

23. An electronic device, characterized in that, include: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the interactive processing method of the virtual scene according to any one of claims 1 to 21.

24. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by the processor, they implement the interactive processing method of the virtual scene according to any one of claims 1 to 21.

25. A computer program product comprising a computer program or computer-executable instructions, characterized in that, When the computer program or computer-executable instructions are executed by the processor, the interactive processing method of the virtual scene as described in any one of claims 1 to 21 is implemented.