Voice processing method and apparatus for virtual scene, information processing method and apparatus for virtual scene, electronic device, storage medium, and computer program product

By automatically switching voice modes in real time based on environmental changes, the problem of voice modes not matching the environment in games is solved, improving the efficiency of voice mode switching and immersive experience, and reducing system resource consumption and latency.

WO2026081682A1PCT designated stage Publication Date: 2026-04-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-08-28
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing technologies cannot achieve precise adaptation between voice modes and virtual scenes in games, resulting in poor immersive perception, high system resource consumption, and frequent latency.

Method used

By sensing environmental changes in real time, the system automatically switches to a voice mode that adapts to the current environment of the virtual object, reducing reliance on preset voice packs and directly matching the voice mode based on the environment, thus reducing the system's processing burden.

Benefits of technology

It improves the efficiency and convenience of voice mode switching, reduces memory usage, lowers latency, enhances the adaptability of voice mode to the environment, and strengthens the immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025117465_23042026_PF_FP_ABST
    Figure CN2025117465_23042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a voice processing method and apparatus for a virtual scene, an information processing method and apparatus for a virtual scene, an electronic device, a storage medium, and a computer program product. The method comprises: displaying a virtual scene, wherein a target account in the virtual scene is in a first voice mode adapted to a first influence factor, and the first influence factor at least comprises a first environment in which a virtual object controlled by the target account is located; and in response to the first influence factor being switched to a second influence factor, switching the target account from the first voice mode to a second voice mode adapted to the second influence factor, wherein the second influence factor at least comprises a second environment in which the virtual object currently controlled by the target account is located.
Need to check novelty before this filing date? Find Prior Art

Description

Voice processing methods, information processing methods, devices, electronic devices, storage media, and computer program products for virtual scenarios

[0001] Cross-reference to related applications

[0002] This application is based on and claims priority to Chinese Patent Application No. 2024114527485, filed on October 16, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of computer technology, and in particular to a voice processing method, information processing method, device, electronic device, storage medium, and computer program product for a virtual scene. Background Technology

[0004] During gameplay, users often need to interact with each other via voice. Related technologies not only support users interacting with their own voices but also allow them to change their voices to disguise their identities and increase the fun of the interaction. However, the voice-changing solutions of these technologies cannot adapt to the complex and ever-changing situations in games, affecting the user's immersive experience.

[0005] Related technologies typically change the voice based on the character's voice pack used in the game, or use a voice corresponding to the emotion expressed in the input content. When the scene in which the user-controlled character is located changes, only the ambient sound in the scene can be changed, resulting in a voice effect that doesn't match the current scene and causing a sense of disconnect for the user during gameplay. Furthermore, these technologies rely on preset character voice packs, requiring the loading of new voice packs when switching scenes, which can easily cause excessive instantaneous memory usage and introduce latency during loading. Matching the voice to the emotion expressed in the input content requires additional text sentiment analysis and other computational steps, increasing the system's processing burden and impacting response efficiency. Summary of the Invention

[0006] In view of this, embodiments of this application provide a voice processing method, information processing method, device, electronic device, storage medium, and computer program product for virtual scenes, which can realize a voice mode that is precisely adapted to the virtual scene, so as to enhance the immersive perception effect when conducting voice interaction in the virtual scene.

[0007] The technical solution of this application embodiment is implemented as follows:

[0008] This application provides a voice processing method for a virtual scene, the method being executed by an electronic device, the method comprising:

[0009] Displaying a virtual scene, wherein the target account in the virtual scene is in a first voice mode adapted to a first influencing factor, and the first influencing factor includes at least a first environment in which the virtual object controlled by the target account is located;

[0010] In response to the first influencing factor switching to the second influencing factor, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second influencing factor includes at least the second environment in which the virtual object currently controlled by the target account is located.

[0011] This application provides an information processing method for a virtual scene, the method being executed by an electronic device, the method comprising:

[0012] Displays a virtual scene based on login to the target account, wherein the virtual scene includes a message editing control;

[0013] In response to a first message input operation in the message editing control, the first input information is displayed;

[0014] In response to the first message output operation, the voice signal of the first input information is played in the virtual scene based on a first voice mode adapted to a first influencing factor of the target account, wherein the first influencing factor includes at least a first environment in which the virtual object controlled by the target account is currently located.

[0015] This application provides a voice processing device for a virtual scene, the device comprising:

[0016] The first display module is configured to display a virtual scene, wherein the target account in the virtual scene is in a first voice mode adapted to a first influencing factor, and the first influencing factor includes at least a first environment in which the virtual object controlled by the target account is located;

[0017] The mode switching module is configured to switch the target account from the first voice mode to a second voice mode adapted to the second voice mode in response to the first influencing factor switching to the second influencing factor, wherein the second influencing factor includes at least the second environment in which the virtual object currently controlled by the target account is located.

[0018] This application provides an information processing device for a virtual scene, the device comprising:

[0019] The second display module is configured to display a virtual scene based on the target account login, wherein the virtual scene includes a message editing control;

[0020] The third display module is configured to display first input information in response to a first message input operation in the message editing control;

[0021] The signal playback module is configured to respond to a first message output operation and play the voice signal of the first input information in the virtual scene based on a first voice mode adapted to a first influencing factor of the target account, wherein the first influencing factor includes at least a first environment in which the virtual object controlled by the target account is currently located.

[0022] This application provides an electronic device, the electronic device comprising:

[0023] Memory is used to store executable instructions or computer programs.

[0024] When a processor executes computer-executable instructions or computer programs stored in the memory, it implements the voice processing method for a virtual scene or the information processing method for a virtual scene provided in the embodiments of this application.

[0025] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the voice processing method or information processing method for a virtual scene provided in this application.

[0026] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the voice processing method for a virtual scene or the information processing method for a virtual scene provided in this application.

[0027] The embodiments of this application have the following beneficial effects:

[0028] When the environment of a virtual object controlled by the target account changes, the system adaptively switches to a voice mode that is at least compatible with the current environment. This achieves automatic voice mode switching, improving efficiency and convenience. Furthermore, by adapting the voice mode to the environment, the perceived effect of the voice signal played based on the voice mode blends with the perceived effect of the environment, facilitating an immersive experience for the player in the virtual environment. Since it does not rely on a large number of preset voice packs, and switches to the corresponding voice mode in real time, it reduces memory usage and latency caused by voice pack loading, ensuring real-time voice mode switching. Simultaneously, directly matching the voice mode based on the environment saves unnecessary intermediate calculations, improving the system's efficiency in adapting the voice mode to the current environment and reducing the system's processing burden. Attached Figure Description

[0029] Figure 1 is a schematic diagram of the architecture of the virtual scene voice processing system 100 provided in an embodiment of this application;

[0030] Figure 2A is a structural schematic diagram of the terminal 400-1 provided in an embodiment of this application;

[0031] Figure 2B is a structural schematic diagram of the terminal 400-2 provided in an embodiment of this application;

[0032] Figure 3A is a schematic diagram of the first process of the voice processing method for a virtual scene provided in an embodiment of this application;

[0033] Figure 3B is a second flowchart illustrating the voice processing method for a virtual scene provided in an embodiment of this application;

[0034] Figure 3C is a schematic diagram of the third process of the voice processing method for a virtual scene provided in the embodiments of this application;

[0035] Figure 3D is a schematic diagram of the fourth process of the voice processing method for a virtual scene provided in the embodiments of this application;

[0036] Figure 3E is a schematic diagram of the fifth process of the voice processing method for virtual scenes provided in the embodiments of this application;

[0037] Figure 3F is a schematic diagram of the sixth process of the voice processing method for a virtual scene provided in the embodiments of this application;

[0038] Figure 3G is a schematic diagram of the seventh process of the voice processing method for a virtual scene provided in the embodiments of this application;

[0039] Figure 4A is a schematic diagram of the first process of the information processing method for a virtual scene provided in an embodiment of this application;

[0040] Figure 4B is a schematic diagram of the second process of the information processing method for a virtual scene provided in an embodiment of this application;

[0041] Figure 4C is a schematic diagram of the third process of the information processing method for a virtual scene provided in an embodiment of this application;

[0042] Figure 4D is a schematic diagram of the fourth process of the information processing method for virtual scenes provided in the embodiments of this application;

[0043] Figure 5A is a schematic diagram illustrating the training principle of the text sentiment recognition model provided in the embodiments of this application;

[0044] Figure 5B is a schematic diagram illustrating the training principle of the text classification model provided in the embodiments of this application;

[0045] Figure 5C is a schematic diagram of the training principle of the machine learning model provided in the embodiments of this application;

[0046] Figure 6A is a schematic diagram of switching voice modes according to environment and role provided in an embodiment of this application;

[0047] Figure 6B is a schematic diagram of switching voice modes according to environment and attributes provided in an embodiment of this application;

[0048] Figure 6C is a schematic diagram of switching voice modes according to environmental and state parameters provided in an embodiment of this application;

[0049] Figure 6D is a schematic diagram of switching voice modes according to environment and input information provided in an embodiment of this application;

[0050] Figure 6E is a first schematic diagram of switching from a third voice mode to a first voice mode according to an embodiment of this application;

[0051] Figure 6F is a second schematic diagram of switching from a third voice mode to a first voice mode provided in an embodiment of this application;

[0052] Figure 6G is a schematic diagram of applying the corresponding voice mode according to the input information provided in an embodiment of this application;

[0053] Figure 7 is a schematic diagram of voice mode switching in the game provided in an embodiment of this application;

[0054] Figure 8 is a schematic diagram of the process for generating speech effects provided in an embodiment of this application.

[0055] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0057] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0058] In the following description, the terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0059] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0060] Unless otherwise specified, "at least one" as used below refers to one or more cases, and "multiple" can refer to two or more cases.

[0061] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for descriptive purposes only and is not intended to limit the scope of this application.

[0062] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0063] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0064] 1) Voice mode: This controls the sound effects used when the user outputs voice signals in a virtual scene. Different voice modes are used to output voice signals with different sound effects, and these different sound effects can be distinguished by the four elements of sound.

[0065] The elements of sound are mainly divided into four basic physical properties, which are:

[0066] 1. Pitch: This refers to the highness or lowness of a sound, and it depends on the frequency of the sound wave. The higher the frequency, the higher the pitch; the lower the frequency, the lower the pitch. For example, small, thin, short, and compact objects vibrate quickly, have a high frequency, and therefore a high pitch.

[0067] 2. Sound intensity: also known as loudness, refers to the strength or volume of a sound, which depends on the amplitude of the sound wave. The larger the amplitude, the greater the sound intensity, and the louder it sounds; the smaller the amplitude, the smaller the sound intensity, and the softer it sounds.

[0068] 3. Duration: This refers to the length of a sound, that is, the duration of the vibration of the sound-producing body. Duration is usually not the primary means of distinguishing meaning in Chinese, but it is an important natural attribute.

[0069] 4. Timbre: Also known as tone color, it refers to the essential characteristics of a sound, the most fundamental feature that distinguishes one sound from others. Timbre depends on the form of the sound waves during pronunciation; different sound wave forms produce different timbres.

[0070] These four elements together determine the physical properties of sound. In daily life, we use these four elements to identify and understand different sounds.

[0071] 2) Influencing factors refer to factors that affect the voice pattern. For example, influencing factors include at least the environment in which the virtual object controlled by the target account is located, and may also include at least one of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the current input information of the target account.

[0072] 3) In response, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.

[0073] 4) Audio effects model, which is used to generate a voice signal that conforms to a specific emotion type based on the input information (which can be in the form of text or speech). The emotion type can include any one of the following: happy, angry, sad, tense, depressed, fearful, calm, curious or disgusted.

[0074] 5) Modulation method refers to the way the speech signal is modulated according to the attributes of the virtual object and the environment. The modulation method is determined by the frequency, amplitude, formants of the speech signal and the simulation of environmental reflections.

[0075] 6) Update method refers to the way input information is updated based on the role characteristics of the virtual object when it moves from the first environment to the second environment. Different roles correspond to different update methods.

[0076] 7) The first attribute of a virtual object refers to the basic characteristics of the virtual object controlled by the target account in the first environment, such as actions, emotions, or states. Different combinations of the first environment and the first attribute correspond to different first voice effects.

[0077] 8) The first state parameter of the target account refers to the parameters of the current state of the target object logged into the target account in the virtual scene, including the target object's actions or expressions. Different combinations of the first environment and the first state correspond to different first voice effects.

[0078] 9) Human-computer interaction interface, which is used to provide human-computer interaction functions / display virtual scenes.

[0079] For example, graphical user interfaces (GUIs) include augmented reality (AR) interfaces, virtual reality (VR) interfaces, voice user interfaces (VUIs), interactive projection interfaces (using projection technology to display information on a flat surface), eye-tracking interfaces (interfaces controlled by detecting the user's gaze), holographic interfaces (three-dimensional holograms formed by projecting images using holographic projection technology, allowing users to see stereoscopic images without wearing special glasses), multimodal interfaces (interfaces that combine multiple interaction methods, such as tactile, visual, and auditory interaction), and brain-machine interfaces (BMIs).

[0080] 10) In response to, used to indicate the conditions or states on which the operation performed depends, when the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay; unless otherwise specified, there is no restriction on the order in which the multiple operations performed are executed.

[0081] When voice chat is enabled in a game, users often want to change the sound of their voice. The voice processing methods of related technologies usually change the voice based on the character's voice pack used by the user in the game, or use a voice that corresponds to the emotion expressed in the input content. When the scene in which the user-controlled character is located changes, only the ambient sound in the scene can be changed. The voice effect does not match the current scene, causing the user to feel disconnected during the game.

[0082] Based on the above analysis, the applicant found that the voice processing methods for virtual scenes in related technologies cannot match the voice pattern with the environment. In order to address the above problem, this application provides a voice processing method for virtual scenes that can achieve a voice pattern that is accurately adapted to the virtual scene, so as to enhance the immersive perception effect when conducting voice interaction in the virtual scene.

[0083] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. The following will describe exemplary applications when the electronic device is implemented as a terminal.

[0084] Referring to Figure 1, which is a schematic diagram of the architecture of a virtual scene voice processing system 100 provided in an embodiment of this application, in order to support a voice processing application for a virtual scene, the terminal 400 connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of both.

[0085] Terminal 400 is used to display a virtual scene on graphical interface 410, wherein the target account in the virtual scene is in a first voice mode adapted to a first influencing factor, the first influencing factor including at least a first environment in which the virtual object controlled by the target account is located; in response to the first influencing factor switching to a second influencing factor, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second influencing factor includes at least a second environment in which the virtual object currently controlled by the target account is located.

[0086] Taking a game scenario as an example, terminal 400 displays a virtual scene on graphical interface 410. In response to the switch from a first influencing factor to a second influencing factor, terminal 400 sends the second influencing factor to server 200, so that server 200 determines a second voice mode based on the second influencing factor and returns the second voice mode to terminal 400. Terminal 400 then plays the voice signal of the target account's input information according to the second voice mode.

[0087] In some embodiments, server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.

[0088] Referring to Figure 2A, which is a schematic diagram of the structure of terminal 400-1 provided in an embodiment of this application, terminal 400-1 is one implementation of the aforementioned terminal 400 for processing voice in a virtual scene. Terminal 400-1 shown in Figure 2A includes: at least one processor 411, a memory 415, at least one network interface 412, and a user interface 413. The various components in terminal 400-1 are coupled together via a bus system 414. It is understood that the bus system 414 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 414 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 414 in Figure 2A.

[0089] Processor 411 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0090] User interface 413 includes one or more output devices 4131 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 413 also includes one or more input devices 4132, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0091] The memory 415 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 415 may optionally include one or more storage devices physically located away from the processor 411.

[0092] Memory 415 may include volatile memory or non-volatile memory, or both. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 415 described in this application embodiment is intended to include any suitable type of memory.

[0093] In some embodiments, memory 415 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0094] Operating system 4151 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks.

[0095] The network communication module 4152 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 412, exemplary network interfaces 412 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0096] Presentation module 4153 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 4131 associated with user interface 413 (e.g., a display screen, a speaker, etc.).

[0097] The input processing module 4154 is used to detect and translate one or more user inputs or interactions from one or more input devices 4132.

[0098] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2A shows a voice processing apparatus 4155 for a virtual scene stored in memory 415. It can be software in the form of programs and plug-ins, including the following software modules: a first display module 41551 and a mode switching module 41552. These modules are logically related and can therefore be arbitrarily combined or further split according to the functions they implement. The functions of each module will be described below.

[0099] Referring to Figure 2B, which is a schematic diagram of the structure of terminal 400-2 provided in an embodiment of this application, terminal 400-2 is one implementation of the aforementioned terminal 400 for processing information in a virtual scene. Terminal 400-2 shown in Figure 2B includes: at least one processor 421, a memory 425, at least one network interface 422, and a user interface 423. The various components in terminal 400-1 are coupled together via a bus system 424. It is understood that the bus system 424 is used to implement communication between these components. In addition to a data bus, the bus system 424 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 424 in Figure 2B.

[0100] Processor 421 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0101] User interface 423 includes one or more output devices 4231 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 423 also includes one or more input devices 4232, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0102] The memory 425 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 425 may optionally include one or more storage devices physically located away from the processor 421.

[0103] Memory 425 may include volatile memory or non-volatile memory, or both. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 425 described in this application embodiment is intended to include any suitable type of memory.

[0104] In some embodiments, memory 425 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0105] Operating system 4251 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0106] The network communication module 4252 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 422, exemplary network interfaces 422 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0107] Presentation module 4253 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 4231 (e.g., a display screen, a speaker, etc.) associated with user interface 423.

[0108] The input processing module 4254 is used to detect and translate one or more user inputs or interactions from one or more input devices 4232.

[0109] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2B shows an information processing apparatus 4255 for a virtual scene stored in memory 425. This apparatus can be software in the form of programs and plugins, including the following software modules: a second display module 42551, a third display module 42552, and a signal playback module 42553. These modules are logically related and can therefore be arbitrarily combined or further divided according to their implemented functions. The functions of each module will be described below.

[0110] In some embodiments, the terminal or server can implement the voice processing method or information processing method for the virtual scene provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as game APPs; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0111] The speech processing method and information processing method for virtual scenes provided in this application will be described in conjunction with exemplary applications and implementations of the terminals provided in the embodiments of this application.

[0112] Referring to Figure 3A, which is a first flowchart of the voice processing method for a virtual scene provided in the embodiments of this application, the steps shown in Figure 3A will be described with the terminal as the main body.

[0113] In step 101, a virtual scene is displayed, wherein the target account in the virtual scene is in a first voice mode adapted to the first influencing factor, and the first influencing factor includes at least the first environment in which the virtual object controlled by the target account is located.

[0114] In some embodiments, in response to a message output operation for the first input information, the voice signal of the first input information is played in a virtual scene based on a first voice mode.

[0115] For example, if the first environment in which the virtual object controlled by the target account is located is a dark and quiet environment, then the first voice mode is a trembling, timid and tense voice effect. If the first input information is output in the virtual scene, then the voice signal of the first input information is played using a trembling, timid and tense voice effect.

[0116] In some embodiments, referring to FIG3B, FIG3B is a second flowchart of the voice processing method for a virtual scene provided in the embodiments of this application. When performing step 101 of FIG3A, or before performing step 101 of FIG3A, steps 201 to 203 of FIG3B can be performed, which are described in detail below.

[0117] In step 201, the influencing factor setting interface is displayed.

[0118] In some embodiments, the influencing factor settings interface includes the environment in which the currently controlled virtual object is located, and the environment is not editable, that is, the environment is always selected.

[0119] For example, the virtual scene includes an entry point to the influencing factors settings interface. In response to a trigger operation on this entry point, the influencing factors settings interface is displayed. The influencing factors settings interface can either completely cover the current virtual scene or float above it with a preset transparency. The size of the influencing factors settings interface can be equal to or smaller than the size of the currently displayed virtual scene.

[0120] In step 202, in response to the input operation in the influencing factor setting interface, at least one dimension of the input is displayed, wherein the at least one dimension is at least one of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the input information of the target account.

[0121] In some embodiments, a dimension input control can be displayed in the influencing factor setting interface. In response to an input operation in the dimension input control, at least one input dimension can be displayed. Alternatively, a selection box can be displayed before each dimension. In response to a selection operation on the selection box corresponding to the at least one dimension, the at least one input dimension can be displayed. For the at least one input dimension, it can be sorted according to a preset priority from high to low, and the at least one dimension can be displayed based on the sorting result. A dimension search control can also be displayed in the influencing factor setting interface. If there are multiple at least one dimension, in response to a search operation on any dimension in the dimension search control, the searched dimensions can be displayed.

[0122] As an example of dimensions, the virtual object controlled by the target account can belong to a variety of roles set in the game, such as a mage, archer, or warrior; the attributes of the virtual object controlled by the target account include at least one of the virtual object's actions, emotions, or states; the state of the target account includes the actions or expressions of the target object (user) logged into the target account; the type of input information of the target account can be voice or text.

[0123] In step 203, at least one dimension is combined with the first environment in which the virtual object currently controlled by the target account is located to form a first influencing factor.

[0124] In some embodiments, the first influencing factor may be any one of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attribute of the virtual object controlled by the target account, the state of the target account, and the input information of the target account, combined with the first environment in which the virtual object currently controlled by the target account is located; the first influencing factor may be any two of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attribute of the virtual object controlled by the target account, the state of the target account, and the input information of the target account, combined with the first environment; the first influencing factor may be any three of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attribute of the virtual object controlled by the target account, the state of the target account, and the input information of the target account, combined with the first environment; or the first influencing factor may be a combination of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attribute of the virtual object controlled by the target account, the state of the target account, and the input information of the target account.

[0125] For example, at least one dimension of the input is "the role (warrior) to which the virtual object belongs" and "the attribute (defense value 90) of the virtual object". The first environment in which the virtual object currently controlled by the target account is located is "mountain battlefield". The role to which the virtual object controlled by the target account belongs, the attribute of the virtual object controlled by the target account, and the first environment are combined into the first influencing factor: {environment: mountain battlefield, role: warrior, attribute: defense value 90}.

[0126] Here, at least one dimension of the first influencing factor, other than the first environment, can be changed to update the first influencing factor.

[0127] This application embodiment allows users to select at least one dimension in the influencing factor setting interface and combine it with the first environment to form a first influencing factor. This enables users to design the first influencing factor according to their own needs and the current virtual scene, which is more in line with users' usage habits and preferences and improves the applicability of the influencing factor.

[0128] Referring again to Figure 3A, in step 102, in response to the first influencing factor being switched to the second influencing factor, the target account is switched from the first voice mode to the second voice mode adapted to the second influencing factor, wherein the second influencing factor includes at least the second environment in which the virtual object currently controlled by the target account is located.

[0129] In some embodiments, the second voice mode may be temporarily generated based on the second influencing factor when the first influencing factor is switched to the second influencing factor; or it may be pre-generated and stored in a database, and when the first influencing factor is switched to the second influencing factor, the database is queried to find the second voice mode that matches the second influencing factor, and the target account is switched from the first voice mode to the second voice mode that is adapted to the second influencing factor.

[0130] Following the example of step 101 above, if the first environment in which the virtual object controlled by the target account is located is a dark and quiet environment, then the first voice mode is a trembling, timid and tense voice effect. If the second environment in which the virtual object controlled by the target account is located is a bright and cheerful environment, then the second voice mode is a happy, relaxed and clear voice effect. When switching from the first environment to the second environment, if the second input information corresponding to the second environment is output in the virtual scene, then the voice effect of the trembling, timid and tense voice will automatically switch to the voice effect of the happy, relaxed and clear voice and the voice signal of the second input information will be played.

[0131] In some embodiments, the first influencing factor further includes a first role to which the virtual object controlled by the target account belongs, and the second influencing factor further includes a second role to which the virtual object controlled by the target account belongs. Step 102 in Figure 3A can be implemented by performing the following process: in response to the role to which the virtual object controlled by the target account belongs switching from the first role to the second role, and the virtual object moving from the first environment to the second environment, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second voice mode is used to output a voice signal adapted to both the second environment and the second role.

[0132] Here, the voice mode will only switch from the first voice mode to the second voice mode, which is adapted to the second influencing factor, when the role of the virtual object controlled by the target account changes, or when the environment changes. Changes in other influencing factors will not cause the voice mode to switch. The second voice mode can also be used to output voice signals that are adapted to both the environmental characteristics of the second environment and the role characteristics of the second character.

[0133] For example, if the first role of the virtual object controlled by the target account is a mage, the mage's corresponding voice mode is a deep and slow voice effect, and the second role of the virtual object controlled by the target account is a traveler, the traveler's corresponding voice effect is a happy and relaxed voice effect, if the role of the virtual object controlled by the target account changes from mage to traveler, and the virtual object moves from the first environment to the second environment, then the voice effect corresponding to both the mage and the first environment will be switched to the voice effect corresponding to both the traveler and the second environment.

[0134] Here, the features of the second character are further broken down, and the first sub-features such as the character's age, skill type, and background are extracted. The features of the second environment are further broken down, and the second sub-features such as the type of environmental sound effects (e.g., wind, rain), space size (e.g., enclosed cabin, open plain), and environmental material (e.g., metal, wood) are extracted. The first voice parameter corresponding to the first sub-feature is queried from the mapping relationship database, and the second voice parameter corresponding to the second sub-feature is queried. The first and second voice parameters are combined into a third voice parameter. Based on the third voice parameter, a voice signal that is suitable for both the second environment and the second character is generated, which is the second voice mode.

[0135] As an example of steps 101 to 102, taking the first influencing factor as the first environment and the first role as an example, refer to Figure 6A. Figure 6A is a schematic diagram of switching voice modes according to the environment and role provided in this application embodiment. In the left side of Figure 6A, a virtual scene 610 is displayed, and an influencing factor setting interface 611 is displayed. The influencing factor setting interface 611 includes at least one dimension, such as the role 612 to which the virtual object controlled by the target account belongs, the attribute 613 of the virtual object controlled by the target account, the status 614 of the target account, and the input information 615 of the target account. The environment 616 where the virtual object is located is always selected. In response to the input operation in the influencing factor setting interface 611, such as the selection operation for the role 612 to which the virtual object controlled by the target account belongs, the role 612 to which the virtual object controlled by the target account belongs and the environment 616 where the virtual object currently controlled by the target account is located, i.e., the first environment, are combined into the first influencing factor. In the middle diagram of Figure 6A, a first environment 617 and a first role 618 are shown. At this time, the target account in the virtual scene is in a first voice mode adapted to the first influencing factor (i.e., the first environment 617 and the first role 618). In response to the first influencing factor switching to a second influencing factor, in the right diagram of Figure 6A, for example, the first environment 617 is switched to the second environment 619 and the first role 618 is switched to the second role 620, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor (i.e., the second environment 619 and the second role 620).

[0136] This application embodiment uses the first environment and the first role as the first influencing factors, and determines the second voice mode by switching between the environment and the role. This makes the switching of the voice mode consistent with the current environment and role, improving the continuity of the switching. Each role has its own specific voice mode, which increases the diversity and complexity of role-playing and enhances the user experience.

[0137] In some embodiments, referring to FIG3C, FIG3C is a third flowchart of the voice processing method for a virtual scene provided in the embodiments of this application. Before the above-mentioned "in response to the virtual object controlled by the target account switching its role from a first role to a second role, and the virtual object moving from a first environment to a second environment, switching the target account from a first voice mode to a second voice mode adapted to the second influencing factor", steps 301 to 304 of FIG3C are executed, which are described in detail below.

[0138] In step 301, the first input information of the target account is detected to obtain the first sentiment type of the first input information.

[0139] In some embodiments, a pre-trained text sentiment recognition model is invoked based on the first input information of the target account to obtain the first sentiment type of the first input information.

[0140] For example, a pre-trained text sentiment recognition model is trained by performing the following processes: obtaining an initialized text sentiment recognition model; obtaining information samples and real sentiment labels, where the real sentiment labels represent the sentiment type expressed by the information samples; calling the initialized text sentiment recognition model based on the information samples to obtain predicted sentiment labels; determining a sentiment loss value based on the real sentiment labels and predicted sentiment labels; and updating the parameters of the initialized text sentiment recognition model based on the sentiment loss value to obtain the pre-trained text sentiment recognition model.

[0141] For example, refer to Figure 5A, which is a schematic diagram of the training principle of the text sentiment recognition model provided in this application embodiment. In Figure 5A, based on information samples, the convolutional and fully connected layers of the initialized text sentiment recognition model can be called to obtain predicted sentiment labels. The sentiment loss value between the real sentiment label and the predicted sentiment label is calculated through a loss function. The sentiment loss value is backpropagated to update the parameters of the initialized text sentiment recognition model. The process of calculating the sentiment loss value and updating the parameters is repeated multiple times until the sentiment loss value no longer increases or decreases, at which point the iteration process stops, forming a pre-trained text sentiment recognition model.

[0142] For example, initializing the representation involves randomly assigning values ​​to the parameters of the text sentiment recognition model, such as assigning all parameters to 0 or all to 1. Backpropagation is implemented using the backpropagation algorithm, calculating the gradient of each neuron from the output layer to the input layer, and updating the neuron's weights and biases based on the gradients. Gradient descent is used to continuously update the parameters, reducing the loss value. Various gradient descent algorithms can be used, such as batch gradient descent, stochastic gradient descent, adaptive gradient descent, and momentum gradient descent. The loss function can be the mean squared error loss function, cross-entropy loss function, multi-label classification loss function, or triplet loss function.

[0143] In step 302, the first audio effect model corresponding to the first emotion type is determined.

[0144] Here, the first audio effect model is used to output a voice signal that matches the first emotion type, that is, a voice signal that matches the first input information.

[0145] In some embodiments, a model database is pre-set, which stores the association between different emotion types and corresponding audio effect models; the model database is used to query the first audio effect model corresponding to the first emotion type. One emotion type can correspond to one audio effect model, or multiple emotion types can be combined to correspond to one audio effect model. Furthermore, for the same first emotion type, different scenarios (such as battle scenarios or social scenarios) can correspond to different audio effect models.

[0146] For example, if the first emotion type is happiness, the model database is queried for the audio effect model corresponding to happiness, and that audio effect model is used as the first audio effect model.

[0147] In step 303, a first modulation method is determined for the first speech signal output by the first audio effect model, wherein the first modulation method includes modulating the first speech signal based on a second environment and a second role.

[0148] In some embodiments, the parameter settings of the first modulation method can be refined, such as adjusting the reverberation duration based on the spatial size of the second environment, adjusting the speech fundamental frequency based on the role age of the second character, and improving the modulation accuracy; a modulation effect feedback mechanism can be established to optimize the rules of the first modulation method based on the user's feedback on the modulated speech signal.

[0149] For example, if the second environment is a bright and sunny environment and the second character is a traveler, then the first modulation method is to add a happy, relaxed and clear voice effect to the first voice signal.

[0150] In step 304, a second speech mode is determined based on the first modulation scheme.

[0151] Here, the first modulation method can be directly used as the second voice mode.

[0152] In some embodiments, the first modulation scheme is decomposed into specific speech parameters, such as pitch adjustment value, reverberation duration, and sound effect overlay type. The decomposed speech parameters are combined with the basic parameters of the first audio effect model to generate a second speech pattern adapted to the second influencing factor, and stored in a speech pattern library for subsequent speech pattern switching. The system supports generating a second speech pattern based on the combination of multiple first modulation schemes. When the first input information includes multiple emotion types and corresponds to multiple first modulation schemes, multiple first modulation schemes can be fused to generate a comprehensive second speech pattern.

[0153] For example, the first modulation method, "increase the tone of voice to enhance penetration, increase the reverberation of the volcanic eruption sound effect, and enhance the roughness of the voice to suit the warrior character," is broken down into voice parameters such as "increase tone by 20%, reverberation duration of 1.2 seconds, and superimpose 50% volcanic eruption sound effect." These parameters are then combined with the basic parameters of the first audio effect model, "rapid rhythm sound effect model," to generate the second voice mode.

[0154] This application embodiment, through emotion detection of input information, can accurately identify the user's emotional state and customize audio effects according to the user's emotional type, providing a more personalized audio experience and making the interaction more closely aligned with the user's emotional state. Audio modulation based on a second environment and a second role allows the voice signal to better adapt to different environments and virtual characters, improving the naturalness and realism of the voice, ensuring consistency between the voice signal and the virtual character, enhancing the user's immersion in the virtual environment, dynamically adjusting audio effects, and providing a more dynamic and richer interactive experience.

[0155] In some embodiments, referring to FIG3D, FIG3D is a fourth flowchart of the voice processing method for a virtual scene provided in the embodiments of this application. Step 304 in FIG3C can be implemented by steps 3041 to 3042 in FIG3D, as described in detail below.

[0156] In step 3041, a first update method for the first input information of the target account is determined, wherein the first update method includes: encoding based on the first input information, the environmental characteristics of the second environment, and the role characteristics of the second role to obtain a first fusion feature, and decoding based on the first fusion feature to obtain the updated first input information.

[0157] In some embodiments, first input information features are extracted from the first input information. Using a multimodal fusion network, feature concatenation techniques, or a deep learning model, the first input information features, the environmental features of the second environment, and the role features of the second character are fused into a unified feature representation, namely the first fused feature. The first fused feature is then processed using a decoder. The decoding process can use a generative model, such as a Generative Adversarial Network (GAN), a Variational Autoencoder (VAE), or other sequence generation models, to decode the first fused feature and obtain the updated first input information.

[0158] For example, the second character's characteristic is using the catchphrase "meow~". If the second environment is a relaxing and peaceful seaside, and the first input information is "I'm so happy today", the catchphrase will be inserted at the end of the first input information or before a punctuation mark in the sentence to update the first input information. For example, the updated first input information could be "I'm so happy today, meow~".

[0159] In step 3042, the first modulation method and the first update method are used as the second voice mode.

[0160] Following the examples of steps 303 and 3041 above, a happy, relaxed, and clear voice effect, along with "meow~" added to the end of "I'm so happy today," is used as the second voice mode. Based on the second voice mode, the updated first input information is played, that is, "I'm so happy today, meow~" is played using the happy, relaxed, and clear voice effect.

[0161] This application embodiment enhances the personalized performance of virtual characters by incorporating specific catchphrases or characteristics of the characters into the voice. By updating the input information and the modulation method based on the environment and the character, the second voice effect can reflect the user's emotional state and improve the fun and uniqueness of the virtual object when corresponding to the character and the environment, so as to facilitate the user's personalized interaction.

[0162] In some embodiments, the first influencing factor further includes a first attribute of the virtual object controlled by the target account, and the second influencing factor includes a second attribute of the virtual object controlled by the target account. Step 102 in Figure 3A can be implemented by performing the following process: in response to the first attribute of the virtual object controlled by the target account switching to the second attribute, and the virtual object moving from the first environment to the second environment, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second voice mode is used to output a voice signal adapted to both the second environment and the second attribute.

[0163] Here, the voice mode will only switch from the first voice mode to the second voice mode, which is adapted to the second influencing factor, when the attributes of the virtual object controlled by the target account change or the environment changes. Changes in other influencing factors will not cause the voice mode to switch. The second voice mode can also be used to output a voice signal that is adapted to both the environmental characteristics of the second environment and the attribute characteristics of the second attribute.

[0164] For example, the attributes of a virtual object include at least one of action, emotion, or state. If the first state of the virtual object controlled by the target account is combat, the voice mode corresponding to the combat state is a tense and high-pitched voice effect. If the second state of the virtual object controlled by the target account is playing, the voice effect corresponding to the playing state is a happy and relaxed voice effect. If the state of the virtual object controlled by the target account switches from combat to playing, that is, from the first attribute to the second attribute, and the virtual object moves from the first environment to the second environment, then the voice effect corresponding to both the combat state and the first environment switches to the voice effect corresponding to both the playing state and the second environment.

[0165] In this embodiment, when both the attributes and environment of the virtual object are switched, the system switches from the first voice mode to the second voice mode. This ensures that the second voice mode after switching is consistent with both the current second environment and the second attributes of the virtual object. This avoids the one-sidedness of adaptation caused by determining the voice mode based on only a single factor (such as only the environment or only the attributes). It enables the voice signal to achieve dual matching with the attribute state and environment of the virtual object in the virtual scene, thereby improving the accuracy of voice processing and enhancing the immersion and realism of the user's voice interaction in the virtual scene.

[0166] In some embodiments, before performing the above-described "in response to the first attribute of the virtual object controlled by the target account switching to the second attribute, and the virtual object moving from the first environment to the second environment, switching the target account from the first voice mode to the second voice mode adapted to the second influencing factor", the following processing may also be performed: detecting the second input information of the target account to obtain the third emotion type of the second input information; determining the third audio effect model corresponding to the third emotion type; determining the third modulation method for the third voice signal output by the third audio effect model, wherein the third modulation method includes modulating the third voice signal based on the second environment and the second attribute; and determining the second voice mode based on the third modulation method.

[0167] Here, the above-mentioned "determining the second voice mode based on the third modulation method" can be achieved by performing the following processing: determining the third update method for the second input information of the target account, wherein the third update method includes: encoding based on the second input information, the environmental features of the second environment, and the attribute features of the second attribute to obtain the third fusion feature, and decoding based on the third fusion feature to obtain the updated second input information; and using the third modulation method and the third update method as the second voice mode.

[0168] As an example of steps 101 to 102, taking the first influencing factor as the first environment and the first attribute as an example, refer to Figure 6B. Figure 6B is a schematic diagram of switching voice modes according to the environment and attribute provided in the embodiment of this application. In the left side of Figure 6B, a virtual scene 610 is displayed, and an influencing factor setting interface 611 is displayed. The influencing factor setting interface 611 includes at least one dimension, such as the role 612 to which the virtual object controlled by the target account belongs, the attribute 613 of the virtual object controlled by the target account, the status 614 of the target account, and the input information 615 of the target account. The environment 616 in which the virtual object is located is always selected. In response to the input operation in the influencing factor setting interface 611, such as the selection operation for the attribute 613 of the virtual object controlled by the target account, the attribute 613 of the virtual object controlled by the target account and the environment 616 in which the virtual object currently controlled by the target account is located, i.e., the first environment, are combined into the first influencing factor. In the middle diagram of Figure 6B, a first environment 617 and a first character 618 are shown. At this time, the first attribute of the first character is walking, and the target account in the virtual scene is in a first voice mode that matches the first influencing factor (i.e., the first attribute of the first environment 617 and the first character 618). In response to the first influencing factor switching to a second influencing factor, in the right diagram of Figure 6B, for example, the first environment 617 is switched to the second environment 619, and the first attribute of the first character 618 is switched to the second attribute, such as from walking to combat, the target account is switched from the first voice mode to the second voice mode that matches the second influencing factor (i.e., the second attribute of the second environment 619 and the first character 618).

[0169] This application embodiment dynamically adjusts the voice mode when the attributes of the virtual object and its environment change, improving the consistency between the voice mode and the attributes of the virtual object and its environment, and enhancing the fit between the voice mode and the environment and attributes. This allows players to use different voice modes when performing different operations with virtual objects in different virtual environments, increasing the fun of the operation and improving the efficiency and convenience of voice mode switching.

[0170] In some embodiments, the first influencing factor further includes a first state parameter of the target account, and the second influencing factor includes a second state parameter of the target account. Step 102 in Figure 3A can be implemented by performing the following process: in response to the target account switching from the first state parameter to the second state parameter, and the virtual object moving from the first environment to the second environment, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second voice mode is used to output a voice signal adapted to both the second environment and the second state parameter.

[0171] Here, the voice mode will only switch from the first voice mode to the second voice mode, which is adapted to the second influencing factor, when the target account's status changes or the environment changes. Changes in other influencing factors will not cause the voice mode to switch. The second voice mode can also be used to output a voice signal that is adapted to both the environmental characteristics of the second environment and the state parameter characteristics of the second state parameter.

[0172] For example, the state of the target account includes at least one of the actions or expressions of the target object logged into the target account in the virtual scene. If the first expression of the virtual object controlled by the target account is anger, the corresponding first voice mode is a rapid and high-pitched voice effect; if the second expression of the virtual object controlled by the target account is happiness, the corresponding second voice effect is a soothing and relaxing voice effect; if the target object's expression changes from anger to happiness, that is, from the first state parameter to the second state parameter, and the virtual object moves from the first environment to the second environment, then the first voice effect corresponding to both anger and the first environment changes to the second voice effect corresponding to both happiness and the second environment.

[0173] In this embodiment, when both the state parameters of the target account and the environment of the virtual object controlled by the target account are switched, the system switches from the first voice mode to the second voice mode. This ensures that the switched voice mode matches both the current second environment of the virtual object and the current second state parameters of the target account. This avoids the one-sidedness of adaptation caused by determining the voice mode based on only a single factor (such as only the environment or only the account state parameters). It enables the voice signal to achieve dual matching with the environment of the virtual object in the virtual scene and the state parameters of the target account, thereby improving the accuracy of voice processing and enhancing the immersion and realism of the user's voice interaction in the virtual scene.

[0174] In some embodiments, before performing the above-described "in response to the target account switching from a first state parameter to a second state parameter, and the virtual object moving from a first environment to a second environment, switching the target account from a first voice mode to a second voice mode adapted to the second influencing factor", the following processing may also be performed: detecting the second input information of the target account to obtain a fourth emotion type of the second input information; determining a fourth audio effect model corresponding to the fourth emotion type; determining a fourth modulation method for the fourth voice signal output by the fourth audio effect model, wherein the fourth modulation method includes modulating the fourth voice signal based on the second environment and the second state parameter; and determining a second voice mode based on the fourth modulation method.

[0175] For example, a pre-trained text sentiment recognition model can be invoked based on the second input information of the target account to obtain the fourth sentiment type of the second input information. The training process of the pre-trained text sentiment recognition model is the same as the example in step 301 above, and will not be repeated here.

[0176] Here, the above-mentioned "determining the second voice mode based on the fourth modulation method" can be achieved by performing the following processing: determining the fourth update method for the second input information of the target account, wherein the fourth update method includes: encoding based on the second input information, the environmental features of the second environment, and the state parameter features of the second state parameters to obtain the fourth fusion feature, and decoding based on the fourth fusion feature to obtain the updated second input information; and using the fourth modulation method and the fourth update method as the second voice mode.

[0177] For example, features of the second input information are extracted from the second input information. Using a multimodal fusion network, feature concatenation techniques, or a deep learning model, the features of the second input information, the environmental features of the second environment, and the state parameter features of the second state parameters are fused into a unified feature representation, i.e., the second fused feature. A decoder is then used to process the second fused feature. The decoding process can use a generative model, such as GAN, VAE, or other sequence generation models, to decode the second fused feature and obtain the updated second input information.

[0178] As an example of steps 101 to 102, taking the first influencing factor as the first environment and the first state parameter as an example, refer to Figure 6C. Figure 6C is a schematic diagram of switching voice modes according to environment and state parameters provided in the embodiment of this application. In the left side of Figure 6C, a virtual scene 610 is displayed, and an influencing factor setting interface 611 is displayed. The influencing factor setting interface 611 includes at least one dimension, such as the role 612 to which the virtual object controlled by the target account belongs, the attribute 613 of the virtual object controlled by the target account, the state 614 of the target account, and the input information 615 of the target account. The environment 616 in which the virtual object is located is always selected. In response to the input operation in the influencing factor setting interface 611, such as the selection operation for the state 614 of the target account, the state 614 of the target account and the environment 616 in which the virtual object currently controlled by the target account is located, i.e., the first environment, are combined into the first influencing factor. In the middle diagram of Figure 6C, a first environment 617 and a first role 618 are shown. The target account in the virtual scene is in a first voice mode that matches the first influencing factor (i.e., the first environment 617 and the first state parameter of the target account). For example, the first state parameter of the target account can be "angry". In response to the first influencing factor switching to a second influencing factor, in the right diagram of Figure 6C, for example, the first environment 617 is switched to the second environment 619, and the first state parameter of the target account is switched to the second state parameter, for example, from "angry" to "happy". The target account is then switched from the first voice mode to the second voice mode that matches the second influencing factor (i.e., the second environment 619 and the second state parameter of the target account).

[0179] This application's embodiments adapt the voice mode to the target account's state parameters and environment, making the interaction more state-aware and reflecting the user's real-time state. Actions or expressions in the virtual scene receive immediate feedback from the voice mode, helping to enhance user engagement and interactivity. Adaptively changing the voice mode to correspond to the environment and the target account's state enhances the fit between the voice mode and the environment and state, allowing players to have an immersive experience in different virtual environments.

[0180] In some embodiments, the first influencing factor further includes first input information of the target account, and the second influencing factor includes second input information of the target account. Step 102 of FIG3A can be implemented by performing the following process: in response to switching from the first input information of the target account to the second input information, and the virtual object moving from the first environment to the second environment, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second voice mode is used to output a voice signal adapted to both the second environment and the second input information.

[0181] Here, the system will only switch from the first voice mode to the second voice mode, which is adapted to the second influencing factor, when the input information of the target account changes or the environment changes. Changes in other influencing factors will not cause the voice mode to switch. The second voice mode can also be used to output a voice signal that is adapted to the environmental characteristics of the second environment and the emotional type reflected by the second input information.

[0182] For example, the types of input information include text and speech. If the emotion expressed by the first input information is anger, the corresponding first speech mode is a rapid and high-pitched speech effect. If the emotion expressed by the second input information is happiness, the corresponding second speech effect is a soothing and relaxing speech effect. If the input information changes from expressing anger to expressing happiness, and the virtual object moves from the first environment to the second environment, the first speech effect, which corresponds to both the emotion of anger and the first environment, changes to the second speech effect, which corresponds to both the emotion of happiness and the second environment.

[0183] As an example of steps 101 to 102, taking the first influencing factor as the first environment and the first input information, and the type of the first input information as text, see Figure 6D. Figure 6D is a schematic diagram of switching voice modes according to the environment and input information provided in this application embodiment. In the left side of Figure 6D, a virtual scene 610 is displayed, and an influencing factor setting interface 611 is displayed. The influencing factor setting interface 611 includes at least one dimension, such as the role 612 of the virtual object controlled by the target account, the attribute 613 of the virtual object controlled by the target account, the status 614 of the target account, and the input information 615 of the target account. The environment 616 where the virtual object is located is always selected. In response to the input operation in the influencing factor setting interface 611, such as the selection operation of the input information 615 of the target account, the input information 615 of the target account and the environment 616 where the virtual object currently controlled by the target account is located, i.e., the first environment, are combined into the first influencing factor. In the middle diagram of Figure 6D, a first environment 617, a first character 618, and a first input message 621 are shown, such as "Ugh, I'm so angry." The target account in the virtual scene is in a first voice mode that matches the first influencing factor (i.e., the first environment 617 and the first input message 621). In response to the first influencing factor switching to a second influencing factor, in the right diagram of Figure 6D, for example, the first environment 617 is switched to the second environment 619, and the first input message 621 is switched to the second input message 622, for example, "Ugh, I'm so angry" is switched to "I'm very happy today." The target account is switched from the first voice mode to a second voice mode that matches the second influencing factor (i.e., the second environment 619 and the second input message 622).

[0184] In this embodiment, when the first input information of the target account is switched to the second input information, and the virtual object moves from the first environment to the second environment, the system switches from the first voice mode to the second voice mode. The second voice mode is adapted to both the current environment of the virtual object and the latest input information of the target account. This avoids the limitations of determining the voice mode solely based on the environment or solely based on the input information. The voice signal can form a dual response with the environmental state of the virtual object in the virtual scene and the input content of the target account. This ensures both the fit between voice processing and scene dynamics and the close association between voice output and user input intent, thereby improving the accuracy of voice interaction and enhancing the user's immersive experience in the virtual scene.

[0185] In some embodiments, before performing the above-described "switching the target account from the first voice mode to the second voice mode adapted to the second influencing factor in response to switching from the first input information of the target account to the second input information and the virtual object moving from the first environment to the second environment", the following processing may also be performed: in response to the first input information of the target account including illegal language, replacing the illegal language with compliant language, wherein the compliant language is adapted to the role to which the virtual object controlled by the target account belongs.

[0186] For example, different virtual objects have different compliant terms corresponding to their attributes. For instance, after performing each action, virtual object A will output the catchphrase "Ya ha ha ha". In response to the first input information of the target object using virtual object A in the virtual scene including the compliant term, the compliant term can be replaced with "Ya ha ha ha" to update the first input information.

[0187] Here, compliance terms do not have to be adapted to the role to which the virtual object belongs. A compliance term database is established, which includes multiple compliance terms. In response to the first input information of the target account, which includes violation terms, any one of the compliance terms is randomly selected to replace the violation term; or all violation terms can be replaced with the same compliance term.

[0188] This application embodiment ensures the coherence and consistency of voice output by replacing non-compliant terms with compliant terms that are appropriate for the role, making the role performance more uniform. It can flexibly replace compliant terms according to different roles, adapt to different virtual objects and environments, and reflect the personalized performance of the role while complying with regulations.

[0189] In some embodiments, before performing the above-mentioned "replacing the non-compliant term with the compliant term", the following processing may also be performed: if the information type of the first input information is text, convert multiple words in the first input information into multiple word feature vectors, and call a pre-trained text classification model based on the multiple word feature vectors to classify the multiple words to obtain the non-compliant term.

[0190] Here, when the information type of the first input information is text, the first input information is segmented into words based on the word segmentation algorithm to obtain multiple words. Each word obtained by word segmentation is mapped to a fixed-dimensional vector space to obtain the word feature vector corresponding to each word. For each word vector feature, a pre-trained text classification model is called to classify the word vector feature to obtain the type of word corresponding to the word vector feature. The word type includes illegal language and compliant language.

[0191] For example, word segmentation algorithms can be used to segment training samples into individual words. Word segmentation algorithms refer to dividing training samples into individual words, such as forward maximum matching, backward maximum matching, string matching-based word segmentation methods, and Hidden Markov Model-based word segmentation algorithms. One-hot encoding, Bag of Words Model, Term Frequency-Inverse Document Frequency (TF-IDF), or N-gram models can be used to obtain the word feature vector corresponding to each word.

[0192] For example, a pre-trained text classification model is trained by performing the following processes: obtaining an initial text classification model; obtaining word feature vector samples and ground truth word labels, where the ground truth word labels represent whether the word corresponding to the word feature vector is an inappropriate word; calling the initial text classification model based on the word feature vector samples to obtain predicted word labels; determining word loss values ​​based on ground truth word labels and predicted word labels; and updating the parameters of the initial text classification model based on the word loss values ​​to obtain the pre-trained text classification model.

[0193] For example, see Figure 5B, which is a schematic diagram of the training principle of the text classification model provided in this application embodiment. In Figure 5B, based on word feature vector samples, the convolutional and fully connected layers of the initialized text classification model can be called to obtain predicted word labels. The word loss value between the true word label and the predicted word label is calculated through a loss function. The word loss value is backpropagated to update the parameters of the initialized text classification model. The process of calculating the word loss value and updating the parameters is repeated multiple times until the word loss value no longer increases or decreases, at which point the iteration process stops, forming a pre-trained text classification model.

[0194] For example, initializing the representation involves randomly assigning values ​​to the parameters of a text classification model, such as assigning all parameters to 0 or all to 1. Backpropagation is implemented using the backpropagation algorithm, calculating the gradient of each neuron from the output layer to the input layer, and updating the neuron's weights and biases based on the gradients. Gradient descent is used to continuously update the parameters, reducing the loss value. Various gradient descent algorithms can be used, such as batch gradient descent, stochastic gradient descent, adaptive gradient descent, and momentum gradient descent. The loss function can be the mean squared error loss function, cross-entropy loss function, multi-label classification loss function, or triplet loss function.

[0195] In this embodiment, when the first input information is text, multiple words in the first input information are converted into multiple word feature vectors. Then, a pre-trained text classification model is called based on these word feature vectors to classify the words, thereby processing the illegal language and accurately identifying the illegal content in the first input information. By accurately representing the text information through word feature vectors and combining this with the mature classification capabilities of the pre-trained text classification model, missed or false judgments of illegal language in the text information are avoided. This ensures that subsequent processing of replacing illegal language with compliant language only targets clearly identified illegal language, guaranteeing the compliance of the first input information processing while avoiding unnecessary intervention in compliant content. This improves the accuracy and efficiency of processing input information for the target account, providing a compliant and accurate input basis for subsequent voice mode adaptation.

[0196] In some embodiments, before performing the above-described "replacing the non-compliant term with the compliant term", the following processing may also be performed: if the information type of the first input information is speech, the first input information is converted into text to be processed, acoustic features are extracted from the first input information, text features are extracted from the text to be processed, and a pre-trained machine learning model is invoked based on the acoustic features and text features to classify multiple words in the text to be processed to obtain the non-compliant term.

[0197] Here, the speech signal of the first input information is preprocessed, and features are extracted from the preprocessed speech signal. The speech waveform is converted into a series of feature vectors representing speech characteristics. The extracted feature vectors are input into the acoustic model, and the acoustic features are converted into phoneme or word representations. The output of the acoustic model is input into the language model, and the next possible word or phrase is predicted based on the context information. The results of the acoustic model and the language model are combined, and a search algorithm is used to determine the word sequence as the text to be processed. The first input information is converted into a digital signal using an analog-to-digital converter. The digital speech signal is pre-emphasized. A window function (such as a Hamming or Hanning window) is applied to each frame of the pre-emphasized digital signal. A Fast Fourier Transform (FFT) is then applied to each frame to convert the time-domain signal into a frequency-domain signal, obtaining the spectral information of the speech signal. Features, such as spectral energy or spectral entropy, are extracted from the spectrum. A Mel filter bank is used to simulate the auditory characteristics of the human ear, dividing the spectrum into multiple frequency bands. The logarithm of the output of each Mel filter is taken to obtain the log-Mel spectrum. Discrete cosine transform (DCT) or other transforms are applied to convert the spectral features into cepstral features. The first and second differences of the cepstral coefficients are calculated to capture the dynamic characteristics of the speech. The extracted features are normalized to obtain the acoustic features of the first input information. Alternatively, the text to be processed can be segmented into words to determine the word feature vector corresponding to each word. The word feature vectors of multiple words are combined to form the text features of the text. Based on the acoustic and text features, a pre-trained machine learning model is called to classify multiple words in the text to obtain the prohibited words.

[0198] For example, preprocessing includes signal enhancement, denoising, silence detection, and endpoint detection. Feature extraction methods include Mel-Frequency Cepstrum Coefficients (MFCCs), Linear Predictive Coding (LPC), or Perceptual Linear Predictive (PLP). Acoustic models can be based on Hidden Markov Models (HMMs), Deep Neural Networks (DNNs), or other machine learning algorithms. Decoding algorithms include the Viterbi Algorithm, dynamic programming, or deep learning-based sequence-to-sequence (Seq2Seq) models.

[0199] For example, a pre-trained machine learning model is trained by performing the following processes: obtaining an initialized machine learning model; obtaining acoustic feature samples and text feature samples, as well as ground truth labels; concatenating the acoustic feature samples and text feature samples to obtain concatenated feature samples, where ground truth labels indicate whether the concatenated feature samples contain prohibited words; calling the initialized machine learning model based on the concatenated feature samples to obtain predicted labels; determining a loss value based on the ground truth labels and predicted labels; updating the parameters of the initialized machine learning model based on the loss value to obtain the pre-trained machine learning model.

[0200] For example, see Figure 5C, which is a schematic diagram of the training principle of the machine learning model provided in this application embodiment. In Figure 5C, based on the concatenated feature samples, the convolutional and fully connected layers of the initialized machine learning model can be called to obtain the predicted label. The loss value between the true label and the predicted label is calculated through a loss function. The loss value is backpropagated to update the parameters of the initialized machine learning model. The process of calculating the loss value and updating the parameters is repeated multiple times until the loss value no longer increases or decreases, at which point the iteration process stops, forming a pre-trained machine learning model.

[0201] For example, initializing the representation involves randomly assigning values ​​to the parameters of a machine learning model, such as setting all parameters to 0 or all to 1. Backpropagation is implemented using the backpropagation algorithm, calculating the gradient of each neuron from the output layer to the input layer, and updating the neuron's weights and biases based on the gradients. Gradient descent is used to continuously update the parameters, reducing the loss value. Various gradient descent algorithms can be used, such as batch gradient descent, stochastic gradient descent, adaptive gradient descent, and momentum gradient descent. The loss function can be the mean squared error loss function, cross-entropy loss function, multi-label classification loss function, or triplet loss function.

[0202] In this embodiment, when the first input information is speech, the first input information is converted into text to be processed. Acoustic features are extracted from the first input information, and text features are extracted from the text to be processed. Then, based on the acoustic and text features, a pre-trained machine learning model is invoked to classify multiple words in the text to be processed, thereby obtaining the processing of illegal language and realizing the early identification of illegal content in the first input information of speech. Combining acoustic and text features, the dual feature dimensions can more comprehensively capture the features of illegal information. Combined with the classification capabilities of the pre-trained machine learning model, it not only ensures the compliance of the processing of the first input information of speech, but also avoids unnecessary intervention in compliant content, thereby improving the accuracy and efficiency of the processing of input information of the target account, and providing a compliant and accurate input basis for subsequent voice mode adaptation.

[0203] In some embodiments, before performing the above-described "switching the target account from the first voice mode to the second voice mode adapted to the second influencing factor in response to switching from the first input information of the target account to the second input information and the virtual object moving from the first environment to the second environment", the following processing may also be performed: if the first input information or the second input information is voice, the first input information or the second input information is captured using the built-in microphone of the electronic device and initially denoised by a noise reduction chip. On this basis, software algorithms may be used for further optimization, such as bandpass filters, to eliminate low-frequency background noise in the environment.

[0204] For example, the information type of the first or second input information is first detected. If the first or second input information is speech, the activation command of the built-in microphone of the electronic device is triggered. The built-in microphone of the electronic device responds to the activation command and enters the audio acquisition state to receive the speech signal (i.e., the first or second input information) input by the target account in real time. The speech signal is converted into audio data and transmitted to the noise reduction chip of the electronic device. The noise reduction chip activates a preset preliminary noise reduction algorithm (such as an adaptive filtering algorithm) to filter out high-frequency burst noise and obvious interference noise in the audio data, and obtains the preliminary noise-reduced audio data. A bandpass filter (or other preset software noise reduction algorithm) is called, the frequency range of the bandpass filter is set, and the preliminary noise-reduced audio data is processed by the bandpass filter to filter out low-frequency background noise (such as ground vibration noise and low-frequency machine roar) with a frequency lower than the preset frequency threshold, and outputs the further optimized audio data.

[0205] In some embodiments, step 102 of FIG3A can be implemented by performing the following process: in response to the virtual object moving from the first environment to the second environment, and at least two other dimensions of the first influencing factor being switched to the second influencing factor, the target account is switched from the first voice mode to the second voice mode adapted to the second influencing factor, wherein the at least two other dimensions are at least two of the following dimensions: the role to which the virtual object currently controlled by the target account belongs, the attributes of the virtual object currently controlled by the target account, the input information of the target account, and the status of the target account.

[0206] For example, the first influencing factor can be any two dimensions from the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the input information of the target account, combined with the first environment; the first influencing factor can be any three dimensions from the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the input information of the target account, combined with the first environment; or the first influencing factor can be the combination of the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the input information of the target account, combined with the first environment.

[0207] This application embodiment adjusts the voice mode through changes in multiple dimensions, ensuring comprehensive adaptability to environmental changes from multiple angles, improving the accuracy of the generated voice mode, making the switching of voice modes more natural, and enhancing user experience and immersion.

[0208] In some embodiments, before performing the above-described "in response to a virtual object moving from a first environment to a second environment, and at least two other dimensions of the first influencing factors being switched to the second influencing factor, switching the target account from the first voice mode to a second voice mode adapted to the second influencing factor", the following processing may also be performed: detecting the second input information of the target account to obtain a fifth emotion type of the second input information; determining a fifth audio effect model corresponding to the fifth emotion type; determining a fifth modulation method for the fifth voice signal output by the fifth audio effect model, wherein the fifth modulation method includes modulating the fifth voice signal based on the second environment, and at least one of the role characteristics of the second role, the attribute characteristics of the second attribute, or the state parameter characteristics of the second state parameter; and determining a second voice mode based on the fifth modulation method.

[0209] Here, the above-mentioned "determining the second voice mode based on the fifth modulation method" can be achieved by performing the following processing: determining the fifth update method for the second input information of the target account, wherein the fifth update method includes: encoding based on at least one of the second input information, the environmental features of the second environment, the role features of the two roles, the attribute features of the second attribute, or the state parameter features of the second state parameter to obtain the fifth fusion feature, and decoding based on the fifth fusion feature to obtain the updated second input information; and using the fifth modulation method and the fifth update method as the second voice mode.

[0210] This application embodiment modulates the fifth voice signal based on at least one of the following: the second environment, the role characteristics of the second character, the attribute characteristics of the second attribute, or the state parameter characteristics of the second state parameter. This determines the second voice mode, allowing it to simultaneously meet the emotional needs of the second input information, the scene requirements of the second environment, and the feature requirements of at least one of the second character, the second attribute, or the second state parameter. This avoids the one-sidedness of the voice mode caused by single-dimensional adaptation. By using the fifth modulation method and the fifth update method as the second voice mode, the fifth modulation method ensures that the fifth voice signal adapts to multi-dimensional factors, while the fifth update method optimizes the second input information itself. This ensures that the second voice mode accurately matches the multi-dimensional features of the second influencing factors, providing a more dynamic basis for subsequent voice mode switching in the virtual scene. This enhances the immersion and accuracy of the target account's voice interaction in the virtual scene, guaranteeing the adaptation effect of voice mode switching under the switching of multi-dimensional influencing factors.

[0211] In some embodiments, referring to FIG3E, FIG3E is a fifth flowchart of the voice processing method for a virtual scene provided in the embodiments of this application. Before step 102, steps 401 to 404 of FIG3E are executed, which are described in detail below.

[0212] In step 401, the second input information is detected to obtain the second emotion type of the second input information.

[0213] In some embodiments, a pre-trained text sentiment recognition model is invoked based on the second input information of the target account to obtain the second sentiment type of the second input information.

[0214] For example, the training process of the text sentiment recognition model can be found in the example of step 301 above, and will not be repeated here. A schematic diagram of the training process of the text sentiment recognition model can be found in Figure 5A.

[0215] In step 402, the second audio effect model corresponding to the second emotion type is determined.

[0216] Here, the second audio effect model is used to output a voice signal that matches the second emotion type, that is, a voice signal that matches the second input information.

[0217] In some embodiments, a model database is pre-set, which stores the association between different emotion types and corresponding audio effect models; the model database is used to query the second audio effect model corresponding to the second emotion type. One emotion type can correspond to one audio effect model, or multiple emotion types can be combined to correspond to one audio effect model. Furthermore, for the same second emotion type, different audio effect models can correspond to different scenarios (such as battle scenarios or social scenarios).

[0218] For example, if the second emotion type is happiness, the model database is queried for the audio effect model corresponding to happiness, and that audio effect model is used as the second audio effect model.

[0219] In step 403, a second modulation method is determined for the second voice signal output by the second audio effect model. The second modulation method includes modulation based on at least one of the following influencing factors: the second environment, the role to which the virtual object currently controlled by the target account belongs, the attributes of the virtual object currently controlled by the target account, and the state of the target account.

[0220] In some embodiments, the parameter settings of the second modulation method can be refined, such as adjusting the reverberation duration based on the spatial size of the second environment, adjusting the speech fundamental frequency based on the role age of the second character, and improving the modulation accuracy; a modulation effect feedback mechanism can be established to optimize the rules of the second modulation method based on the user's feedback on the modulated speech signal.

[0221] For example, if the second environment is a bright and sunny environment, the virtual object currently controlled by the target account belongs to the role of a traveler, the attribute of the virtual object currently controlled by the target account is travel, and the state of the target account is smiling, then the second modulation method is to add a happy, relaxed and clear voice effect to the second voice signal.

[0222] In step 404, a second speech mode is determined based on the second modulation scheme.

[0223] Here, the second modulation method can be directly used as the second voice mode.

[0224] In some embodiments, the first modulation scheme is decomposed into specific speech parameters, such as pitch adjustment value, reverberation duration, and sound effect overlay type. The decomposed speech parameters are combined with the basic parameters of the second audio effect model to generate a second speech pattern adapted to the second influencing factor, and stored in a speech pattern library for subsequent speech pattern switching. The system supports generating a second speech pattern based on the combination of multiple second modulation schemes. When the second input information includes multiple emotion types and corresponds to multiple second modulation schemes, multiple second modulation schemes can be fused to generate a comprehensive second speech pattern.

[0225] For example, the second modulation method, "increase the tone of voice to enhance penetration, increase the reverberation of the volcanic eruption sound effect, and enhance the roughness of the voice to suit the warrior character," is broken down into voice parameters such as "increase tone by 20%, reverberation duration of 1.2 seconds, and superimpose 50% volcanic eruption sound effect." These parameters are then combined with the basic parameters of the second audio effect model, "rapid rhythm sound effect model," to generate the second voice mode.

[0226] This application embodiment, through emotion detection of input information, can accurately identify the user's emotional state and customize audio effects according to the user's emotional type, providing a more personalized audio experience and making the interaction more closely aligned with the user's emotional state. Audio modulation based on influencing factors such as the second environment and second role allows the speech signal to better adapt to different environments and virtual roles, improving the naturalness and realism of the speech, ensuring consistency between the speech signal and influencing factors such as virtual roles, enhancing the user's immersion in the virtual environment, dynamically adjusting audio effects, and providing a more dynamic and richer interactive experience.

[0227] In some embodiments, referring to FIG3F, FIG3F is a sixth flowchart of the voice processing method for a virtual scene provided in the embodiments of this application. Step 404 in FIG3E can be implemented by steps 4041 to 4042 in FIG3F, as described in detail below.

[0228] In step 4041, a second update method for the second input information of the target account is determined. The second update method includes: encoding based on the second input information, the environmental characteristics of the second environment, and the role characteristics of the role to which the virtual object controlled by the target account belongs to obtain a second fusion feature, and decoding based on the second fusion feature to obtain the updated second input information.

[0229] Here, features of the second input information are extracted from the second input information. Using a multimodal fusion network, feature concatenation techniques, or a deep learning model, the features of the second input information, the environmental features of the second environment, and the role features of the second character are fused into a unified feature representation, namely the second fused feature. A decoder is used to process the second fused feature. The decoding process can use a generative model, such as a generative adversarial network, a variational autoencoder, or other sequence generation models, to decode the second fused feature and obtain the updated second input information.

[0230] For example, the second character is characterized by using the catchphrase "haha". If the second environment is a relaxing and peaceful seaside, and the second input information is "apples are delicious", the catchphrase will be inserted at the end of the second input information or before a punctuation mark in the sentence to update the second input information. For example, the updated second input information could be "apples are delicious, haha".

[0231] In step 4042, the second modulation method and the second update method are used as the second speech mode.

[0232] Following the examples of steps 303 and 3041 above, a happy, relaxed, and clear voice effect, along with "haha" added to the end of "apples are delicious," is used as the second voice mode. Based on the second voice mode, the updated second input information is played, that is, "apples are delicious, haha" is played using a happy, relaxed, and clear voice effect.

[0233] This application embodiment determines a second update method and uses the second modulation method and the second update method as a second voice mode. The second modulation method can ensure that the second voice signal is adapted to multi-dimensional influencing factors, and the second update method can optimize the second input information itself, so that the second input information is more closely related to the environmental characteristics of the second environment and the role characteristics of the virtual object. This ensures that the second voice mode can accurately match the multi-dimensional characteristics of the second influencing factors, providing a more dynamic basis for subsequent voice mode switching in the virtual scene. This enhances the immersion and accuracy of the target account's voice interaction in the virtual scene and ensures the adaptation effect of voice mode switching under the switching of multi-dimensional influencing factors.

[0234] In some embodiments, referring to FIG3G, FIG3G is a seventh flowchart of the voice processing method for a virtual scene provided in the embodiments of this application. The above-mentioned "second modulation method" can be implemented through steps 501 to 505 of FIG3G, which will be described in detail below.

[0235] In step 501, the frequency of the second voice signal is determined based on the attributes of the virtual object currently controlled by the target account.

[0236] In some embodiments, the frequency of the second speech signal is determined by analyzing the spectral characteristics of the second speech signal through Fourier transform based on the attributes of the virtual object currently controlled by the target account.

[0237] For example, if the second character's attribute is to cast magic, then the frequency of the second voice signal is lower than the frequency of other second characters' attributes, in order to create a deep voice effect.

[0238] In step 502, the amplitude of the second speech signal is determined based on the second emotion type.

[0239] In some embodiments, the volume of the second speech signal is adjusted using dynamic range compression technology according to the second emotion type to determine the amplitude of the second speech signal.

[0240] For example, if the second emotion type is joy, increase the amplitude of the second speech signal; if the second emotion type is sadness, decrease the amplitude of the second speech signal.

[0241] In step 503, the formant peak of the second voice signal is located based on the attributes of the virtual object currently controlled by the target account.

[0242] In some embodiments, based on the attributes of the virtual objects currently controlled by the target account, the formants of the second speech signal are identified and relocated using linear predictive coding (LPC) or other acoustic models to mimic the sound characteristics of different virtual objects.

[0243] For example, if the attribute of the second character is to cast magic, the formant frequency of the second voice signal is lower than the formant frequency of the attributes of other second characters. By changing the position and intensity of the formants, the pitch and timbre of the voice can be changed to make it closer to the attribute of the character to which the virtual object belongs.

[0244] In step 504, environmental reflection simulation is performed on the second speech signal according to the second environment to obtain the environmental reflection sound of the second speech signal.

[0245] In some embodiments, based on the second environment or the environmental characteristics of the second environment, convolutional reverberation or physical modeling techniques are used to simulate environmental reflections of the second speech signal. The environmental reflection simulation may involve adding echoes and spatial sense to obtain the environmental reflection sound of the second speech signal.

[0246] For example, if the second environment is a cave scene, a long-tailed echo can be added to the second speech signal to obtain the environmental reflection sound of the second speech signal; if the second environment is a grassland, the reflection effect can be reduced to obtain the environmental reflection sound of the second speech signal.

[0247] In step 505, the second speech signal is modulated according to the frequency, amplitude, formant and ambient reflected sound to obtain the third speech signal.

[0248] Here, multiple modulation effects are determined based on frequency, amplitude, formant, and ambient reflected sound. These multiple modulation effects are then combined to obtain a merged effect. The second speech signal is then adjusted based on the merged effect to obtain the third speech signal.

[0249] This application's embodiments adjust the frequency and amplitude of the speech signal based on the attributes and emotional type of the virtual object, generating highly personalized speech output that better meets the user's individual needs. Determining the amplitude based on the second emotional type more realistically reflects the emotional state of the virtual object, enhancing the emotional expression of the speech. By locating formants and simulating reflections according to the environment, a more natural speech signal can be obtained, improving the naturalness and realism of the speech. Environmental reflection simulation of the second speech signal adapts it to different environmental conditions, providing speech output that matches the environment.

[0250] In some embodiments, the virtual scene includes a first voice mode control corresponding to a first voice mode and a third voice mode control corresponding to a third voice mode. The third voice mode is a mode that plays the voice signal of the target account itself. The third voice mode is in the on state by default. Before executing step 102 "in response to the first influencing factor switching to the second influencing factor, switch the target account from the first voice mode to the second voice mode adapted to the second influencing factor", the following processing can be performed: in response to the trigger operation on the first voice mode control, put the first voice mode in the on state and turn off the third voice mode.

[0251] For example, triggering operations include at least one of the following: single click, double click, swipe, long press, gesture trigger, voice trigger, and peripheral device trigger (such as game controller button trigger). Voice mode status prompts can be added. When switching between the first and third voice modes, a pop-up window, sound effect, or text prompt within the virtual scene, such as "First voice mode is enabled, third voice mode is disabled," can be used to inform the user of the current voice mode, avoiding confusion regarding the mode status.

[0252] For example, referring to Figure 6E, Figure 6E is a first schematic diagram of switching from a third voice mode to a first voice mode according to an embodiment of this application. In the left side view of Figure 6E, the virtual scene 610 includes a third voice mode control 623 corresponding to the third voice mode (e.g., original voice mode) and a first voice mode control 624 corresponding to the first voice mode. The third voice mode is the mode for playing the target account's own voice signal, that is, the target account's original voice mode. The third voice mode is in the on state by default. The third voice mode control 623 can be highlighted to indicate that the third voice mode is in the on state, and a first prompt message 625 is displayed, such as "Using original voice mode". In response to the trigger operation of the first voice mode control 624, in the right side view of Figure 6E, the first voice mode is put into the on state, for example, by highlighting the first voice mode control 624, and the third voice mode is turned off, for example, by unhighlighting the third voice mode control 623, and a second prompt message 626 is displayed, such as "Using a new voice mode".

[0253] This application embodiment allows users to select different voice modes through trigger operations, improving the flexibility and autonomy of interaction. Users can enable or disable different voice modes according to their preferences and scenario needs, thereby obtaining a personalized voice experience. Furthermore, the trigger operation of the voice mode control for switching voice modes improves the efficiency and convenience of voice mode switching.

[0254] Here, a voice mode switching control can also be displayed in the virtual scene, wherein the third voice mode is enabled by default; in response to the trigger operation of the voice mode switching control, the first voice mode is enabled and the third voice mode is disabled.

[0255] For example, referring to Figure 6F, Figure 6F is a second schematic diagram of switching from a third voice mode to a first voice mode according to an embodiment of this application. In the left side of Figure 6F, the virtual scene 610 includes a voice mode switching control 627, which is used to switch between the third voice mode and the first voice mode. The third voice mode is the mode that plays the target account's own voice signal, that is, the target account's original voice mode. The third voice mode is on by default. At this time, the first prompt message 625 is displayed, such as "Using the original voice mode". In response to the trigger operation of the voice mode switching control 627, in the right side of Figure 6F, the first voice mode is turned on and the third voice mode is turned off, and the second prompt message 626 is displayed, such as "Using the new voice mode".

[0256] This application embodiment displays a voice mode switching control in a virtual scene, with the third voice mode being enabled by default. Simply triggering this control will directly enable the first voice mode and disable the third voice mode, eliminating the need for separate controls for the first and third voice modes. A single control enables switching between the two voice modes, simplifying the operation and reducing user complexity. Furthermore, the default enable setting of the third voice mode aligns with initial user habits, ensuring that users can directly interact using their own voice signals after entering the virtual scene, reducing operational errors and improving the convenience and efficiency of voice mode control in the virtual scene.

[0257] Referring to Figure 4A, which is a first flowchart of the information processing method for a virtual scene provided in an embodiment of this application, the steps shown in Figure 4A will be described with the terminal as the main focus.

[0258] In step 601, a virtual scene based on the target account login is displayed, wherein the virtual scene includes a message editing control.

[0259] In some embodiments, the display method of the message editing control can be adapted to the type of electronic device. When the type of electronic device is a mobile device, the message editing control can be displayed by floating in the virtual scene; when the type of electronic device is a personal computer (PC), the message editing control can be displayed in a fixed area at the bottom; when the type of electronic device is a virtual reality (VR) device, a three-dimensional (3D) message editing control can be displayed.

[0260] For example, a message editing control could be an input box or a voice input button.

[0261] In step 602, in response to the first message input operation in the message editing control, the first input information is displayed.

[0262] For example, the first input information can be of various types, such as text or voice. If the message editing control is an input box and the type of the first input information is text, the first input information is displayed in response to an editing operation in the message editing control, where the editing operation is the first message input operation; if the message editing control is a voice input button and the type of the first input information is voice, the first input information is displayed in response to a click or long press operation on the voice input button, where the click or long press operation is the first message input operation.

[0263] In step 603, in response to the first message output operation, the voice signal of the first input information is played in the virtual scene based on the first voice mode adapted to the first influencing factor of the target account, wherein the first influencing factor includes at least the first environment in which the virtual object controlled by the target account is currently located.

[0264] Here, based on a first voice mode adapted to the first influencing factor of the target account, a voice signal of the first input information with the target account as the sound source is played in a virtual scene. The first message output operation can be input to a single object, and the listening range of the playback is the single object; the first message output operation can also be output to multiple objects simultaneously, and the listening range of the playback is multiple objects, for example, multiple objects corresponding to multiple accounts in the same camp as the target account.

[0265] As an example of steps 601 to 603, taking text as the first input information, refer to Figure 6G. Figure 6G is a schematic diagram of applying a corresponding voice mode based on the input information according to an embodiment of this application. In the left side of Figure 6G, a virtual scene 610 based on target account login is shown, wherein the virtual scene 610 includes a message editing control 628. In response to a first message input operation in the message editing control 628, the first input information 629 is shown in the right side of Figure 6G. In response to a first message output operation, such as in response to a trigger operation on the send control 630, the voice signal of the first input information 629 is played in the virtual scene 610 based on a first voice mode adapted to a first influencing factor of the target account, wherein the first influencing factor includes at least a first environment 617 in which the virtual object controlled by the target account is currently located.

[0266] This application embodiment displays a virtual scene and message editing controls based on the target account login, allowing users to quickly enter the scene and input information. In response to the first message input operation, the first input information is displayed in real time, ensuring the input content is clear and searchable. In response to the first message output operation, a first voice mode is adapted based on a first influencing factor including the first environment, playing the voice signal of the first input information, ensuring the voice signal accurately matches the environment of the virtual object. The entire process not only ensures the convenience of message input and output in the virtual scene but also enhances the immersive experience of voice interaction through environment-adapted voice modes, avoiding the problem of voice signals being out of sync with the scene. It also supports multi-terminal and multi-scene adaptation, meeting the operational and experience needs of different users and comprehensively improving the information interaction effect of the target account in the virtual scene.

[0267] In some embodiments, referring to FIG4B, FIG4B is a second flowchart of the information processing method for a virtual scene provided in the embodiments of this application. Before step 603 "playing the voice signal of the first input information in the virtual scene", steps 701 to 704 of FIG4B are executed, which are described in detail below.

[0268] In step 701, the first input information is classified to obtain the emotion type of the first input information.

[0269] In some embodiments, a pre-trained text sentiment recognition model is invoked based on the first input information of the target account to obtain the sentiment type of the first input information.

[0270] For example, the training process of the text sentiment recognition model can be found in the example of step 301 above. A schematic diagram of the training process of the text sentiment recognition model can be found in Figure 5A, which will not be repeated here.

[0271] In step 702, the audio effect model corresponding to the emotion type is determined.

[0272] Here, the audio effects model is used to output a speech signal that matches the emotion type, that is, a speech signal that matches the first input information.

[0273] In some embodiments, a model database is pre-set, which stores the association between different emotion types and corresponding audio effect models; the model database is used to query the audio effect model corresponding to an emotion type. One emotion type can correspond to one audio effect model, or multiple emotion types can be combined to correspond to one audio effect model. Alternatively, different audio effect models can be used for the same emotion type in different scenarios (such as battle scenarios or social scenarios).

[0274] For example, if the emotion type is "happy", the model database will be searched for the audio effect model corresponding to "happy", and that audio effect model will be used as the audio effect model corresponding to the emotion type.

[0275] In step 703, the modulation method of the voice signal output by the audio effects model is determined. The modulation method includes modulation based on at least one of the following influencing factors: the environment of the virtual object currently controlled by the target account, the role to which the virtual object currently controlled by the target account belongs, the attributes of the virtual object currently controlled by the target account, and the status of the target account.

[0276] For example, if the first environment is a bright and sunny environment, the virtual object currently controlled by the target account belongs to the role of a traveler, the attribute of the virtual object currently controlled by the target account is travel, and the state of the target account is smiling, then the modulation method is to add a happy, relaxed and clear voice effect to the first voice signal.

[0277] In step 704, a first voice mode is determined based on the modulation method, and the voice signal output by the audio effects mode is applied to the first voice mode to obtain the voice signal of the first input information.

[0278] Here, the modulation method can be directly used as the first speech mode. Audio processing tools and algorithms are then used to apply the effect of the first speech mode to the modulated speech signal.

[0279] This application embodiment, through emotion detection of input information, can accurately identify the user's emotional state and customize audio effects according to the user's emotional type, providing a more personalized audio experience and making the interaction more closely resemble the user's emotional state. Audio modulation based on multi-dimensional influencing factors allows the speech signal to better adapt to different influencing factors, improving the naturalness and realism of the speech, ensuring consistency between the speech signal and various influencing factors, enhancing the user's immersion in the virtual environment, dynamically adjusting audio effects, and providing a more dynamic and richer interactive experience.

[0280] In some embodiments, referring to FIG4C, FIG4C is a third flowchart of the information processing method for a virtual scene provided in the embodiments of this application. Step 704 of FIG4B, "determining the first voice mode based on the modulation method", can be implemented by steps 7041 to 7042 of FIG4C, which will be described in detail below.

[0281] In step 7041, the update method for the first input information of the target account is determined. The update method includes: encoding based on the first input information, the environmental characteristics of the first environment, and the role characteristics of the role to which the virtual object controlled by the target account belongs to obtain the fusion feature, and decoding based on the fusion feature to obtain the updated first input information.

[0282] In some embodiments, features of the first input information are extracted from the first input information. Using a multimodal fusion network, feature concatenation techniques, or a deep learning model, the features of the first input information, the environmental features of the first environment, and the role features of the virtual object controlled by the target account are fused into a unified feature representation, namely the first fused feature. The first fused feature is then processed using a decoder. The decoding process can utilize a generative model, such as a generative adversarial network, a variational autoencoder, or other sequence generation models, to decode the first fused feature and obtain the updated first input information.

[0283] For example, the character to which the virtual object controlled by the target account belongs is characterized by using the catchphrase "Meow~". If the first environment is a relaxing and peaceful beach, and the first input message is "I'm so happy today", the catchphrase will be inserted at the end of the first input message or before a punctuation mark in the sentence to update the first input message. For example, the updated first input message could be "I'm so happy today, meow~".

[0284] In step 7042, the modulation method and update method are used as the first speech mode.

[0285] Following the examples of steps 703 and 7041 above, a happy, relaxed, and clear voice effect, along with "meow~" added to the end of "I'm so happy today," is used as the first voice mode. Based on the first voice mode, the updated first input information is played, that is, "I'm so happy today, meow~" is played using the happy, relaxed, and clear voice effect.

[0286] In some embodiments, referring to FIG4D, FIG4D is a fourth flowchart of the information processing method for a virtual scene provided in the embodiments of this application. After step 603 "playing the voice signal of the first input information in the virtual scene", steps 604 to 606 of FIG4D are executed, which are described in detail below.

[0287] In step 604, in response to the first environment currently occupied by the virtual object controlled by the target account and at least one other dimension of the first influencing factor of the target account being switched to the second influencing factor, the first voice mode adapted to the first influencing factor of the target account is switched to the second voice mode adapted to the second influencing factor. The at least one other dimension is at least one of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the current input information of the target account. The second influencing factor includes at least the second environment occupied by the virtual object controlled by the target account.

[0288] For example, taking the target account's state as an example, if the first environment in which the virtual object controlled by the target account is located is a dark and quiet environment, and the first state parameter of the target account in the first environment is fear, then the first voice mode is a trembling, timid, and tense voice effect. If the second environment in which the virtual object controlled by the target account is located is a bright and cheerful environment, and the second state parameter of the target account in the second environment is happiness, then the second voice mode is a happy, relaxed, and clear voice effect. When switching from the first environment to the second environment, and when the target account's state parameter changes from fear to happiness, the voice effect automatically changes from trembling, timid, and tense to happy, relaxed, and clear.

[0289] In step 605, in response to a second message input operation in the message editing control, the second input information is displayed.

[0290] For example, the second input information can be of various types, such as text or voice. If the message editing control is an input box and the type of the second input information is text, the second input information is displayed in response to an editing operation in the message editing control, where the editing operation is the second message input operation; if the message editing control is a voice input button and the type of the second input information is voice, the second input information is displayed in response to a click or long press operation on the voice input button, where the click or long press operation is the second message input operation.

[0291] In step 606, in response to the second message output operation, the voice signal of the second input information is played in the virtual scene based on the second voice mode adapted to the second influencing factor of the target account.

[0292] Here, based on a second voice mode adapted to the second influencing factor of the target account, a voice signal of the second input information with the target account as the sound source is played in the virtual scene. The second message output operation can be input to a single object, and the listening range of the playback is the single object; the second message output operation can also be output to multiple objects simultaneously, and the listening range of the playback is multiple objects, for example, multiple objects corresponding to multiple accounts in the same camp as the target account.

[0293] Following the example of step 604 above, in response to the second message output operation, the voice signal of the second input information is played in the virtual scene using a happy, relaxed and clear voice effect.

[0294] This application's embodiments can automatically adjust the voice mode according to changes in the virtual object's environment, providing a high degree of environmental sensitivity. Considering changes in multiple dimensions such as role, attributes, state, and input information, the voice mode switching is more comprehensive and accurate. Voice mode switching based on multi-dimensional information makes the interaction of virtual objects more natural and realistic. It allows users to edit and input messages in real time, and the user's input messages can be instantly converted into voice signals for playback, enhancing the real-time nature and responsiveness of the interaction.

[0295] The following will describe exemplary applications of the embodiments of this application in game application scenarios.

[0296] When voice chat is enabled in a game, users often want to change the sound of their voice. The voice processing methods of related technologies usually change the voice based on the character's voice pack used by the user in the game, or use a voice that corresponds to the emotion expressed in the input content. When the scene in which the user-controlled character is located changes, only the ambient sound in the scene can be changed. The voice effect does not match the current scene, causing the user to feel disconnected during the game.

[0297] This application embodiment automatically adjusts the voice and content (i.e., voice mode) based on the selected game character (i.e., the character to which the virtual object controlled by the target account belongs), the current environment (i.e., the environment where the virtual object currently controlled by the target account is located), and the speech content (i.e., the input information of the target account), making it more in line with the actual scene, ensuring the dynamic adaptability of the voice effect, making the voice switching process convenient and quick, and enhancing the fit between the voice mode and the environment.

[0298] Meanwhile, the voice processing method for virtual scenes provided in this application allows for the addition of unique language habits and catchphrases to each game character, thus creating a personalized "voice" for each player. This highly personalized approach goes beyond simple spectrum or waveform modification, delving deeper into the process of person and emotional expression, effectively improving user engagement and satisfaction. Rules can also be set to automatically replace inappropriate or offensive content with other character-related but friendly expressions, maintaining a positive community atmosphere while ensuring fun and immersion.

[0299] In response to a player initiating speech (i.e., responding to the first message input in the message editing control), the system recognizes and analyzes the player's language input (i.e., the first input information of the target account) in real time, while outputting more immersive and context-appropriate voice content. Players only need to enable voice chat in the game to use this function; players can choose their character's voice (i.e., the first voice mode) or enable their original voice (i.e., the third voice mode). The system will adjust the player's voice effects in real time based on the character's characteristics. For example, if character A is silent and concise, the system will automatically recognize and judge the player's speech and output the most concise sentences; if character B likes to use "meow~" as a catchphrase, the system will automatically insert the catchphrase at an appropriate position in the player's speech. When a player uses uncivilized language (i.e., inappropriate language), the system will automatically replace the uncivilized language with the character's catchphrase or other content, such as "hahahaha," with different replacement content for different characters. At the same time, the character's voice effects (i.e., voice modes) will be dynamically adjusted according to the current game context (i.e., the environment in which the virtual object currently controlled by the target account is located) and the content of the user's speech (i.e., the input information of the target account). For example, when encountering a powerful opponent and the user says provocative words, the character's voice will become more powerful and accompanied by a more obvious reverberation effect; when encountering a dangerous situation and the player's chances of winning are very low, the character's voice will automatically become more timid and weak.

[0300] Referring to Figure 7, which is a schematic diagram of voice mode switching in a game according to an embodiment of this application, the user's original voice is used by default in the upper view of Figure 7. In response to switching from the original voice to the character's voice, character A is displayed in the middle view of Figure 7, and the voice mode corresponding to character A is applied. In response to the switching operation for character A, character B is displayed in the lower view of Figure 7, and the voice mode corresponding to character B is applied.

[0301] To achieve the aforementioned voice switching process, an efficient and low-latency distributed framework is established to support real-time voice processing and data transmission. The core components include: a front-end client (e.g., terminal 400 mentioned above), responsible for collecting user voice input and sending it to the back-end server; a back-end server (e.g., server 200 mentioned above), which uses artificial intelligence algorithms to perform voice analysis and processing, and then returns the optimized sound effect (i.e., the second voice mode) to the front-end client; and a database, which stores user personalized settings, role configuration information, and a filter word library.

[0302] Referring to Figure 8, which is a schematic flowchart of the process for generating voice effects according to an embodiment of this application, the steps shown in Figure 8 will be explained using a server (e.g., the server 200 mentioned above) as the executing entity as an example.

[0303] In step 801, voice is acquired in real time and preprocessed.

[0304] In some embodiments, in response to receiving voice input from the user (i.e., input information from the target account), the built-in microphone of the electronic device is used to capture the player's voice and perform initial noise reduction through a noise reduction chip to ensure that the input signal is clear and free of impurities. On this basis, software algorithms can be used for further optimization, such as using a bandpass filter to eliminate low-frequency background noise in the environment.

[0305] For example, the information type of the first or second input information is first detected. If the first or second input information is speech, the activation command of the built-in microphone of the electronic device is triggered. The built-in microphone of the electronic device responds to the activation command and enters the audio acquisition state to receive the speech signal (i.e., the first or second input information) input by the target account in real time. The speech signal is converted into audio data and transmitted to the noise reduction chip of the electronic device. The noise reduction chip activates a preset preliminary noise reduction algorithm (such as an adaptive filtering algorithm) to filter out high-frequency burst noise and obvious interference noise in the audio data, and obtains the preliminary noise-reduced audio data. A bandpass filter (or other preset software noise reduction algorithm) is called, the frequency range of the bandpass filter is set, and the preliminary noise-reduced audio data is processed by the bandpass filter to filter out low-frequency background noise (such as ground vibration noise and low-frequency machine roar) with a frequency lower than the preset frequency threshold, and outputs the further optimized audio data.

[0306] In step 802, the preprocessed speech is converted into text, and natural language processing is performed on the text.

[0307] In some embodiments, step 802 can be implemented by performing the following steps 8021 to 8022.

[0308] In step 8021, emotion recognition is performed on the text, and the audio effects (i.e., audio effect models) corresponding to the recognized emotion categories are adjusted.

[0309] In some embodiments, the emotion recognition model (which may be the pre-trained text emotion recognition model mentioned above) uses the Transformer deep learning algorithm to classify the emotions in the player's speech, such as joy, sadness, anger, etc., and then adjusts the corresponding audio effects accordingly.

[0310] For example, the sentiment recognition model is trained by performing the following processes: obtaining an initialized sentiment recognition model; obtaining information samples and real sentiment labels, where the real sentiment labels represent the sentiment type expressed by the information samples; calling the initialized sentiment recognition model based on the information samples to obtain predicted sentiment labels; determining the sentiment loss value based on the real sentiment labels and predicted sentiment labels; updating the parameters of the initialized sentiment recognition model based on the sentiment loss value to obtain a pre-trained sentiment recognition model.

[0311] In step 8022, the text is subjected to sensitive word detection, and sensitive words (i.e., illegal terms) are replaced with compliant terms.

[0312] In some embodiments, an uncivilized language or special terms (i.e., illegal language) in the text is identified by a sensitive word detection model (which may be the pre-trained text classification model mentioned above), and then a conversion program is called to replace the sensitive words with compliant language to update the text (i.e., input information).

[0313] For example, the text is segmented based on a word segmentation algorithm to obtain multiple words. Each word is mapped to a fixed-dimensional vector space to obtain a word feature vector corresponding to each word. For each word vector feature, a sensitive word detection model is called to classify the word vector feature to obtain the word type corresponding to the word vector feature. The word type includes replacing sensitive words with compliant terms.

[0314] In step 803, speech synthesis and modulation are performed based on the updated text.

[0315] In some embodiments, step 803 can be implemented by performing steps 8031 ​​to 8032 as described below.

[0316] In step 8031, the text is converted into voice effects corresponding to the character.

[0317] In some embodiments, WaveNet neural networks are used to generate fluent and natural human voice samples, with voice effects corresponding to specific attributes of the selected character.

[0318] For example, a mage character uses a deeper and more mysterious tone of voice, while a warrior uses a firm and powerful voice.

[0319] In step 8032, the speech effect is modulated using a modulation algorithm (i.e., modulation method) to obtain the final speech mode (i.e., the second speech mode).

[0320] In some embodiments, a curve varying within a custom parameter range is fitted, including but not limited to frequency, amplitude, and resonance peak position, and physical modeling is combined to recover and simulate the attenuation phenomenon of reflected echoes in the real environment.

[0321] For example, the specific implementation process of the modulation algorithm is as follows: Fourier transform is used to analyze the spectral characteristics of the speech signal, and the basic frequency is adjusted according to the character setting. For example, a magician character may need to lower the basic frequency to create a deep effect; dynamic range compression technology is applied to adjust the volume according to the emotion recognition results, such as increasing the amplitude when happy and decreasing it when sad; formants are identified and relocated through linear predictive coding or other acoustic models to imitate the characteristics of different characters' voices, such as trembling, sharpness and other tone quality changes; convolutional reverberation or physical modeling techniques are used to add echo and spatial sense to make the sound more three-dimensional. For example, long tail echo is added in a cave scene, while reflection effect is reduced in a grassy area.

[0322] The game client (such as terminal 400 mentioned above) collects the player's voice and performs preliminary noise reduction processing. The terminal sends the voice to the server (such as server 200 mentioned above) so that the artificial intelligence model can analyze the content (i.e., input information), identify emotions and detect inappropriate words (i.e., prohibited words). Based on the character characteristics (i.e., character features) and emotional results (i.e., emotional type), the sound effects are adjusted, such as changing the tone and adding reverberation. Relevant configurations, such as the voice characteristics of different characters (i.e., character features), are extracted from the database. After the server completes all processing steps, it transmits the optimized voice (i.e., the second voice mode) back to the game client so that the player hears voice that matches the game atmosphere.

[0323] In this embodiment, when the environment of the virtual object controlled by the target account changes, the system adaptively switches to a voice mode that is at least compatible with the current environment of the virtual object. This achieves automatic voice mode switching, improving the efficiency and convenience of voice mode conversion. Furthermore, by adapting the voice mode to the environment, the perceptual effect of the voice signal played based on the voice mode can be integrated with the user's perception of the environment, facilitating an immersive experience for the player in the virtual environment. The system can automatically adjust the voice mode according to changes in the virtual object's environment, providing high environmental sensitivity. Considering changes in multiple dimensions such as role, attributes, state, and input information, the voice mode switching is more comprehensive and accurate. Voice mode switching based on multi-dimensional information makes the interaction of the virtual object more natural and realistic. Users can edit and input messages in real time, and the user's input messages can be instantly converted into voice signals for playback, enhancing the real-time nature and responsiveness of the interaction. By incorporating specific catchphrases or characteristics of a character into their voice, the personalized performance of virtual characters can be enhanced. Through updating input information and modulating the voice based on the environment and character, the second voice effect can reflect the user's emotional state while also increasing the fun and uniqueness of the virtual object, facilitating personalized interaction. Replacing inappropriate language with compliant language appropriate to the character ensures the coherence and consistency of the voice output, making the character's performance more uniform. Compliant language can be flexibly replaced according to different characters, adapting to different virtual objects and environments, maintaining compliance while showcasing the character's individuality. Different voice modes can be selected through triggering actions, improving the flexibility and autonomy of interaction. Users can turn different voice modes on or off according to their preferences and scenario needs, thus obtaining a personalized voice experience. Furthermore, triggering actions for voice mode controls to switch voice modes improves the efficiency and convenience of voice mode switching.

[0324] The following description continues to illustrate the exemplary structure of the virtual scene voice processing device 4155 provided in this application embodiment as a software module. In some embodiments, as shown in FIG2A, the software module stored in the virtual scene voice processing device 4155 in the memory 415 may include:

[0325] The first display module 41551 is configured to display a virtual scene, wherein the target account in the virtual scene is in a first voice mode adapted to the first influencing factor, and the first influencing factor includes at least the first environment in which the virtual object controlled by the target account is located.

[0326] The mode switching module 41552 is configured to switch the target account from the first voice mode to the second voice mode adapted to the second voice mode in response to the switching of the first influencing factor to the second influencing factor. The second influencing factor includes at least the second environment in which the virtual object currently controlled by the target account is located.

[0327] In some embodiments, the first display module 41551 is further configured to display an influencing factor setting interface; in response to an input operation in the influencing factor setting interface, display at least one dimension of the input, wherein the at least one dimension is at least one of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the input information of the target account; and combine the at least one dimension with the first environment in which the virtual object currently controlled by the target account is located to form a first influencing factor.

[0328] In some embodiments, the first influencing factor further includes the first role to which the virtual object controlled by the target account belongs, and the second influencing factor further includes the second role to which the virtual object controlled by the target account belongs. The mode switching module 41552 is further configured to switch the target account from the first voice mode to the second voice mode adapted to the second influencing factor in response to the role to which the virtual object controlled by the target account switches from the first role to the second role and the virtual object moves from the first environment to the second environment. The second voice mode is used to output a voice signal adapted to both the second environment and the second role.

[0329] In some embodiments, the first influencing factor further includes a first attribute of the virtual object controlled by the target account, and the second influencing factor includes a second attribute of the virtual object controlled by the target account. The mode switching module 41552 is further configured to switch the target account from the first voice mode to a second voice mode adapted to the second influencing factor in response to the first attribute of the virtual object controlled by the target account being switched to the second attribute and the virtual object being moved from the first environment to the second environment. The second voice mode is used to output a voice signal adapted to both the second environment and the second attribute.

[0330] In some embodiments, the mode switching module 41552 is further configured to detect the second input information of the target account to obtain the third emotion type of the second input information; determine the third audio effect model corresponding to the third emotion type; determine the third modulation method for the third speech signal output by the third audio effect model, wherein the third modulation method includes modulating the third speech signal based on the second environment and the second attribute; and determine the second speech mode based on the third modulation method.

[0331] In some embodiments, the mode switching module 41552 is further configured to determine a third update method for the second input information of the target account, wherein the third update method includes: encoding based on the second input information, environmental features of the second environment, and attribute features of the second attribute to obtain a third fusion feature, and decoding based on the third fusion feature to obtain the updated second input information; and using the third modulation method and the third update method as the second voice mode.

[0332] In some embodiments, the first influencing factor further includes a first state parameter of the target account, and the second influencing factor includes a second state parameter of the target account. The mode switching module 41552 is further configured to switch the target account from the first voice mode to a second voice mode adapted to the second influencing factor in response to the target account switching from the first state parameter to the second state parameter and the virtual object moving from the first environment to the second environment. The second voice mode is used to output a voice signal adapted to both the second environment and the second state parameter.

[0333] In some embodiments, the first influencing factor further includes the first input information of the target account, and the second influencing factor includes the second input information of the target account. The mode switching module 41552 is further configured to switch the target account from the first voice mode to the second voice mode adapted to the second influencing factor in response to switching from the first input information of the target account to the second input information and the virtual object moving from the first environment to the second environment. The second voice mode is used to output a voice signal adapted to both the second environment and the second input information.

[0334] In some embodiments, the mode switching module 41552 is further configured to replace the non-compliant term with a compliant term in response to the first input information of the target account including non-compliant terminology, wherein the compliant terminology is adapted to the role to which the virtual object controlled by the target account belongs.

[0335] In some embodiments, the mode switching module 41552 is further configured to, when the information type of the first input information is text, convert multiple words in the first input information into multiple word feature vectors, and classify the multiple words based on the multiple word feature vectors by calling a pre-trained text classification model to obtain the illegal terminology; when the information type of the first input information is speech, convert the first input information into text to be processed, extract acoustic features from the first input information, extract text features from the text to be processed, and classify the multiple words in the text to be processed by calling a pre-trained machine learning model based on the acoustic features and text features to obtain the illegal terminology.

[0336] In some embodiments, the mode switching module 41552 is further configured to switch the target account from the first voice mode to a second voice mode adapted to the second influencing factor in response to the virtual object moving from the first environment to the second environment and at least two other dimensions of the first influencing factor being switched to the second influencing factor, wherein the at least two other dimensions are at least two of the following dimensions: the role to which the virtual object currently controlled by the target account belongs, the attributes of the virtual object currently controlled by the target account, the input information of the target account, and the status of the target account.

[0337] In some embodiments, the mode switching module 41552 is further configured to detect the first input information of the target account to obtain the first emotion type of the first input information; determine the first audio effect model corresponding to the first emotion type; determine the first modulation method of the first voice signal output by the first audio effect model, wherein the first modulation method includes modulating the first voice signal based on the second environment and the second role; and determine the second voice mode based on the first modulation method.

[0338] In some embodiments, the mode switching module 41552 is further configured to invoke a pre-trained text sentiment recognition model based on the first input information of the target account to obtain a first sentiment type of the first input information. The pre-trained text sentiment recognition model is trained by performing the following processes: obtaining an initialized text sentiment recognition model; obtaining information samples and real sentiment labels, wherein the real sentiment labels represent the sentiment type expressed by the information samples; invoking the initialized text sentiment recognition model based on the information samples to obtain predicted sentiment labels; determining a sentiment loss value based on the real sentiment labels and the predicted sentiment labels; and updating the parameters of the initialized text sentiment recognition model based on the sentiment loss value to obtain the pre-trained text sentiment recognition model.

[0339] In some embodiments, the mode switching module 41552 is further configured to determine a first update method for the first input information of the target account, wherein the first update method includes: encoding based on the first input information, environmental features of the second environment, and role features of the second role to obtain a first fusion feature, and decoding based on the first fusion feature to obtain the updated first input information; and using the first modulation method and the first update method as a second voice mode.

[0340] In some embodiments, the mode switching module 41552 is further configured to detect the second input information to obtain a second emotion type of the second input information; determine a second audio effect model corresponding to the second emotion type; determine a second modulation method for the second speech signal output by the second audio effect model, wherein the second modulation method includes modulation based on at least one of the following influencing factors, the at least one influencing factor including: a second environment, the role to which the virtual object currently controlled by the target account belongs, the attribute of the virtual object currently controlled by the target account, and the state of the target account; and determine a second speech mode based on the second modulation method.

[0341] In some embodiments, the mode switching module 41552 is further configured to determine a second update method for the second input information of the target account, wherein the second update method includes: encoding based on the second input information, environmental features of the second environment, and role features of the role to which the virtual object controlled by the target account belongs to obtain a second fusion feature, and decoding based on the second fusion feature to obtain the updated second input information; and using the second modulation method and the second update method as a second voice mode.

[0342] In some embodiments, the mode switching module 41552 is further configured to: determine the frequency of the second voice signal based on the attributes of the virtual object currently controlled by the target account; determine the amplitude of the second voice signal based on the second emotion type; locate the formant of the second voice signal based on the attributes of the virtual object currently controlled by the target account; simulate environmental reflection of the second voice signal based on the second environment to obtain the environmental reflection sound of the second voice signal; and modulate the second voice signal based on the frequency, amplitude, formant, and environmental reflection sound to obtain the third voice signal.

[0343] In some embodiments, the virtual scene includes a first voice mode control corresponding to a first voice mode and a third voice mode control corresponding to a third voice mode. The third voice mode is a mode that plays the voice signal of the target account itself. The third voice mode is in the on state by default. The mode switching module 41552 is also configured to respond to the trigger operation of the first voice mode control, put the first voice mode in the on state, and turn off the third voice mode.

[0344] The following description continues to illustrate the exemplary structure of the virtual scene information processing device 4255 provided in this application embodiment as a software module. In some embodiments, as shown in FIG2B, the software module stored in the virtual scene information processing device 4255 in the memory 425 may include:

[0345] The second display module 42551 is configured to display a virtual scene based on the target account login, wherein the virtual scene includes a message editing control.

[0346] The third display module 42552 is configured to display the first input information in response to the first message input operation in the message editing control.

[0347] The signal playback module 42553 is configured to respond to a first message output operation and play the voice signal of the first input information in a virtual scene based on a first voice mode adapted to a first influencing factor of the target account. The first influencing factor includes at least the first environment in which the virtual object controlled by the target account is currently located.

[0348] In some embodiments, the signal playback module 42553 is further configured to classify the first input information to obtain the emotion type of the first input information; determine the audio effect model corresponding to the emotion type; determine the modulation method of the voice signal output by the audio effect model, wherein the modulation method includes modulation based on at least one of the following influencing factors, the at least one influencing factor including: the environment of the virtual object currently controlled by the target account, the role to which the virtual object currently controlled by the target account belongs, the attribute of the virtual object currently controlled by the target account, and the state of the target account; determine a first voice mode based on the modulation method, and apply the first voice mode to the voice signal output by the audio effect mode to obtain the voice signal of the first input information.

[0349] In some embodiments, the signal playback module 42553 is further configured to determine an update method for the first input information of the target account, wherein the update method includes: encoding based on the first input information, environmental features of the first environment, and role features of the role to which the virtual object controlled by the target account belongs to obtain fusion features, and decoding based on the fusion features to obtain the updated first input information; and using the modulation method and the update method as the first voice mode.

[0350] In some embodiments, the signal playback module 42553 is further configured to, in response to the first environment currently occupied by the virtual object controlled by the target account and at least one other dimension of the first influencing factor of the target account being switched to a second influencing factor, switch the first voice mode adapted to the current first influencing factor of the target account to a second voice mode adapted to the second influencing factor, wherein at least one other dimension is at least one of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the state of the target account, and the current input information of the target account, and the second influencing factor includes at least the second environment occupied by the virtual object controlled by the target account; in response to a second message input operation in the message editing control, display the second input information; and in response to a second message output operation, play the voice signal of the second input information in the virtual scene based on the second voice mode adapted to the second influencing factor of the target account.

[0351] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the virtual scene voice processing method or virtual scene information processing method described above in this application.

[0352] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the virtual scene speech processing method or the virtual scene information processing method provided in this application, such as the virtual scene speech processing method shown in FIG3A, or the virtual scene information processing method shown in FIG4A.

[0353] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0354] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0355] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0356] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0357] In summary, this application's embodiments adaptively switch to a voice mode that is at least compatible with the current environment of the virtual object controlled by the target account when the environment changes. This achieves automatic voice mode switching, improving conversion efficiency and convenience. Furthermore, by adapting the voice mode to the environment, the user's perception of the voice signal played based on the voice mode can be integrated with the perceived environment, facilitating an immersive experience for players in the virtual environment. The voice mode can be automatically adjusted according to changes in the virtual object's environment, providing high environmental sensitivity. Considering changes in multiple dimensions such as role, attributes, status, and input information, the voice mode switching is more comprehensive and accurate. Voice mode switching based on multi-dimensional information makes the interaction of virtual objects more natural and realistic. Users can edit and input messages in real time, and the input messages can be instantly converted into voice signals for playback, enhancing the real-time nature and responsiveness of the interaction. By incorporating specific catchphrases or characteristics of a character into their voice, the personalized performance of virtual characters can be enhanced. Through updating input information and modulating the voice based on the environment and character, the second voice effect can reflect the user's emotional state while also increasing the fun and uniqueness of the virtual object, facilitating personalized interaction. Replacing inappropriate language with compliant language appropriate to the character ensures the coherence and consistency of the voice output, making the character's performance more uniform. Compliant language can be flexibly replaced according to different characters, adapting to different virtual objects and environments, maintaining compliance while showcasing the character's individuality. Different voice modes can be selected through triggering actions, improving the flexibility and autonomy of interaction. Users can turn different voice modes on or off according to their preferences and scenario needs, thus obtaining a personalized voice experience. Furthermore, triggering actions for voice mode controls to switch voice modes improves the efficiency and convenience of voice mode switching.

[0358] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A voice processing method for a virtual scene, the method being executed by an electronic device, the method comprising: Displaying a virtual scene, wherein the target account in the virtual scene is in a first voice mode adapted to a first influencing factor, and the first influencing factor includes at least a first environment in which the virtual object controlled by the target account is located; In response to the first influencing factor switching to the second influencing factor, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second influencing factor includes at least the second environment in which the virtual object currently controlled by the target account is located.

2. The method of claim 1, wherein, The first influencing factor also includes the first role to which the virtual object controlled by the target account belongs, and the second influencing factor also includes the second role to which the virtual object controlled by the target account belongs; The step of switching the target account from the first voice mode to a second voice mode adapted to the second voice mode in response to the first influencing factor switching to the second influencing factor includes: In response to the virtual object controlled by the target account switching its role from the first role to the second role, and the virtual object moving from the first environment to the second environment, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second voice mode is used to output a voice signal adapted to both the second environment and the second role.

3. The method of claim 2, wherein, Before switching the target account from the first voice mode to a second voice mode adapted to the second influencing factor, the method further includes: The first input information of the target account is detected to obtain the first sentiment type of the first input information; Determine the first audio effect model corresponding to the first emotion type; A first modulation method is determined for the first speech signal output by the first audio effect model, wherein the first modulation method includes modulating the first speech signal based on the second environment and the second role; The second voice mode is determined based on the first modulation method.

4. The method of claim 3, wherein, The step of detecting the first input information of the target account to obtain the first sentiment type of the first input information includes: Based on the first input information of the target account, a pre-trained text sentiment recognition model is invoked to obtain the first sentiment type of the first input information. The pre-trained text sentiment recognition model is trained by performing the following processes: obtaining an initialized text sentiment recognition model; obtaining information samples and real sentiment labels, wherein the real sentiment labels represent the sentiment type expressed by the information samples; invoking the initialized text sentiment recognition model based on the information samples to obtain predicted sentiment labels; determining a sentiment loss value based on the real sentiment labels and the predicted sentiment labels; and updating the parameters of the initialized text sentiment recognition model based on the sentiment loss value to obtain the pre-trained text sentiment recognition model.

5. The method of claim 3 or 4, wherein, Determining the second voice mode based on the first modulation scheme includes: A first update method is determined for the first input information of the target account, wherein the first update method includes: encoding based on the first input information, the environmental features of the second environment, and the role features of the second role to obtain a first fusion feature, and decoding based on the first fusion feature to obtain the updated first input information; The first modulation method and the first update method are used as the second voice mode.

6. The method of claim 1, wherein, The first influencing factor also includes the first attribute of the virtual object controlled by the target account, and the second influencing factor includes the second attribute of the virtual object controlled by the target account; The step of switching the target account from the first voice mode to a second voice mode adapted to the second voice mode in response to the switching from the first influencing factor to the second influencing factor includes: In response to the first attribute of the virtual object controlled by the target account being switched to the second attribute, and the virtual object being moved from the first environment to the second environment, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second voice mode is used to output a voice signal adapted to both the second environment and the second attribute.

7. The method of claim 6, wherein, Before switching the first attribute of the virtual object controlled by the target account to the second attribute in response to the virtual object moving from the first environment to the second environment, and switching the target account from the first voice mode to the second voice mode adapted to the second influencing factor, the method further includes: The second input information of the target account is detected to obtain the third sentiment type of the second input information; Determine the third audio effect model corresponding to the third emotion type; A third modulation method is determined for the third speech signal output by the third audio effect model, wherein the third modulation method includes modulating the third speech signal based on the second environment and the second attribute; The second voice mode is determined based on the third modulation method.

8. The method according to claim 7, wherein, Determining the second voice mode based on the third modulation scheme includes: A third update method is determined for the second input information of the target account, wherein the third update method includes: encoding based on the second input information, the environmental features of the second environment, and the attribute features of the second attribute to obtain a third fusion feature, and decoding based on the third fusion feature to obtain the updated second input information; The third modulation method and the third update method are used as the second voice mode.

9. The method of claim 1, wherein, The first influencing factor also includes the first status parameter of the target account, and the second influencing factor includes the second status parameter of the target account; The step of switching the target account from the first voice mode to a second voice mode adapted to the second voice mode in response to the switching from the first influencing factor to the second influencing factor includes: In response to the target account switching from the first state parameter to the second state parameter, and the virtual object moving from the first environment to the second environment, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second voice mode is used to output a voice signal adapted to both the second environment and the second state parameter.

10. The method of claim 1, wherein, The first influencing factor also includes the first input information of the target account, and the second influencing factor includes the second input information of the target account; The step of switching the target account from the first voice mode to a second voice mode adapted to the second voice mode in response to the switching from the first influencing factor to the second influencing factor includes: In response to switching from the first input information of the target account to the second input information, and the virtual object moving from the first environment to the second environment, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the second voice mode is used to output a voice signal adapted to both the second environment and the second input information.

11. The method of claim 10, wherein, Before switching the target account from the first voice mode to a second voice mode adapted to the second influencing factor, the method further includes: The second input information is detected to obtain the second emotion type of the second input information; Determine the second audio effect model corresponding to the second emotion type; A second modulation method is determined for the second speech signal output by the second audio effect model, wherein the second modulation method includes modulation based on at least one of the following influencing factors: the second environment, the role to which the virtual object currently controlled by the target account belongs, the attributes of the virtual object currently controlled by the target account, and the state of the target account; The second voice mode is determined based on the second modulation method.

12. The method of claim 11, wherein, Determining the second voice mode based on the second modulation scheme includes: A second update method for the second input information of the target account is determined, wherein the second update method includes: encoding based on the second input information, the environmental characteristics of the second environment, and the role characteristics of the role to which the virtual object controlled by the target account belongs, to obtain a second fusion feature, and decoding based on the second fusion feature to obtain the updated second input information; The second modulation method and the second update method are used as the second voice mode.

13. The method of claim 11 or 12, wherein, The second modulation method includes: The frequency of the second voice signal is determined based on the attributes of the virtual object currently controlled by the target account; The amplitude of the second voice signal is determined based on the second emotion type; Locate the formant peak of the second voice signal based on the attributes of the virtual object currently controlled by the target account; Based on the second environment, environmental reflection simulation is performed on the second speech signal to obtain the environmental reflection sound of the second speech signal; The second speech signal is modulated according to the frequency, amplitude, resonant peak and ambient reflected sound to obtain the third speech signal.

14. The method according to any one of claims 10 to 13, wherein, Before switching the target account from the first input information to the second input information in response to the virtual object moving from the first environment to the second environment, and switching the target account from the first voice mode to the second voice mode adapted to the second influencing factor, the method further includes: In response to the first input information of the target account including prohibited terms, the prohibited terms are replaced with compliant terms, wherein the compliant terms are adapted to the role to which the virtual object controlled by the target account belongs.

15. The method of claim 14, wherein, Before replacing the non-compliant terms with compliant terms, the method further includes: When the information type of the first input information is text, multiple words in the first input information are converted into multiple word feature vectors. Based on the multiple word feature vectors, a pre-trained text classification model is called to classify the multiple words to obtain the illegal language. If the information type of the first input information is speech, the first input information is converted into text to be processed, acoustic features are extracted from the first input information, text features are extracted from the text to be processed, and a pre-trained machine learning model is called based on the acoustic features and the text features to classify multiple words in the text to be processed, thereby obtaining the illegal language.

16. The method of claim 1, wherein, The step of switching the target account from the first voice mode to a second voice mode adapted to the second voice mode in response to the first influencing factor switching to the second influencing factor includes: In response to the virtual object moving from the first environment to the second environment, and at least two other dimensions of the first influencing factor being switched to the second influencing factor, the target account is switched from the first voice mode to a second voice mode adapted to the second influencing factor, wherein the at least two other dimensions are at least two of the following dimensions: the role to which the virtual object currently controlled by the target account belongs, the attributes of the virtual object currently controlled by the target account, the input information of the target account, and the status of the target account.

17. The method of any one of claims 1 to 16, wherein, The method further includes: Displays the settings interface for influencing factors; In response to an input operation in the influencing factor settings interface, at least one dimension of the input is displayed, wherein the at least one dimension is at least one of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the input information of the target account; The first influencing factor is formed by combining the at least one dimension with the first environment in which the virtual object currently controlled by the target account is located.

18. The method of any one of claims 1 to 17, wherein, The virtual scene includes a first voice mode control corresponding to the first voice mode and a third voice mode control corresponding to the third voice mode. The third voice mode is a mode that plays the voice signal of the target account itself. The third voice mode is enabled by default. Before switching the target account from the first voice mode to a second voice mode adapted to the second influencing factor in response to the first influencing factor switching to the second influencing factor, the method further includes: In response to a trigger operation on the first voice mode control, the first voice mode is turned on, and the third voice mode is turned off.

19. A method for processing information in a virtual scene, the method being executed by an electronic device, the method comprising: Displays a virtual scene based on login to the target account, wherein the virtual scene includes a message editing control; In response to a first message input operation in the message editing control, the first input information is displayed; In response to the first message output operation, the voice signal of the first input information is played in the virtual scene based on a first voice mode adapted to a first influencing factor of the target account, wherein the first influencing factor includes at least a first environment in which the virtual object controlled by the target account is currently located.

20. The method of claim 19, wherein, After playing the voice signal of the first input information in the virtual scene, the method further includes: In response to the first environment currently occupied by the virtual object controlled by the target account, and at least one other dimension of the first influencing factor of the target account being switched to the second influencing factor, the first voice mode adapted to the first influencing factor of the target account is switched to the second voice mode adapted to the second influencing factor. The at least one other dimension is at least one of the following dimensions: the role to which the virtual object controlled by the target account belongs, the attributes of the virtual object controlled by the target account, the status of the target account, and the current input information of the target account. The second influencing factor includes at least the second environment occupied by the virtual object controlled by the target account. In response to a second message input operation in the message editing control, the second input information is displayed; In response to the second message output operation, the voice signal of the second input information is played in the virtual scene based on the second voice mode adapted to the second influencing factor of the target account.

21. The method of claim 19 or 20, wherein, Before playing the voice signal of the first input information in the virtual scene, the method further includes: The first input information is classified to obtain the sentiment type of the first input information; Determine the audio effect model corresponding to the emotion type; Determine the modulation method for the speech signal output by the audio effect model, wherein the modulation method includes modulation based on at least one of the following influencing factors: the environment of the virtual object currently controlled by the target account, the role to which the virtual object currently controlled by the target account belongs, the attributes of the virtual object currently controlled by the target account, and the status of the target account; The first voice mode is determined based on the modulation method, and the voice signal output by the audio effect mode is applied to the first voice mode to obtain the voice signal of the first input information.

22. The method of claim 21, wherein, Determining the first voice mode based on the modulation scheme includes: Determine an update method for the first input information of the target account, wherein the update method includes: encoding based on the first input information, the environmental characteristics of the first environment, and the role characteristics of the role to which the virtual object controlled by the target account belongs, to obtain a fusion feature, and decoding based on the fusion feature to obtain the updated first input information; The modulation method and the update method are used as the first voice mode.

23. A voice processing device for a virtual scene, the device comprising: The first display module is configured to display a virtual scene, wherein the target account in the virtual scene is in a first voice mode adapted to a first influencing factor, and the first influencing factor includes at least a first environment in which the virtual object controlled by the target account is located; The mode switching module is configured to switch the target account from the first voice mode to a second voice mode adapted to the second voice mode in response to the first influencing factor switching to the second influencing factor, wherein the second influencing factor includes at least the second environment in which the virtual object currently controlled by the target account is located.

24. An information processing device for a virtual scene, the device comprising: The second display module is configured to display a virtual scene based on the target account login, wherein the virtual scene includes a message editing control; The third display module is configured to display first input information in response to a first message input operation in the message editing control; The signal playback module is configured to respond to a first message output operation and play the voice signal of the first input information in the virtual scene based on a first voice mode adapted to a first influencing factor of the target account, wherein the first influencing factor includes at least a first environment in which the virtual object controlled by the target account is currently located.

25. An electronic device, the electronic device comprising: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the voice processing method for a virtual scene according to any one of claims 1 to 18, or the information processing method for a virtual scene according to any one of claims 19 to 22.

26. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program, when executed by a processor, implement the voice processing method for a virtual scene according to any one of claims 1 to 18, or the information processing method for a virtual scene according to any one of claims 19 to 22.

27. A computer program product comprising computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program, when executed by a processor, implement the voice processing method for a virtual scene according to any one of claims 1 to 18, or the information processing method for a virtual scene according to any one of claims 19 to 22.

Citation Information

Patent Citations

  • Emotion processing method for input information and electronic equipment

    CN114495988A

  • Audio data processing method, related device, equipment and storage medium

    CN118113249A

  • Method and system for providing interactive personalized immersive content

    US20220335476A1