Voice interaction processing method, device and equipment and computer readable storage medium

By receiving voice recognition results and game logos in game games and automatically selecting and displaying emoticon images, the problem of single voice to text function is solved, and the richness and fun of game interaction is improved.

CN120242498APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410016404.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In game games, the existing voice to text function lacks expression interaction, resulting in a single and boring gameplay that cannot meet the diverse needs of players.

Method used

By receiving the voice recognition results and game logo of the voice recognition server, the expression classification logo is determined, candidate expression images are obtained from the expression image library, and the target expression images are determined based on these images, and sent to the terminal to display, realizing the automatic selection and display of the expression images.

Benefits of technology

Without increasing the complexity of user operations, the efficiency and frequency of expression images selection during gameplay are improved, the richness and fun of game interactions are enhanced, and user viscosity is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120242498A_ABST
    Figure CN120242498A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction processing method, device and equipment, a computer program product and a computer readable storage medium. The method comprises the following steps: receiving a voice recognition result and a game identifier sent by a voice recognition server, wherein the voice recognition result is obtained by the voice recognition server through recognizing voice data collected by a first terminal in a game playing process; determining an expression classification identifier corresponding to the voice recognition result; obtaining at least one candidate expression image corresponding to the expression classification identifier from an expression image library corresponding to the game identifier; determining a target expression image corresponding to the voice recognition result based on the at least one candidate expression image; and sending the target expression image to the first terminal. By means of the method and device, the expression image selection efficiency in the game playing process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to voice processing technology, and in particular to a voice interaction processing method, apparatus, device and computer-readable storage medium. Background Art

[0002] In a game session, players can communicate with their teammates through voice or text. To avoid affecting game operations by typing text, players generally prefer to send voice messages when sending messages. However, when receiving a voice message, the playback duration of the voice is much longer than the duration of viewing text content. Therefore, when conducting instant messaging during a game session, the communication method of converting voice to text has been widely used. The voice-to-text function can not only avoid text input but also save the other party's message viewing time. However, the voice-to-text function can only convert voice into pure text content, lacking expression interaction and having a single gameplay. Summary of the Invention

[0003] Embodiments of this application provide a voice interaction processing method, apparatus and computer-readable storage medium, which can improve the selection efficiency of expression images during a game session.

[0004] The technical solution of the embodiments of this application is implemented as follows:

[0005] Embodiments of this application provide a voice interaction processing method, the method comprising:

[0006] Receiving a voice recognition result and a game identifier sent by a voice recognition server, where the voice recognition result is obtained by the voice recognition server recognizing voice data collected by a first terminal during a game session;

[0007] Determining an expression classification identifier corresponding to the voice recognition result;

[0008] Obtaining at least one candidate expression image corresponding to the expression classification identifier from an expression image library corresponding to the game identifier;

[0009] Determining a target expression image corresponding to the voice recognition result based on the at least one candidate expression image;

[0010] Sending the target expression image to the first terminal.

[0011] Embodiments of this application provide a voice interaction processing method, the method comprising:

[0012] During a game session, in response to a received voice collection instruction, collecting voice data;

[0013] In response to the received voice collection completion instruction, send the collected voice data to the voice recognition server;

[0014] Receive the voice recognition result sent by the voice recognition server, and receive the expression image corresponding to the voice recognition result sent by the expression prediction server, where the target expression image is determined by the expression prediction server based on the voice recognition result sent by the voice recognition server and the game identifier;

[0015] Display the voice recognition result and the target expression image on the information input interface;

[0016] In response to the received message sending instruction, send at least one of the voice recognition result and the target expression image to the second terminal.

[0017] An embodiment of the present application provides a voice interaction processing device, including:

[0018] A first receiving module, configured to receive the voice recognition result and the game identifier sent by the voice recognition server, where the voice recognition result is obtained by the voice recognition server recognizing the voice data collected by the first terminal during the game session;

[0019] A first determining module, configured to determine the expression classification identifier corresponding to the voice recognition result;

[0020] A first obtaining module, configured to obtain at least one candidate expression image corresponding to the expression classification identifier from the expression image library corresponding to the game identifier;

[0021] A second determining module, configured to determine the target expression image corresponding to the voice recognition result based on the at least one candidate expression image;

[0022] A first sending module, configured to send the target expression image to the first terminal.

[0023] An embodiment of the present application provides a voice interaction processing device, including:

[0024] A voice collection module, configured to collect voice data in response to the received voice collection instruction during the game session;

[0025] A second sending module, configured to send the collected voice data to the voice recognition server in response to the received voice collection completion instruction;

[0026] A second receiving module, configured to receive the speech recognition result sent by the speech recognition server, and receive the expression image corresponding to the speech recognition result sent by the expression prediction server, where the target expression image is determined by the expression prediction server based on the speech recognition result sent by the speech recognition server and the game identifier;

[0027] A first display module, configured to display the speech recognition result and the target expression image on the information input interface;

[0028] A second sending module, configured to, in response to a received message sending instruction, send at least one of the speech recognition result and the target expression image to a second terminal.

[0029] An embodiment of the present application provides an electronic device, where the electronic device includes:

[0030] A memory, configured to store computer-executable instructions;

[0031] A processor, configured to implement the speech interaction processing method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0032] An embodiment of the present application provides a computer-readable storage medium, storing a computer program or computer-executable instructions, which are configured to implement the speech interaction processing method provided by the embodiment of the present application when being executed by a processor.

[0033] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, where when the computer program or computer-executable instructions are executed by a processor, the speech interaction processing method provided by the embodiment of the present application is implemented.

[0034] The embodiment of the present application has the following beneficial effects:

[0035] After receiving the voice to be recognized collected by the first terminal during the game session, the voice recognition server performs recognition on the voice to obtain a voice recognition result, and then sends the voice recognition result and the game identifier to the expression prediction server. The expression prediction server determines the expression classification identifier corresponding to the voice recognition result, thereby obtains at least one candidate expression image corresponding to the expression classification identifier from the expression image library corresponding to the game identifier, and finally determines the target expression image corresponding to the voice recognition result based on the at least one candidate expression image, and then sends the target expression image to the first terminal, so that the target expression image matching the voice recognition result is displayed on the first terminal. In this way, the game expression image matching the text content corresponding to the voice data input by the player during the game session can be automatically determined, which can improve the selection efficiency and usage efficiency of the expression image during the game session without increasing the operation complexity of the user, thereby improving the richness and interest of the interaction during the game session and increasing the user stickiness. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic diagram of the network architecture of the game system 100 provided by an embodiment of the present application;

[0037] Figure 2 is a schematic diagram of the structure of the expression prediction server 400 provided by an embodiment of the present application;

[0038] Figure 3A is a schematic diagram of an implementation process of the voice interaction processing method provided by an embodiment of the present application;

[0039] Figure 3B is a schematic diagram of the implementation process of determining the expression classification identifier of the voice recognition result provided by an embodiment of the present application;

[0040] Figure 4A is a schematic diagram of an implementation process of determining the target expression image provided by an embodiment of the present application;

[0041] Figure 4B is a schematic diagram of another implementation process of determining the target expression image provided by an embodiment of the present application;

[0042] Figure 5 is a schematic diagram of another implementation process of the voice interaction processing method provided by an embodiment of the present application;

[0043] Figure 6 is a schematic diagram of still another implementation process of the voice interaction processing method provided by an embodiment of the present application;

[0044] Figure 7 is a schematic diagram of the interface for the player to click the voice-to-text function control to record voice during the game session provided by an embodiment of the present application;

[0045] Figure 8 It is a schematic diagram of the interface for converting speech to text provided by an embodiment of the present application;

[0046] Figure 9 It is a schematic diagram of the interface for displaying game expression images provided by an embodiment of the present application;

[0047] Figure 10 It is another schematic diagram of the implementation process of the voice interaction processing method provided by an embodiment of the present application;

[0048] Figure 11A It is a network structure diagram of the voice interaction processing method provided by an embodiment of the present application;

[0049] Figure 11B It is a schematic diagram of the information interaction between the game client, the ASR server, and the expression configuration server provided by an embodiment of the present application;

[0050] Figure 12A It is a schematic diagram of the structure of the expression classification prediction model provided by an embodiment of the present application;

[0051] Figure 12B It is a schematic diagram of the training data of the expression classification prediction model and the prediction results of the training samples provided by an embodiment of the present application;

[0052] Figure 13 It is a schematic diagram of the information interaction between the expression configuration server, the LSTM inference server, and the game client provided by an embodiment of the present application. Detailed implementation manners

[0053] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0054] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0055] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0056] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the embodiments of this application are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0057] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described, and the nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0058] 1) Expression image. An expression is an external manifestation mode of the subjective experience of emotion. In the context of network applications, an expression image is an artistic expression of daily life and is regarded as the third language besides speech and text.

[0059] 2) Game expression image. An expression image created for game characters in different games;

[0060] 3) Bidirectional Long Short-Term Memory Network (Bi-Directional LSTM). An RNN model developed on the basis of ordinary LSTM. The main idea is that for a sequence of data, when processing a specific element, it not only depends on the data elements before this element but may also depend on the data elements after this element.

[0061] 4) Game ID (GAME ID): The unique identifier of a game in the internal system, which can locate a certain online game.

[0062] The embodiments of this application provide a voice interaction processing method, device, equipment, computer-readable storage medium, and computer program product, which can solve the problems of single voice-to-text function and dull interaction in game applications in related technologies. The following describes the exemplary applications of the electronic devices provided in the embodiments of this application. The devices provided in the embodiments of this application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (such as mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), smart phones, smart speakers, smart watches, smart TVs, in-vehicle terminals, etc., or can also be implemented as servers. The following will describe the exemplary applications when the device is implemented as a server.

[0063] See Figure 1 , Figure 1 is a schematic diagram of the network architecture of the game system 100 provided in the embodiments of this application. To support a game application, terminals (exemplarily shown as the first terminal 200-1 and the second terminal 200-2 in Figure 1 ), through the network ( Figure 1(not shown) connects the speech recognition server 300, the expression prediction server 400, and the game server 500, and the speech recognition server 300 is connected to the expression prediction server 400 and the game server 500 through a network. The network can be a wide area network, a local area network, or a combination of the two. In some embodiments, the speech recognition server 300 and the game server 500 can be the same server, that is, the game server can provide speech recognition functions.

[0064] Various application programs can be installed in the first terminal 200-1 and the second terminal 200-2, such as game application programs, shopping application programs, instant messaging application programs, etc. The first terminal 200-1 and the second terminal 200-2 can start a game session through the game application program, can start a game session through a game applet in other application programs, or can start a game session from a game web page. During the game session, the game session screen is displayed on the graphical interface of the first terminal 200-1. The first terminal 200-1 is also used to receive the player's game operations, send the operation data to the game server 500, and receive the game data sent by the game server 500. The first terminal 200-1 can collect the voice data during the game session and send the voice data to the speech recognition server 300. The speech recognition server 300 recognizes the voice data to obtain a speech recognition result. The speech recognition server 300 sends the speech recognition result to the first terminal 200-1, and the speech recognition server 300 sends the speech recognition result and the game identifier to the expression prediction server 400. The expression prediction server 400 determines a target expression image based on the speech recognition result and the game identifier and sends it to the first terminal 200-1. The first terminal 200-1 sends at least one of the speech recognition result and the target expression image to the second terminal 200-2 via the game server 500 in response to the received message sending instruction.

[0065] In some embodiments, the speech recognition server 300, the expression prediction server 400, and the game server 500 can be independent physical servers, can also be a server cluster or a distributed system composed of multiple physical servers, or can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The first terminal 200-1 and the second terminal 200-2 can be smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, vehicle terminals, etc., but are not limited thereto. The terminals and the servers can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.

[0066] See Figure 2 , Figure 2 which is a schematic structural diagram of the expression prediction server 400 provided by an embodiment of the present application. Figure 2 As shown, the expression prediction server 400 includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the expression prediction server 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 440.

[0067] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0068] The user interface 430 includes one or more output devices 431 capable of presenting media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.

[0069] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located away from the processor 410.

[0070] The memory 450 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0071] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of these data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0072] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks.

[0073] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.

[0074] The presentation module 453 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (such as a display screen, a speaker, etc.).

[0075] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one of one or more input devices 432.

[0076] In some embodiments, the device provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is a voice interaction processing device 455 stored in the memory 450, which can be software in the form of programs and plugins, etc., including the following software modules: a first receiving module 4551, a first determining module 4552, a first obtaining module 4553, a second determining module 4554, and a first sending module 4555. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.

[0077] In some other embodiments, the device provided by the embodiments of the present application may be implemented in a hardware manner. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the voice interaction processing method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0078] The exemplary applications and implementations of the server provided by the embodiments of the present application will be combined to illustrate the voice interaction processing method provided by the embodiments of the present application.

[0079] Next, the voice interaction processing method provided by the embodiments of the present application will be described. As mentioned above, the electronic device implementing the voice interaction processing method of the embodiments of the present application may be a terminal, a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.

[0080] See Figure 3A , Figure 3A is a schematic diagram of an implementation process of the voice interaction processing method provided by the embodiments of the present application, which will be described in combination with Figure 3A the steps shown. Figure 3A The subject of the step is the expression prediction server.

[0081] In step 101, the voice recognition result and the game identifier sent by the voice recognition server are received.

[0082] Among them, the voice recognition result is obtained by the voice recognition server by recognizing the voice data collected during the game session of the first terminal. The voice recognition result is a text in a preset language, and the preset language may be the language type of the collected voice data. Exemplarily, if the collected voice data is Chinese voice, then the voice recognition result is also a Chinese text; if the collected voice data is English voice, then the voice recognition result is also an English text. In some embodiments, the preset language may also be the target language set by the terminal when turning on voice-to-text. For example, the target language may be set to Chinese, and then the voice recognition result is a Chinese text.

[0083] The game identifier is the game ID of the game session currently being played on the terminal. For example, it can be the game name or the game number. The game identifier is unique, that is, different games correspond to different game identifiers, and the unique corresponding game can be determined through the game identifier.

[0084] In step 102, determine the expression classification identifier corresponding to the speech recognition result.

[0085] In some embodiments, the speech recognition result can be subjected to sentiment analysis to obtain a sentiment analysis result, and then the expression classification identifier can be determined based on the sentiment analysis result. The expression classification identifier can indicate the sentiment classification corresponding to the speech recognition result. For example, it can be happy, excited, sad, etc.

[0086] In some embodiments, refer to Figure 3B , step 102 can be implemented through Figure 3B the steps 1021 to 1024 shown below. The following is a specific description.

[0087] In step 1021, obtain a trained expression classification prediction model.

[0088] In some embodiments, the expression classification prediction model can be a neural network model. For example, it can be a convolutional neural network model or a recurrent neural network model. Exemplarily, the expression classification prediction model can be a bidirectional long short-term memory network model. The bidirectional long short-term memory network model is a recurrent neural network model that can capture sequence information. The bidirectional long short-term memory network model is divided into 2 independent long short-term memory networks. The input sequence is input into the 2 long short-term memory networks for feature extraction in forward and reverse order respectively. The word vector formed by splicing the 2 output vectors (i.e., the extracted feature vectors) is used as the final feature expression of the word. The model design concept of the bidirectional long short-term memory network is to make the feature data obtained at time t have information between the past and the future at the same time. This neural network structure model is superior to a single long short-term memory network structure model in terms of text feature extraction efficiency and performance. It should be noted that the network parameters of the 2 long short-term memory networks in the bidirectional long short-term memory network are independent of each other, and they only share the word-embedding word vector list.

[0089] In step 1022, use the embedding layer in the trained expression classification prediction model to process the speech recognition result to obtain the embedding vector corresponding to the speech recognition result.

[0090] In some embodiments, the trained facial expression classification prediction model includes an embedding layer, a fully connected layer, an activation layer, and an output layer. The embedding layer is used to convert the input discrete data (usually vocabulary) into a continuous vector representation. When implementing step 1022, first, the speech recognition result is segmented to obtain each segment, and then the embedding layer in the trained facial expression classification prediction model is used to process each segment to obtain the embedding vector corresponding to the speech recognition result.

[0091] In step 1023, the fully connected layer, activation layer, and output layer in the trained facial expression classification prediction model are used to perform prediction processing on the embedding vector to obtain the facial expression classification identifier corresponding to the speech recognition result.

[0092] In some embodiments, the role of the fully connected layer is to perform feature learning and model decision-making. In the fully connected layer, each node is connected to all nodes in the previous layer. By learning the weights of these connections, useful features can be extracted from the input data and decisions can be made. After performing a fully connected process on the embedding vector using the fully connected layer of the trained facial expression classification prediction model to obtain the fully connected processing result, the activation function in the activation layer is used to perform normalization processing on the fully connected processing result to obtain the probability values that the speech recognition result belongs to each facial expression classification. The probability values are real numbers between 0 and 1. Finally, the output layer in the trained facial expression classification prediction model determines the facial expression classification identifier corresponding to the highest probability value among the probability values as the facial expression classification identifier corresponding to the speech recognition result.

[0093] Continue to refer to Figure 3A , and continue to describe step 102.

[0094] In step 103, at least one candidate facial expression image corresponding to the facial expression classification identifier is obtained from the facial expression image library corresponding to the game identifier.

[0095] In some embodiments, the facial expression prediction server stores facial expression image libraries corresponding to multiple game applications. Each facial expression image library includes multiple facial expression images that can be selected in the game application, and each facial expression image has its own facial expression classification identifier, thereby characterizing the type to which the facial expression image belongs. When implementing step 103, the facial expression image library corresponding to the game identifier can be determined first, and then at least one candidate facial expression image corresponding to the facial expression classification identifier is obtained from this facial expression image library.

[0096] In step 104, the target facial expression image corresponding to the speech recognition result is determined based on at least one candidate facial expression image.

[0097] In some embodiments, when the number of candidate expression images is one, in order to improve the determination efficiency of the target expression image, the candidate expression image can be directly determined as the target expression image. In some embodiments, the image to be synthesized corresponding to the first terminal can also be obtained, where the image to be synthesized is the facial image in the game avatar of the terminal or the facial image of the virtual character adopted by the terminal; then the candidate expression image and the image to be synthesized are subjected to a synthesis process to obtain the target expression image.

[0098] In the actual application process, the game avatar of the terminal can be obtained first, and it is determined whether the game avatar includes a facial image. When the game avatar includes a facial image, the facial image in the game avatar can be determined as the image to be synthesized. When the game avatar does not include a facial image, the facial image of the virtual character adopted by the terminal in the game session can be obtained, and the facial image of the virtual character is determined as the image to be synthesized. When performing the synthesis process on the candidate expression image and the image to be synthesized, facial detection can be performed on the candidate expression image first to determine the facial image in the candidate expression image, and then the facial image in the candidate expression image is replaced with the image to be synthesized to obtain the target expression image. In this way, the facial image in the target expression image is the facial image of the game avatar or the virtual character of the terminal, which can realize the personalized customization of the target expression image and improve the interest of information interaction.

[0099] In some embodiments, when the number of candidate expression images is two or more, step 104 can be implemented through Figure 4A the steps 1041A to 1042A shown below, which will be described in conjunction with Figure 4A below.

[0100] In step 1041A, when the number of candidate expression images is two or more, the popularity value of each candidate expression image is obtained.

[0101] Among them, the popularity value represents the number of times the candidate expression image is selected and used. That is to say, the popularity value of an expression image can reflect the popularity of the expression image. The popularity value of the expression image can be obtained by counting the number of times it is selected and used by players after the game application is launched.

[0102] In step 1042A, the candidate expression image corresponding to the highest popularity value among the respective popularity values is determined as the target expression image.

[0103] Through the above steps 1041A to 1042A, when the number of candidate expression images is two or more, the candidate expression image with the highest popularity value can be determined as the target expression image. In this way, the most popular expression image can be recommended to players, and the usage times of the target expression image can be increased.

[0104] In some embodiments, when the number of candidate expression images is two or more, step 104 may be implemented through Figure 4B steps 1041B to 1043B shown below, which will be described in conjunction with Figure 4B the following.

[0105] In step 1041B, when the number of candidate expression images is two or more, obtain the popularity values of each candidate expression image.

[0106] In step 1042B, obtain the image to be synthesized corresponding to the first terminal.

[0107] Wherein, the image to be synthesized is the facial image in the game avatar or the facial image of the virtual character adopted by the first terminal; in some embodiments, first obtain the game avatar of the terminal, determine whether the game avatar includes a facial image, when the game avatar includes a facial image, the facial image in the game avatar may be determined as the image to be synthesized, and when the game avatar does not include a facial image, the facial image of the virtual character adopted by the terminal in the game session may be obtained and the facial image of the virtual character may be determined as the image to be synthesized.

[0108] In step 1043B, perform a synthesis process on the candidate expression image corresponding to the highest popularity value among the popularity values and the image to be synthesized to obtain the target expression image.

[0109] In some embodiments, when performing the synthesis process on the candidate expression image corresponding to the highest popularity value among the popularity values and the image to be synthesized, facial detection may be first performed on the candidate expression image to determine the facial image in the candidate expression image, and then the facial image in the candidate expression image may be replaced with the image to be synthesized to obtain the target expression image.

[0110] Through the above steps 1041B to 1043B, when the number of candidate expression images is two or more, the candidate expression image with the highest popularity value may be synthesized with the facial image in the game avatar or with the facial image of the virtual character adopted by the terminal to obtain the target expression image. In this way, not only the most popular expression image can be recommended to the player, but also the interest of the target expression image can be improved through the process of synthesizing the images.

[0111] In step 105, send the target expression image to the first terminal.

[0112] In some embodiments, the expression prediction server sends the target expression image to the terminal, so that the target expression image is displayed in the display interface of the terminal, thereby realizing the automatic selection of the target expression image, reducing the operation of the player, and increasing the usage frequency of the game expression image.

[0113] In some embodiments, the speech recognition server may send the speech recognition result and the speech to be recognized to the expression prediction server. After determining the target expression image, the expression prediction server may add a voice playback control to the target expression image and bind the speech to be recognized to the target expression image. The expression prediction server sends the target expression to the first terminal. It can be understood that the target expression image bound with the speech to be recognized is sent to the first terminal. Therefore, when the target expression image is clicked, the speech to be recognized can be played, thus improving the diversity of interactive communication.

[0114] In the speech interaction processing method provided in the embodiments of the present application, after the speech recognition server recognizes the speech to be recognized collected by the first terminal during the game session and obtains the speech recognition result, it sends the speech recognition result and the game identifier to the expression prediction server. The expression prediction server determines the expression classification identifier corresponding to the speech recognition result, and thus obtains at least one candidate expression image corresponding to the expression classification identifier from the expression image library corresponding to the game identifier. Finally, based on the at least one candidate expression image, the target expression image corresponding to the speech recognition result is determined, and the target expression image is sent to the terminal, so that the target expression image matching the speech recognition result is displayed in the terminal. In this way, the game expression image matching the text content corresponding to the voice data input by the player during the game session can be automatically determined, and the selection efficiency and usage efficiency of the expression image during the game session can be improved without increasing the operation complexity of the user, thereby improving the richness and interest of the interaction during the game session and increasing the user stickiness.

[0115] Based on the foregoing embodiments, the embodiments of the present application further provide a speech interaction processing method, which is applied to Figure 1 the network architecture shown in Figure 5 FIG. is another schematic implementation flowchart of the speech interaction processing method provided in the embodiments of the present application. The following will be described in conjunction with Figure 5 this.

[0116] In step 201, the first terminal 200-1 responds to the game session start instruction and starts the game session.

[0117] In some embodiments, after the first terminal 200-1 starts a game application or opens a game webpage, in response to a received game session start instruction, it starts a game session. The first terminal 200-1 displays a virtual scene in the game in the human-computer interaction interface, and in the human-computer interaction interface, the virtual scene can be displayed from the first-person perspective (for example, playing the virtual object in the game from the perspective of the player himself); it can also be displayed from the third-person perspective (for example, the player chasing the virtual object in the game to play); it can also be displayed from an aerial perspective; among them, the above perspectives can be switched arbitrarily. The virtual scene includes virtual objects, virtual buildings, virtual plants, special effect props, etc. The virtual object can be a game character controlled by the user, that is, the virtual object is controlled by the real user. The first terminal 200-1 controls the virtual object to move in the virtual scene in response to the operation of the real user on the controller (including touch screen, voice control switch, keyboard, mouse, and joystick, etc.) during the game session, and sends game data to the game server.

[0118] In step 202, during the game session, the first terminal 200-1 collects voice data in response to a received voice collection instruction.

[0119] In some embodiments, during the game session, a player can send text or voice messages to other players in the same team to communicate game tactics or predict the game tactics of the other party, etc. To avoid the impact of the text input process on the game process, the player can turn on the voice-to-text function. At this time, when communicating with other players during the game session, voice data can be collected using a voice collection device, but the message sent to other players is a text message. In some embodiments, when the first terminal 200-1 receives a touch operation on the voice collection operation control, it determines that a voice collection instruction has been received, and at this time, it collects voice data.

[0120] It should be noted that in the embodiments of the present application, when the terminal collects voice data in response to a voice collection instruction, it needs to obtain the user's permission or consent, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0121] In step 203, the first terminal 200-1 sends the collected voice data to the voice recognition server 300 in response to a received voice collection completion instruction.

[0122] In some embodiments, during the process of the first terminal 200-1 collecting voice data, if it receives a touch operation on the voice collection operation control again, it determines that a voice collection completion instruction has been received, and at this time, it sends the collected voice data to the voice recognition server 300.

[0123] In step 204, the speech recognition server 300 recognizes the received speech data to obtain a speech recognition result.

[0124] In some embodiments, when the speech recognition server 300 recognizes the received speech data, it first performs analog-to-digital conversion, noise reduction processing, speech enhancement processing, endpoint detection, etc. on the speech data, and then extracts features from the processed speech data to obtain speech features. After that, it uses the trained acoustic model and language model to perform encoding and decoding processing on the speech features to obtain a speech recognition result, which is text content.

[0125] In step 205, the speech recognition server 300 sends the speech recognition result to the first terminal 200-1.

[0126] In step 206, the speech recognition server 300 sends the speech recognition result and the game identifier to the expression prediction server 400.

[0127] In some embodiments, the speech recognition result is the text content included in the speech data, and the game identifier is used to uniquely identify the online game application.

[0128] In step 207, the expression prediction server 400 determines the expression classification identifier corresponding to the speech recognition result.

[0129] In step 208, the expression prediction server 400 obtains at least one candidate expression image corresponding to the expression classification identifier from the expression image library corresponding to the game identifier.

[0130] In step 209, the expression prediction server 400 determines the target expression image corresponding to the speech recognition result based on at least one candidate expression image.

[0131] In step 210, the expression prediction server 400 sends the target expression image to the first terminal 200-1.

[0132] It should be noted that the implementation processes of steps 207 to 210 are the same as those of steps 101 to 105 in other embodiments, and the implementation processes of steps 101 to 105 can be referred to.

[0133] In step 211, the first terminal 200-1 displays the speech recognition result and the target expression image on the information input interface.

[0134] In some embodiments, when the terminal displays the speech recognition result and the target expression image on the information input interface, it can display the speech recognition result and the target expression image on the information input interface simultaneously, or display the speech recognition result and the target expression image on the information input interface in sequence.

[0135] In some embodiments, when the target expression image sent by the expression prediction server to the first terminal has a voice playback control and is bound with voice data, when the first terminal receives a click or touch operation on the target expression image, the voice data can be played.

[0136] It should be noted that the simultaneous display of the voice recognition result and the target expression image on the information input interface means that the voice recognition result (text content) and the target expression image are synchronously displayed on the information input interface. The sequential display of the voice recognition result and the target expression image on the information input interface means that the voice recognition result can be first displayed on the information input interface, and after the voice recognition result is sent to the second terminal, the target expression image is then displayed on the information input interface. At this time, the voice recognition result is no longer displayed on the information input interface.

[0137] In step 212, the first terminal 200-1 sends at least one of the voice recognition result and the target expression image to the second terminal 200-2 in response to the received message sending instruction.

[0138] In some embodiments, when the voice recognition result and the target expression image are simultaneously displayed on the information input interface, if the terminal does not receive a message editing instruction (which can be an editing instruction for the voice recognition result or a deletion instruction for deleting the target expression image), and receives a message sending instruction, then the voice recognition result and the target expression image are simultaneously sent to the second terminal, that is, the voice recognition result and the target expression image are carried in an instant messaging message and sent to the second terminal simultaneously.

[0139] In some embodiments, the second terminal 200-2 receives at least one of the voice recognition result and the target expression image sent by the first terminal 200-1 and presents the received information. Since there may be multiple interactive messages received during the game session, the previously received messages may be overwritten by the newly received messages, resulting in message omission. Therefore, when the second terminal 200-2 receives the target expression image, it can determine the output area of the target expression image based on the current position of the virtual object controlled by the second terminal and present the target expression image in this output area, so as to avoid message omission. Exemplarily, the output area of the target expression image can be the area at the upper left of the virtual object and at a first preset distance from the virtual object. Among them, if the target expression image received by the second terminal 200-2 has a voice playback control and is bound with voice data, then when presenting the target expression image, the voice playback control can be presented, and when the second terminal 200-2 receives a click operation or a touch operation on the target expression image, the voice data is played.

[0140] In some embodiments, when the speech recognition result and the target expression image are displayed in sequence in the information input interface, it may be that the speech recognition result is received first. Then, the speech recognition result is first displayed in the information input interface. The information input interface also displays a send control and a cancel control. When a touch operation on the send control is received, it is determined that a message sending instruction is received. In response to the received message sending instruction, the speech recognition result is sent to the second terminal. At this time, the speech recognition result will no longer be displayed in the information input interface. When the target expression image sent by the expression prediction server is received, the target expression image is displayed in the information input interface. Similarly, at this time, the send control and the cancel control are also displayed in the information input interface. When a touch operation on the send control is received, it is determined that a message sending instruction is received. Then, in response to the received message sending instruction again, the target expression image is sent to the second terminal. That is to say, the speech recognition result and the target expression image are sent through two instant messaging messages. At this time, it can not only avoid waiting for the target expression image sent by the expression prediction server, ensure that the speech recognition result can be sent to the second terminal in time, but also realize the automatic selection of the game expression image according to the speech recognition result, improve the usage frequency of the game expression image, and increase the diversity and interest of the game interaction.

[0141] In some embodiments, such as Figure 6 shown, when the speech recognition result and the target expression image are displayed simultaneously in the information input interface, after step 211, the following steps 213 to 216 may also be executed. The following is combined with Figure 6 for specific description.

[0142] In step 213, the first terminal 200-1 deletes the target expression image in response to the received expression image deletion instruction.

[0143] In some embodiments, when the first terminal 200-1 receives a deletion operation on the expression image, it is determined that an expression image deletion instruction is received, and the target expression image is deleted. At this time, only the speech recognition result is displayed in the information input interface.

[0144] In step 214, the first terminal 200-1 sends an expression deletion message to the expression prediction server 400.

[0145] In some embodiments, the first terminal 200-1 sends an expression deletion message to the expression prediction server 400 to enable the expression prediction server to update the heat value of the target expression image based on the deletion message.

[0146] In step 215, the first terminal 200-1 sends the speech recognition result to the second terminal 200-2 in response to the received message sending instruction.

[0147] In step 216, when the expression prediction server 400 receives the expression deletion message sent by the first terminal, it subtracts 1 from the popularity value of the target expression image to obtain the updated popularity value.

[0148] In some embodiments, when the first terminal 200-1 cancels sending the target expression image, it sends an expression deletion message to the expression prediction server, thereby notifying the expression prediction server that the user has not selected to use the target expression image. At this time, the expression prediction server subtracts 1 from the popularity value of the target expression image to obtain the updated popularity value.

[0149] In some embodiments, when the expression prediction server 400 does not receive the expression deletion message sent by the first terminal within the preset duration, it indicates that the first terminal has sent the target expression image to the second terminal, that is, the user has selected to use the target expression image. At this time, the popularity value of the target expression is incremented by 1 to obtain the updated popularity value.

[0150] In some embodiments, when the speech recognition result and the target expression image are displayed in sequence in the information input interface, if the first terminal 200-1 receives a delete expression instruction, then the first terminal 200-1 also sends an expression deletion message to the expression prediction server, so that the expression prediction server updates the popularity value of the target expression image according to the received expression deletion message.

[0151] Through the above steps 213 to 216, after the target expression image is displayed on the first terminal, if a delete expression instruction is received, the first terminal sends an expression deletion message to the expression prediction server, so that the expression prediction server reduces the popularity value of the target expression image based on the expression deletion message, and when the expression deletion message is not received within the preset duration, increases the popularity value of the target expression image, thereby improving the accuracy of the target expression image determined by the expression prediction model.

[0152] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0153] The voice interaction processing method provided by the embodiments of the present application can be applied to a game application. When a game voice-to-text service is provided in the game application, the LSTM inference server infers the emotional type of the chat context based on the text recognized by the speech recognition server, so that the expression configuration server matches the unique expression images in the game according to the emotional type and sends the expression images to the chat box. The player only needs to click OK to send them to the teammates for interaction.

[0154] Figure 7 is a schematic diagram of the interface where the player clicks the voice-to-text function control to record voice during the game session provided by the embodiments of the present application, as Figure 7As shown, when in the process of voice input, "Voice input in progress" will be displayed in the message input box 701, and the input duration will also be displayed. An input completion control and a cancellation control are also displayed in the message input box 701.

[0155] The player speaks to the voice collection device, and the voice collection device will collect the player's voice. When receiving the voice input completion instruction, it will upload the collected voice to the ASR server, and the ASR server will perform voice recognition to obtain the recognized text result. The ASR server will send the recognized text result to the player's terminal, and the player's terminal will display the text result in the message input box, as Figure 8 shown. The text result recognized by the ASR server displayed in the message input box is "Ready to start a team fight". At this time, a send button control 801 and a cancel button control 802 are provided in the message input box. When the player clicks the "Send" button control 801, the terminal will send the text recognition result to the teammate's terminal via the server. The ASR server will send the recognized text result to the LSTM inference server, and the LSTM inference server will determine the expression type based on the text result and send the expression type to the expression configuration server, so that the expression configuration server can filter out the expression images in the game based on the expression type. Continuing with the above example, since the text result is "Ready to start a team fight", after the inference of the LSTM inference server, it is found that it is passionate and full of fighting spirit, and it matches the "Don't give up" expression in the game very well. Therefore, the expression image of "Don't give up" is recommended to the player's terminal, and the expression configuration server will send the expression image to the player's terminal. After the player's terminal receives the expression image, as Figure 9 shown, the expression image is displayed in the message input box.

[0156] Figure 10 is another schematic diagram of the implementation process of the voice interaction processing method provided by the embodiment of the present application. The following will be described in conjunction with Figure 10 for illustration.

[0157] In step 1001, the collected voice data is obtained.

[0158] In step 1002, it is judged whether the voice data is empty.

[0159] In some embodiments, when the voice data is not empty, step 1003 is entered; when the voice data is empty, the process ends.

[0160] In step 1003, it is judged whether the voice-to-text service is successfully executed.

[0161] Among them, if the voice-to-text service is successfully executed, step 1004 is entered; if the voice-to-text service fails, the process ends.

[0162] In step 1005, request the LSTM expression inference service.

[0163] In step 1006, the inference result is mapped to game-specific expressions.

[0164] Figure 11A is the network structure diagram of the voice interaction processing method provided by the embodiments of the present application. As Figure 11A shown, the network structure diagram includes: game client 1101, ASR server 1102, LSTM inference server 1103, and expression configuration server 1104, where:

[0165] The game client 1101 is used to provide games and voice services. The ASR server 1102 is used to provide the function of converting speech to text and send the converted text content to the expression configuration server 1103. The LSTM inference server 1103 is used to infer according to the text content sent by the ASR server 1102 through the trained expression classification prediction model, deduce the expression classification ID that best matches the text, and send the expression classification ID to the expression configuration server 1104. The expression configuration server 1104 is used to determine the game expression image to be sent according to the deduced expression classification ID.

[0166] Figure 11B is the information interaction schematic diagram between the game client, the ASR server, and the expression configuration server provided by the embodiments of the present application. The following combines Figure 11A and Figure 11B to illustrate the voice interaction processing method provided by the embodiments of the present application. The game client takes the PCM voice raw data collected by the voice acquisition device as the request parameter of the speech recognition request and sends it to the ASR server. After receiving the speech recognition request, the ASR server starts to process it. Finally, the game client obtains the expression content finally matched by the expression configuration server and displays the corresponding expression in the chat box for the player to choose whether to send it for interaction.

[0167] The ASR server translates the voice uploaded by the game client into the corresponding text information through the speech-to-text distributed service. After the translation is completed, the text information is sent to the downstream LSTM inference server.

[0168] The LSTM inference server infers according to the text information sent by the ASR, through the trained expression classification prediction model, deduces the most likely emotional expression type of the text (i.e., the expression classification identifier), and then obtains the actual game expression image shown from the game expression configuration server according to the emotional expression type ID as the final result.

[0169] Figure 12A is the structural schematic diagram of the expression classification prediction model provided by the embodiments of the present application. AsFigure 12A As shown, the expression classification prediction model includes an input layer 1201, an embedding layer 1202, a dense layer 1203, an activation layer 1204, and an output layer 1205, where:

[0170] Input Layer 1201: Receives the prediction text as an input sequence.

[0171] Embedding Layer 1202: Used to convert the input discrete data (usually vocabulary) into a continuous vector representation.

[0172] Dense Layer 1203 is used for feature learning and model decision-making. In the dense layer, each node is connected to all nodes in the previous layer. By learning the weights of these connections, useful features can be extracted from the input data and decisions can be made.

[0173] Activation Layer (Softmax activation function) 1204 is the last layer for multi-classification problems. Suppose there are 3 expression categories for decision-making. After passing through the Softmax function, each output value is between 0 and 1, and the sum of the 3 values is 1, which can be regarded as the probability of belonging to each expression category. This characteristic exactly meets the requirements of multi-classification problems, making this function often used in the output layer of classifiers.

[0174] Output Layer 1205 is used to obtain the probabilities of each expression category given by Softmax and take the expression category with the highest probability as the prediction result. The category number here represents the expression ID.

[0175] Figure 12B is a schematic diagram of the training data of the expression classification prediction model provided by the embodiments of the present application and the prediction results of the training samples, where Figure 12B the first column represents the number of the training sample, the second column represents the training sample, the third column represents the label of the training sample, and the fourth column represents the prediction result of the training sample. Figure 12B The training data composed of the training samples and the labels of the training samples in is used to train the expression classification prediction model to obtain a trained expression classification prediction model.

[0176] Figure 13 is a schematic diagram of the information interaction between the expression configuration server, the LSTM inference server, and the game client provided by the embodiments of the present application, as Figure 13As shown in the figure, the LSTM inference server uses GAME ID and Emoj ID as parameters to request the Emoj Config Server to determine the target emoji image. The Emoj Config Server pulls the latest emoji map of the game from the database based on the GAME ID in the parameter (the ID that can locate a specific online game), and then obtains the corresponding game emoji map in the game based on the Emoj ID in the request parameter. Finally, the game emoji map is sent to the game client as the final result. The game client calls out the emoji prompt in the corresponding chat box for the player to choose whether to send it to teammates for interaction. This can improve the playability of the voice-to-text function and expand the diversity of the interactive gameplay of the player battle process.

[0177] It is understandable that in the embodiments of the present application, related data such as user information and voice to be recognized are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0178] The following is a description of an exemplary structure of the voice interaction processing device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules stored in the voice interaction processing device 455 of the memory 450 may include:

[0179] The first receiving module 4551 is used to receive a speech recognition result and a game identifier sent by a speech recognition server, wherein the speech recognition result is obtained by the speech recognition server recognizing speech data collected by the first terminal during the game;

[0180] A first determination module 4552, used to determine the expression classification identifier corresponding to the speech recognition result;

[0181] A first acquisition module 4553 is used to acquire at least one candidate expression image corresponding to the expression classification identifier from the expression image library corresponding to the game identifier;

[0182] A second determination module 4554, configured to determine a target expression image corresponding to the speech recognition result based on the at least one candidate expression image;

[0183] The first sending module 4555 is used to send the target expression image to the first terminal.

[0184] In some embodiments, the second determination module 4554 is further configured to:

[0185] When the number of the candidate expression images is one, determine the candidate expression image as the target expression image; or,

[0186] When the number of the candidate expression images is one, obtain the image to be synthesized corresponding to the first terminal, where the image to be synthesized is the facial image in the game avatar of the first terminal or the facial image of the virtual character adopted by the first terminal;

[0187] Perform a synthesis process on the candidate expression image and the image to be synthesized to obtain the target expression image.

[0188] In some embodiments, the second determination module 4554 is further configured to:

[0189] When the number of the candidate expression images is two or more, obtain the heat value of each candidate expression image, where the heat value represents the number of times the candidate expression image is selected and used;

[0190] Determine the candidate expression image corresponding to the highest heat value among the heat values as the target expression image.

[0191] In some embodiments, the second determination module 4554 is further configured to:

[0192] When the number of the candidate expression images is two or more, obtain the heat value of each candidate expression image;

[0193] Obtain the image to be synthesized corresponding to the first terminal, where the image to be synthesized is the facial image in the game avatar or the facial image of the virtual character adopted by the first terminal;

[0194] Perform a synthesis process on the candidate expression image corresponding to the highest heat value among the heat values and the image to be synthesized to obtain the target expression image.

[0195] In some embodiments, the device further includes:

[0196] A first update module, configured to subtract 1 from the heat value of the target expression image to obtain an updated heat value when receiving an expression deletion message sent by the first terminal;

[0197] A second update module, configured to add 1 to the heat value of the target expression to obtain an updated heat value when not receiving an expression deletion message sent by the first terminal within a preset duration.

[0198] In some embodiments, the first determination module 4552 is further configured to:

[0199] Obtain a trained expression classification prediction model;

[0200] Process the speech recognition result by using the embedding layer in the trained expression classification prediction model to obtain an embedding vector corresponding to the speech recognition result;

[0201] Perform prediction processing on the embedding vector by using the fully connected layer, activation layer, and output layer in the trained expression classification prediction model to obtain an expression classification identifier corresponding to the speech recognition result.

[0202] An embodiment of the present application provides a terminal, which includes: at least one processor, a memory, at least one network interface, and a user interface. Each component in the terminal is coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus. In some embodiments, the speech interaction processing device stored in the memory can be software in the form of programs and plugins, including the following software modules:

[0203] A speech acquisition module, configured to acquire speech data in response to a received speech acquisition instruction during a game session;

[0204] A second sending module, configured to send the acquired speech data to a speech recognition server in response to a received speech acquisition completion instruction;

[0205] A second receiving module, configured to receive the speech recognition result sent by the speech recognition server, and receive the expression image corresponding to the speech recognition result sent by an expression prediction server, where the target expression image is determined by the expression prediction server based on the speech recognition result and a game identifier sent by the speech recognition server;

[0206] A first display module, configured to display the speech recognition result and the target expression image on the information input interface;

[0207] A second sending module, configured to send at least one of the speech recognition result and the target expression image to a second terminal in response to a received message sending instruction.

[0208] In some embodiments, the first display module is further configured to:

[0209] Simultaneously display the speech recognition result and the target expression image on the information input interface;

[0210] Correspondingly, the second sending module is further configured to:

[0211] Send the speech recognition result and the target expression image to the second terminal simultaneously in response to a received message sending instruction.

[0212] In some embodiments, the first display module is further configured to:

[0213] display the speech recognition result and the target expression image in sequence on the information input interface;

[0214] Correspondingly, the second sending module is further configured to:

[0215] after displaying the speech recognition result on the information input interface, in response to a received message sending instruction, send the speech recognition result to the second terminal;

[0216] after displaying the target expression image on the information input interface, in response to a received message sending instruction again, send the target expression image to the second terminal.

[0217] In some embodiments, the apparatus further includes:

[0218] an image deletion module, configured to delete the target expression image in response to a received expression image deletion instruction;

[0219] a third sending module, configured to send an expression deletion message to the expression prediction server, so that the expression prediction server updates the popularity value of the target expression image based on the deletion message;

[0220] a fourth sending module, configured to send the speech recognition result to the second terminal in response to a received message sending instruction.

[0221] An embodiment of the present application provides a computer program product, which includes a computer program or computer executable instructions, and the computer program or computer executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instructions from the computer-readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the speech interaction processing method described above in the embodiments of the present application.

[0222] An embodiment of the present application provides a computer-readable storage medium storing computer executable instructions, where computer executable instructions or a computer program are stored, and when the computer executable instructions or the computer program are executed by a processor, the processor will be caused to execute the speech interaction processing method provided by the embodiments of the present application, for example, as Figure 3A 、 Figure 5 、 Figure 6 shown in the speech interaction processing method.

[0223] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0224] In some embodiments, the computer-executable instructions may be in the form of a program, software, a software module, a script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, a component, a subroutine, or other unit suitable for use in a computing environment.

[0225] As an example, the computer-executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or portions of code).

[0226] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0227] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are all included in the protection scope of the present application.

Claims

1. A voice interaction processing method, characterized in that, The method includes: Receiving a speech recognition result and a game identifier sent by a speech recognition server, where the speech recognition result is obtained by the speech recognition server recognizing speech data collected by a first terminal during a game session; Determining an expression classification identifier corresponding to the speech recognition result; Obtaining at least one candidate expression image corresponding to the expression classification identifier from an expression image library corresponding to the game identifier; Determining a target expression image corresponding to the speech recognition result based on the at least one candidate expression image; Sending the target expression image to the first terminal.

2. The method according to claim 1, wherein The determining a target expression image corresponding to the speech recognition result based on the at least one candidate expression image includes: When the number of candidate expression images is one, determining the candidate expression image as the target expression image; or, When the number of candidate expression images is one, obtaining a to-be-synthesized image corresponding to the first terminal, where the to-be-synthesized image is a facial image in the game avatar of the first terminal or a facial image of a virtual character adopted by the first terminal; Performing a synthesis process on the candidate expression image and the to-be-synthesized image to obtain the target expression image.

3. The method according to claim 2, wherein The determining a target expression image corresponding to the speech recognition result based on the at least one candidate expression image includes: When the number of candidate expression images is two or more, obtaining a popularity value of each candidate expression image, where the popularity value represents the number of times the candidate expression image is selected and used; Determining the candidate expression image corresponding to the highest popularity value among the popularity values as the target expression image.

4. The method according to claim 2, characterized in that, The determining a target expression image corresponding to the speech recognition result based on the at least one candidate expression image includes: When the number of candidate expression images is two or more, obtaining a popularity value of each candidate expression image; Obtaining a to-be-synthesized image corresponding to the first terminal, where the to-be-synthesized image is a facial image in the game avatar or a facial image of a virtual character adopted by the first terminal; Performing a synthesis process on the candidate expression image corresponding to the highest popularity value among the popularity values and the to-be-synthesized image to obtain the target expression image.

5. The method according to any one of claims 1 to 4, characterized in that The method further includes: When receiving an expression deletion message sent by a first terminal, subtracting 1 from the popularity value of the target expression image to obtain an updated popularity value; When not receiving an expression deletion message sent by the first terminal within a preset duration, adding 1 to the popularity value of the target expression to obtain an updated popularity value.

6. The method according to any one of claims 1 to 4, characterized in that, The determining an expression classification identifier corresponding to the speech recognition result includes: Obtaining a trained expression classification prediction model; Processing the speech recognition result by using an embedding layer in the trained expression classification prediction model to obtain an embedding vector corresponding to the speech recognition result; Performing a prediction process on the embedding vector by using a fully connected layer, an activation layer, and an output layer in the trained expression classification prediction model to obtain an expression classification identifier corresponding to the speech recognition result.

7. A method for voice interaction processing, characterized in that, The method includes: During the game session, in response to the received voice collection instruction, collect voice data; In response to the received voice collection completion instruction, send the collected voice data to the voice recognition server; Receive the voice recognition result sent by the voice recognition server, and receive the expression image corresponding to the voice recognition result sent by the expression prediction server, where the target expression image is determined by the expression prediction server based on the voice recognition result sent by the voice recognition server and the game identifier; Display the voice recognition result and the target expression image on the information input interface; In response to the received message sending instruction, send at least one of the voice recognition result and the target expression image to the second terminal.

8. The method according to claim 7, characterized in that, The displaying the voice recognition result and the target expression image on the information input interface includes: Simultaneously display the voice recognition result and the target expression image on the information input interface; Correspondingly, in response to the received message sending instruction, sending at least one of the voice recognition result and the target expression image to the second terminal includes: In response to the received message sending instruction, send the voice recognition result and the target expression image to the second terminal simultaneously.

9. The method according to claim 7, characterized in that, The displaying the voice recognition result and the target expression image on the information input interface includes: Display the voice recognition result and the target expression image on the information input interface in sequence; Correspondingly, in response to the received message sending instruction, sending at least one of the voice recognition result and the target expression image to the second terminal includes: After displaying the voice recognition result on the information input interface, in response to the received message sending instruction, send the voice recognition result to the second terminal; After displaying the target expression image on the information input interface, in response to the received message sending instruction again, send the target expression image to the second terminal.

10. The method according to claim 7, wherein After displaying the voice recognition result and the target expression image on the information input interface, the method further includes: In response to the received expression image deletion instruction, delete the target expression image; Send an expression deletion message to the expression prediction server, so that the expression prediction server updates the heat value of the target expression image based on the deletion message; In response to the received message sending instruction, send the voice recognition result to the second terminal.

11. A voice interaction processing device, characterized in that, The device includes: A first receiving module, configured to receive the voice recognition result and the game identifier sent by the voice recognition server, where the voice recognition result is obtained by the voice recognition server recognizing the voice data collected by the first terminal during the game session; A first determining module, configured to determine the expression classification identifier corresponding to the voice recognition result; A first obtaining module, configured to obtain at least one candidate expression image corresponding to the expression classification identifier from the expression image library corresponding to the game identifier; A second determining module, configured to determine the target expression image corresponding to the voice recognition result based on the at least one candidate expression image; A first sending module, configured to send the target expression image to the first terminal.

12. A voice interaction processing device, characterized in that, The device includes: A voice acquisition module, configured to acquire voice data in response to a received voice acquisition instruction during a game session; A second sending module, configured to send the acquired voice data to a voice recognition server in response to a received voice acquisition completion instruction; A second receiving module, configured to receive a voice recognition result sent by the voice recognition server and receive an expression image corresponding to the voice recognition result sent by an expression prediction server, where the target expression image is determined by the expression prediction server based on the voice recognition result sent by the voice recognition server and a game identifier; A first display module, configured to display the voice recognition result and the target expression image on the information input interface; A second sending module, configured to send at least one of the voice recognition result and the target expression image to a second terminal in response to a received message sending instruction.

13. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the method according to any one of claims 1 to 6 or any one of claims 7 to 10 when executing the computer-executable instructions stored in the memory.

14. A computer-readable storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 6 or any one of claims 7 to 10.

15. A computer program product, comprising computer-executable instructions, characterized in that, The computer-executable instructions, when executed by the processor, implement the method according to any one of claims 1 to 6 or any one of claims 7 to 10.