Data processing method, device, storage medium, and electronic device

By converting the names of virtual game characters into voice data and splicing them with interactive content, the problem of poor interactiveness of virtual game characters is solved, and the immersion and interactivity of the game is improved.

CN113920983BActive Publication Date: 2025-07-25NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111241373.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-25
Publication Date
2025-07-25
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

In the prior art, the names of virtual game characters are usually fed back in text form, resulting in poor interaction with players and lack of voice interaction, which separates players from the identity of virtual game characters.

Method used

By receiving the client's target request, the target text is converted into voice data, and spliced with the voice data other than the name in the interactive content to form complete voice feedback, and use the AI speech generation model and reference audio encoder to ensure style consistency.

Benefits of technology

It improves the name interactivity of virtual game characters, enhances the immersion and visual auditory experience of the game, and enhances the player's sense of identity immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920983B_ABST
    Figure CN113920983B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method, apparatus, storage medium, and electronic device. The method includes: receiving a target request from a client, where the information carried in the target request at least includes a target text, and the target text is used to represent the name of a virtual game character; responding to the target request, converting the target text into first voice data; sending the first voice data to the client so that the client splices the first voice data and second voice data into third voice data, and the second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character. Through the present invention, the technical effect of improving the interactivity of the name of the virtual game character is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular, to a data processing method, apparatus, storage medium, and electronic device. Background Art

[0002] Currently, it is possible to name virtual game characters in a game, and the names of virtual game characters are usually presented in text form and generally displayed in function display interfaces such as a text chat interface, a team formation interface, and personal information.

[0003] In addition, in the conversations of virtual game characters, the names of virtual game characters are usually skipped in the conversation, and only other fixed text content is read, or the virtual game characters are addressed in other fixed ways set by the game.

[0004] Therefore, the name of the virtual game character does not interact with the virtual game character, which makes the player's understanding of the virtual game character and their own identity relatively fragmented, resulting in the technical problem of poor interactivity of the name of the virtual game character.

[0005] Aiming at the technical problem of poor interactivity of the name of the virtual game character in the prior art, no effective solution has been proposed yet. Summary of the Invention

[0006] The main purpose of the present invention is to provide a data processing method, apparatus, storage medium, and electronic device to at least solve the technical problem of poor interactivity of the name of the virtual game character.

[0007] To achieve the above object, according to one aspect of the present invention, a data processing method is provided. The method may include: receiving a target request from a client, where the information carried in the target request includes at least a target text, and the target text is used to represent the name of a virtual game character; in response to the target request, converting the target text into first voice data; and sending the first voice data to the client so that the client concatenates the first voice data and second voice data into third voice data, where the second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character.

[0008] Optionally, the method further includes: obtaining style information of the second voice data, where the style information is used to represent the voice style to which the second voice data belongs; and converting the target text into first voice data includes: converting the style information and the target text into first voice data, where the voice style to which the first voice data belongs is the same as the voice style to which the second voice data belongs.

[0009] Optionally, obtain the style information of the second voice data, including: extracting the first acoustic feature of the second voice data; determining the style information based on the first acoustic feature.

[0010] Optionally, convert the style information and the target text into the first voice data, including: extracting the text feature of the target text; aligning the text feature and the first acoustic feature to obtain an alignment result; converting the style information and the alignment result into the first voice data.

[0011] Optionally, the information carried in the target request further includes the first identification information of the virtual game character, and the method further includes: obtaining the target vector of the virtual game character based on the first identification information, where the target vector is used to represent the timbre of the virtual game character; converting the style information and the alignment result into the first voice data, including: converting the target vector, the style information and the alignment result into the first voice data.

[0012] Optionally, converting the target vector, the style information and the alignment result into the first voice data includes: synthesizing the target vector, the style information and the alignment result into a second acoustic feature; converting the second acoustic feature into the first voice data.

[0013] Optionally, the information carried in the target request further includes the second identification information of the second voice data, and the method further includes: obtaining the second voice data based on the second identification information.

[0014] Optionally, the method further includes: converting the target text into a phoneme data sequence and / or a prosody data sequence; converting the target text into the first voice data, including: converting the phoneme data sequence and / or the prosody data sequence into the first voice data.

[0015] To achieve the above object, according to another aspect of the present invention, another data processing method is further provided. The method may include: when it is detected that the interaction content to be performed by the virtual game character includes the name of the virtual game character, sending a target request to the server, where the information carried in the target request at least includes the target text, and the target text is used to represent the name; obtaining the first voice data, where the first voice data is obtained by the server in response to the target request and converting the target text; splicing the first voice data and the second voice data to obtain the third voice data, where the second voice data is the voice data corresponding to the content other than the name in the interaction content; playing the third voice data.

[0016] To achieve the above object, according to another aspect of the present invention, there is also provided a data processing device. The device may include: a receiving unit, configured to receive a target request from a client, wherein the information carried in the target request at least includes a target text, and the target text is used to represent the name of a virtual game character; a conversion unit, configured to respond to the target request and convert the target text into first voice data; a first sending unit, configured to send the first voice data to the client, so that the client splices the first voice data and second voice data into third voice data, and the second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character.

[0017] To achieve the above object, according to another aspect of the present invention, there is also provided another data processing device. The device may include: a second sending unit, configured to send a target request to a server when it is detected that the interaction content to be performed by the virtual game character includes the name of the virtual game character, wherein the information carried in the target request at least includes a target text, and the target text is used to represent the name; an obtaining unit, configured to obtain first voice data, wherein the first voice data is obtained by the server in response to the target request and converting the target text; a splicing unit, configured to splice the first voice data and second voice data to obtain third voice data, wherein the second voice data is the voice data corresponding to the content other than the name in the interaction content; a playing unit, configured to play the third voice data.

[0018] To achieve the above object, according to another aspect of the present invention, there is provided a computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein when the computer program is run by a processor, it controls the device where the computer-readable storage medium is located to execute the data processing method of one of the embodiments of the present invention.

[0019] To achieve the above object, according to another aspect of the present invention, there is provided an electronic device. The electronic device includes a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the data processing method of one of the embodiments of the present invention.

[0020] In the data processing method of this embodiment, a target request from a client is received. Among them, the information carried in the target request includes at least a target text, and the target text is used to represent the name of a virtual game character; in response to the target request, the target text is converted into first voice data; the first voice data is sent to the client so that the client splices the first voice data and second voice data into third voice data, and the second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character. That is to say, in this application, the target text used to represent the name of the virtual game character is converted into first voice data to be spliced with the second voice data corresponding to the interaction content, achieving the purpose of adding voice feedback to the interaction content containing the name of the virtual game character, avoiding that the name of the virtual game character is usually fed back in text form, and also avoiding skipping the name of the virtual game character in the conversation, thereby solving the technical problem of poor interactivity of the name of the virtual game character and achieving the technical effect of improving the interactivity of the name of the virtual game character. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings that form a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0022] Figure 1 is a hardware structure block diagram of a mobile terminal for a data processing method according to an embodiment of the present invention;

[0023] Figure 2 is a flowchart of a data processing method according to an embodiment of the present invention;

[0024] Figure 3 is a flowchart of another data processing method according to an embodiment of the present invention;

[0025] Figure 4 is a schematic diagram of a text-to-speech server according to an embodiment of the present invention;

[0026] Figure 5 is a schematic diagram of a text-to-speech system with reference audio according to an embodiment of the present invention;

[0027] Figure 6 is a schematic diagram of a data processing device according to an embodiment of the present invention;

[0028] Figure 7 is a schematic diagram of another data processing device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0030] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present application described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] The method embodiment provided by one of the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a data processing method according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only for illustration and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than those shown in Figure 1 the figure, or have a different configuration from that shown in Figure 1 the figure.

[0033] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to a data processing method in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0034] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC for Network Interface Controller), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0035] Next, a data processing method according to one embodiment of the present invention will be introduced from the server (server side).

[0036] Figure 2 is a flowchart of a data processing method according to an embodiment of the present invention. As Figure 2 shown, the method may include the following steps:

[0037] Step S202, receiving a target request from a client, where the information carried in the target request includes at least a target text, and the target text is used to represent the name of a virtual game character.

[0038] In the technical solution provided in step S202 of the present invention, the client can be a game client, and the running game scene includes virtual game characters. For example, the virtual game character is a non-player character (NPC for short). When the interaction content that the virtual game character is about to perform includes the name of the virtual game character, for example, when the interaction content includes the name given by the player to the virtual game character, the server can receive a target request from the above-mentioned client. The information carried in the target request includes at least the target text, which is also the name text to be synthesized and is used to represent the above-mentioned name of the virtual game character. Among them, the interaction content can be the dialogue content carried out by the virtual game character.

[0039] Step S204: In response to the target request, convert the target text into first voice data.

[0040] In the technical solution provided in step S204 of the present invention, after receiving the target request from the client, the server responds to the target request and converts the target text into first voice data.

[0041] In this embodiment, after receiving the target request from the client, the server can identify the target text from the target request, and then input the target text into the voice generation model preset in the server. The voice generation model processes the target text, so as to obtain the first voice data that matches the name of the virtual game character. This stage is also the name synthesis voice stage. Among them, the voice generation model can be an artificial intelligence model (Artificial Intelligence, abbreviated as AI) training model, that is, an AI voice generation model or an AI training model. The first voice data is also the synthesized voice of the name synthesis voice and the name text, which can be reflected by voice waveforms / signals, thus achieving the purpose of feeding back the name of the virtual game character through the first voice data.

[0042] Optionally, the fact that the first voice data in this embodiment matches the name of the virtual game character may mean that the voice content of the first voice data includes the name of the virtual game character, the timbre of the first voice data is the protagonist timbre of the virtual game character, the pitch of the first voice data conforms to the context voice as much as possible, and the length of the first voice data conforms to the overall speaking speed of the protagonist, etc.

[0043] Optionally, in this embodiment, the above target text can be converted into first voice data in the text-to-speech system with reference audio in the server.

[0044] Optionally, the server in this embodiment can set up the voice generation models for each virtual game character corresponding to the game.

[0045] Step S206: Send the first voice data to the client so that the client splices the first voice data and the second voice data into the third voice data, where the second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character.

[0046] In the technical solution provided in step S206 of the present invention above, after responding to the target request and converting the target text into the first voice data, the first voice data can be sent to the client so that the client splices the first voice data and the second voice data into the third voice data, where the second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character.

[0047] In this embodiment, the interaction content to be performed currently includes content other than the name of the virtual game character, and the content other than the name of the virtual game character corresponds to the second voice data, which can be an audio, that is, the audio to be spliced with the first voice data. After the first voice data is received by the client, the above-mentioned first voice data and the second voice data can be spliced by the client. For example, the client splices the first voice data and the second voice data according to the sampling points of the first voice data and the second voice data, where the sampling points can include time information, and the first voice data and the second voice data are spliced through it to obtain the third voice data, which is played by the client.

[0048] The third voice data in this embodiment includes the voice data of the name of the virtual game character, achieving the purpose of adding voice feedback to the interaction content containing the name of the virtual game character. This allows players to more immersively experience the plot and gain a better sense of identity substitution during the game plot and important prompts, thereby enhancing the visual and auditory experience of the game, bringing a stronger sense of immersion to players, and improving the authenticity and interactivity of the virtual world of the game scene.

[0049] Through the above steps S202 to S206 of this application, a target request is received from the client. The information carried in the target request includes at least a target text, which is used to represent the name of a virtual game character. In response to the target request, the target text is converted into first voice data. The first voice data is sent to the client so that the client can splice the first voice data and second voice data into third voice data. The second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character. That is to say, in this embodiment, the target text used to represent the name of the virtual game character is converted into first voice data to be spliced with the second voice data corresponding to the interaction content, achieving the purpose of adding voice feedback to the interaction content containing the name of the virtual game character, avoiding that the name of the virtual game character is usually feedback in text form, and also avoiding skipping the name of the virtual game character in the conversation, thus solving the technical problem of poor interactivity of the name of the virtual game character and achieving the technical effect of improving the interactivity of the name of the virtual game character.

[0050] The above method of this embodiment will be further introduced in detail below.

[0051] As an optional implementation manner, the method further includes: obtaining style information of the second voice data, where the style information is used to represent the voice style to which the second voice data belongs; step S204, converting the target text into first voice data, including: converting the style information and the target text into first voice data, where the voice style to which the first voice data belongs is the same as the voice style to which the second voice data belongs.

[0052] In this embodiment, the style information can be extracted from the second voice data. For example, the style information can be the overall style information of the second voice data, which is used to represent the voice style to which the second voice data belongs. Optionally, the style information of this embodiment can be a style vector, that is, a mathematical vector used to represent the style. This embodiment can encode the above style information to obtain style encoding information. Optionally, the second voice data of this embodiment can be used as a reference audio for the style of the first voice data. This embodiment can convert the style encoding information of the second voice data and the above target text into first voice data, so that the voice style to which the first voice data belongs is the same as the voice style to which the second voice data belongs. That is to say, the style information of the second voice data in this embodiment is used to control the style of the name of the virtual game character when it is played by voice, so that the splicing place of the first voice data synthesized from the target text and the second voice data is natural and coherent, which mainly includes coherent rhythm and consistent channel.

[0053] Optionally, the text-to-speech system with reference audio in this embodiment may include a reference audio encoder. The style information can be extracted from the second speech data through the reference audio encoder, and then the style information is encoded to obtain the above-mentioned style encoding information.

[0054] As another alternative implementation, obtaining the style information of the second speech data includes: extracting the first acoustic feature of the second speech data; determining the style information based on the first acoustic feature.

[0055] In this embodiment, when implementing the acquisition of the style information of the second speech data, the first acoustic feature may be first extracted from the second speech data. The first acoustic feature, that is, the speech acoustic feature, may be a speech feature sequence. Optionally, the above-mentioned text-to-speech system with reference audio in this embodiment may include an acoustic feature extraction module, which can convert the second speech data from a waveform into some information-rich features to obtain the first acoustic feature. Optionally, the first acoustic feature may be a Mel spectrogram.

[0056] After the first acoustic feature is extracted from the second speech data, the style information can be determined based on the first acoustic feature. For example, the reference audio encoder may include a neural network model. The above-mentioned first acoustic feature input is subjected to information extraction and information compression through the neural network model to obtain a style vector. Among them, the neural network model belongs to an unsupervised learning model. After the style vector is obtained, the style vector can be encoded to obtain the style encoding information.

[0057] Optionally, the above-mentioned neural network model may include a convolutional neural network model (Convolutional Neural Network, abbreviated as CNN) and a long short-term memory network model (Long Short-Term Memory, abbreviated as LSTM). That is, the above-mentioned reference audio encoder in this embodiment is implemented based on CNN and LSTM.

[0058] As another alternative implementation, converting the style information and the target text into the first speech data includes: extracting the text feature of the target text; aligning the text feature and the first acoustic feature to obtain an alignment result; converting the style information and the alignment result into the first speech data.

[0059] In this embodiment, when realizing the conversion of style information and target text into first speech data, text features may be first extracted from the target text, and the text features may be a text feature sequence. Optionally, the above text-to-speech system with reference audio in this embodiment may include a text encoder. By inputting the target text into the text encoder, the target text can be processed by the text encoder to obtain text features. Optionally, the text encoder in this embodiment may map the target text to a high-dimensional text feature space through a non-linear transformation for encoding, so as to obtain the above text features.

[0060] In this embodiment, since the data lengths of the text features and the first acoustic features are different. For example, the data length of the first acoustic features is longer than that of the text features. Therefore, after the text features of the target text are extracted, the text features and the first acoustic features can be aligned to obtain an alignment result. Optionally, the above text-to-speech system with reference audio in this embodiment may include an attention mechanism model, and the attention mechanism model can be used to align the text features and the first acoustic features to obtain an alignment result. That is to say, the inputs of the above reference audio encoder and the attention mechanism model in this embodiment are both for the same first acoustic features.

[0061] As another alternative implementation manner, the information carried in the target request further includes the first identification information of the virtual game character, and the method further includes: obtaining a target vector of the virtual game character based on the first identification information, where the target vector is used to represent the timbre of the virtual game character; converting the style information and the alignment result into first speech data, including: converting the target vector, the style information and the alignment result into first speech data.

[0062] In this embodiment, the information carried in the target request sent by the client received by the server may further include the first identification information of the virtual game character, and the first identification information can be used to uniquely identify the virtual game character. For example, it is an identity identification (ID), which can also be referred to as the target speaker ID. This embodiment can obtain the target vector of the virtual game character based on the first identification information. For example, the first identification information is converted into a target vector, and the target vector is also the speaker vector, which can be a vector table and is used to represent the timbre of the virtual game character. Optionally, the above text-to-speech system with reference audio in this embodiment may include a speaker vector table module, which can convert the first identification information into a target vector.

[0063] After obtaining the target vector of the virtual game character based on the first identification information, the server may convert the target vector, the style information, and the alignment result into first voice data, which includes the timbre of the above virtual game character. That is, the above target vector in this embodiment is used to control the timbre of the name of the virtual game character when it is played through voice.

[0064] Optionally, the above target vector and style information in this embodiment may also be input into the attention mechanism module to convert the target vector, the style information, and the alignment result into first voice data.

[0065] As another alternative implementation, converting the target vector, the style information, and the alignment result into first voice data includes: synthesizing the target vector, the style information, and the alignment result into a second acoustic feature; converting the second acoustic feature into first voice data.

[0066] In this embodiment, when implementing the conversion of the target vector, the style information, and the alignment result into first voice data, it may be to collectively call the target vector, the style information, and the alignment result as a second acoustic feature. Optionally, the above text-to-speech system with reference audio in this embodiment may include an acoustic decoder, which can return the text feature, the alignment result obtained by the attention mechanism module with the first acoustic feature, and the target vector and style information to the original voice acoustic feature space through a non-linear transformation to obtain a second acoustic feature, which is also the predicted voice acoustic feature, and then convert it into first voice data. Optionally, the above text-to-speech system with reference audio in this embodiment may include a vocoder, and the second acoustic feature may be converted into a voice waveform / signal through the vocoder to obtain first voice data.

[0067] As another alternative implementation, the information carried in the target request further includes second identification information of the second voice data, and the method further includes: obtaining the second voice data based on the second identification information.

[0068] In this embodiment, the information carried in the target request sent by the server received from the client may further include second identification information, which can be used to uniquely identify the second voice data and can be called the splicing audio ID.

[0069] As another alternative implementation, the method further includes: converting the target text into a phoneme data sequence and / or a prosody data sequence; step S204, converting the target text into first voice data, including: converting the phoneme data sequence and / or the prosody data sequence into first voice data.

[0070] In this embodiment, before converting the target text into the first voice data, the target text may be preprocessed. Optionally, the above-mentioned text-to-speech system with reference audio of this embodiment may include a text preprocessing module, and the target text may be preprocessed by the text preprocessing module. Optionally, this embodiment converts the target text into a corresponding phoneme data sequence and / or prosody data sequence through a text preprocessing module, which may be a series of rule-based or neural network models to convert the target text into a corresponding phoneme data sequence and / or prosody data sequence, and then text features may be extracted from the phoneme data sequence and / or prosody data sequence to align with the first acoustic feature, and then the obtained alignment result, target vector and style information are converted into the first voice data.

[0071] Optionally, this embodiment may combine the above-mentioned phoneme data sequence and prosodic data sequence to obtain a final phoneme prosodic data sequence.

[0072] Optionally, this embodiment can convert the target text into a corresponding phoneme data sequence through a text-to-phoneme model, wherein the text-to-phoneme model can be a neural network model using a CNN+LSTM structure and trained using a cross-entropy loss function.

[0073] Optionally, this embodiment can convert the target text into a corresponding prosodic data sequence through a text-to-prosodic model, wherein the text-to-phoneme model can be a neural network model using an LSTM structure and trained using a cross entropy loss function.

[0074] For example, if the target text input in this embodiment is "I love China", then the phoneme data sequence "w o3 a i3 zh ong1 g uo2" can be obtained by processing it through the text-to-phoneme model; the target text is processed through the text-to-rhythm model, and the prosodic data sequence obtained is "#1#1*#4"=", then the final phoneme prosodic data sequence can be "w o3#1 a i3#1 zh ong1 g uo2#4".

[0075] It should be noted that in this embodiment, the style information of the second speech data can be effectively obtained through the above reference encoder for use in the final acquisition stage of the first speech data. If the reference audio encoder is not used in this embodiment, the first speech data synthesized from the same target text will always have the same style. That is, in the absence of the reference audio encoder to determine the style information, the attention mechanism module only receives the text features output by the text encoder, the first acoustic features output by the acoustic feature extraction module, and the target vector. After adding the reference audio encoder in this embodiment, the attention mechanism module additionally receives the style information (style encoding information) of the second speech data. Since the text features output after the same target text is input to the text encoder are fixed. Therefore, when the first acoustic features and the target vector remain unchanged, the synthesized first speech data are basically of the same style. After adding the style encoding information of the second speech data in this embodiment, the overall style of the synthesized first speech data is affected by the style information of the second speech data. In this way, when different second speech data are used, the style of the first speech data synthesized from the same target text will change.

[0076] In addition, if the reference audio encoder is not used in this embodiment, the style of the first speech data synthesized from the target text may not be consistent in the case of the same second speech data. This is because in the absence of the influence of the style information of the second speech data output by the reference audio encoder, the first speech data synthesized from the target text is generally more affected by the input of the target text. Sometimes, the styles of the first speech data corresponding to different target texts may also be different. However, by introducing the style information of the second speech data output by the reference audio encoder in this embodiment, the style information of the same second speech data can be used when synthesizing different target texts, so that the styles of the first speech data synthesized from different target texts can be kept consistent.

[0077] In this embodiment, in the absence of the style information of the second speech data output by the reference audio encoder, it will cause the splicing of the first speech data and the second speech data synthesized from the target text to be unnatural, which mainly includes discontinuous prosody rhythm and inconsistent channels. And this embodiment uses the style information of the second speech data output by the reference audio encoder to effectively solve this problem, making the splicing of the first speech data and the second speech data synthesized from the target text relatively natural and coherent.

[0078] One of the embodiments of the present invention also provides another data processing method from the client side.

[0079] Figure 3 It is a flowchart of another data processing method according to the embodiment of the present invention. As Figure 3As shown, the method may include the following steps:

[0080] Step S302: When it is detected that the interaction content to be performed by the virtual game character includes the name of the virtual game character, send a target request to the server. The information carried in the target request includes at least target text, and the target text is used to represent the name.

[0081] In the technical solution provided in step S302 of the present invention, the client may be a game client. The game scene run by the game client includes virtual game characters. When the client detects that the interaction content to be performed by the virtual game character includes the name of the virtual game character, for example, when the client detects that the interaction content includes the name given by the player to the virtual game character, the client may send a target request to the server. The information carried in the target request includes at least target text, which is also the name text to be synthesized and is used to represent the above name of the virtual game character. Among them, the interaction content may be the dialogue content carried out by the virtual game character.

[0082] Optionally, the above client may be set on a mobile terminal or on a personal computer (PC) terminal, and no specific limitation is made here.

[0083] Step S304: Obtain first voice data, where the first voice data is obtained by the server in response to the target request and converting the target text.

[0084] In the technical solution provided in step S304 of the present invention, after the client sends a target request to the server, the client obtains first voice data. The target text can be recognized by the server from the target request after receiving the target request, and processed by a voice generation model into first voice data that matches the name of the virtual game character, so as to achieve the purpose of feeding back the name of the virtual game character through the first voice data.

[0085] Optionally, the client of this embodiment may obtain the first voice data returned by the server through a target interface. Optionally, the client obtains a voice stream returned by the server through the target interface, and the voice stream includes byte stream data of the voice.

[0086] Step S306: Concatenate the first voice data and the second voice data to obtain third voice data, where the second voice data is the voice data corresponding to the content other than the name in the interaction content.

[0087] In the technical solution provided in step S306 of the present invention, after the client obtains the first voice data, it may concatenate the first voice data and the second voice data to obtain third voice data.

[0088] In this embodiment, the current interaction content to be carried out includes content other than the name of the virtual game character, and the content other than the name of the virtual game character corresponds to second voice data, which can be audio, that is, the audio to be spliced. After the client receives the first voice data, the client can splice the first voice data and the second voice data. For example, the client splices the first voice data and the second voice data according to the sampling points of the first voice data and the second voice data. For example, the sampling points include time information, and the first voice data and the second voice data can be spliced through it, so as to obtain the third voice data.

[0089] Step S308, play the third voice data.

[0090] In the technical solution provided in step S308 of the present invention, after the client splices the first voice data and the second voice data to obtain the third voice data, the third voice data can be played, that is, the interaction content is played through voice, so as to achieve the purpose of playing the name of the virtual game character in the interaction content by voice when the virtual game character proceeds to the corresponding interaction content.

[0091] It should be noted that since the duration of the target text to be synthesized in this embodiment generally reaches 0.5-1 second, and the waiting duration for the model to generate the first voice data corresponding to this section of the target text basically does not exceed 0.1 second, splicing the first voice data with the second voice data corresponding to the content other than the name of the virtual game character in the interaction content and playing the obtained third voice data will not consume too much time, so players should not feel the interruption of the playback of the entire interaction content.

[0092] When the game proceeds to the interaction content of the corresponding virtual game character and the interaction content contains the name of the virtual game character, this embodiment can call the target interface of the server to return the first voice data matching the name of the virtual game character; the client splices the first voice data returned by the server to the second voice data of the current conversation to obtain the third voice data, and then plays the third voice data. That is to say, this embodiment adds voice feedback to the interaction content containing the name of the virtual game character through the above method, making players more immersive during the game plot and important prompts, improving the authenticity and interactivity of the game virtual world, and optimizing the user experience.

[0093] Next, the technical solution of one of the embodiments of the present invention will be further illustrated by combining preferred embodiments.

[0094] In a related technology, in games on mobile terminals and PC games, after a player names the corresponding virtual game character in the game, the name of the virtual game character is fed back in text.

[0095] In another related technology, after a player names the corresponding virtual game character in the game, the name of the virtual game character will be skipped in the voice of the virtual game character, and only other fixed text content will be read, or it will be fed back in a fixed voice. For example, the player can be addressed in other ways set by the game, such as "Miss", "Girl" or similar addresses.

[0096] However, in the above methods, the name given by the player to the virtual game character generally only appears in function interfaces such as the text chat interface, the team formation interface, and personal information, and there is no interaction with the virtual game character in the game, resulting in poor interactivity of the name of the virtual game character. In addition, the voice data of the related technology does not attach importance to the name given by the player to the virtual game character, and the name of the virtual game character lacks voice interaction, and generally only returns fixed voice content (such as "Miss", "Girl", etc.), resulting in a weak sense of immersion for the player when experiencing the game plot. Even in games with a first-person perspective, the player's self-defined name for the virtual game character is not really used, resulting in a relatively fragmented understanding of the virtual game character and the player's own identity.

[0097] To address the above problems, in this embodiment, the voice of the corresponding virtual game character can be synthesized according to the name given by the player to the virtual game character in the game, and the voice data matching the name of the virtual game character is fed back to the player. The following further introduces this method.

[0098] In this embodiment, the server sets up the AI voice generation model for each virtual game character corresponding to the game; when the game progresses to the dialogue content of the corresponding virtual game character and the dialogue content contains the name of the virtual game character given by the player, the target interface of the server can be called to return the voice data corresponding to the name; the client splices the voice data corresponding to the name returned by the server to the voice of the current dialogue content, and then plays the voice corresponding to the spliced voice data.

[0099] The following introduces the server-side method of converting the text corresponding to the name of the virtual game character into voice data in this embodiment.

[0100] Figure 4 It is a schematic diagram of a text-to-speech server according to an embodiment of the present invention. As Figure 4As shown, the target request sent by the client to the server includes the name text to be synthesized, the spliced audio ID, and the target speaker ID. Among them, the spliced audio can be obtained through the spliced audio ID, and then in the text-to-speech system with reference audio, by processing the name text to be synthesized, the spliced audio, and the target speaker ID, the name synthesis voice corresponding to the name of the virtual game character is obtained and returned to the client through the audio stream.

[0101] The text-to-speech system with reference audio in this embodiment will be further introduced below.

[0102] Figure 5 It is a schematic diagram of a text-to-speech system with reference audio according to an embodiment of the present invention. As Figure 5 shown, the text-to-speech system with reference audio may include: a text preprocessing module 51, an acoustic feature extraction module 52, a speaker vector table module 53, a text encoder 54, a reference audio encoder 55, an attention mechanism module 56, an acoustic decoder 57, and a vocoder 58.

[0103] The text preprocessing module 51 can be used to convert the name text to be synthesized input to the system into a phoneme data sequence and a prosody data sequence corresponding to the name text to be synthesized through a series of rule-based or neural network models.

[0104] In this embodiment, the text preprocessing module 51 may include a text-to-phoneme model, which adopts a neural network model with a CNN+LSTM structure and is trained using a Cross-entropy loss function.

[0105] In this embodiment, the text preprocessing module 51 may include a text-to-prosody model, which adopts a neural network model with an LSTM structure and is trained using a Cross-entropy loss function.

[0106] For example, if the input name text to be synthesized is "I love China", then the name text is converted through the text-to-phoneme model, and the phoneme data sequence "w o3 a i3 zh ong1 g uo2" is output; the name text is converted through the text-to-prosody model, and the prosody data sequence "#1#1*#4" is output. The phoneme data sequence and the prosody data sequence are combined to finally obtain the phoneme prosody data sequence "w o3#1 a i3#1 zh ong1 g uo2#4".

[0107] The acoustic feature extraction module 52 can be used to convert the spliced audio obtained from the spliced audio ID from the waveform into some informative acoustic features, and the acoustic features can be Mel spectrograms.

[0108] The speaker vector table module 53 can be used to convert the target speaker ID into a target speaker vector for controlling the speaker timbre corresponding to the synthesized voice data.

[0109] The text encoder 54 can map the input sequence of phoneme prosody data to a high-dimensional text feature space encoding through a non-linear transformation to obtain a text feature sequence.

[0110] The reference audio encoder 55 can extract the overall style information of the entire audio to be spliced (reference audio) by receiving the acoustic features of the audio to be spliced and encode the overall style information to obtain the overall style encoding information.

[0111] In this embodiment, a reference audio encoder based on CNN and LSTM can be adopted. By performing information extraction and information compression on the acoustic features of the input audio to be spliced, a mathematical vector representing the style, that is, the style vector, can be finally obtained. Among them, the method of performing information extraction and information compression on the acoustic features of the audio to be spliced belongs to unsupervised learning.

[0112] The attention mechanism module 56 can align the text feature sequence and the speech feature sequence (the acoustic features of the audio to be spliced) to obtain an alignment result because the speech feature sequence is longer than the text feature sequence. In addition, the attention mechanism module 56 also receives the overall style encoding information and the target speaker vector from the reference audio encoder 55.

[0113] The acoustic decoder 57 can return the alignment result obtained by aligning the text feature sequence and the speech feature sequence through the attention mechanism module, as well as the overall style encoding information and the target speaker vector, back to the original speech acoustic feature space through a non-linear transformation to return the predicted speech acoustic features.

[0114] The vocoder 58 converts the above predicted speech acoustic features into a speech waveform / signal to obtain the named synthesized speech, which is returned to the client through the speech stream so that the client can splice it on the speech data of the current conversation of the speaker in the game and then play the spliced speech data. Among them, the splicing in this embodiment can refer to the splicing of the audio sampling points corresponding to the speech data.

[0115] In this embodiment, the most important component in the text-to-speech server is the reference audio encoder. Through this reference encoder, the overall style of the audio to be spliced can be effectively encoded to extract style information for the final name synthesis speech stage. If this reference audio encoder is not used in this embodiment, the synthesized speech for the same name text will always have the same style. That is, in the absence of the overall style encoding information of the reference audio output by the reference audio encoder, the attention mechanism module only receives the text feature sequence output by the text encoder, the acoustic features output by the acoustic feature extraction module, and the speaker vector. After adding the reference audio encoder in this embodiment, the attention mechanism module additionally receives the overall style encoding information of the reference audio. Since the text feature sequence output after the same name text is input into the text encoder is fixed. Therefore, when the acoustic features and the speaker vector remain unchanged, the synthesized speech data is basically of the same style. After adding the overall style encoding information of the reference audio in this embodiment, the overall style of the synthesized speech data is affected by the overall style encoding information of the reference audio. Thus, when different reference audios are used, the style of the synthesized speech for the same name text will change.

[0116] In addition, if this reference audio encoder is not used in this embodiment, the style of the synthesized speech for the name text may not be consistent in the case of the same audio to be spliced. This is because in the absence of the overall style encoding information of the reference audio output by the reference audio encoder, the synthesized speech for the name text is generally more affected by the input of the name text. Thus, sometimes the styles of the synthesized speech corresponding to different name texts will also be different. However, by introducing the overall style encoding information of the reference audio output by the reference audio encoder in this embodiment, the overall style encoding information of the same reference audio can be used when synthesizing different name texts, so that the styles of the synthesized speech for different name texts can be kept relatively consistent.

[0117] In this embodiment, in the absence of the overall style encoding information of the reference audio output by the reference audio encoder, it will cause the splicing part between the synthesized speech of the name text and the audio to be spliced to be unnatural, which mainly includes discontinuous prosody rhythm and inconsistent channels. And this problem can be effectively solved by using the overall style encoding information of the reference audio output by the reference audio encoder in this embodiment, making the splicing of the synthesized speech of the name text and the audio to be spliced relatively natural and coherent.

[0118] The above method of this embodiment, which is also a voice feedback method through AI voice recognition, can recognize the name given by the player to the virtual game character and feed back the voice data matching the name of the virtual game character to the player. That is, this embodiment achieves the purpose of adding voice feedback to the conversation containing the player's name, improving the authenticity and interactivity of the game virtual world. In this way, during the game plot and important prompts, the player can have a stronger sense of immersion, be more immersed in experiencing the plot, thus obtaining a better sense of identity substitution, and further enhancing the visual and auditory experience of the game.

[0119] One embodiment of the present invention further provides a data processing device. It should be noted that the data processing device of this embodiment can be used to execute the data processing method shown in the embodiments of the present invention. Figure 2 The data processing method shown.

[0120] Figure 6 It is a schematic diagram of a data processing device according to an embodiment of the present invention. As Figure 6 shown, the data processing device 60 may include: a receiving unit 61, a conversion unit 62, and a first sending unit 63.

[0121] The receiving unit 61 is configured to receive a target request from the client. Among them, the information carried in the target request at least includes a target text, and the target text is used to represent the name of the virtual game character.

[0122] The conversion unit 62 is configured to convert the target text into first voice data in response to the target request.

[0123] The first sending unit 63 is configured to send the first voice data to the client, so that the client splices the first voice data and second voice data into third voice data. The second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character.

[0124] One embodiment of the present invention further provides another data processing device. It should be noted that the data processing device of this embodiment can be used to execute the data processing method shown in the embodiments of the present invention. Figure 3 The data processing method shown.

[0125] Figure 7 It is a schematic diagram of another data processing device according to an embodiment of the present invention. As Figure 7 shown, the data processing device 70 may include: a second sending unit 71, an obtaining unit 72, a splicing unit 73, a splicing unit 73, and a playing unit 74.

[0126] A second sending unit 71, configured to send a target request to a server when it is detected that the interaction content to be performed by the virtual game character includes the name of the virtual game character, where the information carried in the target request includes at least target text, and the target text is used to represent the name.

[0127] An obtaining unit 72, configured to obtain first voice data, where the first voice data is obtained by the server in response to the target request and converting the target text.

[0128] A splicing unit 73, configured to splice the first voice data and second voice data to obtain third voice data, where the second voice data is the voice data corresponding to the content other than the name in the interaction content.

[0129] A playing unit 74, configured to play the third voice data.

[0130] In the data processing device of this embodiment, the target text used to represent the name of the virtual game character is converted into first voice data to be spliced with the second voice data corresponding to the interaction content, achieving the purpose of adding voice feedback to the interaction content including the name of the virtual game character, avoiding that the name of the virtual game character is usually feedback in text form, and also avoiding skipping the name of the virtual game character in the conversation, thereby solving the technical problem of poor interactivity of the name of the virtual game character and achieving the technical effect of improving the interactivity of the name of the virtual game character.

[0131] One of the embodiments of the present invention further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium. When the computer program is run by a processor, it controls the device where the computer-readable storage medium is located to execute the data processing method of the embodiment of the present invention.

[0132] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store computer programs.

[0133] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0134] Optionally, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0135] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to be implemented. In this way, the present invention is not limited to any specific combination of hardware and software.

[0136] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data processing method, characterized in that, Including: Receiving a target request from a client, where the information carried in the target request includes at least a target text, and the target text is used to represent the name of a virtual game character; Responding to the target request and converting the target text into first voice data; Sending the first voice data to the client so that the client concatenates the first voice data and second voice data into third voice data, where the second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character; Among them, converting the target text into first voice data includes: extracting text features of the target text; performing waveform conversion on the second voice data to obtain first acoustic features; using an attention mechanism model to align the text features and the first acoustic features to obtain an alignment result; converting the style information of the second voice data and the alignment result into the first voice data; Among them, the method further includes: determining the style information based on the first acoustic features.

2. The method according to claim 1, wherein The information carried in the target request further includes first identification information of the virtual game character. The method further includes: obtaining a target vector of the virtual game character based on the first identification information, where the target vector is used to represent the timbre of the virtual game character; Converting the style information and the alignment result into the first voice data includes: converting the target vector, the style information, and the alignment result into the first voice data.

3. The method according to claim 2, wherein Converting the target vector, the style information, and the alignment result into the first voice data includes: Synthesizing the target vector, the style information, and the alignment result into second acoustic features; Converting the second acoustic features into the first voice data.

4. The method according to claim 3, wherein The method further includes: Encoding the style information to obtain style encoding information; Synthesizing the target vector, the style information, and the alignment result into second acoustic features includes: synthesizing the target vector, the style encoding information, and the alignment result into the second acoustic features.

5. The method according to claim 1, wherein The information carried in the target request further includes second identification information of the second voice data, and the method further includes: Obtaining the second voice data based on the second identification information.

6. The method according to any one of claims 1 to 5, characterized in that Extracting the text features of the target text includes: Converting the target text into a phoneme data sequence and / or a prosody data sequence; Extracting the text features from the phoneme data sequence and / or the prosody data sequence.

7. A data processing method, characterized in that, Including: When it is detected that the interaction content to be performed by the virtual game character includes the name of the virtual game character, sending a target request to the server, where the information carried in the target request includes at least a target text, and the target text is used to represent the name; Obtain first voice data, where the first voice data is obtained by the server in response to the target request and converting the target text, the first voice data is determined by the alignment result, the alignment result is obtained by using an attention mechanism model to align the text features extracted from the target text and the first acoustic features obtained by waveform conversion of the second voice data, the first voice data is obtained by converting the style information of the second voice data and the alignment result, and the style information is determined based on the first acoustic features; Concatenate the first voice data and the second voice data to obtain third voice data, where the second voice data is the voice data corresponding to the content other than the name in the interaction content; Play the third voice data.

8. A data processing device, characterized in that, Includes: A receiving unit, configured to receive a target request from a client, where the information carried in the target request includes at least a target text, and the target text is used to represent the name of a virtual game character; A conversion unit, configured to respond to the target request and convert the target text into first voice data; A first sending unit, configured to send the first voice data to the client, so that the client concatenates the first voice data and the second voice data into third voice data, and the second voice data is the voice data corresponding to the content other than the name in the interaction content to be performed by the virtual game character; Wherein, the conversion unit is configured to convert the first voice data by performing the following steps: extract the text features of the target text; perform waveform conversion on the second voice data to obtain first acoustic features; use an attention mechanism model to align the text features and the first acoustic features to obtain an alignment result; convert the style information of the second voice data and the alignment result into first voice data; Wherein, the device is further configured to perform the following steps: determine the style information based on the first acoustic features.

9. A data processing device, characterized in that, Includes: A second sending unit, configured to send a target request to the server when it is detected that the interaction content to be performed by the virtual game character includes the name of the virtual game character, where the information carried in the target request includes at least a target text, and the target text is used to represent the name; An obtaining unit, configured to obtain first voice data, where the first voice data is obtained by the server in response to the target request and converting the target text, the first voice data is determined by the alignment result, the alignment result is obtained by using an attention mechanism model to align the text features extracted from the target text and the first acoustic features obtained by waveform conversion of the second voice data, the first voice data is obtained by converting the style information of the second voice data and the alignment result, and the style information is determined based on the first acoustic features; A splicing unit, configured to splice the first voice data and the second voice data to obtain third voice data, where the second voice data is voice data corresponding to the content other than the name in the interaction content; A playing unit, configured to play the third voice data.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where when the computer program is run by a processor, it controls the device where the computer-readable storage medium is located to execute the method described in any one of claims 1 to 7.

11. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech synthesis method and device

    CN109584859A

  • Voice synthetic method and device, dictionary constructional method and computer ready-read medium

    CN1282017A