Voice Processing Method, Apparatus and Electronic Device

Through the voice generation plug-in to generate and play audio data during game running, it solves the cumbersome and error problems caused by manual import of voice resources in the prior art, and improves game testing and development efficiency.

CN114360487BActive Publication Date: 2025-07-18NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111639353.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-18
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

In the prior art, it is necessary to manually import the game engine after generating game voice resources, resulting in cumbersome operations and error-proneness, and reducing game testing and development efficiency.

Method used

The voice generation plug-in generates voice resource packages based on preset voice parameters and playback rules, and directly plays audio data when the game is running, avoiding the manual import step.

Benefits of technology

It simplifies the operation process, reduces time costs, improves game testing and development efficiency, and reduces the probability of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360487B_ABST
    Figure CN114360487B_ABST
Patent Text Reader

Abstract

The present invention provides a voice processing method, apparatus and electronic device. A voice generation plug-in generates a voice resource packet according to preset voice parameters and pre-configured playback rules; in response to a first operation on a processing control for a target game, obtains the voice parameters and playback rules from the voice resource packet; generates audio data according to the voice parameters; and plays the audio data during the operation of the target game according to the playback rules. In this way, by integrating the voice generation plug-in into the audio engineering process of the game engine, it is possible to directly generate the audio data required for the game according to the preset voice parameters during the game operation and play it according to the playback rules, without pre-generating the audio data and manually importing it into the game engine, reducing the time cost, with a simple operation process and not easily making mistakes, and improving the game testing efficiency and development efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of game voice, and in particular to a voice processing method, device and electronic device. Background Art

[0002] Game voice, as a tool to convey the game world view and shape game characters, is an important component module in the game audio system. The game voice usually needs to be recorded by professionals, and then voice designers apply the recorded voice resources to the game through programs. However, due to the long recording cycle and time of voice, generally, during the process of testing the game voice function, it is not necessary to use the recorded voice resources for game testing. In order to improve the testing efficiency, third-party tools are usually used to generate voice resources to provide temporary voice resources for game testing.

[0003] In the related art, after generating voice resources through a third-party tool, voice designers need to manually import the voice resources into the audio engine of the game before they can perform game voice function testing based on the voice resources. However, the third-party tool can only generate individual audio resources. Therefore, voice designers need to manually import each audio resource into the audio engine of the game one by one before they can provide the corresponding voice resources for game voice function testing. This method is very time-consuming, the operation process is cumbersome and error-prone, resulting in a reduction in game testing efficiency and affecting game development efficiency. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a voice processing method, device and electronic device to simplify the voice generation process, and at the same time improve game testing efficiency and development efficiency.

[0005] In a first aspect, an embodiment of the present invention provides a voice processing method, which is applied to a voice generation plugin; the voice generation plugin provides a game development interface, and the game development interface includes a processing control for a target game; the method includes: generating a voice resource package corresponding to voice parameters according to preset voice parameters and pre-configured playback rules; in response to a first operation on the processing control for the target game, obtaining the voice parameters and playback rules from the voice resource package; generating audio data corresponding to the voice parameters according to the voice parameters; and playing the audio data during the operation of the target game according to the playback rules.

[0006] Further, the editing interface includes a parameter input control and a serialization control; the step of generating a voice resource package corresponding to voice parameters according to preset voice parameters and pre-configured playback rules includes: in response to a voice parameter input operation on the parameter input control, obtaining the voice parameters; in response to a second operation on the serialization control, serializing the voice parameters and the pre-configured playback rules according to a preset encoding format to generate a voice resource package.

[0007] Further, the step of obtaining voice parameters and playback rules from the voice resource package includes: deserializing the voice resource package according to a preset decoding format to obtain voice parameters and playback rules.

[0008] Further, the step of generating audio data corresponding to the voice parameters according to the voice parameters includes: obtaining a target voice resource matching the voice parameters from a preset voice resource library through a preset voice interface; wherein, the voice parameters at least include: the target text; synthesizing audio data of the target text through the target voice resource.

[0009] Further, after the step of generating audio data corresponding to the voice parameters according to the voice parameters, the method further includes: caching the audio data into a first specified memory by sampling; wherein, the size of the first specified memory is determined according to the sampling rate and the specified audio stream channel number during the initial sampling process; writing the memory address of the first specified memory where the audio data is cached into a preset pointer variable.

[0010] Further, the above method further includes: converting the data structure of the audio data into a specified format through a data signal processing module.

[0011] Further, the above method further includes: when the audio data playback is completed or stopped, deleting the cached audio data and releasing the first specified memory.

[0012] Further, the step of playing the audio data during the operation of the target game according to the playback rules includes: obtaining the target audio data from the first specified memory according to the preset pointer variable; outputting the target audio data to the audio stream channel corresponding to the voice generation plugin to play the target audio data during the operation of the target game according to the playback rules.

[0013] Further, the editing interface further includes a voice audition control; after the step of obtaining the voice parameters in response to the voice parameter input operation of the parameter input control, the above method further includes: responding to a third operation on the voice audition control, obtaining the voice parameters from the cache space; generating audio data corresponding to the voice parameters according to the voice parameters, caching the audio data into a second specified memory, and playing the audio data in the editing interface.

[0014] In a second aspect, an embodiment of the present invention provides a voice processing device, which is disposed in a voice generation plugin; the voice generation plugin provides a game development interface, and the game development interface includes processing controls for a target game; the device includes: a resource package generation module, configured to generate a voice resource package corresponding to voice parameters according to preset voice parameters and preconfigured playback rules; an acquisition module, configured to, in response to a first operation on the processing control for the target game, acquire the voice parameters and the playback rules from the voice resource package; a data generation module, configured to generate audio data corresponding to the voice parameters according to the voice parameters; and a playback module, configured to play the audio data according to the playback rules when the target game is running.

[0015] In a third aspect, an embodiment of the present invention provides an electronic device, which is characterized by including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the voice processing method according to any one of the first aspect.

[0016] In a fourth aspect, an embodiment of the present invention provides a machine-readable storage medium, which is characterized in that the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the voice processing method according to any one of the first aspect.

[0017] The embodiments of the present invention bring the following beneficial effects:

[0018] The present invention provides a voice processing method, device and electronic device. The voice generation plugin generates a voice resource package according to preset voice parameters and preconfigured playback rules; in response to a first operation on the processing control for the target game, the voice parameters and the playback rules are acquired from the voice resource package; audio data is generated according to the voice parameters; and the audio data is played according to the playback rules when the target game is running. In this way, by integrating the voice generation plugin into the audio engineering process of the game engine, the audio data required for the game can be directly generated according to the preset voice parameters during the game operation and played according to the playback rules, without the need to pre-generate the audio data and manually import it into the game engine, reducing the time cost, with a simple operation process and not easily making mistakes, improving the game testing efficiency and development efficiency.

[0019] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.

[0020] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given below and, in conjunction with the accompanying drawings, are described in detail as follows. Brief Description of the Drawings

[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a flowchart of a voice processing method provided by an embodiment of the present invention;

[0023] Figure 2 It is a schematic diagram of an editing interface in a voice processing method provided by an embodiment of the present invention;

[0024] Figure 3 It is a schematic diagram of the interface integration structure in a voice processing method provided by an embodiment of the present invention;

[0025] Figure 4 It is a schematic diagram of the structure of a voice processing device provided by an embodiment of the present invention;

[0026] Figure 5 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0027] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention.

[0028] As a tool for conveying the game world view and shaping game characters, in-game voice is an important component of the game audio system. The in-game voice usually needs to be recorded by professionals, and then voice designers apply the recorded voice resources to the game through programs. However, since the probability of using temporary voice resources in the final release version of the game is very low, and the iteration of the text is very frequent during the process of testing the audio function, there will be a large waste of recording costs. Moreover, since recording voice requires selecting actors and arranging schedules, and finding suitable venues and equipment, each recording cycle will be relatively long. And the voice system must use temporary voice resources as a reference to start building, which means that the functional developers of the voice system also have to wait for the same long time period before they can start working. Since the development of the voice system itself is a process that requires repeated running-in and iteration, in the case where resource output and function development cannot be carried out synchronously, the development cycle will increase significantly, resulting in low development efficiency.

[0029] To improve the testing efficiency, third-party tools are usually used to generate voice resources to provide temporary voice resources for game testing. In related technologies, after generating voice resources through third-party tools, voice designers need to manually import the voice resources into the game's audio engine before they can perform in-game voice function testing based on the voice resources. However, third-party tools can only generate individual audio resources. Therefore, voice designers need to manually import each audio resource into the game's audio engine one by one before they can provide the corresponding voice resources for in-game voice function testing. This method is very time-consuming, the operation process is cumbersome and error-prone, resulting in a decrease in game testing efficiency and affecting game development efficiency.

[0030] To facilitate the understanding of this embodiment, first, a voice processing method disclosed in an embodiment of the present invention will be introduced in detail. This method is applied to a voice generation plug-in; the voice generation plug-in provides a game development interface, and the game development interface includes processing controls for the target game;

[0031] Wwise is a game audio solution launched by Audiokinetic Inc., and Wwise Plugin SDK is an open plug-in development framework provided by Wwise. Developers provide extended functions for the editing module by implementing each interface of the plug-in. The above-mentioned voice generation plug-in (Source Plugin) is used to generate audio resources and output them to the audio workflow of game development. The above-mentioned voice generation plug-in is integrated in Wwise (i.e., the audio workflow) of the game engine. Usually, the above-mentioned voice generation plug-in includes two modules, an editing module and a sound engine module. The sound engine module is integrated into the game development application in the form of a plug-in. As Figure 1 shown, the method includes:

[0032] Step S102: Generate a voice resource package corresponding to the voice parameters according to the preset voice parameters and the pre-configured playback rules;

[0033] The above-mentioned preset voice parameters can be set according to actual needs and may include the text to be converted into voice, the gender of the speaker, the voice rate, and the language type. For example, the text can be converted into voice using English, Chinese, Korean, Japanese, German, etc., and of course, it can also be languages from various regions. Additionally, the above voice parameters usually need to be input by voice designers according to the game requirements.

[0034] The above-mentioned pre-configured playback rules are pre-configured into the voice generation plug-in and usually include information such as playback logic and signal routing, which need to be implemented by voice designers through the existing Wwise framework. Specifically, after voice designers set up the playback logic and routing structure in the Wwise editing module, they add an instance of the voice generation plug-in to this structure. In terms of playback logic, Wwise provides the following functions: Containers, including random containers, switch containers, mix containers, etc.; A random container will randomly play the content inside; A switch container switches to play the content inside according to the Switch field sent by the game; A mix container mixes the content inside according to the RTPC (continuously changing real-time data) sent by the game. Playback control, including loop playback, delayed playback, etc. Parameter control, including modulators such as Envelope, LFO (Low Frequency Oscillator), Time, etc., which can be used to modulate a series of playback parameters such as volume, pitch, and delay duration. In terms of signal routing, Wwise provides two main tools: Audio Bus and Aux Bus. Users can customize the output bus or auxiliary bus of the container structure and separately control the input audio signals at the bus end. In the playback rules section, Wwise provides a priority system. An object with a high priority can interrupt the playback of an object with a low priority; The interrupted object can be specified whether to continue playing at a certain future moment or stop playing.

[0035] In actual implementation, voice designers can input voice parameters in the editing interface of the voice generation plug-in and then click the generation control. The editing module generates a voice resource package based on the preset generation scheme. This voice resource package can be represented as a Soundbank, which is the basic unit for loading audio assets during game operation. The purpose of generating the voice resource package is to support the writing and reading of various languages to convert voice parameters into corresponding voices.

[0036] Step S104: In response to a first operation on the processing control for the target game, obtain the voice parameters and playback rules from the voice resource package;

[0037] This step is completed in the game development interface provided by the voice generation plugin. Specifically, game developers can click on the processing control in the game development interface, and through the sound engine module of the voice generation plugin, obtain the above-generated voice resource package from the editing module, and then read the voice parameters and playback rules in the resource package according to the preset reading method. The purpose is to correctly read the voice parameters input by the voice designer and convert the voice parameters into audio data.

[0038] Step S106, generate audio data corresponding to the voice parameters according to the voice parameters;

[0039] Since the voice parameters set various attributes of the audio data, such as gender, rate, language, etc., it is necessary to obtain the sound resources that can convert the text into audio data with corresponding attributes according to the voice parameters, and then convert the text in the voice parameters into voice. For example, the text included in the voice parameters is "Hello", the gender is female, the language is Chinese, and the speech rate is "0" (i.e., normal speech rate). The audio data generated according to this voice parameter includes the Chinese "Hello" spoken in a female voice at a normal speech rate. Usually, the text in the above voice parameters can include multiple, and each text can set different voice parameters. Usually, the game voice of a game will include multiple audio data for playing in different scenarios.

[0040] Step S108, play the audio data during the operation of the target game according to the playback rules.

[0041] Specifically, the corresponding audio data can be played in real time during the game operation according to the playback rules to develop the voice interaction function during the game operation. For example, during the game operation, if the virtual character kills the enemy virtual character, the first audio data will be played. If the virtual character is killed, the second audio data will be played, etc. Specifically, it is necessary to play the generated audio data in real time according to the playback rules. It should be noted that the above-generated audio data is only used for developing the voice function of the game. In actual game applications, the temporary audio data during development is rarely used. Usually, the recorded voice resources need to replace the above audio data.

[0042] A voice processing method provided by an embodiment of the present invention generates a voice resource package through a voice generation plug-in according to preset voice parameters and pre-configured playback rules; in response to a first operation on a processing control for a target game, obtains the voice parameters and playback rules from the voice resource package; generates audio data according to the voice parameters; and plays the audio data during the operation of the target game according to the playback rules. In this way, by integrating the voice generation plug-in into the audio engineering process of the game engine, the audio data required by the game can be directly generated according to the preset voice parameters during the game operation and played according to the playback rules, without pre-generating the audio data and manually importing it into the game engine, reducing the time cost, the operation process is simple and not prone to errors, and the game testing efficiency and development efficiency are improved.

[0043] The above editing interface includes a parameter input control and a serialization control; as Figure 2 shown in the editing interface, the parameter input control therein includes controls that need to be input, such as "Hello" at the "Text to speak" in the figure, and also includes controls that need to be selected. For example, "Female" and "Chinese" at the "Gender" and "Language" in the figure, that is, the text of gender and various languages can be input; among them, the gender can be selected in the drop-down box; the speech rate (corresponding to "Voice Rate" in the figure) can be a controllable component, and different rates can be selected by adjusting the controllable component.

[0044] In addition, the above editing interface is implemented by setting a graphical user interface. Specifically, a series of interface-related attributes of each parameter are defined in an XML file, including: the name and display name of the parameter, the type of the parameter (integer, floating-point, string, etc.), the default value of the parameter, the maximum and minimum value intervals of the parameter, the enumeration type and enumeration values of the parameter. Different from the parameter types defined in the XXXPluginSourceParams class in the sound engine part, the parameter attributes defined in the XML file are used to specify the input control form of the parameter in the editing interface. For example, for a parameter defined as a string type in the XML, the input control of this parameter in the editing interface will be a text input box; for a floating-point parameter with a defined value range, the input control of this parameter in the editing interface will be a Slider (slider) plus a text input box. The user sets the value of this parameter by dragging the Slider, and the sliding range of the Slider will be between the value ranges defined by the parameter;

[0045] For parameters that define an enumeration type (such as a language code), assuming the parameter itself is of string type, the corresponding parameter input control is no longer a text input box but becomes a dropdown menu: Each option in the dropdown menu is an enumeration value defined in the XML file (for example, Chinese / English / French, etc.). Selecting one of the enumeration values means that the value of this string variable (language code) is a definite value. The correspondence between this definite value and the enumeration value is also defined in the XML file (Chinese corresponds to 804, English corresponds to 409, French corresponds to 400, etc.).

[0046] The following describes the steps to generate a voice resource package corresponding to voice parameters according to preset voice parameters and pre-configured playback rules. A possible implementation: Respond to a voice parameter input operation on the parameter input control to obtain the voice parameters; Respond to a second operation on the serialization control, and serialize the voice parameters and the pre-configured playback rules according to a preset encoding format to generate a voice resource package.

[0047] The above serialization process can be understood as an agreement, that is, the data written in accordance with the preset encoding format during serialization should be read in the same format during deserialization. For example, for string data, different character encoding / decoding formats interpret the same string data differently. In this embodiment, the process of storing and retrieving strings and the process of encoding and decoding strings are set in different modules respectively. Serialization is performed in the editing module, and deserialization is performed in the sound engine.

[0048] In order to support the correct writing and reading of text in multiple languages, the voice generation plugin in this embodiment uses variable-length UTF-8 encoding to record the strings in the voice parameters. In the serialization interface (GetBankParameters), the voice parameters input by the user (i.e., the string) are directly passed to the DataWriter in the form of a char array, and the string is written into the voice resource package Soundbank "as is".

[0049] In the above method, the voice generation plugin provides an editing interface for the user. The user can set voice parameters in the Wwise editing interface. At the same time, the voice generation plugin is tightly integrated with Wwise, enabling voice designers to modify the voice text and other parameters in the game in real time for rapid iteration. Through the preset encoding format, the string can be serialized into the voice resource package and correctly deserialized during the game runtime, and the input string can be text in languages from all over the world without additional settings.

[0050] In order to read the voice parameters and playback rules in the voice resource package, in the sound engine module, the voice resource package is deserialized according to a preset decoding format to obtain the voice parameters and playback rules.

[0051] The above deserialization process is completed in the sound engine. For example, in the deserialization interface (SetParamsBlock), the char array stored in the voice resource package Soundbank can be directly encapsulated into a string variable. When the game-side plugin gets this string variable, it will convert the string variable into a wide string of wstring type according to the "agreed" variable-length UTF-8 encoding, and then decode the wide string according to the preset decoding format to obtain the above voice parameters and playback rules. In this way, through the process of serialization and deserialization, the voice parameters input by the voice designer in the editing interface can be read into the sound engine to generate corresponding audio data through the deserialized voice parameters.

[0052] In the entire serialization and deserialization process of the above strings, the voice generation plugin in this embodiment stores and retrieves the string data as a simple char array: that is, the actual encoding type of the string is not considered during the storage and retrieval process. After retrieving the string data, the string is decoded according to the agreed decoding format to obtain the correct "interpretation" of the string.

[0053] The following describes the steps of generating audio data corresponding to voice parameters according to voice parameters. Specifically, through a preset voice interface, a target voice resource matching the voice parameters is obtained from a preset sound resource library; wherein, the voice parameters at least include: target text; the audio data of the target text is synthesized through the target voice resource.

[0054] In fact, the above voice plugin also integrates the Microsoft Speech API (SAPI), that is, the above preset voice interface is a set of voice generation tools provided by Microsoft, including multiple functional modules such as text-to-speech, speech recognition, and speech synthesis. In this embodiment, in order to enable the text in the voice parameters to be converted into voice in real time, by integrating the voice interface in the voice generation plugin, the voice engine function of the current system is opened to developers in a simple form. In addition, the above voice parameters can also include language type, audio gender, language rate, etc.

[0055] Specifically, the voice resources can be requested from the current operating system through a preset voice interface. First, it is necessary to obtain the first resources that can be converted to the gender according to the gender attribute in the deserialized voice parameters. It is also necessary to obtain the second resources that can be converted to the language type according to the language type attribute in the deserialized voice parameters, such as English. Then, according to the obtained voice resources, the target text is converted into audio data according to the target text and speech rate in the deserialized voice parameters.

[0056] It can be understood that the above voice resources refer to the system resources provided by the current operating system, and the search operation can be performed on them to obtain audio data. It can be understood as a "person who can speak". Only by finding this "person" can it be made to speak the input voice parameters, that is, the steps of synthesizing the audio data of the target text through the target voice resources.

[0057] As Figure 3 shown in the figure, the figure includes various interfaces integrated with the voice generation plug-in. Among them, Application is an application, which can be a game development application. By integrating the voice interface API and SAPI in the voice generation plug-in, and through the automatic collection of channel data (Distributors Data Integration, DDI), the voice resources are obtained from the voice conversion engine of the current operating system, and the target text is converted into audio data in the current operating system. In addition, Figure 3 the Recognition Engine in it is a voice recognition engine. In this embodiment, this module is not applied, and only the TTS Engine (voice conversion engine) is applied.

[0058] In the above method, by using the voice resources provided by the current operating system, voice designers can design localized voices in up to hundreds of countries / regions, and can set the gender and speech rate to achieve the purpose of voice conversion in the current operating system. At the same time, the deployment volume of the plug-in is very small.

[0059] In addition, it should be noted that the path to obtain the voice resources through the voice interface is fixed, and the voice resources are managed in the registry. The voice packs of various countries in the world can be installed through the settings of the current operating system (such as the windows system). However, by default, these resources cannot be read by the voice interface. In this embodiment, this problem is solved by modifying the registry item, and the same number of voice resources of countries or regions as that of the windows narrator can be provided for the voice generation plug-in.

[0060] After obtaining the sound resource in the current operating system, by default, the synthesized audio data is directly played through the audio playback device of the current operating system; that is to say, it will not pass through the audio channel of the game application. Therefore, in this embodiment, after the step of synthesizing the audio data of the target text through the target sound resource, the audio data is cached in the first specified memory by sampling; wherein, the size of the first specified memory is determined according to the sampling rate and the specified number of audio channels during the initial sampling process; the memory address of the first specified memory where the audio data is cached is written into a preset pointer variable.

[0061] Since the device can only store digital audio formats, but the audio data actually obtained by the operating system is an analog signal, it is necessary to sample the generated audio data and cache the audio data in the digital audio format in the first specified memory of the speech generation plug-in, which can be the cache memory of the sound engine module. This is to facilitate the real-time acquisition of audio data during game operation and play it on the audio playback device of the game development application.

[0062] In addition, to address the problem that the large volume of audio data causes the game package size to increase, in this embodiment, the first specified memory for caching the audio data is flexibly determined according to the volume of the Audio Buffer of the audio data, that is, calculated from the sampling rate and the number of channels. Therefore, there is no problem of insufficient or excessive allocation of the first specified memory. Among them, the higher the sampling rate, the more audio data information is retained. For example, if the sampling rate is 48,000 Hertz, that is, the audio data is sampled 48,000 times in one second. If the specified number of audio channels is two, then multiply by 2. If each sampling point is 16 bits, then multiply by 2 again to obtain the memory size of the currently sampled data.

[0063] The above-mentioned preset pointer variable is a variable defined within the program. Its purpose is to save the address information of the first specified memory where the audio data is cached and at the same time indicate the currently played audio data, so that the target audio data can be obtained from the corresponding first specified memory according to the pointer variable during game operation. In fact, the memory address of the first specified memory where the audio data is cached can be written into the specified variable pData to record the starting position of the first specified memory. In the above method, the first specified memory is allocated in real time for the synthesized audio data, and the audio data is cached in the specified content, so that the new voice will not increase the game package size.

[0064] Since the data structure of the audio data synthesized through the voice interface is different from the data structure of the audio data inside Wwise, and the data structure of the audio data generated by the voice interface is defaulted to the 48000Hz / 24bit PCM format, while Wwise internally defaults to the 32Bit floating-point encoding format for processing. Therefore, in this embodiment, the data structure of the audio data is converted to a specified format through the data signal processing module; that is, the format conversion from PCM to FP32 is achieved. This method ensures the normal playback of the audio data through the format conversion of the audio data.

[0065] In addition, in order not to affect the size of the game package, when the audio data playback is completed or stopped, the cached audio data is deleted and the first specified memory is released. In fact, when the voice resource generated by the plug-in finishes playing, or the plug-in playback is stopped, the sound engine will call the Term function of the plug-in, and at this time, the first specified memory temporarily allocated is released; at this time, the plug-in no longer occupies the memory space of the audio media.

[0066] The following describes the steps of playing audio data during the operation of the target game according to the playback rules, including: obtaining the target audio data from the first specified memory according to the preset pointer variable; outputting the target audio data to the audio stream channel corresponding to the voice generation plug-in to play the target audio data during the operation of the target game according to the playback rules.

[0067] In order to play audio data during the operation of the game, it is necessary to output the audio data to be played to the audio stream channel corresponding to the voice generation plug-in, that is, the audio channel of Wwise, so that the target audio data can be played during the operation of the target game. In fact, the audio data can be obtained from the first specified memory according to the memory address indicated by the pointer variable pData and output to the audio stream channel of the game development application for playback. This method outputs the audio data to the processing channel of the Wwise sound engine, thereby allowing the user to apply functions such as 3D positioning attenuation and effect processing of Wwise to the voice sound effects.

[0068] The above editing interface also includes a voice audition control; to facilitate the user to audition the input voice parameters, in response to the third operation on the voice audition control, obtain the voice parameters from the cache space; generate the audio data corresponding to the voice parameters according to the voice parameters, cache the audio data in the second specified memory, and play the audio data in the editing interface.

[0069] The steps of generating the audio data are the same as the aforementioned process of generating audio data, both of which are generated in the current operating system through the voice interface, and will not be elaborated here. In this embodiment, the voice parameters can be obtained from the cache space of the editing module. In addition, in the editing module, the generated audio data can be saved in the second specified memory and then output to the audio playback channel of the editing module to audition the above audio data.

[0070] In addition, in this embodiment, the voice designer can modify the voice parameters in the editing interface, and the voice parameters will be updated when the game is running. The corresponding audio data will be regenerated according to the updated voice parameters. It realizes the real-time change of the text, country / region, gender and speech rate of the voice in the game for rapid function iteration. Since the present invention generates audio data in real time instead of generating fixed voice files offline, when the user modifies the voice parameters of the plugin instance, the next time the plugin is triggered to play, the audio data will be generated with the new parameters during the plugin initialization process and then played during the game operation. The game operation will call the Execute interface.

[0071] Corresponding to the above method embodiment, an embodiment of the present invention provides a voice processing device, which is arranged in the voice generation plugin; the voice generation plugin provides a game development interface, and the game development interface includes a processing control for the target game; as Figure 4 shown, the device includes:

[0072] A resource package generation module 41, configured to generate a voice resource package corresponding to the voice parameters according to the preset voice parameters and the pre-configured playback rules;

[0073] An acquisition module 42, configured to, in response to a first operation on the processing control for the target game, acquire the voice parameters and the playback rules from the voice resource package;

[0074] A data generation module 43, configured to generate audio data corresponding to the voice parameters according to the voice parameters;

[0075] A playback module 44, configured to play the audio data when the target game is running according to the playback rules.

[0076] A voice processing device provided by an embodiment of the present invention generates a voice resource package through a voice generation plug-in according to preset voice parameters and pre-configured playback rules; in response to a first operation on a processing control for a target game, obtains the voice parameters and playback rules from the voice resource package; generates audio data according to the voice parameters; and plays the audio data during the operation of the target game according to the playback rules. In this way, by integrating the voice generation plug-in into the audio engineering process of the game engine, it is possible to directly generate the audio data required for the game according to the preset voice parameters during the game operation and play it according to the playback rules, without pre-generating the audio data and manually importing it into the game engine, reducing the time cost, with a simple operation process and not easily making mistakes, and improving the game testing efficiency and development efficiency.

[0077] Further, the above voice generation plug-in also provides an editing interface; the editing interface includes a parameter input control and a serialization control; the above resource package generation module is further configured to: in response to a voice parameter input operation on the parameter input control, obtain the voice parameters; in response to a second operation on the serialization control, serialize the voice parameters and the pre-configured playback rules according to a preset encoding format to generate a voice resource package.

[0078] Further, the above obtaining module is further configured to: deserialize the voice resource package according to a preset decoding format to obtain the voice parameters and the playback rules.

[0079] Further, the above data generation module is further configured to: through a preset voice interface, obtain a target sound resource matching the voice parameters from a preset sound resource library; where the voice parameters at least include: a target text; and synthesize the audio data of the target text through the target sound resource.

[0080] Further, the above device further includes: a caching module, configured to cache the audio data into a first specified memory by sampling; where the size of the first specified memory is determined according to the sampling rate and the specified audio stream channel number during the initial sampling process; a writing module, configured to write the memory address of the first specified memory where the audio data is cached into a preset pointer variable.

[0081] Further, the above device further includes: a format conversion module, configured to convert the data structure of the audio data into a specified format through a data signal processing module.

[0082] Further, the above device further includes: a memory release module, configured to delete the cached audio data and release the first specified memory when the audio data playback is completed or stopped.

[0083] Further, the above-mentioned playback module is further configured to: obtain target audio data from the first specified memory according to a preset pointer variable; output the target audio data to the audio stream channel corresponding to the voice generation plugin, so as to play the target audio data during the operation of the target game according to the playback rule.

[0084] Further, the above-mentioned editing interface further includes a voice audition control; the above-mentioned device further includes an audition module, configured to: in response to a third operation on the voice audition control, obtain voice parameters from the cache space; generate audio data corresponding to the voice parameters according to the voice parameters, cache the audio data in the second specified memory, and play the audio data in the editing interface.

[0085] The voice processing device provided by the embodiment of the present invention has the same technical features as the voice processing method provided by the above-mentioned embodiment, so it can also solve the same technical problems and achieve the same technical effects.

[0086] This embodiment further provides an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned voice processing method. The electronic device can be a server or a terminal device.

[0087] See Figure 5 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the above-mentioned voice processing method.

[0088] Further, Figure 5 the electronic device shown further includes a bus 102 and a communication interface 103, and the processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.

[0089] Among them, the memory 101 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 5 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0090] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0091] This embodiment also provides a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the above voice processing method.

[0092] The computer program product of the voice processing method, device, electronic device and system provided by the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For the specific implementation, reference can be made to the method embodiments, which will not be elaborated herein.

[0093] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.

[0094] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0095] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0096] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0097] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A voice processing method, characterized in that, The method is applied to a voice generation plugin; The voice generation plugin provides a game development interface, and the game development interface includes processing controls for a target game; the method includes: Generating a voice resource package corresponding to the voice parameters according to preset voice parameters and pre-configured playback rules; In response to a first operation on the processing control for the target game, obtaining the voice parameters and the playback rules from the voice resource package; Generating audio data corresponding to the voice parameters according to the voice parameters; Playing the audio data during the operation of the target game according to the playback rules; The voice generation plugin further provides an editing interface; the editing interface includes parameter input controls and serialization controls; The step of generating a voice resource package corresponding to the voice parameters according to preset voice parameters and pre-configured playback rules includes: In response to a voice parameter input operation on the parameter input control, obtaining the voice parameters; In response to a second operation on the serialization control, serializing the voice parameters and the pre-configured playback rules according to a preset encoding format to generate the voice resource package.

2. The method according to claim 1, wherein The step of obtaining the voice parameters and the playback rules from the voice resource package includes: Deserializing the voice resource package according to a preset decoding format to obtain the voice parameters and the playback rules.

3. The method according to claim 1, characterized in that, The step of generating audio data corresponding to the voice parameters according to the voice parameters includes: Through a preset voice interface, obtaining a target sound resource matching the voice parameters from a preset sound resource library; wherein, the voice parameters at least include: target text; Synthesizing audio data of the target text through the target sound resource.

4. The method according to claim 3, wherein After the step of generating audio data corresponding to the voice parameters according to the voice parameters, the method further includes: Caching the audio data into a first specified memory by sampling; wherein, the size of the first specified memory is determined according to the sampling rate and the specified audio stream channel number during the initial sampling process; Writing the memory address of the first specified memory where the audio data is cached into a preset pointer variable.

5. The method according to claim 4, wherein The method further includes: Converting the data structure of the audio data into a specified format through a data signal processing module.

6. The method according to claim 4, characterized in that, The method further includes: When the audio data is played or stopped, deleting the cached audio data and releasing the first specified memory.

7. The method according to claim 1, characterized in that The step of playing the audio data during the operation of the target game according to the playback rules includes: Obtaining target audio data from the first specified memory according to a preset pointer variable; Outputting the target audio data to the audio stream channel corresponding to the voice generation plugin to play the target audio data during the operation of the target game according to the playback rules.

8. The method according to claim 1, characterized in that The editing interface further includes a voice audition control; after the step of obtaining the voice parameters in response to a voice parameter input operation on the parameter input control, the method further includes: In response to a third operation on the voice audition control, obtaining the voice parameters from the cache space; Generate audio data corresponding to the voice parameters according to the voice parameters, cache the audio data in a second specified memory, and play the audio data in the editing interface.

9. A voice processing device, characterized in that, The device is arranged in a voice generation plugin; The voice generation plugin provides a game development interface, and the game development interface includes processing controls for a target game; The device includes: A resource package generation module, configured to generate a voice resource package corresponding to the voice parameters according to preset voice parameters and preconfigured playback rules; An acquisition module, configured to, in response to a first operation on a processing control for a target game, acquire the voice parameters and the playback rules from the voice resource package; A data generation module, configured to generate audio data corresponding to the voice parameters according to the voice parameters; A playback module, configured to play the audio data during the operation of the target game according to the playback rules; The voice generation plugin further provides an editing interface; The editing interface includes a parameter input control and a serialization control; The resource package generation module is further configured to: in response to a voice parameter input operation on the parameter input control, acquire the voice parameters; in response to a second operation on the serialization control, serialize the voice parameters and the preconfigured playback rules according to a preset encoding format to generate the voice resource package.

10. An electronic device, characterized in that, It includes a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the voice processing method according to any one of claims 1-8.

11. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the voice processing method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Audio binding method and device, equipment and storage medium

    CN112717395A

  • Generating Expressive Speech Audio From Text Data

    US20210151029A1