A voice information processing method, a terminal and a storage medium
By categorizing, storing, and retrieving auxiliary voice data within the game, the problem of users being unable to communicate when it is inconvenient to speak is solved, thus improving the gaming experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2026-03-24
AI Technical Summary
In situations where speaking is inconvenient, users cannot communicate with other players in the game via voice, which reduces the game's fun and user experience.
By categorizing and storing auxiliary voice data according to the names of preset games, the system can identify whether the currently running application is a preset game and whether the voice assistance function is enabled, and then call up the pre-stored auxiliary voice data to communicate with other players.
In scenarios where it is inconvenient for users to speak, they can communicate with other players by calling up pre-stored auxiliary voice data, thereby enhancing the user's gaming experience.
Smart Images

Figure CN115920416B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of game voice processing, and particularly relates to a voice information processing method, a terminal and a storage medium. BACKGROUND
[0002] With the development of games and the improvement of living standards, users are willing to invest more money or time in their favorite games. Users can play games anytime and anywhere, but in inconvenient speaking scenarios such as high-speed trains, airplanes, boring parties, and physical discomfort, users cannot communicate with other players through voice, which reduces the interest of the game and affects the user's game experience. SUMMARY
[0003] Therefore, the purpose of the embodiments of the present application is to provide a voice information processing method, a terminal and a storage medium to solve the problem that users cannot communicate with other players through voice in inconvenient speaking scenarios.
[0004] The technical solutions adopted by the present application to solve the above technical problems are as follows:
[0005] According to an aspect of the embodiments of the present application, a voice information processing method is provided, which comprises:
[0006] storing auxiliary voice data according to the name of a preset game;
[0007] identifying whether the currently running application is a preset game and the voice assistance function is in an open state;
[0008] if yes, calling the pre-stored auxiliary voice data to communicate with other players.
[0009] Optionally, the storing of the auxiliary voice data according to the name of the preset game comprises:
[0010] obtaining original voice data of user communication with other players in a preset time period during the running of the preset game;
[0011] generating the auxiliary voice data based on the original voice data and storing the auxiliary voice data according to the game name.
[0012] Optionally, the generating of the auxiliary voice data based on the original voice data and the storing of the auxiliary voice data according to the game name comprises at least one of the following ways:
[0013] collecting a preset number of words with the highest frequency spoken by the user in the original voice data of the preset time period, and if the words are irrelevant to the current game scenario of the application and / or the current game operation of the user, saving the voice data corresponding to the words as first auxiliary voice data;
[0014] identify a game scene corresponding to the original voice data, and save the voice data of the user in the original voice data as second auxiliary voice data corresponding to the game scene;
[0015] identify a game operation of the user corresponding to the original voice data, and save the voice data of the user in the original voice data as third auxiliary voice data corresponding to the game operation.
[0016] Optionally, the generating the auxiliary voice data based on the original voice data and storing the auxiliary voice data according to game names further comprises:
[0017] collecting a word with the highest frequency of occurrence in the original voice data of the preset time length, and saving voice data corresponding to the word as a speech accent word; and / or,
[0018] collecting a voice tone of the user and / or other players in the original voice data of the preset time length, and saving the voice tone in association with a speaker corresponding thereto.
[0019] Optionally, the identifying a game scene corresponding to the original voice data and saving the voice data of the user in the original voice data as second auxiliary voice data corresponding to the game scene comprises:
[0020] identifying a game scene corresponding to the original voice data, combining the voice data of the user in the original voice data with the first auxiliary voice data, a speech accent word and / or a voice tone, and saving the combined voice data as second auxiliary voice data corresponding to the game scene;
[0021] The identifying a game operation of the user corresponding to the original voice data and saving the voice data of the user in the original voice data as third auxiliary voice data corresponding to the game operation comprises:
[0022] identifying a game operation of the user corresponding to the original voice data, combining the voice data of the user in the original voice data with the first auxiliary voice data, a speech accent word and / or a voice tone, and saving the combined voice data as third auxiliary voice data corresponding to the game operation.
[0023] Optionally, the calling the pre-stored auxiliary voice data to communicate with other players comprises at least one of the following manners:
[0024] outputting a first auxiliary voice corresponding to the pre-stored first auxiliary voice data according to a first preset rule;
[0025] identifying a current game scene of the application, and outputting second auxiliary voice corresponding to second auxiliary voice data corresponding to the current game scene;
[0026] identifying a current game operation of the user, and outputting third auxiliary voice corresponding to third auxiliary voice data corresponding to the current game operation.
[0027] Optionally, when the first auxiliary voice data comprises a plurality of pieces, the outputting of the first auxiliary voice corresponding to the pre-stored first auxiliary voice data according to the first preset rule comprises: selecting one from the pre-stored plurality of first auxiliary voice data according to a second preset rule, and outputting the first auxiliary voice corresponding to the selected first auxiliary voice data according to the first preset rule;
[0028] When the second auxiliary voice data comprises a plurality of pieces, the identifying of the current game scene of the application, and the outputting of the second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene comprises: identifying the current game scene of the application, selecting one from the pre-stored plurality of second auxiliary voice data according to a third preset rule, and outputting the second auxiliary voice corresponding to the selected second auxiliary voice data;
[0029] When the third auxiliary voice data comprises a plurality of pieces, the identifying of the current game operation of the user, and the outputting of the third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation comprises: identifying the current game operation of the user, selecting one from the pre-stored plurality of third auxiliary voice data according to a fourth preset rule, and outputting the third auxiliary voice corresponding to the selected third auxiliary voice data.
[0030] Optionally, after the identifying of the current game scene of the application, and the outputting of the second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene; or the identifying of the current game operation of the user, and the outputting of the third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation, the method further comprises:
[0031] If a reply of another player is received, fourth auxiliary voice corresponding to fourth auxiliary voice data corresponding to the reply is outputted.
[0032] According to another aspect of the embodiments of the present application, a terminal is provided, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor; the computer program, when executed by the processor, implements the steps of the voice information processing method.
[0033] According to still another aspect of the embodiments of the present application, a computer readable storage medium is provided, and the computer readable storage medium has stored thereon a voice information processing program which, when executed by a processor, implements the steps of the voice information processing method.
[0034] The voice information processing method, the terminal and the storage medium provided by the embodiments of the present application can store the auxiliary voice data according to the name of the preset game, identify whether the currently running application is the preset game and the voice auxiliary function is in the open state, and call the pre-stored auxiliary voice data to communicate with other players if yes. By using the technical scheme, the user can communicate with other players by calling the pre-stored auxiliary voice data in the scene where the user is inconvenient to speak, and the game experience of the user is improved. BRIEF DESCRIPTION OF DRAWINGS
[0035] The present application will be further described below with reference to the accompanying drawings and embodiments, in which:
[0036] Figure 1 is a schematic diagram of a hardware structure of a mobile terminal related to the present application;
[0037] Figure 2 is a flowchart of a voice information processing method provided by the embodiments of the present application;
[0038] Figure 3 is a schematic diagram of a structure of a terminal provided by the embodiments of the present application. DETAILED DESCRIPTION
[0039] It should be understood that the specific embodiments described herein merely exemplify the present application and should not be used to limit the present application.
[0040] In the following description, the suffixes used for elements, such as "module", "part", or "unit", are used only for convenience of explanation of the present application, and have no specific meaning by themselves. Thus, "module", "part", or "unit" can be mixedly used.
[0041] The terminal can be implemented in various forms. For example, the terminal described in the present application can include a mobile terminal such as a mobile phone, a tablet computer, a notebook computer, a palmtop computer, a Personal Digital Assistant (PDA), a Portable Media Player (PMP), a navigation device, a wearable device, a smart band, a pedometer, and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like.
[0042] In the following description, a mobile terminal will be exemplified, and it will be understood by those skilled in the art that the configuration according to the embodiments of the present application can be applied to a fixed type terminal, except for elements particularly used for mobile purposes.
[0043] Referring to Figure 1 is a hardware structure diagram of a mobile terminal for implementing various embodiments of the present application. The mobile terminal 100 can include an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (audio / video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will appreciate that the mobile terminal structure shown in FIG. 1 is not intended to limit the scope of the present application, and the mobile terminal can include more or less components, or some components can be combined, or different components can be arranged, than those shown in the drawing. Figure 1 The mobile terminal structure shown in FIG. 1 is not intended to limit the scope of the present application, and the mobile terminal can include more or less components, or some components can be combined, or different components can be arranged, than those shown in the drawing.
[0044] Hereinafter, the components of the mobile terminal will be described in detail. Figure 1 The components of the mobile terminal will be described in detail.
[0045] The RF unit 101 can be used for receiving and transmitting signals in the information receiving or communication process. Specifically, the RF unit 101 receives downlink information from a base station and transmits uplink data to the base station. Generally, the RF unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc. In addition, the RF unit 101 can communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System for Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), and TDD-LTE (Time Division Duplexing-Long Term Evolution), etc.
[0046] WiFi belongs to a short-range wireless transmission technology, and the mobile terminal can help users to send and receive e-mails, browse web pages, and access streaming media, etc. through the WiFi module 102, which provides users with wireless broadband Internet access. Although Figure 1 The WiFi module 102 is shown, but it is understood that it does not belong to the necessary components of the mobile terminal, and can be omitted as needed without changing the essence of the invention.
[0047] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in a call signal reception mode, a call mode, a recording mode, a voice recognition mode, a broadcast reception mode, etc. Moreover, the audio output unit 103 can also provide audio output related to a particular function performed by the mobile terminal 100 (e.g., a call signal reception sound, a message reception sound, etc.). The audio output unit 103 can include a speaker, a buzzer, etc.
[0048] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 can include a graphics processing unit (GPU) 1041 and a microphone 1042, the graphics processing unit 1041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frame can be displayed on the display unit 106. The image frame processed by the graphics processing unit 1041 can be stored in the memory 109 (or other storage medium) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) via the microphone 1042 in a telephone call mode, a recording mode, a voice recognition mode, etc. and can process such sound into audio data. The processed audio (voice) data can be converted into a format transmittable to a mobile communication base station in the case of a telephone call mode and outputted. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to cancel (or suppress) noise or interference generated in the process of receiving and transmitting audio signals.
[0049] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 1061 according to the brightness of ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is moved to the ear. As one of the motion sensors, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and when at rest, can detect the magnitude and direction of gravity, and can be used for applications such as identifying the posture of the mobile phone (such as switching between horizontal and vertical screens, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometers, tapping), and the like. As for the fingerprint sensor, pressure sensor, iris sensor, molecular sensor, gyroscope, barometer, hygrometer, thermometer, infrared sensor and other sensors that can be configured on the mobile phone, they will not be described here.
[0050] The display unit 106 is configured to display information input by a user or information provided to the user. The display unit 106 can include a display panel 1061, which can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0051] The user input unit 107 can be configured to receive input digital or character information, and to generate key signal inputs related to user settings and function controls of the mobile terminal. Specifically, the user input unit 107 can include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can collect a user's touch operation (such as the user's operation on or near the touch panel 1071 using a finger, a stylus, or any suitable object or accessory) and drive the corresponding connection device according to the pre-set program. The touch panel 1071 can include two parts, a touch detection device and a touch controller. The touch detection device detects the user's touch position and detects the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch coordinates, and sends it to the processor 110, and can also receive commands from the processor 110 and execute them. In addition, the touch panel 1071 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 can also include other input devices 1072. Specifically, the other input devices 1072 can include one or more of a physical keyboard, function keys (such as volume control buttons, on / off buttons, etc.), trackballs, mice, joysticks, and the like, without limitation.
[0052] Further, the touch panel 1071 can cover the display panel 1061, and when the touch panel 1071 detects a touch operation thereon or thereabout, transmit the same to the processor 110 to determine the type of the touch event, and then the processor 110 provides a corresponding visual output on the display panel 1061 according to the type of the touch event. Although in the above description, the touch panel 1071 and the display panel 1061 are implemented as two independent components to realize the input and output functions of the mobile terminal, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal, which is not limited herein. Figure 1 Further, the touch panel 1071 can cover the display panel 1061, and when the touch panel 1071 detects a touch operation thereon or thereabout, transmit the same to the processor 110 to determine the type of the touch event, and then the processor 110 provides a corresponding visual output on the display panel 1061 according to the type of the touch event. Although in the above description, the touch panel 1071 and the display panel 1061 are implemented as two independent components to realize the input and output functions of the mobile terminal, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal, which is not limited herein.
[0053] The interface unit 108 serves as an interface through which at least one external device can be connected with the mobile terminal 100. For example, the external device can include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device having an identification module, an audio input / output (I / O) port, a video I / O port, an earphone port, or the like. The interface unit 108 can be used as a path through which input data can be input to the mobile terminal 100 or through which the mobile terminal 100 can output data to an external device.
[0054] The memory 109 is operable to store various data used by at least one component of the mobile terminal 100. The memory 109 can include an internal memory and / or external memory. The internal memory can include at least one of a volatile memory (e.g., a random access memory (RAM), a dynamic RAM (DRAM), or a static RAM (SRAM)) and a non-volatile memory (e.g., a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically EPROM (EEPROM), a flash memory, a hard disk drive, or a solid state drive).
[0055] The processor 110 is a control center of the mobile terminal 100 and controls a plurality of constituent components connected thereto by using various programs stored in the memory 109 and / or data stored in the memory 109, and performs various functions of the mobile terminal 100 by processing the data. The processor 110 can include one or more processors. Preferably, the processor 110 can include an application processor and a modem processor. The application processor mainly processes an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication. It is understood that the modem processor can not be integrated into the processor 110.
[0056] The mobile terminal 100 can further include a power supply 111 (such as a battery) for powering the various components of the mobile terminal 100. In addition, the power supply 111 can be preferably logically connected to the processor 110 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption management, etc.
[0057] Although Figure 1 Not shown, the mobile terminal 100 can further include a Bluetooth module, etc., which will not be described here.
[0058] Based on the terminal hardware structure described above, various embodiments of the method of the present application are proposed.
[0059] Embodiment one
[0060] Figure 2 is a flowchart of a voice information processing method provided by an embodiment of the present application. The method of this embodiment is automatically run by the terminal, wherein each step can be performed in the order shown in the flowchart when running, or multiple steps can be performed simultaneously according to actual conditions, which is not limited here. The voice information processing method provided by the present application includes:
[0061] Step S1000, storing auxiliary voice data according to the name of a preset game;
[0062] Step S2000, identifying whether the currently running application is a preset game and the voice assistance function is in an open state;
[0063] Step S3000, if yes, calling the pre-stored auxiliary voice data to communicate with other players.
[0064] Through the above implementation, first, the auxiliary voice data is stored according to the name of a preset game; then, it is confirmed whether the gray mean value is greater than a preset threshold; then, it is identified whether the currently running application is a preset game and the voice assistance function is in an open state; finally, if yes, the pre-stored auxiliary voice data is called to communicate with other players.
[0065] In this embodiment, first of all, it needs to be pointed out that, considering the problem in the prior art that in a scenario where it is inconvenient to speak, the user cannot communicate with other players through voice. Therefore, in this embodiment, in order to solve the above technical problem, the auxiliary voice data is stored according to the name of a preset game; it is identified whether the currently running application is a preset game and the voice assistance function is in an open state; if yes, the pre-stored auxiliary voice data is called to communicate with other players. By using this voice information processing method, in a scenario where it is inconvenient for the user to speak, the user can communicate with other players by calling the pre-stored auxiliary voice data, thereby improving the user's game experience.
[0066] The above steps will be described in detail in conjunction with the specific embodiments.
[0067] In step S1000, the auxiliary voice data is stored according to the name of a preset game.
[0068] Specifically, the terminal stores the auxiliary voice data according to the name of a preset game, so as to subsequently call the auxiliary voice data to communicate with other players. The preset game refers to a game with a voice assistance function, and the preset game can include one or more games.
[0069] In an embodiment, the auxiliary voice data is stored according to the name of a preset game, and the method comprises the following steps:
[0070] In step S1100, original voice data of user communication with other players in a preset time period during the running of a preset game is obtained.
[0071] In step S1200, the auxiliary voice data is generated based on the original voice data, and the auxiliary voice data is stored according to the game name.
[0072] Specifically, the terminal can record original voice data of user voice communication with other players in a preset time period during the game when the user uses the preset game for the first time, then generate the auxiliary voice data based on the original voice data, and store the auxiliary voice data according to the game name. The terminal can also record original voice data of user voice communication with other players in a preset time period during the game when the user uses the preset game subsequently, then generate a plurality of auxiliary voice data based on the original voice data, and store the plurality of auxiliary voice data according to the game name and de-duplicate, so as to enrich the content of the auxiliary voice data.
[0073] In an embodiment, the auxiliary voice data is generated based on the original voice data, and the auxiliary voice data is stored according to the game name, which comprises at least one of the following methods:
[0074] In step S1210, a preset number of words with the highest frequency spoken by the user in the original voice data in the preset time period are collected, and if the words are irrelevant to the current game scene of the application and / or the current game operation of the user, the voice data corresponding to the words is saved as first auxiliary voice data.
[0075] In step S1220, the game scene corresponding to the original voice data is identified, and the voice data of the user in the original voice data is saved as second auxiliary voice data corresponding to the game scene.
[0076] In step S1230, the game operation of the user corresponding to the original voice data is identified, and the voice data of the user in the original voice data is saved as third auxiliary voice data corresponding to the game operation.
[0077] Specifically, the terminal generates the auxiliary voice data based on the original voice data, and stores the auxiliary voice data according to the game name classification in one or more of the following ways. For example, the terminal collects the preset number of words with the highest frequency spoken by the user in the original voice data of the preset time length. If the words are not associated with the current game scene of the application and / or the current game operation of the user, it indicates that the words are the user's verbal habits. At this time, the voice data corresponding to the words is saved as first auxiliary voice data. The terminal determines whether the speech of the user in the original voice data corresponds to the current game scene through image recognition. If yes, the voice data of the user in the original voice data is saved as second auxiliary voice data corresponding to the game scene. The terminal determines whether the speech of the user in the original voice data corresponds to the current game operation of the user through image recognition. If yes, the voice data of the user in the original voice data is saved as third auxiliary voice data corresponding to the game operation. In order to subsequently call the first auxiliary voice data, the second auxiliary voice data and the third auxiliary voice data according to different needs to communicate with other players. The specific value of the preset number can be set according to actual needs, and the specific value is not limited in this embodiment.
[0078] For example, the terminal collects the three words with the highest frequency spoken by the user in the original voice data of the preset time length. If the words are not associated with the current game scene of the application and / or the current game operation of the user, it indicates that the words are the user's verbal habits. At this time, the voice data corresponding to the words is saved as first auxiliary voice data.
[0079] For example, the terminal identifies and records a specific game scene, detects the position, the enemy and me situation, and records the different language of the user when a new term appears in the game, to form an auxiliary voice database file of the game scene, which is convenient to call. For example, in the game of King Glory, when the voice mentions fighting the dragon, the teammates gather to the dragon king, and the position of fighting the dragon and the voice conversation during the process are recorded through image recognition. Generally, when the enemy comes, according to the number of people, combined with the number of people on the current board, whether to require teammates to retreat, when our side is 5 people and the other side is 2 people, and the number of people on the board is more than the opponent, the user voice informs the teammates to continue to fight the dragon "not to panic, can fight", at this time, "not to panic, can fight" is saved as the second auxiliary voice data associated with the game scene. On the contrary, when fighting the dragon, the number of our side is less than the other side, and the board is lost to the other side, the user voice informs the teammates to retreat "everyone, don't fight, quickly retreat", at this time, "everyone, don't fight, quickly retreat" is saved as the second auxiliary voice data associated with the game scene. For example, in the game of Peace Elite, the terminal finds that the user drives a car and encounters enemy firing in a certain place on the map, and the user's visual angle shows that the firepower is strong, which may be that many people are in ambush and well equipped, and the user voice informs the teammates "the situation is unknown, quickly drive away", at this time, "the situation is unknown, quickly drive away" is saved as the second auxiliary voice data associated with the game scene. When the player encounters sporadic firepower, the user voice informs the teammates "not many people, fight directly", at this time, "not many people, fight directly" is saved as the second auxiliary voice data associated with the game scene. In this way, the auxiliary voice database file of a certain game scene can be formed.
[0080] In an embodiment, the generating the auxiliary voice data based on the original voice data and storing the auxiliary voice data according to the game name classification further comprises:
[0081] Step S1240, collecting the highest frequency word in the original voice data of the preset time length, and saving the voice data corresponding to the word as a pronunciation word; and / or,
[0082] Step S1250, collecting the tone of the user and / or other players in the original voice data of the preset time length, and saving the tone associated with the corresponding speaker.
[0083] Specifically, the terminal can collect the highest frequency single word in the original voice data of the preset time length, and save the voice data corresponding to the word as a pronunciation word, so as to enrich the auxiliary voice data by superimposing the pronunciation word. In addition, the terminal can also collect the tone of the user and / or other players in the original voice data of the preset time length, and save the tone associated with the corresponding speaker, so as to enrich the auxiliary voice data by superimposing the tone and increase the interest of the game.
[0084] For example, the "er" sound is used in "jiner" (today), "zhier" (here), and "zaier" (what's wrong). The single word with the highest frequency generally appears at the end of a coherent sentence and is saved as a dialect word. For example, because the game has multiple players, the terminal can record the voice tones of the user and other players and correspond the voice tones to the user and other players based on the frequency of the voice tones. Based on this, the terminal can randomly or according to other preset orders intelligently simulate the voice tones of other players to issue auxiliary voice, thereby increasing the interest of the game.
[0085] In an embodiment, the identifying the game scene corresponding to the original voice data and saving the voice data of the user in the original voice data as second auxiliary voice data corresponding to the game scene includes:
[0086] Step S1221, identifying the game scene corresponding to the original voice data, combining the voice data of the user in the original voice data with the first auxiliary voice data, dialect words, and / or voice tones, and saving the combined voice data as second auxiliary voice data corresponding to the game scene;
[0087] The identifying the game operation of the user corresponding to the original voice data and saving the voice data of the user in the original voice data as third auxiliary voice data corresponding to the game operation includes:
[0088] Step S1231, identifying the game operation of the user corresponding to the original voice data, combining the voice data of the user in the original voice data with the first auxiliary voice data, dialect words, and / or voice tones, and saving the combined voice data as third auxiliary voice data corresponding to the game operation.
[0089] Specifically, the terminal can enrich the content of the second auxiliary voice data by combining the voice data corresponding to the game scene with the first auxiliary voice data, dialect words, and / or voice tones, thereby avoiding repeating the same second auxiliary voice for the same game scene. The terminal can enrich the content of the third auxiliary voice data by combining the voice data corresponding to the game operation of the user with the first auxiliary voice data, dialect words, and / or voice tones, thereby avoiding repeating the same third auxiliary voice for the same game operation, thereby improving the intelligence and interest of the user in communicating with other players through auxiliary voice.
[0090] For example, a game scene corresponding to the original voice data is identified, the voice data of the user in the original voice data is combined with the first auxiliary voice data and the voice tone of the other player selected at random, and the combined voice data is saved in association as second auxiliary voice data corresponding to the game scene; or a game operation of the user corresponding to the original voice data is identified, the voice data of the user in the original voice data is combined with the accent word of the user, and the combined voice data is saved in association as third auxiliary voice data corresponding to the game operation.
[0091] In step S2000, it is identified whether the currently running application is a preset game and the voice assistance function is in an open state.
[0092] Specifically, the terminal can identify whether the currently running application is a preset game through system identification of the package name, and of course, the terminal can also identify whether the currently running application is a preset game through other manners. Alternatively, when multiple applications are running in the background, the application running in the focus window is identified as the currently running application.
[0093] In step S3000, if yes, the pre-stored auxiliary voice data is called to communicate with other players.
[0094] Specifically, the terminal calls the pre-stored auxiliary voice data to communicate with other players when it is identified that the currently running application is a preset game and the voice assistance function is in an open state. Thus, in a scenario where the user is inconvenient to speak, the user can communicate with other players through the pre-stored auxiliary voice data, thereby improving the user's game experience. In a scenario where the user is convenient to speak, the user can choose not to open the voice assistance function of the preset game and directly communicate with other players through speaking.
[0095] In an embodiment, the calling of the pre-stored auxiliary voice data to communicate with other players includes at least one of the following manners:
[0096] In step S3100, first auxiliary voice corresponding to the pre-stored first auxiliary voice data is output according to a first preset rule.
[0097] In step S3200, a game scene of the application is identified, and second auxiliary voice corresponding to second auxiliary voice data corresponding to the current game scene is output.
[0098] In step S3300, a game operation of the user is identified, and third auxiliary voice corresponding to third auxiliary voice data corresponding to the current game operation is output.
[0099] Specifically, the first preset rule can be to output the first auxiliary voice randomly or according to other preset time intervals. The manner in which the terminal communicates with other players by invoking the pre-stored auxiliary voice data can include one or more of the following manners. For example, when the terminal identifies that the currently running application is a preset game and the voice assistance function is in an open state, the terminal communicates with other players by outputting the first auxiliary voice corresponding to the pre-stored first auxiliary voice data according to the first preset rule in the game process, or outputting the first auxiliary voice corresponding to the pre-stored first auxiliary voice data according to a preset time interval; or, the terminal communicates with other players by outputting the second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene of the application by image recognition; or, the terminal communicates with other players by outputting the third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation of the user by image recognition.
[0100] In an embodiment, when the first auxiliary voice data includes multiple, the outputting the first auxiliary voice corresponding to the pre-stored first auxiliary voice data according to the first preset rule includes:
[0101] Step S3110, selecting one from the pre-stored multiple first auxiliary voice data according to a second preset rule and outputting the first auxiliary voice corresponding to the selected first auxiliary voice data according to the first preset rule;
[0102] When the second auxiliary voice data includes multiple, the identifying the current game scene of the application and outputting the second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene includes:
[0103] Step S3210, identifying the current game scene of the application, selecting one from the pre-stored multiple second auxiliary voice data according to a third preset rule and outputting the second auxiliary voice corresponding to the selected second auxiliary voice data;
[0104] When the third auxiliary voice data includes multiple, the identifying the current game operation of the user and outputting the third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation includes:
[0105] Step S3310, identifying the current game operation of the user, selecting one from the pre-stored multiple third auxiliary voice data according to a fourth preset rule and outputting the third auxiliary voice corresponding to the selected third auxiliary voice data.
[0106] Specifically, the second preset rule, the third preset rule and the fourth preset rule can be randomly selected or selected according to other preset sequences. When the first auxiliary voice data, the second auxiliary voice data and / or the third auxiliary voice data include multiple, by selecting different auxiliary voice data, it can avoid always repeating the first auxiliary voice, or always repeating the same second auxiliary voice for the same game scene, or always repeating the same third auxiliary voice for the same game operation, thereby improving the intelligence of auxiliary voice communication.
[0107] In an embodiment, after identifying the current game scene of the application, outputting the second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene; or, identifying the current game operation of the user, outputting the third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation, the method further comprises:
[0108] Step S3400, if the reply of other players is received, output the fourth auxiliary voice corresponding to the fourth auxiliary voice data corresponding to the reply.
[0109] Specifically, the terminal identifies the current game scene of the application when identifying that the currently running application is a preset game and the voice assistance function is in an open state, and outputs the second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene; or, identifies the current game operation of the user, and outputs the third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation, and then, if the reply of other players is received, outputs the fourth auxiliary voice corresponding to the fourth auxiliary voice data corresponding to the reply, to realize a certain voice automatic dialogue mechanism.
[0110] For example, in the game, when the user requires “team battle”, the other players reply “aggregation”, and the user continues to reply “first attack the one with less blood”, to further communicate and realize a certain voice automatic dialogue mechanism.
[0111] In the embodiment of the application, the auxiliary voice data is stored according to the name of the preset game; whether the currently running application is a preset game and the voice assistance function is in an open state is identified; if yes, the pre-stored auxiliary voice data is called to communicate with other players. By using the voice information processing method, in the scene where the user is inconvenient to speak, the pre-stored auxiliary voice data can be called to communicate with other players, thereby improving the game experience of the user.
[0112] Embodiment two
[0113] Figure 3is a structural schematic diagram of a terminal 300 provided by an embodiment of the present application. The terminal 300 comprises a memory 301, a processor 302, and a computer program (not shown in the figure) stored on the memory 301 and capable of running on the processor 302, which, when executed by the processor 302, implements the steps of the voice information processing method as described in Embodiment One above.
[0114] The terminal of the embodiment of the present application and the voice information processing method of Embodiment One above belong to the same concept, and the specific implementation process is described in the corresponding method embodiment, and the technical features in the method embodiment are all applicable to the present terminal embodiment, which will not be described here.
[0115] Embodiment Three
[0116] The embodiment of the present application also provides a computer readable storage medium, characterized in that the computer readable storage medium has a voice information processing program stored thereon, and the voice information processing program, when executed by a processor, implements the steps of the voice information processing method as described in Embodiment One above.
[0117] The computer readable storage medium of the embodiment of the present application and the method of Embodiment One above belong to the same concept, and the specific implementation process is described in the corresponding method embodiment, and the technical features in the method embodiment are all applicable to the present computer readable storage medium embodiment, which will not be described here.
[0118] The corresponding technical features in each of the above embodiments can be used with each other on the premise that it does not cause contradiction or unimplementable scheme.
[0119] It should be noted that in this document, the term "comprising" or "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or apparatus including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of another identical element in the process, method, article or apparatus including the element.
[0120] The above embodiment numbers of the present application are only for description, not representing the pros and cons of the embodiments.
[0121] Those skilled in the art can clearly understand the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, also can be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application essentially or say the part which contributes to the prior art can be embodied in the form of software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a plurality of instructions to make a terminal (may be a mobile phone, computer, server, air conditioner, or network equipment, etc.) execute the method described in various embodiments of the present application.
[0122] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative, not limiting, and those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.
Claims
1. A method for processing voice information, characterized in that, The method includes: Assistive voice data is stored according to the preset game names; It can identify whether the currently running application is a preset game and whether the voice assistance function is turned on. If so, then use the pre-stored auxiliary voice data to communicate with other players; The storage of auxiliary voice data categorized according to the preset game names includes: Acquire raw voice data of user interactions with other players during a preset game run for a preset duration; The auxiliary voice data is generated based on the original voice data, and the auxiliary voice data is stored according to the game name. The step of generating the auxiliary voice data based on the original voice data and storing the auxiliary voice data according to game name includes: Collect the most frequently spoken words from the original voice data of the preset duration. If the words are not related to the current game scene of the application and / or the current game operation of the user, the voice data corresponding to the words is saved as the first auxiliary voice data. Identify the game scene corresponding to the original voice data, and save the user's voice data in the original voice data as a second auxiliary voice data corresponding to the game scene; Identify the user's game operation corresponding to the original voice data, and save the user's voice data in the original voice data as a third auxiliary voice data corresponding to the game operation.
2. The voice information processing method according to claim 1, characterized in that, The process of generating the auxiliary voice data based on the original voice data and storing the auxiliary voice data according to game name further includes: Collect the most frequently occurring character from the raw speech data of the preset duration, and save the speech data corresponding to that character as an accent word; and / or, Collect the user's and / or other players' voice timbre from the original voice data of the preset duration, and associate and save the voice timbre with its corresponding speaker.
3. The voice information processing method according to claim 2, characterized in that, The step of identifying the game scene corresponding to the original voice data and associating and saving the user's voice data in the original voice data as second auxiliary voice data corresponding to the game scene includes: Identify the game scene corresponding to the original voice data, combine the user's voice data in the original voice data with the first auxiliary voice data, accent words and / or timbre, and save the combined voice data as the second auxiliary voice data corresponding to the game scene; The step of identifying the user's game operation corresponding to the original voice data and associating and saving the user's voice data in the original voice data as third auxiliary voice data corresponding to the game operation includes: Identify the user's game operation corresponding to the original voice data, combine the user's voice data in the original voice data with the first auxiliary voice data, accent words and / or timbre, and save the combined voice data as the third auxiliary voice data corresponding to the game operation.
4. The voice information processing method according to claim 1, characterized in that, The method of calling pre-stored auxiliary voice data to communicate with other players includes at least one of the following: The first auxiliary voice corresponding to the pre-stored first auxiliary voice data is output according to the first preset rule; Identify the current game scene of the application and output the second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene; Identify the user's current game operation and output the third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation.
5. The voice information processing method according to claim 4, characterized in that, When the first auxiliary voice data includes multiple data, the step of outputting the first auxiliary voice corresponding to the pre-stored first auxiliary voice data according to the first preset rule includes: selecting one data from the multiple pre-stored first auxiliary voice data according to the second preset rule and outputting the first auxiliary voice corresponding to the selected first auxiliary voice data according to the first preset rule. When the second auxiliary voice data includes multiple data, the step of identifying the current game scene of the application and outputting the second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene includes: identifying the current game scene of the application, selecting one data from the multiple pre-stored second auxiliary voice data according to a third preset rule, and outputting the second auxiliary voice corresponding to the selected second auxiliary voice data. When the third auxiliary voice data includes multiple data, the step of recognizing the user's current game operation and outputting the third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation includes: recognizing the user's current game operation, selecting one data from the multiple pre-stored third auxiliary voice data according to the fourth preset rule, and outputting the third auxiliary voice corresponding to the selected third auxiliary voice data.
6. The voice information processing method according to claim 4 or 5, characterized in that, After identifying the current game scene of the application and outputting second auxiliary voice corresponding to the second auxiliary voice data corresponding to the current game scene; or, identifying the user's current game operation and outputting third auxiliary voice corresponding to the third auxiliary voice data corresponding to the current game operation, the method further includes: If a response is received from another player, the fourth auxiliary voice corresponding to the fourth auxiliary voice data corresponding to the response is output.
7. A terminal, characterized in that, The terminal includes: a memory, a processor, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a voice information processing program, which, when executed by a processor, implements the steps of the voice information processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Game voice matching generation method and device and computer readable storage medium
CN115475387A