Information processing apparatus and method, and computer readable storage medium
By selecting sound elements related to scene features from the sounds, establishing a correspondence library and generating customized sounds, the problem of single game commentary in the existing technology is solved, and personalized and interactive audio production is achieved.
Patent Information
- Application Number
- CN201910560709.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-26
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2039-06-26
AI Technical Summary
In existing audio production technology, game commentary audio files are single and pre-recorded, resulting in a boring player experience that lacks personalization and interactivity.
Through information processing equipment and methods, sound elements related to scene features are selected from the sound, corresponding relationships are established and stored in a corresponding relationship library, and customized personalized sounds are generated based on the scene features.
It generates unique personalized game commentary audio, increases player participation and convenience of information interaction, and enriches the content and form of game commentary.
Smart Images

Figure CN112233647B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information processing, in particular to an information processing device and method capable of generating a self-defined personalized sound and a corresponding computer readable storage medium. BACKGROUND
[0002] In the existing audio production technology, only the voice content inherent to the system can be used to produce an audio file, which is easy to make the user feel boring. For example, in the game platform scene, only the pre-recorded commentary audio file in the game can be used to realize the game commentary, which is easy to make the player feel monotonous and boring. SUMMARY
[0003] A brief summary of the disclosure is given in the following to provide a basic understanding of some aspects of the disclosure. It should be understood that this summary is not an exhaustive overview of the disclosure. It is not intended to identify key or important parts of the disclosure, nor is it intended to limit the scope of the disclosure. Its purpose is only to give some concepts in a simplified form as a prelude to the more detailed description discussed later.
[0004] According to an aspect of the present application, an information processing device is provided, comprising: processing circuitry configured to: select a sound element related to a scene feature during sound emission from the sound; establish a correspondence relationship, the correspondence relationship comprising a first correspondence relationship between the scene feature and the sound element and between each sound element, and store the scene feature and the sound element and the correspondence relationship in a correspondence relationship library in association; and generate a sound to be reproduced based on the reproduced scene feature and the correspondence relationship library.
[0005] According to another aspect of the present application, an information processing method is provided, comprising: selecting a sound element related to a scene feature during sound emission from the sound; establishing a correspondence relationship, the correspondence relationship comprising a first correspondence relationship between the scene feature and the sound element and between each sound element, and storing the scene feature and the sound element and the correspondence relationship in a correspondence relationship library in association; and generating a sound to be reproduced based on the reproduced scene feature and the correspondence relationship library.
[0006] According to another aspect of the present application, there is provided an information processing apparatus including: a manipulation device for a user to manipulate the information processing apparatus; a processor; and a memory including instructions readable by the processor, and the instructions, when read by the processor, cause the information processing apparatus to perform the following processing: selecting, from a sound, a sound element related to a scene feature during a period when the sound is emitted; establishing a correspondence relation including a first correspondence relation between the scene feature and the sound element and between the respective sound elements, and storing the scene feature and the sound element and the correspondence relation in a correspondence relation library in association with each other; and generating a sound to be reproduced based on a re-scene feature and the correspondence relation library.
[0007] According to other aspects of the present disclosure, there are also provided computer program codes and computer program products for implementing the above information processing method and a computer readable storage medium having the computer program codes for implementing the above information processing method recorded thereon.
[0008] These and other advantages of the present disclosure will become more apparent from the following detailed description of the preferred embodiments thereof, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to further illustrate the above and other advantages and features of the present disclosure, a specific embodiment thereof will be further described in detail below with reference to the accompanying drawings. Such drawings are included herein and form a part of the specification. The same reference numbers in different drawings identify the same elements. It should be understood that these drawings are not necessarily to scale as the emphasis is generally placed upon illustrating the principles of the disclosure. In the drawings:
[0010] Figure 1 Fig. 1 shows a functional module block diagram of an information processing device according to an embodiment of the present disclosure;
[0011] Figure 2 Fig. 2 is a flowchart showing a flow example of an information processing method according to an embodiment of the present disclosure;
[0012] Figure 3 Fig. 3 is a block diagram showing an exemplary structure of a general personal computer in which the method and / or device according to an embodiment of the present disclosure can be implemented; and
[0013] Figure 4 Fig. 4 shows a structural block diagram of an information processing apparatus according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0014] Exemplary embodiments of the present disclosure will be described hereinafter with reference to the accompanying drawings. In the description, specific terminology and descriptions are set forth in order to provide a thorough understanding of the present disclosure. However, it should be understood that the present disclosure can be practiced without the specific details. Furthermore, it should be understood that the terminology and the descriptions are the matters of design and selection and are the matters of judgment of the present disclosure. Accordingly, it is not intended to confine the present disclosure to the well-known forming techniques and technologies. It is to be understood that the aspects of the present disclosure that are functional, are implemented to achieve the benefits described below, and that the conceptual aspects of the present disclosure apply to, and can be implemented in conjunction with, hardware and software.
[0015] It is also to be noted that, in the drawings, only the structures and / or processing steps that are closely related to the solutions according to the present disclosure are shown, and other details that are not closely related to the present disclosure are omitted, in order to avoid obscuring the present disclosure with unnecessary details.
[0016] Figure 1 A functional module block diagram of the information processing device 100 according to an embodiment of the present disclosure is shown in FIG. 1. Figure 1 As shown in FIG. 1, the information processing device 100 includes a sound element selection unit 101, a correspondence relationship establishment unit 103, and a generation unit 105.
[0017] The sound element selection unit 101, the correspondence relationship establishment unit 103, and the generation unit 105 can be implemented by one or more processing circuits, which can be implemented as a chip, a processor, for example. It should be understood that the sound element selection unit 101, the correspondence relationship establishment unit 103, and the generation unit 105 are logical modules divided according to the specific functions they implement, and are not intended to limit the specific implementation. Figure 1 The various functional units shown in FIG. 1 are logical modules divided according to the specific functions they implement, and are not intended to limit the specific implementation.
[0018] For the convenience of description, the information processing device 100 according to an embodiment of the present disclosure is described below in the application scenario of a game entertainment platform. However, the information processing device 100 according to an embodiment of the present disclosure can not only be applied to a game entertainment platform, but also be applied to a live broadcast of a sports game, a documentary, or other audio-visual products with a commentary, etc.
[0019] The sound element selection unit 101 can be configured to select a sound element from the sound, which is related to the scene characteristics during the sound is emitted.
[0020] As an example, the sound includes the speech of a speaker (e.g., the speech of a game player). As an example, the sound can also include at least one of clapping, cheering, cheering, music, etc.
[0021] As an example, the sound element selection unit 101 can perform sound processing on external sounds collected in real time during game system startup and during game play, so as to recognize the speech of the game player, for example, to recognize the comments of the game player during game play. The sound element selection unit 101 can also recognize sound information such as clapping, cheering, cheering, music, etc. through sound processing.
[0022] As an example, the scene features include at least one of game content, game character name (e.g. player name), game action, game or match nature, real-time game scene, game scene description. It can be seen that the scene features can include various characteristics or attributes related to the scene in which the sound is located.
[0023] As an example, the sound elements include information for describing the scene features and / or information for expressing emotions, which includes the tone of the sound and / or the rhythm of the sound.
[0024] As an example, the sound element selection unit 101 performs comparative analysis on the sound according to predetermined rules to select sound elements in the sound that are related to the scene features during the sound is emitted. The predetermined rules at least define the correspondence between the sound elements and the scene features, and the correspondence between the sound elements. For example, the predetermined rules can be designed by referring to at least part of the original voice commentary information of the game. For example, the predetermined rules can be designed by cutting and converting the sound into text, and then performing semantic analysis. For example, if it is judged that the name "Messi" is the name of a new player, the sound element "Messi" can be recorded, and the corresponding scene feature is marked as "player name", and more sound elements and scene features can be recorded according to the context, such as the voice "Messi's shot is too cool" will also record: the sound element "shot" corresponds to the scene feature "game action", and since it is judged that Messi is usually related to shooting, the correspondence between the sound elements "Messi" and "shot" is recorded (in this example, "Messi" is the subject, and "shot" is the action, so the correspondence between "Messi" and "shot" is subject + action). The above recorded information is used as the predetermined rules. As an example, the correspondence between the sound elements can also be specified in combination with a grammar model (e.g. "subject + predicate", "subject + predicate + object", "subject + attributive", "subject + adverbial", etc.).
[0025] As an example, the sound element selection unit 101 filters out the sound elements in the sound that are not related to the scene features during the sound is emitted.
[0026] As an example, the sound element selection unit 101 can be deployed locally on the game device, or can be implemented using cloud platform resources.
[0027] As can be seen from the above description, the sound element selection unit 101 can analyze and identify and finally screen out effective sound elements.
[0028] The correspondence relationship establishing unit 103 can be configured to establish a correspondence relationship including a first correspondence relationship between scene features and sound elements and between various sound elements, and store the scene features and the sound elements and the correspondence relationship in the correspondence relationship library in association.
[0029] The correspondence relationship establishing unit 103 labels the sound elements selected by the sound element selection unit 101 and the corresponding scene features, and establishes the correspondence relationship between the scene features and the sound elements and between various sound elements by referring to the above predetermined rules, for example, by machine learning (for example, neural network). Taking the voice "C Ronaldo scored a goal, really amazing" as an example, the correspondence relationship establishing unit 103 establishes the correspondence relationship between the sound element "C Ronaldo" and the scene feature "player name", and the correspondence relationship between the sound element "goal" and the scene feature "game action". Since it is judged by machine learning that C Ronaldo is usually related to goal, the correspondence relationship between the sound element "C Ronaldo" and the sound element "shot" is also established. If the above scene features and sound elements are not stored in the correspondence relationship library, the above scene features, sound elements and correspondence relationship are stored in the correspondence relationship library in association.
[0030] In addition, the above predetermined rules can also be stored in the correspondence relationship library. As the sound elements and scene features in the correspondence relationship library become more and more, the correspondence between the sound elements and the scene features and the correspondence between various sound elements will also become more and more complex. The predetermined rules are updated with the update of the correspondence between the sound elements and the scene features and the correspondence between various sound elements.
[0031] As an example, the correspondence relationship library can be continuously expanded and improved by machine learning (for example, neural network).
[0032] The correspondence relationship library can be stored in a local or remote platform (network space or cloud storage space).
[0033] The correspondence relationship can be stored in the form of a correspondence relationship matrix, a mapping diagram, etc.
[0034] The generating unit 105 can be configured to generate the sound to be reproduced based on the re-creation scene feature and the correspondence relationship library. Specifically, the generating unit 105 can generate the sound to be reproduced according to the correspondence relationship between the scene features and the sound elements in the correspondence relationship library, and between the sound elements, based on the re-creation scene feature and the correspondence relationship library. With the continuous updating of the scene features, sound elements, and correspondence relationships in the correspondence relationship library, the sound to be reproduced can be continuously updated, optimized, and enriched. As an example, under the scene trigger with the re-creation scene feature in the game, the generating unit 105 can generate a brand-new game commentary audio information file according to the player's voice stored in the correspondence relationship library, which will contain the comments of the game player during the game process, etc., so as to make the game commentary audio information more personalized and become a unique audio commentary information file for the game player. This personalized audio commentary information can be shared on the platform, thereby increasing the convenience of information interaction.
[0035] As an example, the generating unit 105 can save the generated sound to be reproduced in the form of a file (for example, an audio commentary information file) in a dedicated area in a local or remote platform (network space or cloud storage space). In addition, the file can be displayed in a customized way (for example, in Chinese, English, Japanese, etc.) in the UI of the game system for the game player to select and use.
[0036] As can be known from the above description, the information processing device 100 according to the embodiments of the present disclosure can generate a customized and personalized sound according to the correspondence relationship between the scene features and the sound elements in the correspondence relationship library, and between the sound elements, based on the re-creation scene feature, thereby solving the defect that only the sound content pre-recorded by the system can be used to make an audio file in the existing audio production technology. For the game entertainment platform, the existing game commentary is single and fixed, however, the information processing device 100 according to the embodiments of the present disclosure can generate a customized and personalized game commentary according to the player's voice stored in the correspondence relationship library.
[0037] Preferably, the information processing device 100 according to the embodiments of the present disclosure can further include a sound collecting unit configured to collect sound via a sound collecting device. The general game system platform at present does not have an external sound collecting device and function. The sound collecting unit according to the embodiments of the present disclosure configures a recording function on the peripheral device. The sound collecting device can be, for example, installed on a gamepad, a mouse, a camera device, a PS Move, a headset, a computer, or a display device such as a television, etc.
[0038] Preferably, the sound collecting unit can collect the sound of each speaker via a sound collecting device set up corresponding to each speaker respectively, and can distinguish the sound of different speakers collected according to the ID of the sound collecting device. Preferably, the ID of the sound collecting device can also be included in the correspondence library. For example, when multiple people participate in the game at the same time, the voices of multiple game players can be recorded simultaneously through the microphone of each game handle and / or the microphone of other game peripherals, and the voices of different players can be distinguished through the ID of the microphone. Preferably, the ID of the microphone can also be included in the correspondence library. For example, player A and friend B play football game at the same time, and the sound collecting unit collects the voices of player A and friend B via the microphones of player A and friend B at the same time, and distinguishes the voices of player A and friend B through the ID of the microphone.
[0039] Preferably, the sound collecting unit can collect the sound of each speaker via a sound collecting device set up corresponding to each speaker respectively, and can distinguish the sound of different speakers collected according to the ID of the sound collecting device. Preferably, the ID of the sound collecting device can also be included in the correspondence library. For example, when multiple people participate in the game at the same time, the voices of multiple game players can be recorded simultaneously through the microphone of each game handle and / or the microphone of other game peripherals, and the voices of different players can be distinguished through the ID of the microphone. Preferably, the ID of the microphone can also be included in the correspondence library. For example, player A and friend B play football game at the same time, and the sound collecting unit collects the voices of player A and friend B via the microphones of player A and friend B at the same time, and distinguishes the voices of player A and friend B through the ID of the microphone.
[0040] The above two sound collecting schemes (i.e., collecting the sound of each speaker via a sound collecting device corresponding to each speaker respectively and collecting the sound of each speaker via a centralized sound collecting device) can be configured to be used respectively, or can be configured to be used simultaneously. For example, the voices of a part of speakers are collected through the respective sound collecting devices, and the voices of a part of users are collected through the centralized sound collecting device. Alternatively, the respective sound collecting device and the centralized sound collecting device can be configured to be used simultaneously, and the sound collecting scheme to be used can be determined according to the actual situation.
[0041] Preferably, the sound collecting unit can collect the sound of each speaker via a sound collecting device, and can distinguish the sound of different speakers by performing sound ray analysis on the collected sound. As an example, during the game, the sound collecting unit can collect the voices of player A and friends B, C collectively via one microphone or can collect the voices of A, B, C respectively via the microphones of A, B, C, and perform sound ray analysis on the collected voices, thereby identifying the voices of player A and friends B, C. As an example, the system can record the real-time position information of the game player (e.g. the relative position of the game player to the game handle or host). Because the collected sound effect can be different due to the difference in the relative position of the same player to the handle, such position information helps to eliminate the sound difference of the sound due to the difference in position, thereby being able to more accurately identify the voices of different players.
[0042] Preferably, the correspondence further includes a second correspondence between the sound and the scene feature and the sound element. For example, the correspondence can further include a second correspondence between the whole sound and the scene feature and the sound element. Taking the above whole voice "Messi's shot is too cool" as an example, the correspondence can further include a second correspondence between the whole voice "Messi's shot is too cool" and the scene feature "player name" and "game action", and between the sound element "Messi" and "shot". Preferably, the correspondence establishing unit 103 can be configured to store the whole sound in association with the scene feature and the sound element and the second correspondence in the correspondence library, and the generating unit 105 can be configured to find the whole sound or the sound element related to the re-scene feature from the correspondence library according to the correspondence, and generate the sound to be reproduced by using the found whole sound or sound element. As an example, if the above whole sound is not stored in the correspondence library, the above whole sound is stored in the correspondence library in association with the scene feature and the sound element and the second correspondence. As an example, the generating unit 105 dynamically and intelligently finds the sound or the sound element from the correspondence library. For example, in the case where there are multiple whole sounds or combinations of multiple sound elements related to the re-scene feature in the correspondence library, one whole sound is dynamically and intelligently selected from the multiple whole sounds, or a combination of sound elements is dynamically and intelligently selected from the combinations of multiple sound elements, and the sound to be reproduced is generated by using the selected whole sound or combination of sound elements.
[0043] Generating the sound to be reproduced by using the found whole sound or sound element can enrich the content of the sound to be reproduced, thereby being able to generate personalized voices.
[0044] For the sake of brevity of description, the "whole sound" is sometimes also referred to as "sound" below.
[0045] As an example, the correspondence establishing unit 103 can periodically analyze the usage of the sound elements and the scene features stored in the correspondence library in generating the sound to be reproduced, and if there are sound elements and scene features in the correspondence library that have not been used to generate the sound to be reproduced for a long time, the correspondence establishing unit 103 can rejudge these sound elements and scene features as invalid information, and further delete them from the correspondence library, thereby saving storage space and improving processing efficiency. For example, the correspondence establishing unit 103 can also delete the entire sound that has not been used to generate the sound to be reproduced for a long time from the correspondence library.
[0046] Preferably, the correspondence further includes a third correspondence between the ID information of the speaker who uttered the sound and the scene features and the sound elements, and the correspondence establishing unit 103 can be configured to further store the ID information of the speaker who uttered the sound in the correspondence library in association with the scene features and the sound elements and the third correspondence. Through the third correspondence between the ID information of the speaker who uttered the sound and the scene features and the sound elements, the generating unit 105 can determine which speaker the found sound element belongs to, and thus the generating unit 105 can generate the sound to be reproduced including the entire sound or the sound element of the desired speaker, thereby improving the user experience.
[0047] Although the first correspondence, the second correspondence, and the third correspondence are described above, the present disclosure is not limited to the correspondence that can only include the first correspondence, the second correspondence, and the third correspondence. Other correspondences can also be generated when analyzing and processing the sound, the sound elements, and the scene features, and the correspondence establishing unit 103 can be configured to also store the other correspondences in the correspondence library.
[0048] Preferably, the generating unit 105 can be configured to, in a case where the re- scene features completely match the scene features in the correspondence library, find the entire sound related to the scene features that completely match the re-scene features, and generate the sound to be reproduced using the found entire sound. Generating the sound to be reproduced using the found entire sound can generate the sound that completely corresponds to the re-scene features.
[0049] As an example, in a case where the re-scene features completely match the scene features corresponding to the voice "Messi's shot is too awesome", the generating unit 105 can find the entire voice "Messi's shot is too awesome" from the correspondence library, and generate the sound to be reproduced using the found entire voice "Messi's shot is too awesome".
[0050] Preferably, the sound is the speaker's voice, and the generating unit 105 can be configured to add the found whole sound in the form of text or audio to the sound information base of the original speaker (e.g., the original commentator in the game), and generate the sound to be reproduced based on the sound information base, so as to render the sound to be reproduced in the pronunciation line of the original speaker, thereby increasing the flexibility of the commentary audio synthesis. In this way, the generating unit 105 adds the found whole sound to the sound information base of the original speaker, and constantly enriches and expands the sound information base of the original speaker. As an example, the generating unit 105 can combine the found whole sound with the voice in the sound information base of the original speaker, and synthesize the sound to be reproduced in the pronunciation line of the original speaker. For the game entertainment platform, under the real-time scene trigger of the game, the generating unit 105 can combine the found whole voice of the player with the original commentary, and synthesize the sound to be reproduced in the pronunciation line of the original commentator in the game, as part of the new game commentary audio.
[0051] Preferably, the generating unit 105 can be configured to generate the sound to be reproduced by using the found whole sound in the form of text or audio, so as to render the sound to be reproduced in the pronunciation line of the speaker who speaks the found whole sound, thereby maximizing the expression of the intonation and rhythm in the found sound. In this way, the generating unit 105 directly saves the found whole sound as a voice file. As an example, the generating unit 105 can directly generate the sound to be reproduced in the pronunciation line of the speaker who speaks the found whole voice. For the game entertainment platform, under the real-time scene trigger of the game, the generating unit 105 can synthesize the found whole voice of the player in the pronunciation line of the found player, as part of the new game commentary audio.
[0052] Preferably, the generating unit 105 can be configured to, in the case that neither the reenacted scene feature nor the scene feature in the correspondence library is completely matched, find the sound elements respectively matched with each part of the reenacted scene feature, and generate the sound to be reproduced by combining the found sound elements. As an example, the generating unit 105 divides the reenacted scene feature into different parts, finds the scene features respectively matched with each part of the reenacted scene feature from the correspondence library, and finds the sound elements "Messi", "shot", "too cool" respectively matched with the matched scene features, and finally generates the sound to be reproduced "Messi's shot is too cool" by combining the found sound elements. By combining the found sound elements respectively matched with the reenacted scene feature, the sound to be reproduced corresponding to the reenacted scene feature can be generated.
[0053] Preferably, the sound is the voice of a speaker, the generating unit 105 can be configured to add the found sound elements into a sound library of the original speaker in the form of text or audio, and generate the sound to be reproduced based on the sound library, so as to render the sound to be reproduced according to the pronunciation line of the original speaker, thereby increasing the flexibility of the commentary audio synthesis. In this way, the generating unit 105 adds the found sound elements into the sound library of the original speaker, and constantly enriches and expands the sound library of the original speaker. As an example, the generating unit 105 can combine the found sound elements with the voice in the sound library of the original speaker, and synthesize the sound to be reproduced according to the pronunciation line of the original speaker. For a game entertainment platform, under the triggering of a real-time scene of the game, the generating unit 105 can combine the found sound elements of the player with the original commentary, and synthesize the commentary according to the pronunciation line of the original game commentator, as part of the new game commentary audio.
[0054] Preferably, the generating unit 105 can be configured to generate the sound to be reproduced by using the found sound elements, so as to render the sound to be reproduced according to the pronunciation line of the speaker who speaks the found sound elements, thereby increasing the participation of the speaker. In this way, the generating unit 105 directly saves the combination of the found sound elements as a voice file. As an example, the generating unit 105 can directly generate the sound to be reproduced by using the combination of the found sound elements according to the pronunciation line of the speaker who speaks the found voice. For a game entertainment platform, under the triggering of a real-time scene of the game, the generating unit 105 can combine the found voice of the player according to the pronunciation line of the found player, and synthesize the commentary as part of the new game commentary audio.
[0055] As an example, in the case that each part of the reproduced scene feature does not match the scene feature in the corresponding relationship library, the sound elements related to the scene feature similar to the reproduced scene feature can be selected according to the similarity between the reproduced scene feature and the scene feature in the corresponding relationship library, and combined to generate the sound to be reproduced.
[0056] Preferably, the generating unit 105 can attach the found whole voice or voice element to the sound in the form of sound barrage to generate the sound to be reproduced. As an example, in the initial stage of collecting the player audio information, when the information collection is not rich enough, the found whole voice or voice element of the game player can be attached to the original commentary audio in the form of "sound barrage" to form a unique audio rendering mode. In this case, the original commentary audio is not changed and remains the same, and only in some specific scenes (such as a goal, a foul, a red or yellow card, etc.), the game will play the found whole voice or voice element of the game player in the form of "sound barrage", thereby enriching the form of commentary audio reproduction.
[0057] The sound to be reproduced generated according to the above processing can be played or reproduced immediately after being generated, or can be cached for subsequent playback or reproduction.
[0058] Preferably, the information processing device 100 according to the embodiments of the present disclosure further comprises a reproduction unit (not shown in the figure). The reproduction unit can be configured to reproduce the sound to be reproduced in a scene with a reproduction scene feature. As an example, the reproduction unit can analyze the game real-time scene in real time according to the original design logic of the game, and trigger the sound to be reproduced (for example, the game commentary audio information file generated according to the above processing) in a scene with a reproduction scene feature. With the increase and continuous enrichment of the voice information collected by the sound collecting unit, the design logic of the game can also be continuously optimized, so as to be able to reproduce the related sound to be reproduced (for example, the game commentary audio information file generated according to the above processing) generated more accurately and more richly according to the real-time scene of the game. Therefore, the reproduction unit can more user-friendlyly present the sound to be reproduced.
[0059] Preferably, the reproduction unit can render the sound to be reproduced according to the pronunciation tone of the original speaker. Specifically, the reproduction unit can analyze the scene of the game in real time according to the original design logic of the game, and in the case that the generating unit 105 adds the found voice element or whole voice to the voice information library of the original speaker as described above, the reproduction unit presents the sound to be reproduced according to the pronunciation tone of the original speaker, so as to continuously enrich and expand the original commentary content information, and make the commentary content have personalized features. In addition, the addition of new voice elements and scene features in the correspondence library will change or more finely enrich the trigger logic and design of the original commentary audio of the game.
[0060] Preferably, the reproduction unit can render the sound to be reproduced according to the pronunciation voice of the speaker who speaks the sound elements or the whole sound that are found. Specifically, when the generation unit 105 directly saves the combination of sound elements or the whole sound that are found as a voice file as described above, the reproduction unit reproduces the sound to be reproduced according to the pronunciation voice of the speaker who speaks the sound elements or the whole sound that are found. For example, when the sound elements or the whole voice of the game player are found, the reproduction unit can display the game commentary audio in the player's own voice according to the design logic of the original game and the real-time scene of the game, and the continuously increasing sound elements and scene features will increase the triggering of the game scene, so that the commentary audio information is more accurate and vividly displayed. In addition, the commentary audio that comes with the original game can also be rendered in the voice of the game player, especially at the beginning when the player's own voice information is not rich enough.
[0061] Preferably, the information processing device 100 according to an embodiment of the present disclosure further includes a communication unit (not shown). The communication unit can be configured to communicate with an external device or network platform via wireless or wired means to transmit information to the external device or network platform. For example, the communication unit can transmit the sound to be reproduced generated by the generation unit 105 to the network platform in the form of a file for sharing between users.
[0062] The above describes the information processing device 100 according to the embodiment of the present disclosure using the application scenario of a game platform, especially sports games (E-Sports) as an example. However, the information processing device 100 according to the embodiment of the present disclosure can also be applied to other similar application scenarios.
[0063] As an example, the information processing device 100 according to an embodiment of the present disclosure can also be applied to live sports broadcasts on television. In this application scenario, the information processing device 100 collects the broadcaster's voice information in real time, performs detailed analysis, and saves the relevant entire sound and / or sound elements, scene features, and their corresponding relationships. This allows the device to automatically generate commentary audio tailored to the real-time scene of the game and in the voice of the target commentator, thus achieving "automatic commentary" in future games.
[0064] As an example, the information processing device 100 according to an embodiment of the present disclosure can also implement "automatic narration" in documentaries or other audio and video products with narration. Specifically, by recording the commentary of a famous announcer, performing voice analysis, and saving the relevant entire sound and / or sound elements, scene features, and their corresponding relationships, it can automatically generate commentary for real-time scenes in other documentaries in the voice of the recorded announcer, thus achieving the generation and playback of "automatic narration."
[0065] Corresponding to the above-described embodiments of the information processing device, the present disclosure also provides embodiments of an information processing method. Figure 2 is a flowchart showing an example of a flow of an information processing method according to an embodiment of the present disclosure. As shown in Figure 2 The information processing method 200 according to an embodiment of the present disclosure includes a sound element selection step S201, a correspondence relationship establishment step S203, and a generation step S205.
[0066] In the sound element selection step S201, a sound element related to a scene feature during a sound emission period is selected from the sound.
[0067] As an example, the sound includes a speaker's voice (e.g., a game player's voice). As an example, the sound can also include at least one of applause, cheers, chants, music, etc.
[0068] As an example, in the sound element selection step S201, sound processing can be performed on external sound collected in real time during a game system startup and during a game process, so as to identify the game player's voice, for example, to identify the game player's comments during the game process. In the sound element selection step S201, sound information such as applause, cheers, chants, music, etc. can also be identified through sound processing.
[0069] As an example, the scene feature includes at least one of game content, game character name (e.g., player name), action in the game, game or match nature, real-time game scene, game scene description. It can be seen that the scene feature can include various characteristics or attributes related to the scene in which the sound is located.
[0070] As an example, the sound element includes information for describing the scene feature and / or information for expressing emotion, and the information for expressing emotion includes tone of voice and / or rhythm of voice.
[0071] As an example, in the sound element selection step S201, the sound is analyzed according to a predetermined rule to select a sound element in the sound related to a scene feature during a sound emission period. The predetermined rule at least defines the correspondence between the sound element and the scene feature, and the correspondence between each sound element.
[0072] Examples of the predetermined rule can be found in the foregoing description of the sound element selection unit 101 in the information processing device embodiments, which will not be repeated here.
[0073] As an example, in the sound element selection step S201, sound elements in the sound that are not related to the scene feature during the sound emission period are filtered out.
[0074] As can be seen from the above description, in the sound element selection step S201, the effective sound elements can be analyzed, identified and finally screened out.
[0075] In the correspondence relationship establishment step S203, the correspondence relationship including the first correspondence relationship between the scene features and the sound elements and between the sound elements can be established, and the scene features and the sound elements and the correspondence relationship are stored in the correspondence relationship library in association.
[0076] In the correspondence relationship establishment step S203, the sound elements selected in the sound element selection step S201 and the corresponding scene features are labeled, and the correspondence relationship between the scene features and the sound elements and between the sound elements is established by referring to the predetermined rules, for example, by machine learning (for example, neural network). If the above-mentioned scene features and sound elements are not stored in the correspondence relationship library, the above-mentioned scene features, sound elements and correspondence relationship are stored in the correspondence relationship library in association.
[0077] For examples of establishing the correspondence relationship, please refer to the description of the correspondence relationship establishment unit 103 in the aforementioned information processing device embodiment, which will not be repeated here.
[0078] In addition, the above-mentioned predetermined rules can also be stored in the correspondence relationship library. As the sound elements and scene features in the correspondence relationship library become more and more, the correspondence between the sound elements and scene features and the correspondence between the sound elements become more and more complex. The predetermined rules are updated with the update of the correspondence between the sound elements and scene features and the correspondence between the sound elements.
[0079] As an example, the correspondence relationship library can be continuously expanded and improved by machine learning (for example, neural network).
[0080] The correspondence relationship library can be stored in a local or remote platform (network space or cloud storage space).
[0081] The correspondence relationship can be stored in the form of correspondence relationship matrix, mapping diagram, etc.
[0082] In the generating step S205, the sound to be reproduced can be generated based on the re-creation scene feature and the correspondence relationship library. Specifically, in the generating step S205, the sound to be reproduced can be generated based on the re-creation scene feature and the correspondence relationship library according to the correspondence relationship between the scene features and the sound elements in the correspondence relationship library, and between each sound element. With the continuous updating of the scene features, sound elements and correspondence relationships in the correspondence relationship library, the sound to be reproduced can be continuously updated, optimized and enriched. As an example, in the generating step S205, a brand new game commentary audio information file can be generated according to the player's voice stored in the correspondence relationship library under the triggering of the scene with the re-creation scene feature in the game. The file will contain the comments of the game player during the game process, etc., so as to make the game commentary audio information more personalized and become a unique audio commentary information file for the game player. This personalized audio commentary information can be shared on the platform, thereby increasing the convenience of information interaction.
[0083] As an example, in the generating step S205, the generated sound to be reproduced can be saved in the form of a file (for example, an audio commentary information file) in a dedicated area in a local or remote platform (network space or cloud storage space). In addition, the file can be displayed in a customized way (for example, in Chinese, English, Japanese, etc.) in the UI of the game system for the game player to choose to use.
[0084] As can be seen from the above description, the information processing method 200 according to the embodiments of the present disclosure can generate a customized personalized sound based on the re-creation scene feature according to the correspondence relationship between the scene features and the sound elements in the correspondence relationship library, and between each sound element, thereby solving the defect that only the sound content pre-recorded by the system can be used to make an audio file in the existing audio production technology. For a game entertainment platform, the existing game commentary is single and fixed, however, the information processing method 200 according to the embodiments of the present disclosure can generate a customized personalized game commentary according to the player's voice stored in the correspondence relationship library.
[0085] Preferably, the information processing method 200 according to the embodiments of the present disclosure can further include a sound collecting step, in which sound is collected via a sound collecting device. The sound collecting device can be, for example, installed on a gamepad, a mouse, a camera device, a PS Move, a headset, a computer, or a display device such as a television, etc.
[0086] Preferably, in the sound collecting step, the sound of each speaker can be collected via a sound collecting device respectively arranged corresponding to each speaker, and the sound of different speakers collected can be distinguished according to the ID of the sound collecting device. Preferably, the ID of the sound collecting device can also be included in the correspondence relationship library.
[0087] Preferably, in the sound collecting step, the sound of each speaker can be collected via a sound collecting device, and the sound of different speakers collected can be distinguished according to the position information and / or sound ray information of the speakers. In addition, the above-mentioned position information can be saved for future use in other applications, such as 3D audio rendering, etc. Preferably, the above-mentioned position information can also be included in the correspondence library.
[0088] Preferably, in the sound collecting step, the sound of each speaker can be collected via a sound collecting device, and the sound of different speakers can be distinguished by sound ray analysis of the collected sound.
[0089] Preferably, the correspondence further includes a second correspondence between the whole sound and the scene feature and the sound element, in the correspondence establishing step S203, the whole sound and the scene feature and the sound element and the second correspondence are also stored in the correspondence library in association with the correspondence, and in the generating step S205, the whole sound or the sound element related to the re-creation scene feature is searched from the correspondence library according to the correspondence, and the whole sound or the sound element searched is used to generate the sound to be reproduced. As an example, the sound or the sound element is dynamically and intelligently searched from the correspondence library. For example, in the case where there are multiple whole sounds or combinations of multiple sound elements related to the re-creation scene feature in the correspondence library, one whole sound is dynamically and intelligently selected from the multiple whole sounds, or a combination of sound elements is dynamically and intelligently selected from the combinations of multiple sound elements, and the whole sound or the combination of sound elements selected is used to generate the sound to be reproduced.
[0090] As an example, in the correspondence establishing step S203, the use of the sound elements and the scene features stored in the correspondence library in generating the sound to be reproduced is periodically analyzed, if there are sound elements and scene features that have not been used to generate the sound to be reproduced for a long time in the correspondence library, these sound elements and scene features are re-identified as invalid information, and then they are deleted from the correspondence library. For example, in the correspondence establishing step S203, the whole sound that has not been used to generate the sound to be reproduced for a long time is also deleted from the correspondence library.
[0091] Preferably, the correspondence further includes a third correspondence between the ID information of the speaker who uttered the sound and the scene feature and the sound element, and in the correspondence establishing step S203, the ID information of the speaker is also stored in the correspondence database in association with the scene feature and the sound element and the third correspondence. By the third correspondence between the ID information of the speaker and the scene feature and the sound element, in the generating step S205, it can be determined which speaker the found sound element belongs to, and thus the sound to be reproduced can be generated including the whole sound or the sound element of the desired speaker.
[0092] Preferably, in the generating step S205, in the case where the re- scene feature completely matches the scene feature in the correspondence database, the whole sound related to the scene feature that completely matches the re- scene feature is searched for, and the sound to be reproduced is generated using the searched whole sound. Generating the sound to be reproduced using the searched whole sound can generate the sound that completely corresponds to the re- scene feature.
[0093] Preferably, in the generating step S205, the searched whole sound can be added to the sound information database of the original speaker in the form of text or audio, and the sound to be reproduced is generated based on the sound information database so as to render the sound to be reproduced in the pronunciation tone of the original speaker, thereby being able to increase the flexibility of the commentary audio synthesis. In this way, in the generating step S205, the searched whole sound is added to the sound information database of the original speaker, and the sound information database of the original speaker is constantly enriched and expanded.
[0094] Preferably, in the generating step S205, the searched whole sound in the form of text or audio can be used to generate the sound to be reproduced so as to render the sound to be reproduced in the pronunciation tone of the speaker who uttered the searched whole sound, thereby being able to maximize the expression of the intonation and rhythm in the searched sound. In this way, in the generating step S205, the searched whole sound is directly saved as a voice file.
[0095] Preferably, in the generating step S205, in the case where neither the re- scene feature completely matches the scene feature in the correspondence database, sound elements related to the scene features that respectively match each part of the re- scene feature are searched for, and the sound to be reproduced is generated by combining the searched sound elements. By combining the searched sound elements related to the re- scene feature, the sound to be reproduced that corresponds to the re- scene feature can be generated.
[0096] Preferably, in the generating step S205, the found sound elements can be added to the sound information library of the original speaker in the form of text or audio, and the sound to be reproduced is generated based on the sound information library, so as to render the sound to be reproduced according to the pronunciation line of the original speaker, thereby increasing the flexibility of commentary audio synthesis. In this way, in the generating step S205, the found sound elements are added to the sound information library of the original speaker, and the sound information library of the original speaker is constantly enriched and expanded.
[0097] Preferably, in the generating step S205, the sound to be reproduced can be generated by using the found sound elements, so as to render the sound to be reproduced according to the pronunciation line of the speaker who speaks the found sound elements, thereby increasing the participation of the speaker. In this way, in the generating step S205, the combination of the found sound elements is directly saved as a speech file.
[0098] As an example, in the case where each part of the reproduced scene feature does not match the scene feature in the corresponding relationship library, the sound elements related to the scene feature with high similarity to the reproduced scene feature can be selected according to the similarity between the reproduced scene feature and the scene feature in the corresponding relationship library, and combined into the sound to be reproduced.
[0099] Preferably, in the generating step S205, the found whole sound or sound elements can be attached to the sound in the form of sound barrage, so as to generate the sound to be reproduced. As an example, in the initial stage of collecting the player audio information, when the information collection is not rich enough, the found whole speech or sound elements of the game player can be attached to the original commentary audio in the form of "sound barrage", thereby forming a unique audio rendering mode. In this case, the original commentary audio is not changed and remains the same, and only in some specific scenes (such as a goal, a foul, a red or yellow card, etc.), the game will play the found whole speech or sound elements of the game player in the form of "sound barrage", thereby enriching the form of commentary audio reproduction.
[0100] The sound to be reproduced generated according to the above processing can be played or reproduced immediately after being generated, or can be cached for subsequent playback or reproduction when needed.
[0101] Preferably, the information processing method 200 according to the embodiments of the present disclosure further comprises a reproducing step, in which the sound to be reproduced can be reproduced in a scene with a re-creation scene characteristic. As an example, in the reproducing step, the real-time scene of the game can be analyzed in real time according to the original design logic of the game, and the sound to be reproduced (e.g., the game commentary audio information file generated according to the above processing) can be triggered in a scene with a re-creation scene characteristic. With the increase and continuous enrichment of the voice information collected in the sound collecting step, the design logic of the game can also be continuously optimized, so that the related sound to be reproduced (e.g., the game commentary audio information file generated according to the above processing) can be reproduced more accurately and more richly according to the real-time scene of the game. Therefore, in the reproducing step, the sound to be reproduced can be more user-friendly.
[0102] Preferably, in the reproducing step, the sound to be reproduced can be rendered according to the pronunciation tone of the original speaker. Specifically, in the reproducing step, the scene of the game can be analyzed in real time according to the original design logic of the game; in the case that the sound elements or the whole sound found are added to the sound information library of the original speaker as described above in the generating step S205, the sound to be reproduced can be presented according to the pronunciation tone of the original speaker in the reproducing step, so that the original commentary content information is continuously enriched and expanded, and the commentary content has personalized characteristics. In addition, the addition of new sound elements and scene characteristics in the correspondence library will change or more finely enrich the triggering logic and design of the original commentary audio of the game.
[0103] Preferably, in the reproducing step, the sound to be reproduced can be rendered according to the pronunciation tone of the speaker who speaks the sound elements or the whole sound found. Specifically, in the case that the combination of the sound elements or the whole sound found is directly saved as a voice file in the generating step S205 as described above, the sound to be reproduced can be reproduced according to the pronunciation tone of the speaker who speaks the sound elements or the whole sound found in the reproducing step. For example, in the case that the sound elements or the whole voice of the game player are found, the game commentary audio can be presented in the player's own tone in the reproducing step according to the original design logic of the game combined with the real-time scene of the game, and the continuously increasing sound elements and scene characteristics will increase the triggering of the game scene, so that the commentary audio information can be more accurately and vividly presented. In addition, the original commentary audio of the game can also be rendered in the player's tone, especially at the beginning when the player's own voice information is not rich enough.
[0104] Preferably, the information processing method 200 according to the embodiments of the present disclosure further comprises a communication step, in which the information can be communicated to an external device or a network platform through wireless or wired manner. For example, in the communication step, the generated sound to be reproduced can be transmitted to the network platform in the form of a file, so as to facilitate sharing among users.
[0105] The information processing method 200 according to the embodiments of the present disclosure is described above by taking the application scenario of game platform, especially the game of sports (E-Sports) as an example. As an example, the information processing method 200 according to the embodiments of the present disclosure can also be applied to the application scenario of live broadcast of sports games. As an example, the information processing method 200 according to the embodiments of the present disclosure can also realize "automatic realization of commentary" and playing in documentary or other audio-visual products with commentary.
[0106] It should be noted that although the functional configuration and operation of the information processing device and method according to the embodiments of the present disclosure are described above, this is only an example and not a limitation, and those skilled in the art can modify the above embodiments according to the principles of the present disclosure, for example, the functional modules and operations in each embodiment can be added, deleted or combined, etc., and such modifications all fall within the scope of the present disclosure.
[0107] In addition, it should also be noted that the method embodiments herein correspond to the above-mentioned device embodiments, and therefore the contents not described in detail in the method embodiments can be referred to the description of the corresponding parts in the device embodiments, which will not be described here again.
[0108] Moreover, the present disclosure also proposes a program product storing machine-readable instruction codes. The instruction codes are read and executed by a machine, and can execute the above-mentioned method according to the embodiments of the present disclosure.
[0109] Correspondingly, the storage medium for carrying the above-mentioned program product storing machine-readable instruction codes is also included in the disclosure of the present disclosure. The storage medium includes but is not limited to floppy disk, optical disk, magneto-optical disk, memory card, memory stick, etc.
[0110] In the case of implementing the present disclosure by software or firmware, the programs constituting the software are installed from the storage medium or the network to the computer with a special hardware structure (for example Figure 3 The general-purpose computer 300 shown), which can perform various functions when various programs are installed.
[0111] In Figure 3In the illustrated example, a central processing unit (CPU) 301 performs various processes in accordance with a program stored in a read only memory (ROM) 302 or a program loaded from a storage section 308 to a random access memory (RAM) 303. In the RAM 303, data required when the CPU 301 performs various processes and the like is also stored as necessary. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output interface 305 is also connected to the bus 304.
[0112] The following components are connected to the input / output interface 305: an input section 306 (including a keyboard, a mouse, and the like), an output section 307 (including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like), a storage section 308 (including a hard disk and the like), a communication section 309 (including a network interface card such as a LAN card, a modem, and the like). The communication section 309 performs communication processing via a network such as the Internet. As necessary, a drive 310 can also be connected to the input / output interface 305. A removable medium 311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is mounted on the drive 310 as necessary, so that a computer program read therefrom is installed in the storage section 308 as necessary.
[0113] In a case where the above-described series of processes are implemented by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 311.
[0114] It should be understood by those skilled in the art that such storage media are not limited to the Figure 3 The removable medium 311 shown is one in which a program is stored, and is distributed separately from an apparatus to provide the program to a user. Examples of the removable medium 311 include a magnetic disk (including a floppy disk (registered trademark)), an optical disk (including a compact disc read only memory (CD-ROM) and a digital versatile disk (DVD)), a magneto-optical disk (including a mini disk (MD) (registered trademark)), and a semiconductor memory. Alternatively, the storage medium can be the ROM 302, a hard disk included in the storage section 308, and the like, in which a program is stored, and is distributed to the user together with the apparatus including them.
[0115] It should also be noted that, in the apparatus and the method of the present disclosure, each unit or each step can be decomposed and / or recombined. These decompositions and / or recombination should be considered as equivalents of the present disclosure. Also, the steps of performing the above-described series of processes can naturally be executed in time series in the order of explanation, but do not necessarily have to be executed in time series. Some steps can be executed in parallel or independently of each other.
[0116] Furthermore, the present disclosure also provides a computer program capable of implementing the above-described embodiments according to the present disclosure (for example, as described above, for example, the series of processes described in the flowcharts of FIGS. 8 to 10, and the like) and a non-transitory computer readable medium (for example, a computer-readable medium such as a magnetic recording device, an optical and semiconductor memory, and the like) storing the computer program. Figure 1An information processing apparatus 400 having the functions of an information processing device (as shown). Figure 4 Schematically shows a structural block diagram of an information processing device 400 according to an embodiment of the present disclosure. Figure 4 As shown, the information processing apparatus 400 according to an embodiment of the present disclosure includes: an operating device 401, a processor 402, and a memory 403. The operating device 401 is used for the user to operate the information processing apparatus 400. The processor 402 may be a central processing unit (CPU) or a graphics processing unit (GPU), etc. The memory 403 includes instructions readable by the processor 402, and the instructions, when read by the processor 402, enable the information processing apparatus 400 to perform the following processing: select sound elements related to the scene features during the period when the sound is emitted from the sound; establish a correspondence, which includes a first correspondence between the scene features and the sound elements and between each sound element, and store the scene features and the sound elements and the correspondence in a correspondence library in association with each other; and generate the sound to be reproduced based on the reproduced scene features and the correspondence library. For examples of the information processing apparatus 400 performing the above-mentioned processing, please refer to the aforementioned information processing apparatus embodiments (for example, Figure 1 The description in FIG. 1 is shown in FIG. 2 and will not be repeated here.
[0117] It should be noted that although Figure 4 The manipulation device 401 is illustrated as being separate from and connected to the processor 402 and the memory 403 through lines, but the manipulation device 401 may be implemented to be integrated with the processor 402 and the memory 403 .
[0118] In a specific embodiment, the information processing device may be configured as a gaming device, wherein the operating device may be a wired game controller or a wireless game controller, and the gaming device may be operated by the game controller.
[0119] The gaming device according to this embodiment can generate customized and personalized game commentary based on the player's voice stored in the correspondence library, thereby solving the problem of single and solidified game commentary in the prior art.
[0120] During operation of the game device, as an example, the memory, the processor, and the manipulation device can be connected to a display device through an HDMI (High-Definition Multimedia Interface) line. The display device can be a television, a projector, a computer display, or the like. In addition, as an example, the game device according to the embodiment can also include a power supply, an input / output interface, an optical drive, or the like. Furthermore, as an example, the game device can be configured as a PlayStation (PS) game machine series. In this configuration scenario, the game device according to the embodiment of the present disclosure can also include a PlayStation Move (somatic controller) or a PlayStation camera, or the like, for acquiring relevant information of a user (e.g., a game player), such as including voice, video images, or the like, of the user.
[0121] Finally, it should be noted that the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, no element is introduced to the process, method, article, or apparatus by the mere inclusion of such element in a list of elements.
[0122] Although the embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, it is to be understood that the above-described implementations are merely for illustration of the present disclosure and do not constitute a limitation of the present disclosure. Various modifications and changes can be made to the above-described implementations by those skilled in the art without departing from the spirit and scope of the present disclosure. Therefore, the scope of the present disclosure is only defined by the appended claims and their equivalent meanings.
[0123] The present technology can also be configured as follows.
[0124] (1) An information processing device, comprising:
[0125] processing circuitry configured to:
[0126] selecting, from the sound, a sound element related to a scene feature during emission of the sound;
[0127] establishing a correspondence including a first correspondence between the scene feature and the sound element and between individual sound elements, and storing the scene feature and the sound element and the correspondence in a correspondence library in association with each other; and
[0128] generating a sound to be reproduced based on a re-scene feature and the correspondence library.
[0129] (2) The information processing device according to (1), in which
[0130] The correspondence further includes a second correspondence between the sound and the scene feature and the sound element; and
[0131] The processing circuitry is configured to:
[0132] store the sound in association with the scene feature and the sound element and the second correspondence in the correspondence library, and
[0133] find a sound or a sound element related to the reproduced scene feature from the correspondence library according to the correspondence, and generate the sound to be reproduced using the found sound or sound element.
[0134] (3) The information processing device according to (2), in which the processing circuitry is configured to:
[0135] in a case where the reproduced scene feature completely matches a scene feature in the correspondence library, find a sound related to the scene feature that completely matches the reproduced scene feature, and generate the sound to be reproduced using the found sound.
[0136] (4) The information processing device according to (3), in which
[0137] the sound is a speaker's voice, and
[0138] the processing circuitry is configured to:
[0139] add the found sound in the form of text or audio to a sound information library of an original speaker, and generate the sound to be reproduced based on the sound information library so that the sound to be reproduced is rendered in a pronunciation tone of the original speaker; or
[0140] generate the sound to be reproduced using the found sound in the form of text or audio so that the sound to be reproduced is rendered in a pronunciation tone of a speaker who spoke the found sound.
[0141] (5) The information processing device according to (2), in which the processing circuitry is configured to:
[0142] in a case where neither the reproduced scene feature completely matches a scene feature in the correspondence library, find sound elements related to scene features that respectively match parts of the reproduced scene feature, and generate the sound to be reproduced by combining the found sound elements.
[0143] (6), the information processing device according to (5), in which
[0144] the sound is a speaker's voice, and
[0145] the processing circuitry is configured to:
[0146] add the found sound element to a sound information library of the original speaker in the form of text or audio, and generate the sound to be reproduced based on the sound information library so as to render the sound to be reproduced in the sound line of pronunciation of the original speaker; or
[0147] generate the sound to be reproduced using the found sound element so as to render the sound to be reproduced in the sound line of pronunciation of the speaker who spoke the found sound element.
[0148] (7), the information processing device according to any one of (1) to (6), in which
[0149] the processing circuitry is configured to collect the sound of each speaker via a sound collection device set up corresponding to each speaker respectively, and distinguish the sound of different speakers collected according to the ID of the sound collection device.
[0150] (8), the information processing device according to any one of (1) to (7), in which
[0151] the processing circuitry is configured to collect the sound of each speaker centrally via one sound collection device, and distinguish the sound of different speakers collected according to the position information and / or sound line information of the speaker.
[0152] (9), the information processing device according to any one of (1) to (8), in which the processing circuitry is configured to collect the sound of each speaker via a sound collection device, and distinguish the sound of different speakers by sound line analysis on the collected sound.
[0153] (10), the information processing device according to any one of (1) to (9), in which
[0154] the correspondence further includes a third correspondence between the ID information of the speaker who uttered the sound and the scene feature and the sound element, and
[0155] the processing circuitry is configured to further store the ID information of the speaker in the correspondence library in association with the scene feature and the sound element and the third correspondence.
[0156] (11), the information processing device according to any one of (1) to (10), in which
[0157] The processing circuit is configured to define the correspondence between the sound elements and the scene features and between the sound elements by using predetermined rules, and to update the predetermined rules as the correspondence between the sound elements and the scene features and the correspondence between the sound elements is updated.
[0158] (12), The information processing device according to any one of (1) to (11), wherein the sound element includes information for describing the scene feature and / or information for expressing emotion, the information for expressing emotion including tone of voice and / or rhythm of voice.
[0159] (13), The information processing device according to any one of (1), (2), (3), and (5), wherein,
[0160] The sound includes at least one of applause, cheers, cheers, and music.
[0161] (14), The information processing device according to (2), wherein,
[0162] The processing circuit is configured to attach the found sound or sound element to the sound in the form of a sound barrage to generate the sound to be reproduced.
[0163] (15), The information processing device according to any one of (1) to (14), wherein,
[0164] The processing circuit is configured to delete sound elements and scene features that have not been used to generate the sound to be reproduced for a long time from the correspondence relationship library.
[0165] (16), The information processing device according to any one of (1) to (15), wherein,
[0166] The processing circuit is configured to reproduce the sound to be reproduced in a scene with the reproduced scene feature.
[0167] (17), The information processing device according to any one of (1) to (16), wherein,
[0168] The processing circuit is configured to communicate with an external device or a network platform by wireless or wired means to transmit information to the external device or the network platform.
[0169] (18), The information processing device according to (8), wherein,
[0170] The position information is used for 3D audio rendering.
[0171] (19) The information processing device according to (2), in which
[0172] dynamically and intelligently finding the sound or the sound element from the correspondence library.
[0173] (20) An information processing method, comprising:
[0174] selecting a sound element related to a scene feature during emission of the sound from the sound;
[0175] establishing a correspondence including a first correspondence between the scene feature and the sound element and between each sound element, and storing the scene feature and the sound element and the correspondence in association with each other in a correspondence library; and
[0176] generating a sound to be reproduced based on a reproduced scene feature and the correspondence library.
[0177] (21) A computer-readable storage medium having stored thereon computer-executable instructions that, when executed, perform a method, the method comprising:
[0178] selecting a sound element related to a scene feature during emission of the sound from the sound;
[0179] establishing a correspondence including a first correspondence between the scene feature and the sound element and between each sound element, and storing the scene feature and the sound element and the correspondence in association with each other in a correspondence library; and
[0180] generating a sound to be reproduced based on a reproduced scene feature and the correspondence library.
[0181] (22) An information processing apparatus, comprising:
[0182] a manipulation device for a user to manipulate the information processing apparatus;
[0183] a processor; and
[0184] a memory including instructions readable by the processor, and the instructions, when read by the processor, cause the information processing apparatus to perform the following processing:
[0185] selecting a sound element related to a scene feature during emission of the sound from the sound;
[0186] establishing a correspondence comprising a first correspondence between the scene feature and the sound elements and between individual sound elements, and storing the scene feature and the sound elements in association with the correspondence in a correspondence library; and
[0187] generating a sound to be reproduced on the basis of the reestablished scene feature and the correspondence library.
Claims
1. An information processing device for a gaming platform, comprising: The processing circuit is configured to: Selecting, from the sounds, sound elements related to scene features during the period when the sounds are emitted, wherein the sounds include a voice of a game player and the scene features include actions in the game; Establishing a correspondence relationship, the correspondence relationship including a first correspondence relationship between the scene feature and the sound element and between each sound element, and storing the scene feature, the sound element and the correspondence relationship in a correspondence relationship library in association with each other; and Based on the reproduction scene features and the corresponding relationship library, generating the sound to be reproduced, The processing circuit is configured to collect the voice of each game player via a sound collection device respectively provided corresponding to each game player, and distinguish the collected voices of different game players according to the ID of the sound collection device; and / or, the processing circuit is configured to centrally collect the voice of each game player via one sound collection device, and distinguish the collected voices of different game players according to the position information and voice line information of the game players.
2. The information processing device according to claim 1, wherein The correspondence relationship further includes a second correspondence relationship between the sound and the scene feature and the sound element; and The processing circuit is configured to: storing the sound in the correspondence library in association with the scene feature, the sound element, and the second correspondence; and According to the corresponding relationship, the sound or sound element related to the reproduction scene feature is searched from the corresponding relationship library, and the sound to be reproduced is generated using the searched sound or sound element.
3. The information processing device according to claim 2, wherein The processing circuit is configured to: In the case that the reproduction scene feature completely matches the scene feature in the correspondence library, a sound related to the scene feature completely matching the reproduction scene feature is searched, and the sound to be reproduced is generated using the searched sound.
4. The information processing device according to claim 3, wherein The processing circuit is configured to: Adding the found sound to a sound information library of the original game player in the form of text or audio, and generating the sound to be reproduced based on the sound information library, so as to render the sound to be reproduced according to the pronunciation voice of the original game player; or The sound to be reproduced is generated by using the found sound in the form of text or audio, so that the sound to be reproduced is rendered according to the pronunciation voice of the game player who speaks the found sound.
5. The information processing device according to claim 2, wherein The processing circuit is configured to: In the case that the reproduced scene feature does not completely match the scene feature in the correspondence library, sound elements related to the scene features that match each part of the reproduced scene feature are searched, and the sound to be reproduced is generated by combining the found sound elements. The information processing device according to claim 5 , wherein: The processing circuit is configured to: Adding the found sound elements to a sound information library of the original game player in the form of text or audio, and generating the sound to be reproduced based on the sound information library, so as to render the sound to be reproduced according to the pronunciation voice of the original game player; or The sound to be reproduced is generated by using the found sound elements, so that the sound to be reproduced is rendered according to the pronunciation voice of the game player who speaks the found sound elements.
7. An information processing method for a gaming platform, comprising: Selecting, from the sounds, sound elements related to scene features during the period when the sounds are emitted, wherein the sounds include a voice of a game player and the scene features include actions in the game; Establishing a correspondence relationship, the correspondence relationship including a first correspondence relationship between the scene feature and the sound element and between each sound element, and storing the scene feature, the sound element and the correspondence relationship in a correspondence relationship library in association with each other; and Based on the reproduction scene features and the corresponding relationship library, generating the sound to be reproduced, The voice of each game player is collected via a sound collection device respectively provided corresponding to each game player, and the collected voices of different game players are distinguished according to the ID of the sound collection device; and / or the voice of each game player is centrally collected via a sound collection device, and the collected voices of different game players are distinguished according to the position information and voice line information of the game player.
8. A computer-readable storage medium having computer-executable instructions stored thereon, which, when executed, perform a method for a gaming platform, the method comprising: Selecting, from the sounds, sound elements related to scene features during the period when the sounds are emitted, wherein the sounds include a voice of a game player and the scene features include actions in the game; Establishing a correspondence relationship, the correspondence relationship including a first correspondence relationship between the scene feature and the sound element and between each sound element, and storing the scene feature, the sound element and the correspondence relationship in a correspondence relationship library in association with each other; and Based on the reproduction scene features and the corresponding relationship library, generating the sound to be reproduced, The voice of each game player is collected via a sound collection device respectively provided corresponding to each game player, and the collected voices of different game players are distinguished according to the ID of the sound collection device; and / or the voice of each game player is centrally collected via a sound collection device, and the collected voices of different game players are distinguished according to the position information and voice line information of the game player.
9. An information processing device for a gaming platform, comprising: A manipulation device, used for a user to manipulate the information processing apparatus; processor; as well as a memory including instructions readable by the processor, and the instructions, when read by the processor, causing the information processing device to execute the following processing: Selecting, from the sounds, sound elements related to scene features during the period when the sounds are emitted, wherein the sounds include a voice of a game player and the scene features include actions in the game; Establishing a correspondence relationship, the correspondence relationship including a first correspondence relationship between the scene feature and the sound element and between each sound element, and storing the scene feature, the sound element and the correspondence relationship in a correspondence relationship library in association with each other; and Based on the reproduction scene features and the corresponding relationship library, generating the sound to be reproduced, The voice of each game player is collected via a sound collection device respectively provided corresponding to each game player, and the collected voices of different game players are distinguished according to the ID of the sound collection device; and / or the voice of each game player is centrally collected via a sound collection device, and the collected voices of different game players are distinguished according to the position information and voice line information of the game player.
Citation Information
Patent Citations
System and method for authoring and providing information relevant to a physical world
US20030155413A1
Audio Diarization System that Segments Audio Input
US20170359666A1