Immersive script antivirus effect generation method and system based on deep learning
By using deep learning technology to obtain environmental and player position information and dynamically adjust the status of the audio unit, the problem of unbalanced sound effects in the script-killing sound effect system is solved, the immersion and sound field optimization are enhanced, and a personalized sound experience is achieved.
Patent Information
- Application Number
- CN202511009219.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-22
AI Technical Summary
The existing script-killing sound effect system is unable to adjust the sound output according to the player's real-time physical position, resulting in problems of sound being too loud or too soft, and insufficient sound field optimization, affecting the immersion and experience comfort.
Through deep learning technology, the environmental space and player position information are obtained, and the sound field control unit is used to dynamically adjust the working status of the audio unit, including volume, surround sound effects, etc., to optimize the sound field according to the player's real-time position and feedback.
It realizes the dynamic adjustment of the sound output effects according to the player's position, improves the immersive feeling and experience comfort of immersive script-killing, and optimizes the uniformity and clarity of the sound field.
Smart Images

Figure CN120676308A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of sound effect control technology, and more specifically, to a method and system for generating immersive script-killing sound effects based on deep learning. Background Art
[0002] The most common method for implementing sound effects in script-based killing scenarios currently available on the market relies primarily on a pre-set script flow. At specific plot points or time points, staff or a central control system trigger and play fixed audio files. This is a linear playback model that relies heavily on the established structure of the script. In this sound effect model, the sound effect content (such as ambient sounds, clue prompts, ambient music, and character-specific sounds), as well as their playback order, volume, and duration, are set before the game begins. Throughout the game, the sound effects are strictly executed according to the script requirements and fixed steps. This lacks the ability to respond to the actual dynamic changes in the game scene in real time. This operating method is essentially a "playlist-style" execution, lacking intelligent adaptation to the audio environment.
[0003] The aforementioned script-killing sound effect mode suffers from the defect of being unable to automatically adjust the volume of the sound effects output by each speaker based on the player's real-time physical location. Regardless of the player's location in the room or how far away from the speaker unit, the preset volume is always played. This can easily lead to the player hearing the sound too loud when close to the speaker, or not being able to hear key sound effects when far away from the speaker, seriously affecting the game's immersion and comfort. At the same time, existing systems generally lack the optimization of the sound field and the ability to fine-tune the coordinated work of multiple speakers, making it difficult to create a uniform, clear, and accurately positioned three-dimensional sound field environment in the space. The sounds of different speakers may interfere with each other, overlap, or form blind spots, making it impossible for players to accurately perceive the direction and distance of the sound source. These defects seriously restrict the highly immersive experience and real sense of interaction pursued by script-killing games. Summary of the Invention
[0004] The purpose of this application is to provide an immersive script-killing sound effect generation method and system based on deep learning, which solves the technical problems of being unable to adjust the sound effects output by each speaker according to the player's real-time physical position and insufficient optimization of the sound field, and achieves the technical effect of adjusting the sound effects output by each speaker according to the player's real-time physical position and efficiently optimizing the sound field.
[0005] An embodiment of the present application provides an immersive script-killing sound effect generation method based on deep learning, the method comprising: obtaining environmental space information corresponding to a target environment, and obtaining position information of multiple audio units and position information of multiple players within the target environment; wherein the environmental space information comprises three-dimensional space information corresponding to the target environment, and one audio unit corresponds to one audio unit position information; through a sound field control unit, the audio unit control information of each audio unit is determined according to the environmental space information, the position information of multiple audio units and the position information of multiple players; and the working state of each audio unit is controlled according to the audio unit control information of each audio unit.
[0006] In one possible implementation, the method further includes: obtaining sound effect optimization information fed back by players corresponding to multiple player position information respectively; wherein the sound effect optimization information includes volume adjustment suggestions; determining, through the sound field control unit, the audio unit control information of each audio unit based on the environmental space information, multiple audio unit position information, multiple player position information and the sound effect optimization information corresponding to the multiple player position information respectively; and controlling the working status of each audio unit according to the audio unit control information of each audio unit.
[0007] In another possible implementation, the method also includes: determining, through a script parsing model, scene information and scene transition information corresponding to the immersive script-killing game based on the script information of the immersive script-killing game, where the scene transition information corresponds to the transition information between different scene information; determining, through a sound field adjustment unit, audio unit adjustment information of each audio unit based on environmental space information, multiple audio unit position information, multiple player position information, scene information and scene transition information; adjusting the working status of each audio unit based on the audio unit adjustment information of each audio unit; wherein the audio unit adjustment information includes volume adjustment information and surround sound adjustment information.
[0008] In another possible implementation, the method further includes: obtaining sound effect preference information fed back by players corresponding to multiple player position information respectively; wherein the sound effect preference information includes a volume mutation preference value, a horror preference value and a low-frequency tolerance value, the volume mutation preference value represents the player's preference for volume mutation, the horror preference value represents the player's preference for horror sound effects, and the low-frequency tolerance value represents the player's tolerance for low-frequency sound effects; through the sound field adjustment unit, determining a limit threshold of the audio unit adjustment information according to the sound effect preference information fed back by players corresponding to multiple player position information respectively; and adjusting the audio unit adjustment information of each audio unit according to the limit threshold of the audio unit adjustment information.
[0009] In another possible implementation, the method further includes: determining the number of player position information as player number information; determining, through the sound field transition adjustment unit, a sudden sound effect rising speed and a soundscape switching speed corresponding to the audio unit adjustment information based on the player number information; and adjusting the working state of each audio unit based on the audio unit adjustment information of each audio unit; wherein, the greater the player number information, the greater the sudden sound effect rising speed and the soundscape switching speed.
[0010] In another possible implementation, the method further includes: obtaining a communication volume preference value set by the player in the game, the communication volume preference value in the game representing the preferred size of the sound effect volume for the player's communication in the game; determining, through the sound field transition adjustment unit, the sudden sound effect rising speed, sound scene switching speed, and sound effect mutation amplitude corresponding to the audio unit adjustment information according to the player number information and the communication volume preference value in the game; adjusting the working state of each audio unit according to the audio unit adjustment information of each audio unit; wherein, the greater the communication volume preference value in the game, the greater the sudden sound effect rising speed, sound scene switching speed, and sound effect mutation amplitude.
[0011] In another possible implementation, the method also includes: obtaining player interaction information in an immersive script-killing game, where the player interaction information includes communication information between multiple players; determining the player's interaction activity factor based on the player quantity information and player interaction information through an activity monitoring model; wherein the interaction activity factor is between 0.8 and 1.5; determining the product of the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude corresponding to the audio unit adjustment information and the interaction activity factor, so as to adjust the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude corresponding to the audio unit adjustment information.
[0012] In another possible implementation, the method also includes: obtaining the volume value of the game NPC in the immersive script-killing game, and obtaining the voice clarity evaluation value of the game NPC by the players in the immersive script-killing game; through the game NPC voice adjustment unit, according to the volume value, interactive activity factor and voice clarity evaluation value of the game NPC, determine the volume adjustment value of the game NPC.
[0013] In another possible implementation, the method also includes: obtaining the interval duration of the same clue sound effect in the immersive script-killing game when it is repeatedly played, as the sound effect repetition interval duration; determining the product of the sound effect repetition interval duration and the interactive activity factor, as the sound effect repetition interval adjustment duration; adjusting the interval duration of the same clue sound effect in the immersive script-killing game when it is repeatedly played according to the sound effect repetition interval adjustment duration.
[0014] An embodiment of the present application also provides an immersive script-killing sound effect generation system based on deep learning, including a unit for executing any of the methods described above.
[0015] Compared with the prior art, the embodiments of the present application have the following beneficial effects: The embodiment of the present application provides an immersive script-killing sound effect generation method based on deep learning, the method comprising: obtaining environmental space information corresponding to the target environment, and obtaining multiple audio unit position information and multiple player position information within the target environment; wherein the environmental space information includes three-dimensional space information corresponding to the target environment, and one audio unit corresponds to one audio unit position information; through a sound field control unit, the audio unit control information of each audio unit is determined according to the environmental space information, the multiple audio unit position information and the multiple player position information; and the working state of each audio unit is controlled according to the audio unit control information of each audio unit. The embodiment of the present application can avoid the problem that the player's voice is too loud when close to the audio, or that the key sound effects cannot be heard clearly when far away from the audio, while improving the user's personalized experience in use and improving the working effect of the immersive script-killing sound effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 A schematic diagram of a flow chart of a first method for generating immersive script-killing sound effects based on deep learning provided in an embodiment of the present application; Figure 2 A schematic diagram of the working process of the first deep learning-based immersive script-killing sound effect generation method provided in an embodiment of the present application; Figure 3 A flowchart of a second method for generating immersive script-killing sound effects based on deep learning provided in an embodiment of the present application; Figure 4 A schematic flow chart of a third method for generating immersive script-killing sound effects based on deep learning provided in an embodiment of the present application; Figure 5 A schematic diagram of the logical structure of an immersive script-killing sound effect generation system based on deep learning provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0019] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0020] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0021] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0022] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0023] The aforementioned script-killing sound effect mode suffers from the defect of being unable to automatically adjust the volume of the sound effects output by each speaker based on the player's real-time physical location. Regardless of the player's location in the room or how far away from the speaker unit, the preset volume is always played. This can easily lead to the player hearing the sound too loud when close to the speaker, or not being able to hear key sound effects when far away from the speaker, seriously affecting the game's immersion and comfort. At the same time, existing systems generally lack the optimization of the sound field and the ability to fine-tune the coordinated work of multiple speakers, making it difficult to create a uniform, clear, and accurately positioned three-dimensional sound field environment in the space. The sounds of different speakers may interfere with each other, overlap, or form blind spots, making it impossible for players to accurately perceive the direction and distance of the sound source. These defects seriously restrict the highly immersive experience and real sense of interaction pursued by script-killing games.
[0024] Based on the above reasons, an embodiment of the present application provides an immersive script-killing sound effect generation method based on deep learning, the method comprising: obtaining environmental space information corresponding to the target environment, and obtaining multiple audio unit position information and multiple player position information within the target environment; wherein the environmental space information includes three-dimensional space information corresponding to the target environment, and one audio unit corresponds to one audio unit position information; through the sound field control unit, the audio unit control information of each audio unit is determined according to the environmental space information, multiple audio unit position information and multiple player position information; and the working state of each audio unit is controlled according to the audio unit control information of each audio unit. The embodiment of the present application can avoid the problem that the player hears too loud sounds when close to the audio, or cannot hear key sound effects clearly when far away from the audio, while improving the personalized experience of the user in use and improving the working effect of the immersive script-killing sound effect.
[0025] In some scenarios, an immersive script-killing sound effect generation method based on deep learning in an embodiment of the present application can be applied to short-duration (1-2 hours) immersive script-killing games combined with AR / VR technology, which can improve the player's gaming experience in existing immersive script-killing games.
[0026] The following is a detailed description of an immersive script-killing sound effect generation method based on deep learning provided in an embodiment of the present application with reference to specific examples.
[0027] Figure 1 The flowchart of the first method for generating immersive script-killing sound effects based on deep learning provided in the embodiment of the present application is as follows: Figure 1 As shown, the existing deep learning-based immersive script-killing sound effect generation method includes S110 to S120, and S110 to S120 are described in detail below.
[0028] S110: Acquire environmental space information corresponding to the target environment, and acquire position information of multiple audio units and position information of multiple players within the target environment. The environmental space information includes three-dimensional space information corresponding to the target environment, and each audio unit corresponds to one audio unit position information.
[0029] In this implementation, the environmental space information corresponding to the target environment can be obtained, and the position information of multiple audio units and multiple players in the target environment can be obtained. The environmental space information includes three-dimensional space information corresponding to the target environment. The three-dimensional space information can be used to accurately describe the spatial structure characteristics of the game venue.
[0030] In this implementation, each audio unit has corresponding audio unit position information. The audio unit position information can be entered into the system when the three-dimensional spatial information corresponding to the target environment is arranged. The three-dimensional spatial information corresponding to the target environment can be stored together with the audio unit position information.
[0031] For example, in an immersive script-killing game, a laser scanner can be used to obtain the three-dimensional point cloud data of the game room to track the position coordinates of each player in real time, and the three-dimensional spatial information and audio unit position information corresponding to the target environment can be obtained through the database. These data together constitute the basic information for sound effect control.
[0032] S120: Determine, by the sound field control unit, audio unit control information for each audio unit based on the ambient space information, the position information of the multiple audio units, and the position information of the multiple players, and control the operating state of each audio unit based on the audio unit control information.
[0033] Figure 2 The working process diagram of the first method for generating immersive script-killing sound effects based on deep learning provided in the embodiment of the present application is as follows: Figure 2 As shown, after obtaining the above information, the sound field control unit can determine the sound unit control information of each sound unit according to the environmental space information, the position information of the multiple sound units and the position information of the multiple players.
[0034] In this implementation, the sound field control unit can analyze the spatial relationship between the player and each audio unit and calculate the optimal sound propagation path.
[0035] Exemplarily, the sound field control unit can be a deep learning model based on a convolutional neural network. The sound field control unit is trained through a large number of acoustic propagation sample data in game scenarios, and can accurately predict the optimal sound parameters in different spatial positions. By inputting environmental space information, multiple audio unit position information and multiple player position information into the sound field control unit, the audio unit control information of each audio unit can be output.
[0036] Exemplarily, the sound field control unit can also be a Transformer-based deep learning model, which takes the environmental space information, audio unit position information, and player position information as input, learns the spatial dependencies (distance, relative orientation) between them through the self-attention mechanism, and obtains the audio unit control information of each audio unit.
[0037] After determining the audio unit control information, the working state of each audio unit can be controlled according to this information, and then by adjusting the sound parameters of each audio unit, the reasonable distribution of sound effects in the space can be ensured.
[0038] For example, when a player is close to a certain audio unit, the volume of the audio unit can be appropriately lowered to avoid the sound being too loud; when the player is far away from a certain audio unit, the volume of the unit can be appropriately increased and sound delay compensation can be added to ensure that the sound is clear and discernible. This dynamic adjustment can achieve balanced propagation of sound in space.
[0039] The beneficial effect of the above implementation method is that by controlling the working status of each audio unit through environmental space information, the player's position, and the position of the audio unit, it can avoid the problem of players hearing too loud sounds when close to the audio, or not being able to hear key sound effects clearly when far away from the audio, thereby improving the user's immersion and experience comfort in the game.
[0040] The beneficial effect of the above-mentioned implementation method is that, compared with the traditional fixed audio and sound effect control method, the working status of each audio unit is controlled by the player's position, and the sound effect is dynamically adjusted according to the player's position, which improves the user's personalized experience during use and improves the working effect of the immersive script-killing sound effect.
[0041] In some implementations, the above method further includes S130 to S140, which are described in detail below.
[0042] S130: Acquire sound effect optimization information fed back by players corresponding to the plurality of player position information, wherein the sound effect optimization information includes volume adjustment suggestions.
[0043] In this implementation, in an immersive script-killing game, sound effect optimization information feedback from players corresponding to multiple player position information can be continuously obtained. The sound effect optimization information may include parameters such as volume adjustment suggestions, reflecting the subjective feelings and optimization needs of players at different positions for the current sound effects.
[0044] For example, during a script-killing game, players in the corner of the room may complain that the volume is too low, while players close to the speakers may complain that the volume is too high. By collecting this sound optimization information, the auditory experience of players in each position can be adjusted in a targeted manner.
[0045] S140: Determine, by the sound field control unit, audio unit control information for each audio unit based on the ambient space information, the position information of the plurality of audio units, the position information of the plurality of players, and the sound effect optimization information corresponding to the plurality of player positions. Control the operating state of each audio unit based on the audio unit control information.
[0046] After obtaining the sound effect optimization information, the sound field control unit can determine the sound unit control information of each sound unit based on the environmental space information, the position information of multiple sound units, the position information of multiple players, and the sound effect optimization information corresponding to the multiple player position information, so that the sound field control unit can comprehensively consider the spatial layout and player feedback to calculate the optimal sound adjustment plan.
[0047] For example, the sound field control unit can analyze the reflection characteristics of sound waves in the ambient space information and calculate the sound field distribution in combination with the player's position information. The sound field control unit can be trained and optimized through historical sound effect data and player feedback data.
[0048] After determining the audio unit control information of each audio unit, the working status of each audio unit can be accurately controlled according to the control information. The audio unit control information can further include volume, equalizer settings, delay parameters, etc., to ensure that each player position can obtain the best auditory effect.
[0049] For example, for player positions where the feedback volume is too low, the output power of the audio unit in the corresponding direction can be increased; for player positions where the feedback volume is too high, the volume of the corresponding audio can be appropriately lowered. This precise control ensures that players at each player position can get a comfortable listening experience.
[0050] The beneficial effect of the above implementation method is to obtain sound effect optimization information corresponding to multiple player position information, so as to obtain the special optimization requirements of the player for sound effects corresponding to each player position information, and further optimize the user experience of the player corresponding to each player position information.
[0051] Figure 3 A flow chart of the second method for generating immersive script-killing sound effects based on deep learning provided in the embodiment of the present application is shown as follows: Figure 3 As shown, the above method further includes S210 to S220, and S210 to S220 are described in detail below.
[0052] S210. Determine the scene information and scene conversion information corresponding to the immersive script-killing game based on the script information of the immersive script-killing game through the script parsing model. The scene conversion information corresponds to conversion information between different scene information.
[0053] In this implementation, during the immersive script-killing game process, the script information can be intelligently analyzed through the script parsing model to further improve the game sound effect control effect. The script parsing model can identify the scene descriptions and scene transition nodes in the script, and extract key scene information and the transition relationship between scenes. The scene information includes the environmental characteristics and plot elements of the current scene, and the scene transition information includes the transition method, transition time point and transition duration when the scene switches.
[0054] For example, the script parsing model can be a deep learning model built using natural language processing technology. The script parsing model is trained with a large number of script samples and can accurately identify scene division signs and transition prompts in the script. The scene information output by the script parsing model can include features such as scene type and environmental atmosphere, and the scene transition information can include parameters such as transition method, transition time point, and transition duration.
[0055] S220: Determine, by the sound field adjustment unit, audio unit adjustment information for each audio unit based on the ambient space information, the position information of the multiple audio units, the position information of the multiple players, the scene information, and the scene transition information. Adjust the operating state of each audio unit based on the audio unit adjustment information. The audio unit adjustment information includes volume adjustment information and surround sound adjustment information.
[0056] In this implementation, the sound field adjustment unit can generate audio unit adjustment information for the position of each audio unit and the player position based on the scene information and scene conversion information output by the script parsing model, combined with the environmental space information, the position information of multiple audio units, and the position information of multiple players. The audio unit adjustment information includes parameters such as volume adjustment and surround sound configuration.
[0057] For example, in the castle investigation scene, the sound field adjustment unit can automatically lower the background music volume according to the scene information to enhance the surround feeling of the environmental sound effects. When the scene switches to an outdoor chase scene, it can smoothly transition the sound effects and gradually increase the volume of the main audio unit to create a tense atmosphere. Each audio unit will receive specific adjustment parameters based on the environmental space information, the position information of multiple audio units, and the position information of multiple players to achieve precise sound field control.
[0058] The beneficial effect of the above-mentioned implementation method is that it can parse scene information and scene transition information based on the script information of the immersive script-killing game, and automatically adjust the sound effect information in the scene based on the environmental space information, multiple audio unit position information, multiple player position information, scene information and scene transition information to optimize the sound effects in the script-killing game.
[0059] The beneficial effect of the above implementation is that when the working state of each audio unit is controlled according to the audio unit control information of each audio unit, the sound effect can be further adjusted according to the audio unit adjustment information, thereby improving the personalization of the audio during operation.
[0060] In some implementations, the above method further includes S230 to S240, and S230 to S240 are described in detail below.
[0061] S230: Acquire sound effect preference information reported by players corresponding to the multiple player position information. The sound effect preference information includes a volume mutation preference value, a horror preference value, and a low-frequency tolerance value. The volume mutation preference value indicates the player's preference for volume mutations, the horror preference value indicates the player's preference for horror sound effects, and the low-frequency tolerance value indicates the player's tolerance for low-frequency sound effects.
[0062] In this implementation, during the immersive script-killing game, the sound effect preference information corresponding to each player's position can be obtained. The volume mutation preference value can reflect the player's acceptance of sudden volume changes, the horror preference value can measure the player's preference for horror sound effects, and the low-frequency tolerance value can evaluate the player's physiological tolerance to low-frequency sound effects. The volume mutation preference value, horror preference value and low-frequency tolerance value can be used to further optimize the game sound effects of the immersive script-killing game.
[0063] For example, these preference information can be obtained through a questionnaire survey before the game starts or a real-time interactive interface. This preference information helps to understand each player's personalized needs for sound effects and provides a data basis for subsequent sound effect adjustments.
[0064] S240: Determine, by the sound field adjustment unit, a limit threshold for the audio unit adjustment information based on the sound effect preference information fed back by the players corresponding to the respective player position information, and adjust the audio unit adjustment information of each audio unit based on the limit threshold for the audio unit adjustment information.
[0065] After obtaining the sound effect preference information, the sound field adjustment unit may further determine a limit threshold of the audio unit adjustment information according to the sound effect preference information of each player.
[0066] Exemplarily, the sound field adjustment unit can be a decision-making system based on a rule engine. The sound field adjustment unit can calculate the optimal sound effect parameter limit suitable for the current player group by analyzing the combined relationship between each player's volume mutation preference value, horror preference value and low-frequency tolerance value. The sound field adjustment unit can comprehensively consider the preference information of all players to avoid extreme sound effect settings.
[0067] After determining the limit thresholds, the volume, frequency and other parameters of each audio unit can be adjusted according to these thresholds. This adjustment can be performed in real time to ensure that every player can get a sound experience that suits their preferences.
[0068] For example, in a horror-themed script-killing game, for player positions with low horror preference values, the intensity of the horror sound effects in that area can be reduced. At the same time, for player positions with low low-frequency tolerance values, the output power of the low-frequency sound effects can be reduced.
[0069] The beneficial effect of the above-mentioned implementation method is that the limit threshold of the audio unit adjustment information is determined through the sound effect preference information feedback by players corresponding to multiple player position information, so as to ensure that the sound effects in the script-killing game meet the user's usage preferences and improve the user's usage experience.
[0070] The beneficial effect of the above implementation is that the audio unit adjustment information of each audio unit is adjusted according to the limit threshold of the audio unit adjustment information, which can optimize the working effect of the audio based on the audio unit adjustment information and improve the user experience.
[0071] Figure 4 A flow chart of a third method for generating immersive script-killing sound effects based on deep learning provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the above method further includes S310 to S320, and S310 to S320 are described in detail below.
[0072] S310: Determine the number of player position information as player quantity information.
[0073] In this implementation, the number of player position information can be determined as player quantity information. The player quantity information reflects the total number of players currently participating in the game. The player quantity information can be used to preliminarily assess the noise level of the on-site environment and provide a basis for subsequent sound adjustments.
[0074] S320: The sound field transition adjustment unit determines, based on the player count information, a burst sound effect ramp-up speed and a soundscape switching speed corresponding to the audio unit adjustment information. The operating state of each audio unit is adjusted based on the audio unit adjustment information. The greater the player count information, the greater the burst sound effect ramp-up speed and the soundscape switching speed.
[0075] After determining the number of players, the sound field transition adjustment unit can be used to determine the sudden sound effect rising speed and sound scene switching speed corresponding to the audio unit adjustment information based on the number of players. The sudden sound effect rising speed indicates the response time of the audio unit from a silent state to the maximum sound effect intensity, and the sound scene switching speed indicates the transition time between different scene sound effects.
[0076] It should be noted that when determining the speed at which sudden sound effects rise and soundscape switching occur, the greater the number of players, the greater the speed at which sudden sound effects rise and soundscape switching occur. This adjustment method can adapt to changes in ambient noise under different numbers of players, ensuring that sound effects can effectively penetrate ambient noise.
[0077] Exemplarily, the sound field transition adjustment unit may be a deep learning model based on a neural network, and the sound field transition adjustment unit may be trained through historical player number information and corresponding optimal sound effect parameters.
[0078] In this implementation, after obtaining the sudden sound effect rising speed and the sound scene switching speed, the working status of each audio unit can be adjusted according to the audio unit adjustment information of each audio unit. By adjusting the working status of each audio unit in real time, precise sound field control can be achieved.
[0079] For example, in a script-killing game with 10 players, due to the loud conversations of the players, the system will automatically increase the rising speed of sudden sound effects and the switching speed of soundscapes, so that key sound effects can stand out quickly; while in a 3-player game, the system will reduce these parameters to maintain a smooth transition of sound effects.
[0080] The beneficial effect of the above implementation method is that it can adjust the switching speed of sound effects according to the number of players. When the number of players is large and the average ambient noise increases, by increasing the rising speed of sudden sound effects and the soundscape switching speed, a faster transient response is used to break through the critical point of ambient noise, avoiding the ambient sound effect time being too long and increasing the noisiness of the on-site sound effects, which can improve the user experience in scenarios with a large number of players.
[0081] The beneficial effect of the above implementation method is that when the number of players is small, by reducing the rising speed of sudden sound effects and the switching speed of soundscapes, the immersive atmosphere can be maintained, and the sound effect transition can be smoothed to maintain a delicate feeling, thereby improving the player's experience.
[0082] In some implementations, the above method further includes S330 to S340, and S330 to S340 are described in detail below.
[0083] S330: Obtain a communication volume preference value set by the player in the game, where the communication volume preference value represents the preferred volume of the sound effect for the player's communication in the game.
[0084] In this implementation, the game system interface can further obtain the communication volume preference value pre-set by the player in the game. The communication volume preference value can reflect the player's personalized demand for the volume of the communication sound effect during the game.
[0085] For example, the communication volume preference value in the game can be set to a continuous numerical range of 1-10. The smaller the communication volume preference value in the game, the smaller the impact of the game sound effects on communication between players. The larger the communication volume preference value in the game, the greater the impact of the game sound effects on communication between players.
[0086] S340. The sound field transition adjustment unit determines, based on the number of players and the in-game communication volume preference value, a burst sound effect rise speed, a soundscape switching speed, and a sound effect mutation amplitude corresponding to the audio unit adjustment information. The operating state of each audio unit is adjusted based on the audio unit adjustment information. The higher the in-game communication volume preference value, the greater the burst sound effect rise speed, soundscape switching speed, and sound effect mutation amplitude.
[0087] After obtaining the communication volume preference value in the game, the sound field transition adjustment unit can be used to determine the sudden sound effect rising speed, sound scene switching speed and sound effect mutation amplitude corresponding to the audio unit adjustment information according to the player number information and the communication volume preference value in the game, so as to further adjust the audio parameters according to the number of players and the sound effect volume requirements in the game communication.
[0088] Exemplarily, the sound field transition adjustment unit can adopt a sound effect control model based on deep learning. The sound field transition adjustment unit can be trained by analyzing the correspondence between the number of players, communication volume preference values and optimal sound effect parameters in historical game scenes. The sound field transition adjustment unit can intelligently calculate the audio unit adjustment information suitable for the current game atmosphere based on the current number of players information and the communication volume preference value in the game.
[0089] It's important to note that audio unit adjustment information includes three key parameters: burst sound effect ramp-up speed, soundscape switching speed, and sound effect mutation amplitude. When the preferred communication volume value is high in-game, the audio control system can increase the values of these three parameters to make the sound effect changes more rapid and noticeable. For example, when the player sets a high communication volume preference, the system will speed up the transition from background music to dialogue scenes. At this time, the high burst sound effect ramp-up speed, soundscape switching speed, and sound effect mutation amplitude may significantly affect the game sound effects and communication between players.
[0090] After obtaining the audio unit adjustment information, the working status of each audio unit can be adjusted in real time based on the calculated audio unit adjustment information, and then the sound intensity, frequency response and spatial positioning effect of each audio unit can be dynamically controlled to ensure that the sound effect changes meet the communication needs of players in the current game.
[0091] For example, when a startling sound effect suddenly appears during the reasoning phase, the amplitude of the sound effect's mutation will be controlled according to the parameter settings to ensure that the sound effect changes meet the communication needs of players in the current game.
[0092] The beneficial effect of the above implementation method is that by obtaining the in-game communication volume preference value that represents the sound effect volume preference of the player's communication in the game, when the communication volume preference value in the game is large, a fast and strong sound effect pulse is generated, so that the volume, sound effect switching speed and sound effect change amplitude can better adapt to the player's overall demand for maintaining a large sound effect volume when communicating in the game, which can better maintain the interactive atmosphere in the game and improve the player's usage experience in the game.
[0093] The beneficial effect of the above implementation method is that when the communication volume preference value in the game is small, the volume, sound effect switching speed and sound effect change range are smaller, which can better meet the communication needs of thinking and quiet players in the game. By maintaining a smooth transition of game sound effects, the user experience of players who prefer quietness is improved.
[0094] The beneficial effect of the above implementation method is that by obtaining the in-game communication volume preference value that represents the sound effect volume preference of the player's communication in the game, and combining it with the number of players, the volume, sound effect switching speed and sound effect change amplitude are comprehensively adjusted, thereby improving the player's in-game experience.
[0095] In some implementations, the above method further includes S410 to S420, and S410 to S420 are described in detail below.
[0096] S410: Obtain player interaction information from the immersive script-killing game. The player interaction information includes communication information between multiple players. Using an activity monitoring model, determine the player's interaction activity factor based on the player count information and player interaction information. The interaction activity factor is between 0.8 and 1.5.
[0097] During the immersive script-killing game, player interaction information can be obtained. Player interaction information includes communication information between multiple players. This communication information can reflect the frequency and depth of interaction between players. By analyzing this interaction information, we can understand the level of player participation in the game.
[0098] After obtaining the communication information between multiple players, the player interaction activity factor can be determined through the activity monitoring model based on the player number information and player interaction information.
[0099] Exemplarily, the activity monitoring model may be a deep learning model based on a neural network, which predicts the interactive activity of the current game by analyzing features such as the number of players, interaction frequency, and interaction content in historical game data.
[0100] For example, the activity monitoring model can be trained using sample player quantity information, sample player interaction information, and sample interaction activity factors. The activity monitoring model can identify the activity levels corresponding to different interaction modes, such as the impact of quick conversations, heated debates, or in-depth discussions on the activity factor.
[0101] It should be noted that the interactive activity factor is between 0.8 and 1.5, and the interactive activity factor can quantify the overall activity level of the player group.
[0102] S420. Determine the product of the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude corresponding to the audio unit adjustment information and the interactive activity factor, so as to adjust the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude corresponding to the audio unit adjustment information.
[0103] After determining the interactive activity factor, the product of the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude corresponding to the audio unit adjustment information can be calculated. Through this multiplication operation, the change rate and intensity of the sound effect can be dynamically adjusted according to the actual activity level of the player. When the interactive activity factor is greater than 1, the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude can be amplified; when the interactive activity factor is less than 1, the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude can be reduced.
[0104] For example, in a detective-themed script-killing game, when players are enthusiastically discussing clues, the Interaction Activity Factor might reach 1.3, and the speed of sound effects changes will increase by 30%, creating a more tense atmosphere. Conversely, when players fall into a state of contemplation, with the Interaction Activity Factor perhaps reaching 0.8, the speed of sudden sound effects, the speed of soundscape switching, and the amplitude of sudden sound effects changes will remain steady, giving players ample time to think.
[0105] The beneficial effect of the above implementation method is that, based on the player quantity information and player interaction information, the player's interactive activity factor is determined, and the sudden sound effect rising speed, sound scene switching speed and sound effect mutation amplitude corresponding to the audio unit adjustment information are adjusted according to the interactive activity factor. The game rhythm can be adjusted according to the player's interactive activity in the actual game, realizing a positive cycle in which the challenge becomes more compact when the player team is more active.
[0106] In some implementations, the above method further includes S430 to S440, and S430 to S440 are described in detail below.
[0107] S430. Obtain the volume value of the game NPC in the immersive script-killing game, and obtain the voice clarity evaluation value of the game NPC by the players in the immersive script-killing game.
[0108] During the immersive script-killing game, the volume value of the game NPC can be continuously obtained, and the player's subjective evaluation value of the NPC's voice clarity can be simultaneously collected. The volume value of the game NPC reflects the actual sound pressure level of the NPC's voice output, and the voice clarity evaluation value quantifies the player's perception of whether the voice is clear.
[0109] S440. Determine the volume adjustment value of the game NPC according to the volume value, interactive activity factor and voice clarity evaluation value of the game NPC through the game NPC voice adjustment unit.
[0110] After obtaining the volume value and speech clarity evaluation value of the game NPC, the game NPC speech adjustment unit can comprehensively analyze the three key parameters of volume value, interactive activity factor and speech clarity evaluation value, and determine the specific volume adjustment value based on these three input parameters.
[0111] For example, in a horror-themed script-killing game, when players' nervousness leads to a decrease in interactive activity, the NPC volume can be automatically raised to ensure the delivery of key clues; while during the deduction phase, when players engage in heated discussions, the NPC volume can be appropriately lowered to avoid disrupting communication. This dynamic adjustment ensures a balanced sound experience across different game stages.
[0112] The beneficial effect of the above implementation method is that when the interactive activity factors of players in the game are different, the volume value of the NPC may affect the effect of the immersive script-killing game. By obtaining the voice clarity evaluation value of the game NPC from user feedback, the volume value of the NPC is adjusted according to the interactive activity factor and the voice clarity evaluation value to improve the control effect of the NPC's volume value, thereby avoiding the interruptive impact of the NPC's volume on the player's interactive activity, and improving the NPC's volume control effect.
[0113] The beneficial effect brought about by the above implementation method is that it takes into account both objective volume measurement values and subjective clarity evaluations, so that the volume adjustment conforms to both acoustic characteristics and player perception needs, thereby improving the naturalness and realism of sound interaction.
[0114] The beneficial effect brought about by the above implementation method is that the introduction of the interactive activity factor enables the system to perceive changes in the game situation, automatically reduce NPC interference when players are focused on discussion, enhance NPC presence when guidance is needed, and realize intelligent volume scene adaptation.
[0115] In some implementations, the above method further includes S450 to S460, which are described in detail below.
[0116] S450. Obtain the duration of the repeated playback of the same clue sound effect in the immersive script-killing game as the duration of the sound effect repetition interval.
[0117] During an immersive script-killing game, the duration of the interval between repeated playbacks of the same clue sound effect can be obtained, that is, the sound effect repetition interval duration. The sound effect repetition interval duration reflects the time interval between repeated playbacks of the clue sound effect. By monitoring this parameter, the repetition frequency of the clue sound effect in the current game process can be understood.
[0118] For example, when players explore a secret room, the ticking sound of a clock played in a loop in the background serves as a key time clue. By recording the time difference between two playbacks of the sound effect as the duration of the sound effect repetition interval, the collection of this timing data can be achieved through the game engine's built-in audio management system, providing basic data support for subsequent dynamic adjustments.
[0119] S460: Determine the product of the sound effect repetition interval duration and the interactive activity factor as the sound effect repetition interval adjustment duration. Adjust the interval duration of repeated playback of the same clue sound effect in the immersive script-killing game based on the sound effect repetition interval adjustment duration.
[0120] After determining the duration of the sound effect repetition interval, we can determine the product of the sound effect repetition interval and the interactive activity factor to obtain the adjusted duration of the sound effect repetition interval. By combining the sound effect repetition interval with the interactive activity factor, which reflects the player's level of engagement, we achieve intelligent matching of the clue sound effect playback rhythm with the actual game progress. When players are actively interacting, the repetition interval is appropriately shortened to maintain the continuity of clue prompts; when players are focused on solving the puzzle, the interval is appropriately extended to avoid distractions.
[0121] Exemplarily, the interactive activity factor may be between 0.7 and 1.8, and the duration of the sound effect repetition interval may be increased or shortened by the interactive activity factor.
[0122] The beneficial effect of the above-mentioned implementation method is that the repetition interval of the clue sound effects in the immersive script-killing game will affect the continuity of the game. By adjusting the repetition time interval of the clue sound effects through the interactive activity factor, the adaptability of the game clue sound effects and the player's activity in the game is further improved, thereby improving the player's gaming experience.
[0123] An embodiment of the present application also provides an immersive script-killing sound effect generation system based on deep learning, including a unit for executing any of the methods described above.
[0124] Figure 5 A logical structure diagram of an immersive script-killing sound effect generation system based on deep learning provided in an embodiment of the present application is shown as follows: Figure 5As shown, the system 1 of this embodiment includes a processing unit 11, a storage unit 12, and a transceiver unit 13. The processing unit 11 is used to process data, the storage unit 12 is used to store data, and the transceiver unit 13 is used to send and receive data. The processing unit 11, the storage unit 12, and the transceiver unit 13 cooperate with each other to implement the above method. The beneficial effects of the embodiment of the present application have been described in the above method and will not be repeated here.
[0125] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.
[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0127] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0128] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0129] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0130] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0131] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0132] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for generating immersive script-killing sound effects based on deep learning, characterized in that: The method comprises: Obtaining environmental spatial information corresponding to a target environment, and obtaining position information of multiple audio units and position information of multiple players within the target environment; wherein the environmental spatial information includes three-dimensional spatial information corresponding to the target environment, and each audio unit corresponds to one audio unit position information; The sound field control unit determines the sound unit control information of each sound unit according to the environmental space information, the position information of multiple sound units and the position information of multiple players; and controls the working state of each sound unit according to the sound unit control information of each sound unit.
2. The method according to claim 1, wherein The method further comprises: Obtaining sound effect optimization information fed back by players corresponding to multiple player position information; wherein the sound effect optimization information includes volume adjustment suggestions; Through the sound field control unit, the audio unit control information of each audio unit is determined according to the environmental space information, the position information of multiple audio units, the position information of multiple players and the sound effect optimization information corresponding to the multiple player position information; and the working state of each audio unit is controlled according to the audio unit control information of each audio unit.
3. The method according to claim 2, wherein The method further comprises: Through the script parsing model, according to the script information of the immersive script-killing game, the scene information and scene conversion information corresponding to the immersive script-killing game are determined, and the scene conversion information corresponds to the conversion information between different scene information; Through the sound field adjustment unit, audio unit adjustment information of each audio unit is determined based on environmental space information, multiple audio unit position information, multiple player position information, scene information and scene transition information; the working status of each audio unit is adjusted according to the audio unit adjustment information of each audio unit; wherein the audio unit adjustment information includes volume adjustment information and surround sound adjustment information.
4. The method according to claim 3, wherein The method further comprises: Acquiring sound effect preference information fed back by players corresponding to multiple player position information; wherein the sound effect preference information includes a volume mutation preference value, a horror preference value, and a low-frequency tolerance value. The volume mutation preference value represents the player's preference for volume mutations, the horror preference value represents the player's preference for horror sound effects, and the low-frequency tolerance value represents the player's tolerance for low-frequency sound effects. The sound field adjustment unit determines a limit threshold of the audio unit adjustment information based on the sound effect preference information fed back by players corresponding to the position information of multiple players respectively, and adjusts the audio unit adjustment information of each audio unit based on the limit threshold of the audio unit adjustment information.
5. The method according to claim 4, wherein The method further comprises: Determine the number of player position information as player quantity information; Through the sound field transition adjustment unit, according to the player number information, the sudden sound effect rising speed and the sound scene switching speed corresponding to the audio unit adjustment information are determined; the working state of each audio unit is adjusted according to the audio unit adjustment information of each audio unit; among which, when the player number information is greater, the sudden sound effect rising speed and the sound scene switching speed are greater.
6. The method according to claim 5, wherein The method further comprises: Get the in-game communication volume preference value set by the player. The in-game communication volume preference value represents the preferred volume of the sound effects for the player's communication in the game. Through the sound field transition adjustment unit, according to the player number information and the communication volume preference value in the game, the sudden sound effect rising speed, sound scene switching speed and sound effect mutation amplitude corresponding to the audio unit adjustment information are determined; the working state of each audio unit is adjusted according to the audio unit adjustment information of each audio unit; among which, when the communication volume preference value in the game is greater, the sudden sound effect rising speed, sound scene switching speed and sound effect mutation amplitude are greater.
7. The method according to claim 6, wherein The method further comprises: Obtain player interaction information in immersive script-killing games, including communication information between multiple players; determine the player interaction activity factor based on the player number information and player interaction information through an activity monitoring model; the interaction activity factor is between 0.8 and 1.5; Determine the product of the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude corresponding to the audio unit adjustment information and the interactive activity factor, so as to adjust the sudden sound effect rising speed, soundscape switching speed, and sound effect mutation amplitude corresponding to the audio unit adjustment information.
8. The method according to claim 7, wherein The method further comprises: Get the volume value of the game NPC in the immersive script-killing game, and get the voice clarity evaluation value of the game NPC by the players in the immersive script-killing game; The game NPC voice adjustment unit determines the volume adjustment value of the game NPC according to the volume value, interactive activity factor and voice clarity evaluation value of the game NPC.
9. The method according to claim 8, wherein The method further comprises: Get the duration between repeated playbacks of the same clue sound effect in the immersive script-killing game as the sound effect repetition interval duration; Determine the product of the sound effect repetition interval duration and the interactive activity factor as the sound effect repetition interval adjustment duration; adjust the interval duration of the same clue sound effect during repeated playback in the immersive script-killing game based on the sound effect repetition interval adjustment duration.
10. An immersive script-killing sound effect generation system based on deep learning, characterized in that: Comprising means for performing the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Sound effect control method and system for multi-modal story content output
CN109542389A
Sound effect adjustment method and device, electronic equipment and storage medium
CN112738705A
Sound parameter adjusting method
CN116095569A
Intelligent acousto-optic control method, system and equipment for dispatching room
CN118411971A
Sound quality intelligent regulation and control system
CN118764777A