Sound field generation method, device, equipment and medium
Through large-scale language models and audio generation technology, sound content and position information matching the scene are generated, which solves the problems of high hardware cost and limited applicability of existing sound field restoration technology, and realizes efficient and realistic sound field reproduction in complex environments.
Patent Information
- Application Number
- CN202411344211.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-09-25
AI Technical Summary
Existing sound field restoration technology relies on multi-channel audio systems and physical acoustic models. It has high hardware costs and difficulty in handling dynamic scene changes, which limits its application in complex environments.
By inputting scene description information into a large language model, sound content information and location information are generated. Combined with the audio generation model and spatial sound restoration technology, the propagation characteristics of sound in space can be accurately reproduced.
It reduces dependence on complex hardware, is applicable to various complex scenes, can respond to scene changes in real time, and provide a realistic dynamic sound environment.
Smart Images

Figure CN119110238B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a sound field generation method, device, equipment and medium. Background Art
[0002] Sound field restoration technology simulates the propagation characteristics of sound in space, achieving spatial positioning, directionality, and environmental reflection, allowing users to experience realistic three-dimensional sound effects. This technology is widely used in virtual reality, augmented reality, immersive audio systems, and other fields to provide users with an immersive auditory experience.
[0003] However, existing sound field restoration technologies rely primarily on multi-channel audio systems and physical acoustic models, often requiring extensive hardware and complex computing resources. Furthermore, these technologies struggle to handle dynamic scene changes and are unable to flexibly address diverse sound field requirements, limiting their application in complex environments. Summary of the Invention
[0004] The present invention provides a sound field generation method, device, equipment and medium, which are used to solve the defects of high cost and limited application range of sound field restoration hardware in related technologies.
[0005] The present invention provides a sound field generation method, comprising the following steps:
[0006] Inputting the scene description information into a large language model to obtain the sound content information and sound position information of the scene corresponding to the scene description information output by the large language model;
[0007] Based on the sound content information and the sound position information, a scene sound field corresponding to the scene description information is generated.
[0008] According to the sound field generation method provided by the present invention, the scene description information includes scene environment description information, and the scene environment description information is used to describe the environment in which the scene is located.
[0009] According to the sound field generation method provided by the present invention, the scene description information further includes scene reference sound information, and the scene reference sound information is used to describe the content and / or position of the candidate sound in the scene.
[0010] According to the sound field generation method provided by the present invention, generating the scene sound field corresponding to the scene description information based on the sound content information and the sound position information includes:
[0011] Acquire audio corresponding to the sound content information;
[0012] Based on the audio and the sound position information, a scene sound field corresponding to the scene description information is generated.
[0013] According to the sound field generation method provided by the present invention, the step of obtaining audio corresponding to the sound content information includes:
[0014] In a case where the sound content information includes human voice content, generating audio corresponding to the human voice content based on a speech synthesis model;
[0015] Furthermore, when the sound content information includes non-human voice content, audio corresponding to the non-human voice content is generated based on a sound effect generation model.
[0016] According to the sound field generation method provided by the present invention, generating the scene sound field corresponding to the scene description information based on the audio and the sound position information includes:
[0017] Based on the sound position information, spatial sound restoration is performed on the audio to obtain a scene sound field corresponding to the scene description information.
[0018] The present invention also provides a sound field generating device, comprising the following modules:
[0019] A scene parsing unit, configured to input the scene description information into a large language model, and obtain sound content information and sound position information of the scene corresponding to the scene description information output by the large language model;
[0020] A sound field generating unit is configured to generate a scene sound field corresponding to the scene description information based on the sound content information and the sound position information.
[0021] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned sound field generation methods is implemented.
[0022] The present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned sound field generation methods when executed by a processor.
[0023] The present invention further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned sound field generation methods.
[0024] The sound field generation method, device, equipment, and medium provided by the present invention use a large-scale language model to analyze scene description information input by the user and generate corresponding sound content information and sound location information, ensuring a high degree of match between the sound and the scene. The sound field is generated by combining the sound content information and sound location information, accurately reproducing the propagation characteristics of sound in space. This invention can reduce reliance on complex hardware and is applicable to a variety of complex application scenarios. Furthermore, it can respond to scene changes in real time, providing a more realistic and dynamic sound environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0026] Figure 1 It is a flow chart of the sound field generating method provided by the present invention.
[0027] Figure 2 It is a structural schematic diagram of the sound field generating device provided by the present invention.
[0028] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0029] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0030] Acoustic field restoration technology is a key technology used to reproduce the propagation characteristics of sound within a specific space. By simulating the spatial positioning, direction, distance, and environmental reflection and absorption of sound, it provides a highly immersive auditory experience. This technology has widespread applications in virtual reality, augmented reality, immersive gaming, and high-fidelity audio systems. For example, in a virtual reality environment, acoustic field restoration technology enables users to accurately perceive the direction and distance of a sound source, enhancing the sense of immersion. In high-fidelity audio systems, acoustic field restoration technology can accurately reproduce the acoustics of a live music concert, providing a realistic audio experience.
[0031] Traditional sound field restoration methods typically rely on multi-channel audio systems and physical acoustic models. These systems require multiple speakers installed in a space and use complex algorithms to simulate sound propagation paths and environmental effects. These methods are not only demanding on hardware and complex to deploy, but also require significant computing resources, making them difficult to implement in real-time applications. Furthermore, these technologies struggle with dynamically changing scenes and cannot flexibly adapt to the user's changing perspective and position in the virtual space, thus limiting the application scenarios for sound field restoration.
[0032] In view of the above problems, an embodiment of the present invention provides a sound field generating method. Figure 1 Schematic diagram of the flow of the sound field generation method provided by the present invention, such as Figure 1 As shown, the method includes:
[0033] Step 110: Input the scene description information into a large language model to obtain the sound content information and sound position information of the scene corresponding to the scene description information output by the large language model.
[0034] Specifically, scene description information refers to a specific scene described in natural language, including the environmental background of the scene, possible sound elements, and the dynamic characteristics of these sound elements. The specific scene here refers to the scene where the sound field to be generated is located. In an embodiment of the present invention, in order to generate a sound field corresponding to the scene description information, the sound content information and sound position information of the scene corresponding to the scene description information can be first generated based on a large language model. That is, the scene description information can help the large model understand the overall atmosphere and sound layout of the scene, thereby generating more accurate sound information.
[0035] Here, a large language model (LLM), also known as a large model, refers to a natural language processing (NLP) model with a large number of parameters. The number of model parameters and / or the complexity of the model structure exceed a preset threshold. The model processes large amounts of text data during training and has the ability to understand and generate natural language. For example, a large language model may include the Spark large model.
[0036] Furthermore, the large language model here can be a large language model in a general field, or a large language model in the code field obtained by fine-tuning data in the sound field generation field. For example, the large language model can be obtained by fine-tuning based on a pre-trained large language model using various sample scene description information and the sound content information and sound position information of the scenes corresponding to the sample scene description information. The embodiment of the present invention does not make specific limitations on this.
[0037] The scene description is fed into the large language model, which uses its powerful natural language understanding and generation capabilities to parse the input. By deeply understanding the scene described, the large language model can predict key sound elements that may be present in the scene and generate corresponding sound content and location information.
[0038] Sound content information can include the types of sounds within the scene, such as wind, rain, human voices, and footsteps. Specifically, it can be represented by sound elements, such as the type of bird song, the nature of human voices (e.g., whispers or shouts), and the intensity and duration of wind sounds. Sound location information describes the specific spatial locations of these sounds. This location can be a static location or a dynamic motion path, such as the direction the sound comes from and whether it is stationary or moving within the scene. For example, the input scene description might be "a quiet park, with distant birdsong, nearby whispers, and the occasional rustle of leaves in the breeze." In this example, the model uses its language understanding capabilities to identify sound elements such as "birdsong," "conversation," and "wind." Through its generation capabilities, it outputs sound content information and location information corresponding to these elements, such as distant birdsong, nearby whispers, and surrounding wind sounds, accurately assigning these sounds to their corresponding spatial locations.
[0039] Step 120: Generate a scene sound field corresponding to the scene description information based on the sound content information and the sound position information.
[0040] Specifically, the sound content information obtained in step 110 includes the specific characteristics of each sound element in the scene, such as the sound type, frequency, pitch, and loudness. These characteristics describe the physical properties of the sound and serve as the basic material for generating the sound. The sound position information provides the spatial distribution and dynamic changes of these sounds, such as the direction of the sound source, the distance from the listener, and the sound's movement path within the scene. Combining these two types of information generates a specific, perceptible three-dimensional sound field—that is, a scene sound field that corresponds to the scene description.
[0041] The core of sound field generation is to convert the sound elements contained in the sound content information into audio. The audio of each sound element is then spatialized based on its location information. This spatialization process accurately represents the propagation characteristics of the sound elements in three-dimensional space, creating a realistic audio environment. This process ensures accurate sound directionality and positioning, allowing listeners to perceive the direction of the sound's source. Spatialization also simulates the sense of distance, adjusting the loudness and frequency characteristics of sound to make distant sounds appear blurrier and nearby sounds clearer. Spatialization also simulates environmental reflections and reverberation. For example, in an open space, the echo and delay of sound are accurately simulated, enhancing the realism of the sound. Spatialization also involves simulating dynamic changes and the movement of sound sources. As a sound source moves within the scene, the system adjusts the sound's location, direction, and intensity in real time to accurately represent the dynamic changes. For example, if a car moves from far away to near and then back away, spatialization allows listeners to perceive the car's sound as it shifts from distant to near and back again.
[0042] The sound field generation method provided by an embodiment of the present invention uses a large language model to parse scene description information input by the user and generate corresponding sound content information and sound location information, ensuring a high degree of match between the sound and the scene. The sound field is generated by combining the sound content information and the sound location information to accurately reproduce the propagation characteristics of sound in space. This method can reduce reliance on complex hardware and is applicable to a variety of complex application scenarios. Furthermore, the method can respond to scene changes in real time, providing a more realistic and dynamic sound environment.
[0043] Based on the above embodiment, in step 110, the scene description information includes scene environment description information, and the scene environment description information is used to describe the environment in which the scene is located.
[0044] Specifically, the scene environment description information is used to describe in detail the environment in which the scene is located. It is a key component of the scene description information and provides a basis for generating a sound environment that matches it.
[0045] The scene environment description covers all aspects of the scene, including its geographic location, environment type, weather conditions, and time of day. For example, the scene environment description might describe whether the scene is a city street, a park, an indoor room, or a rural field. These details of the location and environment type directly affect the sound characteristics, such as sound reverberation, attenuation, and propagation through different materials and spaces.
[0046] Weather conditions also play an important role in describing the scene environment, such as sunny, rainy, windy, or snowstorms. These conditions affect the background sounds in the environment, as well as the texture and intensity of the sounds. For example, on a rainy day, the environment may have a persistent sound of raindrops, while the sound of wind will be more prominent in an open environment.
[0047] Time is also an important element in describing scene environments. Different times of day, such as daytime, nighttime, or dusk, can affect the sound characteristics and atmosphere of an environment. For example, a night scene might have a quieter background sound, occasionally accompanied by a distant barking dog or the sound of wind, while a daytime city scene might be filled with traffic noise and the hustle and bustle of people.
[0048] Through detailed scene environment description information, we can better understand and simulate the natural characteristics of sound in the scene, so that the generated sound environment is highly consistent with the actual situation of the scene, enhancing the user's immersion and sense of reality in applications such as virtual reality and augmented reality.
[0049] Based on the above embodiment, in step 110, the scene description information further includes scene reference sound information, and the scene reference sound information is used to describe the content and / or position of the candidate sound in the scene.
[0050] Specifically, scene-referenced sound information refers to specific information within the scene description that further refines the scene's sound features. This information describes the content and / or location of candidate sounds within the scene. By providing scene-referenced sound information, large-scale language models can more accurately understand which sound elements should appear in a scene, as well as their specific location or dynamics within the scene.
[0051] Specifically, scene-referenced sound information may include descriptions of specific sounds, such as "there should be distant birdsong in the scene" or "there should be the sound of flowing water nearby." These descriptions clarify the content of the candidate sounds, enabling the system to generate sound materials that meet user expectations based on these prompts. In addition, scene-referenced sound information can also include spatial positioning information of these sounds, such as "the birdsong comes from the tree in the front left" or "the sound of water flows from the right." These details help to accurately locate the direction and distance of the sound source when generating the sound field.
[0052] By incorporating scene-specific sound information, the sound field generation process can more accurately reflect the user's intent, not only matching the sound content to the scene requirements but also achieving realistic positioning in spatial layout. This refined information input improves the accuracy of the generated sound environment, making the resulting sound field more closely aligned with the actual scene.
[0053] Based on the above embodiment, in step 120, generating the scene sound field corresponding to the scene description information based on the sound content information and the sound position information includes:
[0054] Acquire audio corresponding to the sound content information;
[0055] Based on the audio and the sound position information, a scene sound field corresponding to the scene description information is generated.
[0056] Specifically, there are multiple methods for obtaining audio corresponding to the sound content information, depending on the requirements of the scene description information and the available resources.
[0057] For example, a pre-trained audio generation model can be used to generate audio. The audio generation model can be a literary sound model, such as a speech synthesis model, or a sound effects synthesis model, which is not specifically limited in the present embodiment. The audio generation model can automatically generate audio corresponding to each sound content that matches the scene based on the sound content information used as input. The generated audio can meet highly customized requirements and accurately reflect specific sound elements in the scene, such as ambient sounds, human voices, or natural sounds.
[0058] Furthermore, audio materials corresponding to the sound content can be retrieved from an audio database. Audio databases typically contain a large number of high-quality pre-recorded or pre-generated sound samples, such as birdsong, flowing water, and wind in nature, as well as traffic background sounds and crowd sounds in urban environments. Based on the sound content, selecting the ready-made audio materials that best meet the scene requirements can effectively improve the efficiency and quality of sound field generation.
[0059] In addition to the two methods mentioned above, you can also use professional recording equipment to directly record audio in real scenes, or use audio synthesis software to generate sound materials with specific effects. The choice of these methods depends on the specific needs and application scenarios of the audio.
[0060] After acquiring appropriate audio, a scene sound field corresponding to the scene description information can be generated based on this audio and sound position information. This sound position information includes details such as the specific location, direction, distance, and dynamic characteristics of each sound in three-dimensional space. By combining this position information, the audio can be spatialized to accurately simulate the propagation characteristics of sound at different locations. For example, the effect of sound coming from a distance or moving through space can be simulated, giving the listener a sense of direction and distance.
[0061] Based on the above embodiment, obtaining the audio corresponding to the sound content information includes:
[0062] In a case where the sound content information includes human voice content, generating audio corresponding to the human voice content based on a speech synthesis model;
[0063] Furthermore, when the sound content information includes non-human voice content, audio corresponding to the non-human voice content is generated based on a sound effect generation model.
[0064] Specifically, the process of obtaining audio corresponding to the sound content information can be achieved through a text-based sound model, and in this process, different situations can be divided into different types of text-based sound models for processing, and the specific method depends on the type of sound contained in the sound content information.
[0065] When the sound types are specifically divided into human voice and non-human voice, the corresponding literary sound models can be a speech synthesis model and a sound effect generation model respectively.
[0066] When the sound content information includes human voice, corresponding human voice audio is generated based on a speech synthesis model. A speech synthesis model is a technology that simulates and generates natural human voices, typically trained through deep learning and a large amount of speech data. It can generate human voices with different speech characteristics, such as those of different genders, ages, emotional states, and languages. This allows for highly realistic human voice audio to be generated based on the vocal content specified in the scene description, such as the speaker's tone, language, and speaking rate. The speech synthesis model not only ensures the fluency of the generated vocals but also allows for personalized adjustments based on specific scene requirements, ensuring that the vocal portion closely matches the overall atmosphere of the scene.
[0067] When the sound content information contains non-human voice content, such as ambient sounds, natural sound effects, or specific background sounds, the audio corresponding to these non-human voice contents is generated based on the sound effect generation model. The sound effect generation model can generate natural sounds (such as birdsong, wind, rain, etc.) or artificial environmental sound effects (such as traffic noise, mechanical operation sounds, etc.). Based on the sound effect generation model, various realistic sound effects can be generated according to the input sound description information to simulate different sound environments. For example, if the scene description information mentions "birds singing in the forest" or "the sound of flowing water in the distance", the sound effect generation model can generate these specific natural sound effects so that the generated audio matches the actual environmental sound field.
[0068] Based on the above embodiment, generating the scene sound field corresponding to the scene description information based on the audio and the sound position information includes:
[0069] Based on the sound position information, spatial sound restoration is performed on the audio to obtain a scene sound field corresponding to the scene description information.
[0070] Here, sound position information provides detailed information about each audio frequency in three-dimensional space, including its specific location, direction, and distance. Sound position information is the foundation for spatial sound restoration, determining the sound's location in space and its propagation path.
[0071] Spatial sound restoration uses sound position information to simulate the propagation characteristics of sound in space, so that the generated sound field can accurately reflect the direction of the sound source, the sense of distance, and the reflection, refraction and attenuation characteristics of sound in different environments.
[0072] Specific spatial sound restoration methods can be implemented using a variety of technologies. For example, physical modeling is a technique based on the principles of physical acoustics. It simulates the propagation path of sound in spaces of various materials and shapes, calculating the reflection, refraction, and absorption effects of sound. This method is particularly suitable for high-precision sound field simulation, such as simulating the reverberation effect of sound in a cathedral or concert hall.
[0073] Software algorithms based on acoustic simulation are also a viable approach for spatial sound restoration. These algorithms utilize mathematical models and physical laws to simulate the propagation and behavior of sound in various environments. For example, Finite Difference Time Domain (FDTD) and ray tracing techniques can simulate the propagation paths of sound waves, calculate sound reflections and attenuation, and recreate the characteristics of sound in complex environments.
[0074] Furthermore, data-driven machine learning methods are an effective means of spatial sound restoration. By training deep learning models, they can learn and predict the propagation characteristics of sound in different environments. This approach can handle complex acoustic conditions and diverse scene variations, simulating the spatial characteristics of sound by generating audio effects that match real-world scenarios. This approach is particularly suitable for applications requiring real-time audio rendering and dynamic sound field generation, and can flexibly handle complex sound sources and nonlinear acoustic environments.
[0075] These spatial sound restoration methods, combined with sound position information and audio generated based on sound content, can generate a scene sound field that is highly consistent with the scene description information. The generated sound field not only accurately reproduces the spatial positioning and propagation characteristics of sound, but also provides a realistic sense of direction and distance.
[0076] The sound field generating device provided by the present invention is described below. The sound field generating device described below and the sound field generating method described above can be referenced to each other. Figure 2 Schematic diagram of the structure of the sound field generating device provided by the present invention, such as Figure 2 As shown, the device includes:
[0077] The scene parsing unit 210 is configured to input the scene description information into a large language model, and obtain the sound content information and sound position information of the scene corresponding to the scene description information output by the large language model.
[0078] The sound field generating unit 220 is configured to generate a scene sound field corresponding to the scene description information based on the sound content information and the sound position information.
[0079] The sound field generation device provided in an embodiment of the present invention uses a large language model to parse scene description information input by the user and generate corresponding sound content information and sound location information, ensuring a high degree of match between the sound and the scene. The sound field is generated by combining the sound content information and the sound location information, accurately reproducing the propagation characteristics of sound in space. This invention can reduce reliance on complex hardware and is applicable to a variety of complex application scenarios. Furthermore, it can respond to scene changes in real time, providing a more realistic and dynamic sound environment.
[0080] Based on any of the above embodiments, in the scene parsing unit:
[0081] The scene description information includes scene environment description information, and the scene environment description information is used to describe the environment in which the scene is located.
[0082] The scene description information further includes scene reference sound information, where the scene reference sound information is used to describe the content and / or position of candidate sounds in the scene.
[0083] Based on any of the above embodiments, the sound field generating unit is specifically configured to:
[0084] Acquire audio corresponding to the sound content information; and generate a scene sound field corresponding to the scene description information based on the audio and the sound position information.
[0085] When the sound content information includes human voice content, audio corresponding to the human voice content is generated based on a speech synthesis model; and when the sound content information includes non-human voice content, audio corresponding to the non-human voice content is generated based on a sound effect generation model.
[0086] Based on the sound position information, spatial sound restoration is performed on the audio to obtain a scene sound field corresponding to the scene description information.
[0087] Figure 3 An example of a physical structure diagram of an electronic device is shown below. Figure 3As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the sound field generation method, which includes:
[0088] Inputting the scene description information into a large language model to obtain the sound content information and sound position information of the scene corresponding to the scene description information output by the large language model;
[0089] Based on the sound content information and the sound position information, a scene sound field corresponding to the scene description information is generated.
[0090] Furthermore, the logic instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the relevant art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0091] On the other hand, the present invention further provides a computer program product, comprising a computer program. The computer program may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the sound field generation method provided by the above methods, the method comprising:
[0092] Inputting the scene description information into a large language model to obtain the sound content information and sound position information of the scene corresponding to the scene description information output by the large language model;
[0093] Based on the sound content information and the sound position information, a scene sound field corresponding to the scene description information is generated.
[0094] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the sound field generation method provided by the above methods, the method comprising:
[0095] Inputting the scene description information into a large language model to obtain the sound content information and sound position information of the scene corresponding to the scene description information output by the large language model;
[0096] Based on the sound content information and the sound position information, a scene sound field corresponding to the scene description information is generated.
[0097] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0098] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A sound field generation method, characterized in that: include: Inputting scene description information into a large language model to obtain sound content information and sound location information of a scene corresponding to the scene description information output by the large language model; the scene description information includes scene environment description information, which is used to describe the environment in which the scene is located; the scene description information also includes scene reference sound information, which is used to describe the content and / or location of candidate sounds in the scene; Generating a scene sound field corresponding to the scene description information based on the sound content information and the sound position information includes: The sound content information includes the specific characteristics of each sound element in the scene, and the sound position information includes the distribution and dynamic changes of each sound element in space. Based on the sound position information, the audio of each sound element is spatialized to accurately represent the propagation characteristics of the audio of each sound element in three-dimensional space. The spatialization processing includes ensuring the direction and positioning of the sound, and simulating the sense of distance of the sound by adjusting the loudness and frequency characteristics of the sound. It also includes simulating environmental reflections and reverberation effects, as well as dynamic changes and movement of the sound source.
2. The sound field generation method according to claim 1, characterized in that: The generating, based on the sound content information and the sound position information, a scene sound field corresponding to the scene description information includes: Acquire audio corresponding to the sound content information; Based on the audio and the sound position information, a scene sound field corresponding to the scene description information is generated.
3. The sound field generation method according to claim 2, characterized in that: The obtaining of audio corresponding to the sound content information includes: In a case where the sound content information includes human voice content, generating audio corresponding to the human voice content based on a speech synthesis model; Furthermore, when the sound content information includes non-human voice content, audio corresponding to the non-human voice content is generated based on a sound effect generation model.
4. The sound field generation method according to claim 2, characterized in that: The generating, based on the audio and the sound position information, a scene sound field corresponding to the scene description information includes: Based on the sound position information, spatial sound restoration is performed on the audio to obtain a scene sound field corresponding to the scene description information.
5. A sound field generating device, characterized in that: include: A scene parsing unit, configured to input scene description information into a large language model, and obtain sound content information and sound location information of a scene corresponding to the scene description information output by the large language model; the scene description information includes scene environment description information, which is used to describe the environment in which the scene is located; the scene description information also includes scene reference sound information, which is used to describe the content and / or location of candidate sounds in the scene; A sound field generation unit is configured to generate a scene sound field corresponding to the scene description information based on the sound content information and the sound position information, including: the sound content information including specific characteristics of each sound element in the scene, the sound position information including the distribution and dynamic changes of each sound element in space, and spatializing the audio of each sound element based on the sound position information to accurately represent the propagation characteristics of the audio of each sound element in three-dimensional space; the spatialization processing includes ensuring the directionality and positioning of the sound, and simulating the sense of distance of the sound by adjusting the loudness and frequency characteristics of the sound; and also includes simulating environmental reflections and reverberation effects, as well as dynamic changes and movement of the sound source.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the sound field generating method according to any one of claims 1 to 4 is implemented.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the sound field generating method according to any one of claims 1 to 4 is implemented.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the sound field generating method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Method and system for creating object-based audio content
CN112334973A
Audio book automatic generation method based on multi-modal large language model
CN116821410A