Scene perception few-sample acoustic modeling method and device based on map guidance, equipment and medium
Through a map-guided scene perception method, semantic features in visual features are extracted and fused, scene feature maps are generated, and predicted RIR is calculated using Gaussian distance weighting, which solves the performance limitations of the acoustic modeling task in the existing technology, and achieves higher prediction accuracy and spatial information understanding capabilities.
Patent Information
- Application Number
- CN202411814239.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The lack of sufficient understanding and utilization of the provided small amount of sample data in the small sample acoustic modeling task limits the performance of acoustic learning, especially in understanding spatial information and generalizing to unknown scenarios.
A scene-aware method based on map guidance is adopted to generate a scene feature map by extracting semantic features in visual features and using depth mapping and pose information to align and fusion. Then, using the Gaussian distance weighted calculation method, query features are obtained from the scene feature map, and input into the autoencoder with audio features to predict the target RIR.
It significantly improves the model's prediction accuracy of RIR, enhances its ability to understand spatial information such as house structure, layout, object semantics, and can better generalize to unknown scenarios.
Smart Images

Figure CN119942543A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual modeling, and in particular to a map-guided scene-aware few-sample acoustic modeling method, device, equipment and medium. Background Art
[0002] The task of few-shot acoustic modeling aims to predict the acoustic characteristics of unknown postures in unknown areas of unknown scenes using a small number of environmental audio-visual observation samples. It has wide application value in fields such as VR and AR. The core goal of this task is to understand the entire acoustic space and then predict the room impulse response function (RIR) at any specified location, and provide powerful spatial perception capabilities for downstream tasks.
[0003] Existing technical solutions cannot fully understand and utilize the small amount of example data provided, which limits the performance of acoustic learning. Either the environment is perceived from the original RGB-D image and the target RIR is predicted in an end-to-end manner, or the local features are implicitly modeled through Nerf to directly predict the target RIR. The former weakens the pose information through sinusoidal position encoding and ignores the connection between each position, while the latter cannot utilize the complete visual information, especially the semantic information, which limits the model's ability to understand the space and cannot generalize well to unknown scenes. Summary of the invention
[0004] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the object of the present invention is to provide a map-guided scene-aware few-sample acoustic modeling method, device, equipment and medium.
[0005] The first technical solution adopted by the present invention is:
[0006] A map-guided scene-aware few-sample acoustic modeling method comprises the following steps:
[0007] Acquire image data, extract visual features according to the image data, and extract scene semantic features from the visual features;
[0008] According to the extracted scene semantic features, the scene features obtained from different postures are aligned and fused to obtain a scene feature map;
[0009] The query coordinates are obtained, query features are obtained according to the query coordinates and the scene feature map, and the target RIR is obtained according to the query features.
[0010] Furthermore, extracting scene semantic features from visual features includes:
[0011] The U-Net model pre-trained on the semantic segmentation task is used to extract semantic features related to acoustics. The U-Net model outputs the last layer of latent features that fuse low-level and high-level features, thereby providing a comprehensive representation of environmental details:
[0012] v=f V (V)
[0013] In the formula, V represents the RGB image in the image data, f V represents the pre-trained U-Net feature extraction model, and v represents the extracted pixel-level acoustic semantic features that are consistent with the input image in length and width.
[0014] Furthermore, the scene features acquired from different postures are aligned and fused according to the extracted scene semantic features to obtain a scene feature map, including:
[0015] With the help of depth mapping technology and the depth map in the image data, the semantic features are mapped to the top view corresponding to the features:
[0016] T=D P (v,D)
[0017] Where T represents the top view corresponding to the feature, D P represents the depth map, D represents the depth map in the observation information, and v represents the extracted pixel-level acoustic semantic features that are consistent with the input image in length and width;
[0018] With the help of the posture information of each observation point, the top view is rotated and translated so that it has a unified world coordinate system:
[0019] T ′ =T R (T,P)
[0020] Where, T ′ It represents the top view corresponding to each observed feature after alignment, T R represents the rotation and translation operation, P represents the posture information corresponding to the observation point (including coordinates and orientation information), and T represents the top view corresponding to each observation feature;
[0021] The aligned top views are fused to retain the results with the most significant features (i.e., the largest eigenvalues), thereby obtaining the feature map of the entire scene:
[0022] M=F(T i )
[0023] In the formula, M represents the top view after fusion, F represents the fusion operation, and T ′ It represents the top view corresponding to each observed feature after alignment.
[0024] Further, the obtaining of query coordinates, obtaining query features according to the query coordinates and the scene feature map, and obtaining a target RIR according to the query features include:
[0025] Get the query coordinates and calculate the Gaussian distance between the query coordinates and each point on the scene feature map;
[0026] The Gaussian distance value of each point is used as the weight to perform weighted summation on the feature values to obtain the query feature under the query coordinates;
[0027] The query features are combined with the audio features and input into the autoencoder to predict the RIR of the query coordinates.
[0028] Furthermore, the calculation formula of the Gaussian distance is:
[0029]
[0030] In the formula, G represents the Gaussian distance, (x, y) represents the coordinates of each point on the scene feature map, (x q ,y q ) represents the query coordinates and σ represents a trainable parameter.
[0031] Furthermore, the calculation formula of the query feature is:
[0032]
[0033] In the formula, v q represents the query feature under the query coordinates, k represents the number of points on the scene map, G represents the Gaussian distance, (x i ,y i ) represents the coordinates of the i-th point on the scene feature map, (x q ,y q ) represents the query coordinates, M i Represents the features of the i-th point on the scene feature map.
[0034] The second technical solution adopted by the present invention is:
[0035] A map-guided scene-aware few-sample acoustic modeling device, comprising:
[0036] A semantic feature extraction module is used to obtain image data, extract visual features based on the image data, and extract scene semantic features from the visual features;
[0037] The scene feature alignment module is used to align and fuse the scene features obtained from different postures according to the extracted scene semantic features to obtain a scene feature map;
[0038] The prediction module is used to obtain the query coordinates, obtain the query features according to the query coordinates and the scene feature map, and obtain the target RIR according to the query features.
[0039] The third technical solution adopted by the present invention is:
[0040] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a map-guided scene-aware few-sample acoustic modeling method as described above.
[0041] The fourth technical solution adopted by the present invention is:
[0042] A computer-readable storage medium, wherein at least one instruction, at least one program, code set or instruction set is stored in the storage medium, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement a map-guided scene-aware few-sample acoustic modeling method as described above.
[0043] The fifth technical solution adopted by the present invention is:
[0044] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.
[0045] The beneficial effects of the present invention are as follows: the present invention proposes a map-guided scene perception solution, which obtains feature information of the entire scene by extracting acoustic-related semantic features from different visual observations and performing feature alignment and fusion with the help of depth mapping, thereby enhancing the model's understanding of spatial information such as house structure, house layout, and object semantics. In addition, the present invention makes full use of the features of each point in the scene feature map by means of Gaussian distance weighted calculation, significantly improving the model's prediction accuracy for RIR. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0047] Figure 1 is a flowchart of a method for scene-aware few-sample acoustic modeling based on map guidance in an embodiment of the present invention;
[0048] Figure 2 Schematic diagram of a map-guided scene-aware few-sample acoustic modeling algorithm in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0050] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0051] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0052] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0053] Example 1
[0054] like Figure 1 and Figure 2 As shown, this embodiment provides a scene-aware few-sample acoustic modeling method based on map guidance. In the virtual modeling fields such as AR and VR, this method is used to solve the problem of weak alignment ability of existing audio-visual feature representations and enhance the model's understanding of spatial information such as house structure, house layout, and object semantics. The method specifically includes the following steps:
[0055] S1. Acquire image data, extract visual features based on the image data, and extract scene semantic features from the visual features.
[0056] As an implementation method, a U-Net model pre-trained on a semantic segmentation task is used to extract acoustic-related semantic features, where the model outputs the last layer of latent features that fuse low-level and high-level features, thereby providing a comprehensive representation of environmental details:
[0057] v=f V (V)
[0058] Where V represents the RGB image in visual observation, f V represents the pre-trained U-Net feature extraction model, and v represents the extracted pixel-level acoustic semantic features that are consistent with the input image in length and width.
[0059] S2. Based on the extracted scene semantic features, the scene features obtained from different postures are aligned and fused to obtain a scene feature map.
[0060] See also Figure 2 In some embodiments, step S2 specifically includes the following steps:
[0061] S2-1: Using depth mapping technology and depth maps from visual observations, semantic features are mapped to the top view:
[0062] T=D P (v,D)
[0063] Where T represents the top view corresponding to the feature, D P represents the depth map, D represents the depth map in the observation information, and v represents the extracted pixel-level acoustic semantic features that are consistent with the input image in length and width.
[0064] S2-2: With the help of the posture information of each observation point, the top view is rotated and translated so that it has a unified world coordinate system:
[0065] T ′ =T R (T,P)
[0066] Among them, T ′ Represents the top view after alignment with the world coordinates, T R represents the rotation and translation operation, P represents the posture information corresponding to the observation point (including coordinates and orientation information), and T represents the top view corresponding to each observation feature.
[0067] S2-3: Fuse the aligned top views and retain the result with the most significant features (i.e. the largest eigenvalue) to obtain the feature map of the entire scene:
[0068] m=F(T ′ )
[0069] Among them, m represents the fused top view, F represents the fusion operation, and T ′ It represents the top view corresponding to each observed feature after alignment.
[0070] S3. Obtain query coordinates, obtain query features according to the query coordinates and the scene feature map, and obtain target RIR according to the query features.
[0071] See also Figure 2 In some embodiments, step S3 specifically includes the following steps:
[0072] S3-1: Calculate the Gaussian distance between the query coordinates and each point on the scene feature map:
[0073]
[0074] Among them, G represents the Gaussian distance, (x, y) represents the coordinates of each point on the scene feature map, (x q ,y q ) represents the query coordinates and σ represents a trainable parameter.
[0075] S3-2: Use the Gaussian distance value of each point as the weight, perform weighted summation on the feature values, and obtain the query feature under the query coordinates:
[0076]
[0077] Among them, v q represents the query feature under the query coordinates, k represents the number of points on the scene map, G represents the Gaussian distance, (x i ,y i ) represents the coordinates of the i-th point on the scene feature map, (x q ,y q ) represents the query coordinates, M i Represents the features of the i-th point on the scene feature map.
[0078] S3-3: Input the query features and audio features into the autoencoder to predict the RIR of the query coordinates.
[0079] In general, existing methods still cannot fully understand and utilize the small amount of example data provided, which limits the performance of acoustic learning. Either the environment is perceived from the original RGB-D image and the target RIR is predicted in an end-to-end manner, or the local features are implicitly modeled through Nerf to directly predict the target RIR. The former weakens the posture information through sinusoidal position encoding and ignores the connection between each position, while the latter cannot utilize complete visual information, especially semantic information, which limits the model's ability to understand space and cannot be well generalized to unknown scenes. In contrast, the present invention innovatively proposes a map-guided scene perception solution, which extracts acoustically related semantic features from different visual observations, and uses deep mapping to align and fuse features to obtain feature information of the entire scene, thereby enhancing the model's understanding of spatial information such as house structure, house layout, and object semantics. In addition, the present invention makes full use of the features of each point in the scene feature map with the help of Gaussian distance weighted calculation method, which significantly improves the model's prediction accuracy for RIR.
[0080] In summary, through feature extraction fusion and sampling, the present invention can significantly improve the RIR prediction accuracy in the few-sample acoustic modeling task, especially in the feature sparse area. The model can simultaneously utilize the features of the entire scene map, making the prediction results more accurate and comprehensive.
[0081] Example 2
[0082] This embodiment provides a map-guided scene-aware few-sample acoustic modeling device, including:
[0083] A semantic feature extraction module is used to obtain image data, extract visual features based on the image data, and extract scene semantic features from the visual features;
[0084] The scene feature alignment module is used to align and fuse the scene features obtained from different postures according to the extracted scene semantic features to obtain a scene feature map;
[0085] The prediction module is used to obtain the query coordinates, obtain the query features according to the query coordinates and the scene feature map, and obtain the target RIR according to the query features.
[0086] Since the device is a map-guided scene-aware few-sample acoustic modeling device of an embodiment of the present invention, and the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0087] Example 3
[0088] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A map-guided scene-aware few-shot acoustic modeling method is shown.
[0089] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0090] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.
[0091] Since the electronic device is an electronic device corresponding to a map-guided scene-aware few-sample acoustic modeling method of an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0092] Example 4
[0093] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A map-guided scene-aware few-shot acoustic modeling method is shown.
[0094] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0095] Since the storage medium is a storage medium corresponding to a map-guided scene-aware few-sample acoustic modeling method of an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0096] Example 5
[0097] In some possible implementations, various aspects of the method of the embodiment of the present invention may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of a map-guided scene-aware few-sample acoustic modeling method according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing various embodiments can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0098] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0099] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0100] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.
Claims
1. A map-guided scene-aware few-sample acoustic modeling method, characterized in that: The following steps are involved: Acquire image data, extract visual features according to the image data, and extract scene semantic features from the visual features; According to the extracted scene semantic features, the scene features obtained from different postures are aligned and fused to obtain a scene feature map; the query coordinates are obtained, the query features are obtained according to the query coordinates and the scene feature map, and the target RIR is obtained according to the query features.
2. The map-guided scene-aware few-sample acoustic modeling method according to claim 1, characterized in that: The step of extracting scene semantic features from visual features includes: The U-Net model pre-trained on the semantic segmentation task is used to extract semantic features related to acoustics. The U-Net model outputs the last layer of latent features that fuse low-level and high-level features, thereby providing a comprehensive representation of environmental details: v=f v (V) In the formula, V represents the RGB image in the image data, f v represents the pre-trained U-Net feature extraction model, and v represents the extracted pixel-level acoustic semantic features that are consistent with the input image in length and width.
3. The map-guided scene-aware few-sample acoustic modeling method according to claim 1, characterized in that: The step of aligning and fusing scene features obtained from different postures according to the extracted scene semantic features to obtain a scene feature map includes: With the help of depth mapping technology and the depth map in the image data, the semantic features are mapped to the top view corresponding to the features: T=D P (v,D) Where T represents the top view corresponding to the feature, D P represents the depth map, D represents the depth map in the observation information, and v represents the extracted pixel-level acoustic semantic features that are consistent with the input image in length and width; With the help of the posture information of each observation point, the top view is rotated and translated so that the top view has a unified world coordinate: T ′ =T R (T,P) Where, T ′ It represents the top view corresponding to each observed feature after alignment, T R represents the rotation and translation operation, P represents the posture information corresponding to the observation point, and T represents the top view corresponding to each observation feature; The aligned top views are fused to retain the most significant results, thus obtaining a feature map of the entire scene: M=F(T ′ ) In the formula, M represents the top view after fusion, F represents the fusion operation, and T ′ It represents the top view corresponding to each observed feature after alignment.
4. The map-guided scene-aware few-sample acoustic modeling method according to claim 1, characterized in that: The obtaining of query coordinates, obtaining query features according to the query coordinates and the scene feature map, and obtaining a target RIR according to the query features include: Get the query coordinates and calculate the Gaussian distance between the query coordinates and each point on the scene feature map; The Gaussian distance value of each point is used as the weight to perform weighted summation on the feature values to obtain the query feature under the query coordinates; The query features are combined with the audio features and input into the autoencoder to predict the RIR of the query coordinates.
5. The map-guided scene-aware few-sample acoustic modeling method according to claim 4, characterized in that: The calculation formula of the Gaussian distance is: In the formula, G represents the Gaussian distance, (x, y) represents the coordinates of each point on the scene feature map, (x q ,y q ) represents the query coordinates and σ represents a trainable parameter.
6. The map-guided scene-aware few-sample acoustic modeling method according to claim 4, characterized in that: The calculation formula of the query feature is: In the formula, v q represents the query feature under the query coordinates, k represents the number of points on the scene map, G represents the Gaussian distance, (x i ,y i ) represents the coordinates of the i-th point on the scene feature map, (x q ,y q ) represents the query coordinates, M i Represents the features of the i-th point on the scene feature map.
7. A map-guided scene-aware few-sample acoustic modeling device, characterized in that: include: A semantic feature extraction module is used to obtain image data, extract visual features based on the image data, and extract scene semantic features from the visual features; The scene feature alignment module is used to align and fuse the scene features obtained from different postures according to the extracted scene semantic features to obtain a scene feature map; The prediction module is used to obtain the query coordinates, obtain the query features according to the query coordinates and the scene feature map, and obtain the target RIR according to the query features.
8. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Long-distance periodic GIS / GIL line breakdown sound pressure level positioning and point distribution method
CN113221414A
Visual language navigation method and device based on open scene map and medium
CN116499471A
Estimation of acoustic parameters for audio system based on stored information about acoustic model
US11598962B1