An intelligent sound control method and system based on sound field adaptive adjustment
By combining real-time audio analysis and visual sensing data of the smart speaker system, the sound field is dynamically adjusted, solving the problem that the smart speaker system cannot adapt to changes in user location and environment, thus improving the sound quality and user experience of the speaker system.
Patent Information
- Application Number
- CN202510995158.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-07-18
AI Technical Summary
Smart speaker systems cannot accurately adjust the sound field according to the user's real-time location and environmental changes, which means they cannot meet the user's personalized sound quality and sound effects needs in different locations or environments.
By extracting real-time audio streams based on a preset audio window for time-frequency analysis, matching real-time sound field templates, and combining binocular visual sensing data for user spatial pose coordinate positioning, a four-dimensional sound image space vector is output. Sound field mapping calculations and low-latency updates are then performed to achieve dynamic adjustment of the sound field.
It enables precise sound field adjustment of the smart speaker system according to changes in user location and environment, enhancing the user's immersive experience and personalized sound quality needs.
Smart Images

Figure CN120848190B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio control technology, specifically to an intelligent audio control method and system based on adaptive sound field adjustment. Background Technology
[0002] Smart speaker systems typically cannot precisely adjust the sound field based on the user's real-time location and environmental changes. In traditional speaker systems, sound field adjustment is usually fixed and cannot adapt to changes in the user's spatial posture or the dynamic changes in the room's acoustic environment. This lack of real-time adaptability limits the immersive experience of the speaker system in different scenarios and makes it difficult to meet the personalized sound quality and effects needs of users in different locations or environments. Summary of the Invention
[0003] This application provides an intelligent audio control method and system based on adaptive sound field adjustment, which addresses the technical problem in the prior art that it is impossible to adjust the sound field in real time according to changes in user position and environment.
[0004] In view of the above problems, this application provides an intelligent audio control method and system based on sound field adaptive adjustment.
[0005] The first aspect of this application provides a smart speaker control method based on adaptive sound field adjustment, the method comprising:
[0006] After capturing the real-time audio stream based on a preset audio window, a real-time sound field template is obtained by performing time-frequency analysis on the real-time audio stream. The user's real-time spatial pose coordinates are located based on binocular visual sensing data. Using the real-time spatial pose as a reference, sound wave propagation attenuation analysis is performed in the listening area to output a four-dimensional sound image space vector. The four-dimensional sound image space vector is used to perform spatial sound field mapping operation on the real-time sound field template to obtain the real-time scene sound field. After loading the real-time scene sound field into the smart speaker, the sound field parameters are pre-calculated based on trajectory prediction according to the spatial pose sequence obtained by the camera skeleton time-series tracking, and the real-time scene sound field is updated with low latency based on the calculation results.
[0007] A second aspect of this application provides an intelligent audio control system based on adaptive sound field adjustment, the system comprising:
[0008] The system includes a time-frequency analysis module for capturing real-time audio streams through a preset audio window and then performing time-frequency analysis to obtain a real-time sound field template; a coordinate positioning module for locating the user's real-time spatial pose coordinates based on binocular visual sensing data; a spatial vector positioning module for performing sound wave propagation attenuation analysis in the listening area based on the real-time spatial pose and outputting a four-dimensional sound image spatial vector; a mapping operation module for performing spatial sound field mapping operation on the real-time sound field template using the four-dimensional sound image spatial vector to obtain the real-time scene sound field; and a low-latency update module for loading the real-time scene sound field into the smart speaker, performing pre-calculation of sound field parameters based on trajectory prediction according to the spatial pose sequence obtained by time-series tracking of the camera skeleton, and performing low-latency updates of the real-time scene sound field based on the calculation results.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0010] This application captures a real-time audio stream through a preset audio window, performs time-frequency analysis on the real-time audio stream, and matches it to obtain a real-time sound field template. It then locates the user's real-time spatial pose coordinates based on binocular visual sensing data. Using the real-time spatial pose as a reference, it performs sound wave propagation attenuation analysis in the listening area, outputting a four-dimensional sound image space vector. The four-dimensional sound image space vector is used to perform spatial sound field mapping calculations on the real-time sound field template to obtain the real-time scene sound field. After loading the real-time scene sound field into the smart speaker, it performs pre-calculation of sound field parameters based on trajectory prediction according to the spatial pose sequence obtained from the camera skeleton time-series tracking, and updates the real-time scene sound field with low latency based on the calculation results. This invention solves the technical problem in the prior art that it cannot adjust the sound field in real time according to changes in user position and environment. By combining real-time audio stream analysis, spatial pose positioning, and four-dimensional sound image space mapping, it achieves the technical effect of precise and dynamic adjustment of the sound field. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A schematic flowchart of an intelligent speaker control method based on adaptive sound field adjustment provided in an embodiment of this application;
[0013] Figure 2 This is a schematic diagram of an intelligent audio control system based on adaptive sound field adjustment, provided as an embodiment of this application.
[0014] Figure labeling: Time-frequency analysis module 11, coordinate positioning module 12, spatial vector positioning module 13, mapping operation module 14, low-latency update module 15. Detailed Implementation
[0015] This application provides an intelligent speaker control method and system based on adaptive sound field adjustment, which addresses the technical problem in the prior art that it is impossible to adjust the sound field in real time according to changes in user position and environment. By combining real-time audio stream analysis, spatial pose positioning and four-dimensional sound image spatial mapping, it achieves the technical effect of precise and dynamic adjustment of the sound field.
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0017] It should be noted that any variation of the terms "comprising" and "having" is intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products, or devices.
[0018] Example 1, as Figure 1 As shown, this application provides an intelligent speaker control method based on adaptive sound field adjustment, the method comprising:
[0019] Step S100: After capturing the real-time audio stream based on the preset audio window, a real-time sound field template is obtained by performing time-frequency analysis on the real-time audio stream.
[0020] In this embodiment, the real-time audio stream is first segmented into multiple segmented audio streams by using 1 / K of a preset audio window as the sliding segmentation window and setting a sliding step size. Next, multi-dimensional feature vectors are extracted from each segmented audio stream, and these feature vectors are compared with a music type vector library to obtain confidence values for multiple music types. Based on the temporal relationship of the audio streams, temporal fitting analysis is performed on these music type confidence values to output the real-time music type. Finally, a real-time sound field template is generated by calling the corresponding standard sound field template based on the real-time music type.
[0021] Furthermore, in the method provided in the application embodiments, after capturing the real-time audio stream based on a preset audio window, a real-time sound field template is obtained by performing time-frequency analysis on the real-time audio stream and matching it. The method further includes:
[0022] Using 1 / K of the preset audio window as a sliding segmentation window; constrained by a preset sliding step size, the real-time audio stream is segmented into multiple segmented audio streams using the sliding segmentation window; multiple multi-dimensional segmentation feature vectors of the multiple segmented audio streams are extracted; the multiple multi-dimensional segmentation feature vectors are compared with a music type vector library to obtain multiple sets of music type confidence values; based on the temporal relationship of the multiple segmented audio streams, a music type temporal fitting analysis is performed on the multiple sets of music type confidence values to output the real-time music type; a standard sound field template bound to the real-time music type is called as the real-time sound field template.
[0023] In this embodiment, the real-time audio stream is first segmented by using 1 / K of a preset audio window as the sliding segmentation window and constrained by a preset sliding step size. Here, 1 / K represents the proportion of the audio window to the total length of the audio stream, and K is a fixed constant value to ensure that the audio segments obtained from each segmentation are relatively evenly sized. The sliding segmentation window continuously divides the audio stream into multiple smaller segments. The sliding step size refers to the time interval between each window slide. By using a preset step size, overlapping portions between the segmented audio segments can be ensured, thereby capturing continuous changes in the audio signal. Through this step, the real-time audio stream is divided into multiple segmented audio streams.
[0024] Next, multiple multidimensional segmented feature vectors are extracted from the segmented audio streams. Specifically, a parallel feature processing framework is pre-constructed, including a temporal dynamic feature extraction unit, a spectral structure feature extraction unit, a harmonic structure feature extraction unit, and a rhythmic texture feature extraction unit. Each segmented audio stream is input into this framework, and after parallel processing by these four units, features of the audio signal in the time domain, frequency domain, and rhythm are extracted respectively. Through the processing of these parallel units, information including the frequency components, volume fluctuations, timbre, and rhythm of the audio is extracted from each audio segment, resulting in multiple multidimensional segmented feature vectors.
[0025] The extracted multi-dimensional segmented feature vectors are then compared with a music genre vector library to obtain multiple sets of confidence values for each music genre. The comparison process is accomplished by calculating the similarity between the feature vector of each segmented audio stream and the music genre template. Cosine similarity or other similarity metrics are used to quantify the degree of matching between the audio segment and the music genre template. After obtaining the similarity between each segmented audio stream and the music genre template, these similarity values are converted into confidence values. Confidence values represent the degree of matching between the audio stream and a specific music genre, and are obtained by normalizing the similarity values. Through this process, multiple sets of confidence values for each music genre are obtained.
[0026] Then, based on the temporal relationship of multiple segmented audio streams, a temporal fitting analysis of music genre confidence values is performed on multiple sets of music genres. The temporal fitting analysis constructs a time-dimensional matrix, where the horizontal axis represents time and the vertical axis represents each music genre. Each value in the matrix is the music genre confidence value at the corresponding time point. Specifically, first, a temporal matrix is constructed, filling it with multiple confidence values for each audio segment. The horizontal axis represents the time step, and the vertical axis represents the music genre. Each element in the matrix represents the confidence value of a specific music genre at a given time point. For example, assuming there are 5 time steps and 3 music genres (rock, jazz, classical), the matrix size is 5×3. Each column represents the confidence value of different music genres at each time point, and the rows represent the confidence values of different music genres at each time point. Next, a fitting analysis is performed, analyzing each column of the temporal matrix. Through simple linear fitting, the changing trend of each music genre on the time axis is captured. Linear fitting fits the relationship between the time step and the confidence value into a straight line, thereby predicting the music genre of the audio stream at future time points. For example, if the confidence value for rock music gradually increases over the past few time points, it predicts that the following audio segment will still belong to rock music. Finally, through time series fitting analysis, based on the changes in the confidence value of each music genre on the time axis, the real-time music genre of the current audio stream is output. If the confidence value of a certain genre shows strong continuity and stability in the time series analysis, it is identified as the real-time music genre of the current audio stream. For example, if the confidence value for rock music is stable and high throughout the entire time series matrix, then the audio stream is determined to belong to rock music.
[0027] Finally, based on the results of the time-series fitting analysis, the real-time music type is output. For example, if the time-series fitting analysis determines that the audio stream belongs to rock music, then rock music is output as the real-time music type. Subsequently, the corresponding standard sound field template is called according to this real-time music type to adjust the sound field of the audio equipment. The standard sound field template is a parameter configuration designed for different music types, which can adjust the audio system according to the characteristics of each music type (such as low frequencies, timbre, etc.) to ensure that the sound effect is optimally optimized when playing a specific type of music.
[0028] Furthermore, in the method provided in the application embodiments, extracting multiple multidimensional segmented feature vectors of the multiple segmented audio streams further includes:
[0029] A pre-constructed parallel feature processing framework is provided, comprising a temporal dynamic feature extraction unit, a spectral structure feature extraction unit, a harmonic structure feature extraction unit, and a rhythmic texture feature extraction unit. A first segmented audio stream is input into the parallel feature processing framework. After the temporal dynamic feature extraction unit, spectral structure feature extraction unit, harmonic structure feature extraction unit, and rhythmic texture feature extraction unit perform feature extraction operations in parallel, the feature extraction results are normalized to generate a first multidimensional segmented feature vector. This process is repeated to extract the multiple multidimensional segmented feature vectors of the multiple segmented audio streams using the parallel feature processing framework.
[0030] In this embodiment, a parallel feature processing framework is first pre-constructed, including a temporal dynamic feature extraction unit, a spectral structure feature extraction unit, a harmonic structure feature extraction unit, and a rhythmic texture feature extraction unit. These feature extraction units are used to extract features of the audio signal from different dimensions. The temporal dynamic feature extraction unit utilizes the instantaneous changes in the audio signal to extract features such as volume changes and waveform fluctuations, mainly focusing on the dynamic characteristics of the audio signal over time; the spectral structure feature extraction unit converts the audio signal into the frequency domain, analyzes the frequency components and spectral structure of the audio signal, and extracts the frequency features of the audio signal; the harmonic structure feature extraction unit extracts information such as pitch and chords based on the harmony of the audio signal, focusing on the harmonic structure of the audio; the rhythmic texture feature extraction unit extracts rhythmic features from the audio signal, identifies rhythmic patterns and changes, and analyzes the rhythmic structure of the audio. These extraction units work in parallel to ensure that the audio signal is analyzed simultaneously from multiple dimensions.
[0031] After the first segment of the audio stream is input into the parallel feature processing framework, four feature extraction units analyze the audio signal in parallel, extracting time-domain, frequency-domain, harmonic structure, and rhythmic features respectively. Each feature extraction unit generates corresponding feature data, comprehensively reflecting the different characteristics of the audio segment. This process outputs multiple feature vectors of the first segment of the audio stream, which describe the time-domain waveform, frequency components, harmonic structure, and rhythmic pattern of the audio signal. Next, the extracted feature results are normalized. Normalization maps the values of all features to a uniform range (e.g., 0 to 1) to ensure consistency in the numerical ranges of different features. Through this normalization process, the first multi-dimensional segmented feature vector is obtained.
[0032] Finally, a parallel feature processing framework is used to process multiple segmented audio streams sequentially. Each segmented audio stream undergoes parallel processing by four feature extraction units, ultimately yielding multiple multidimensional segmented feature vectors for the multiple segmented audio streams.
[0033] Step S200: Determine the user's real-time spatial pose coordinates based on binocular vision sensor data.
[0034] In this embodiment, to obtain the user's real-time spatial pose coordinates, a binocular vision sensor first uses two cameras to capture the same scene from different angles, obtaining two images. Then, by comparing the pixel differences in the same areas of the two images, the depth information of each point, i.e., the distance of that point from the camera, is calculated. Then, combining the camera's position and orientation, a geometric method is used to calculate the three-dimensional coordinates of each point. Through this method, the user's position in space is tracked in real time, and the user's position and orientation are calculated, thereby obtaining the user's real-time spatial pose coordinates.
[0035] Step S300: Based on the real-time spatial pose, perform sound wave propagation attenuation analysis in the listening area and output a four-dimensional sound image space vector.
[0036] In this embodiment, when performing sound wave propagation attenuation analysis in the listening area based on real-time spatial pose, the user's real-time head coordinates are first extracted from the real-time spatial pose data. Then, the acoustic parameter set of the room (such as wall absorption coefficient, air absorption coefficient, etc.) and the coordinate information of the speaker array are called. Based on the user's head coordinates and acoustic parameters, spatial sound propagation attenuation analysis is performed, and the gradient distribution of sound pressure level and sound image width angle are output.
[0037] Subsequently, based on the coordinates of the speaker array and the user's real-time head coordinates, a direct sound vector is constructed, and the head's horizontal and vertical pitch angles are calculated using this vector as a reference. Finally, the sound pressure level gradient distribution, head horizontal angle, vertical pitch angle, and sound image width angle are integrated to generate a four-dimensional sound image space vector.
[0038] Furthermore, in the method provided in the application embodiment, based on the real-time spatial pose, sound wave propagation attenuation analysis in the listening area is performed to output a four-dimensional sound image space vector, which further includes:
[0039] Real-time head coordinates are extracted from the real-time spatial pose; the room acoustic parameter set and speaker array coordinates are locally invoked; spatial sound propagation attenuation analysis is performed based on the real-time head coordinates and room acoustic parameter set, and the sound pressure level gradient distribution and sound image width angle are output; after constructing the direct sound vector based on the speaker array coordinates and real-time head coordinates, the head horizontal direction angle and head vertical pitch angle are calculated and output based on the direct sound vector; the sound pressure level gradient distribution, head horizontal direction angle, head vertical pitch angle and sound image width angle are integrated to output the four-dimensional sound image space vector.
[0040] In this embodiment, the real-time head coordinates, i.e., the user's current position and orientation, are first extracted from the real-time spatial pose data. This step is achieved using a binocular vision sensor, and the user's head position is determined by combining the sensor's coordinate system.
[0041] Next, the room acoustic parameter set and speaker array coordinates are retrieved locally. The room acoustic parameter set includes wall sound absorption coefficients, air temperature and humidity, etc. These parameters are predetermined. Similarly, the speaker array coordinates are also pre-calibrated.
[0042] Next, spatial sound propagation attenuation analysis is performed based on real-time head coordinates and room acoustic parameter sets. Specifically, firstly, by analyzing room acoustic parameters such as wall absorption coefficients and air temperature and humidity, the reflected sound path is fitted, and the attenuation of reflected sound waves is calculated. Then, based on the coordinates of the speaker array and real-time head coordinates, the length of the direct sound path is calculated, and a spherical wave diffusion attenuation model is applied to calculate the attenuation of the direct sound pressure level along this path. Next, the attenuation of the direct sound pressure level and the attenuation of the reflected sound pressure level are fused to generate a sound pressure level gradient distribution, which reflects the change in sound intensity in space. Finally, the ratio of the reflected sound pressure level attenuation to the direct sound pressure level attenuation is calculated, and this ratio is mapped to the sound image width angle.
[0043] Next, a direct sound vector is constructed based on the coordinates of the speaker array and the real-time head coordinates. Specifically, this vector is determined by calculating the straight-line distance and direction between the speaker array and the user's head. This step is accomplished through simple geometric calculations. The straight-line distance between the two points is calculated using the Euclidean distance formula, and the direction vector from the speaker array to the user's head is calculated based on the coordinate difference between the two points. This direction vector is the direct sound vector.
[0044] Then, using the direct sound vector as a reference, the head horizontal angle and head vertical pitch angle are calculated. The horizontal angle and vertical pitch angle represent the user's head position relative to the speaker array in the horizontal and vertical directions, respectively. Using vector calculation methods, the angle between the direct sound vector and the horizontal direction (x-axis) is first calculated to obtain the head horizontal angle. Then, trigonometric functions are used to calculate the angle between the direct sound vector and the vertical direction (y-axis) to obtain the head vertical pitch angle.
[0045] Finally, by integrating the sound pressure level gradient distribution, head horizontal angle, head vertical pitch angle, and sound image width angle, a four-dimensional sound image space vector is output.
[0046] Furthermore, in the method provided in the application embodiment, spatial sound propagation attenuation analysis is performed based on the real-time head coordinates and room acoustic parameter set to output the sound pressure level gradient distribution and sound image width angle, which further includes:
[0047] Based on the room acoustic parameter set, the reflected sound path is fitted, and the reflected sound pressure level attenuation is output. After calculating the output direct sound path length according to the coordinates of the speaker array and the real-time head coordinates, the direct sound pressure level attenuation of the direct sound path length is calculated using a spherical wave diffusion attenuation model. The direct sound pressure level attenuation and the reflected sound pressure level attenuation are fused to generate the sound pressure level gradient distribution. After calculating the proportion of reflected sound pressure level to the direct sound pressure level attenuation and the reflected sound pressure level attenuation, the proportion of reflected sound pressure level is mapped to the sound image width angle.
[0048] In this embodiment, when fitting the reflected sound path based on the room acoustic parameter set, a pre-stored room acoustic parameter set is first extracted from the local parameter library. This parameter set includes a wall absorption coefficient table, an air absorption coefficient function, and a room geometric model. Next, the coordinates of the speaker array and the real-time head coordinates are projected onto the room geometric model for collaborative reflection region analysis, and the dominant reflection region is located at the main reflection surface. Based on the main reflection surface, the speaker array coordinates are converted into an equivalent virtual sound source array using the mirror method, and then the weighted center coordinates of this array are calculated as the center of the virtual sound source. Then, the wall absorption coefficient table and the air absorption coefficient function are loaded as sound absorption interference terms into the room geometric model to calculate the average reflection path length. Based on this reflection path length, the spherical wave diffusion attenuation is calculated, reflecting the sound propagation attenuation in space. Then, the main absorption coefficient is calculated based on the regional material of the dominant reflection region and the wall absorption coefficient table, thereby obtaining the additional wall absorption attenuation. Simultaneously, the air absorption attenuation is calculated based on the average reflection path length and the air absorption coefficient function. Finally, the attenuation of spherical wave diffusion, air absorption, and wall sound absorption are superimposed to obtain the attenuation of reflected sound pressure level.
[0049] Next, the direct sound path length is calculated based on the speaker array coordinates and the real-time head coordinates. This process uses the known speaker array position and the user's head position, employing the Euclidean distance formula to calculate the straight-line distance from the speaker array to the user's head; this straight-line distance is the direct sound path length. Then, the spherical wave diffusion attenuation model is applied to calculate the direct sound pressure level attenuation along the direct sound path. The spherical wave diffusion attenuation model describes the characteristics of sound propagation in three-dimensional space, assuming that the sound originates from a point source and spreads along a sphere; the sound pressure level attenuates with increasing propagation distance. By substituting the direct sound path length between the speaker array and the user's head into the spherical wave diffusion attenuation model, the sound attenuation is calculated based on the propagation distance, thus obtaining the direct sound pressure level attenuation.
[0050] The direct sound pressure level attenuation and the reflected sound pressure level attenuation are then fused to generate a sound pressure level gradient distribution. This process combines the attenuation of reflected and direct sound to generate a distribution map representing changes in sound intensity. The fusion process uses a simple summation method, adding the attenuation of direct and reflected sound to obtain the overall sound pressure level gradient distribution.
[0051] Next, the proportion of reflected sound pressure level to the direct sound pressure level attenuation and the reflected sound pressure level attenuation is calculated. The reflected sound pressure level proportion represents the intensity ratio of the reflected sound wave to the direct sound wave, reflecting the contribution of each sound wave to the total sound intensity. The reflected sound pressure level proportion is obtained by subtracting the sum of the direct and reflected sound pressure level attenuations from the reflected sound pressure level attenuation. Finally, the reflected sound pressure level proportion is mapped to the sound image width angle.
[0052] Furthermore, in the method provided in the application embodiment, the method of fitting the reflected sound path based on the room acoustic parameter set and outputting the reflected sound pressure level attenuation also includes:
[0053] The pre-stored room acoustic parameter set is extracted from the local parameter library. This set includes a wall absorption coefficient table, an air absorption coefficient function, and a room geometric model. The coordinates of the speaker array and the real-time head coordinates are projected onto the room geometric model for collaborative reflection region analysis to locate the dominant reflection region, which is located on the primary reflection surface. Based on the primary reflection surface, the speaker array coordinates are mirrored into an equivalent virtual sound source array. The weighted center coordinates of the equivalent virtual sound source array are then calculated as the virtual sound source center. The wall absorption coefficient table and air absorption coefficient function are then used to analyze the room acoustic parameters. After the sound absorption interference is loaded into the room geometry model as a sound absorption interference term, the average reflection path length is calculated based on the virtual sound source center; the spherical wave diffusion attenuation is calculated based on the average reflection path length; the wall sound absorption coefficient table is indexed using the regional material of the dominant reflection area to obtain the main sound absorption coefficient, and the additional wall sound absorption attenuation is calculated based on the main sound absorption coefficient; the air absorption attenuation is calculated based on the average reflection path length and the air absorption coefficient function; and the superposition value of the spherical wave diffusion attenuation, the air absorption attenuation, and the additional wall sound absorption attenuation is used as the reflected sound pressure level attenuation.
[0054] In this embodiment, a room acoustic parameter set is first extracted from a local parameter library. This set includes a wall absorption coefficient table, an air absorption coefficient function, and a room geometric model. The wall absorption coefficient table lists the sound absorption characteristics of different surface materials (such as walls and ceilings) within the room, describing their ability to absorb sound waves. The air absorption coefficient function represents the energy loss caused by air friction, temperature, humidity, and other factors during sound wave propagation in the air. The room geometric model provides the room's physical structure data, such as the room's dimensions, wall locations, and their material properties.
[0055] Next, the coordinates of the speaker array and the real-time head coordinates are projected onto the room geometry model for collaborative reflection region analysis. In this process, the acoustic focusing center is first calculated based on the inverse square relationship between the speaker array coordinates and the real-time head coordinates. Then, a virtual reference ray pointing to the real-time head coordinates is constructed, and using this as a guiding direction, joint ray cluster projection is performed on the room geometry model based on the speaker array coordinates, generating multiple ray cluster arrays to simulate the sound propagation path. Next, point cloud collision spatial cluster analysis is performed on the generated ray cluster arrays. By analyzing the collision between the rays and the room's interior surfaces, the dominant reflection region is located on the primary reflection surface.
[0056] Subsequently, based on the primary reflector, the coordinates of the speaker array are mirrored to form an equivalent virtual sound source array. This process simulates the source of reflected sound waves by geometrically mirroring the speaker array relative to the primary reflector, calculating the mirrored position of the speaker array relative to the primary reflector. At this point, the position of the mirrored speaker array represents the source of the reflected sound waves and is used as the equivalent virtual sound source array. Next, the weighted center coordinates of the equivalent virtual sound source array are calculated, serving as the virtual sound source center. In this step, the coordinates of the virtual sound source center are obtained by performing an equal-weighted average calculation on the positions of all equivalent virtual sound sources.
[0057] Next, the wall sound absorption coefficient table and air absorption coefficient function are loaded as sound absorption interference terms into the room geometry model. Through physical modeling methods, these sound absorption interference terms are integrated into the room geometry model to ensure that the environmental sound absorption effect is considered when simulating sound wave propagation. Then, the average reflection path length is calculated based on the virtual sound source center. This step is accomplished by calculating the average length of the reflection path from the virtual sound source center to the user's head. Specifically, the average value of all possible reflection paths is calculated, as these paths propagate through various reflective surfaces in the room. The average reflection path length is obtained through geometric calculations.
[0058] Next, the spherical wave diffusion attenuation is calculated based on the average reflection path length. The spherical wave diffusion attenuation model is used to simulate the propagation of sound in space, where the intensity of the sound wave decreases as the propagation distance increases. The attenuation is calculated using the inverse square of the path length, as the attenuation is inversely proportional to the square of the distance. Therefore, based on the average reflection path length, the attenuation along that path is calculated, yielding the spherical wave diffusion attenuation.
[0059] Then, the primary sound absorption coefficient is obtained by indexing the wall sound absorption coefficient table using the regional material index of the dominant reflection area. The primary sound absorption coefficient reflects the sound wave absorption capacity of the primary reflection surface material. By consulting the wall sound absorption coefficient table, the corresponding sound absorption coefficient is obtained based on the material properties of the primary reflection area. The additional sound absorption attenuation of the wall is calculated using the primary sound absorption coefficient, which is the extra attenuation caused by the wall material absorbing sound. The additional sound absorption attenuation of the wall is calculated by combining the primary sound absorption coefficient with the length of the sound wave propagation path.
[0060] Simultaneously, based on the average reflection path length and the air absorption coefficient function, the air absorption attenuation is calculated. Air absorption attenuation represents the energy loss of sound waves propagating in the air due to factors such as air friction and scattering. By combining the path length with the air absorption coefficient function, the effect of air absorption on sound waves is calculated, yielding the air absorption attenuation.
[0061] Finally, the attenuation of spherical wave diffusion, air absorption, and wall sound absorption are superimposed to obtain the attenuation of reflected sound pressure level.
[0062] Furthermore, in the method provided in the application embodiment, projecting the coordinates of the speaker array and the real-time head coordinates onto the room geometry model for collaborative reflection area analysis to locate the dominant reflection area, further includes:
[0063] Based on the inverse square distance ratio array of the sound array coordinates and the real-time head coordinates, a weighted acoustic focusing center is calculated; a virtual reference ray is constructed pointing from the acoustic focusing center to the real-time head coordinates; using the virtual reference ray as the guiding direction, a joint ray cluster projection is performed on the room geometry model according to the sound array coordinates to generate a ray cluster array; point cloud collision spatial clustering analysis is performed on the ray cluster array to locate the dominant reflection region, wherein the dominant reflection region is located on the main reflection surface.
[0064] In this embodiment, the acoustic focusing center is first calculated using a weighted average based on the inverse square distance ratio between the speaker array coordinates and the real-time head coordinates. During this process, the weight of each propagation path is calculated using the inverse square distance method, based on the distance relationship between the speaker array coordinates and the real-time head coordinates. Since the path weight is inversely proportional to the square of the distance, the initial weight of each path is obtained by calculating the distance between the speaker array and the user's head. To ensure the reasonableness of the weights, the weights of all paths are normalized, that is, the weight value of each path is adjusted so that the sum of all weights is 1. Then, the coordinates of each reflection path are weighted and calculated using a weighted average method to finally obtain an acoustic focusing center.
[0065] Next, based on the calculated acoustic focus center, a virtual reference ray pointing to the real-time head coordinates is constructed. This virtual reference ray represents the path of sound propagation from the focus center to the user's head, providing the direction of sound propagation. When constructing this ray, the relative position between the acoustic focus center and the real-time head coordinates is first determined. Through geometric modeling, the straight-line distance and direction from the acoustic focus center to the user's head are calculated, and this direction is defined as the virtual reference ray.
[0066] Then, guided by a virtual reference ray, a joint ray cluster projection is performed on the room's geometric model to generate a ray cluster array. This process is accomplished using ray tracing, emitting multiple rays from the sound array and simulating their propagation along reflective surfaces within the room to generate a ray cluster array. The ray cluster array reflects sound propagation paths in different directions, comprehensively simulating the propagation characteristics of sound, including multiple reflections and refractions.
[0067] Next, point cloud collision spatial clustering analysis is performed on the ray cluster array. A collision detection algorithm is used to calculate the intersection points of each ray with various reflective surfaces in the room (such as walls, ceilings, and floors), generating point cloud data. The point cloud data is the collection of all points where rays intersect with reflective surfaces. Then, spatial clustering analysis is performed on these points, for example, using the K-means clustering algorithm. By analyzing spatial distance relationships, the point cloud data is grouped to identify the areas with the strongest sound wave reflection, i.e., the dominant reflection regions. These dominant reflection regions are located on the primary reflective surface and are where sound reflection is strongest, directly affecting the sound propagation effect.
[0068] Step S400: Use the four-dimensional sound image space vector to perform spatial sound field mapping calculation on the real-time sound field template to obtain the real-time scene sound field.
[0069] In this embodiment, when performing spatial sound field mapping calculations on a real-time sound field template using a four-dimensional sound image space vector, the sound pressure level gradient distribution, head horizontal angle, head vertical pitch angle, and sound image width angle are first extracted from the four-dimensional sound image space vector. Then, the sound pressure level gradient distribution is used to dynamically compensate the frequency response curve of the real-time sound field template, outputting a corrected frequency response curve. Next, based on the head horizontal angle and head vertical pitch angle, the sound source coordinates of the real-time sound field template are spatially remapped, thereby outputting calibrated sound image coordinates. Next, the sound field widening parameter in the real-time sound field template is mapped using a widening coefficient based on the sound image width angle, thereby outputting an extended sound field distribution. Finally, the corrected frequency response curve, calibrated sound image coordinates, and extended sound field distribution are fused to obtain the real-time scene sound field.
[0070] Furthermore, in the method provided in the application embodiment, the method uses the four-dimensional sound image space vector to perform spatial sound field mapping operation on the real-time sound field template to obtain the real-time scene sound field, and further includes:
[0071] The sound pressure level gradient distribution, head horizontal angle, head vertical pitch angle, and sound image width angle are extracted from the four-dimensional sound image space vector. The sound pressure level gradient distribution is used to dynamically compensate the frequency response curve of the real-time sound field template, and a corrected frequency response curve is output. Spatial remapping is performed on the sound source coordinates of the real-time sound field template according to the head horizontal angle and head vertical pitch angle, and calibrated sound image coordinates are output. The sound field widening parameter widening coefficient of the real-time sound field template is mapped according to the sound image width angle, and an extended sound field distribution is output. The corrected frequency response curve, calibrated sound image coordinates, and extended sound field distribution are fused to reconstruct the real-time scene sound field.
[0072] In this embodiment of the application, the sound pressure level gradient distribution, head horizontal direction angle, head vertical pitch angle and sound image width angle are first extracted from the four-dimensional sound image space vector.
[0073] Next, the extracted sound pressure level gradient distribution is used to dynamically compensate the frequency response curve of the real-time sound field template. In this step, the sound pressure response at different locations and frequency bands in the real-time sound field template is analyzed. Based on the gradient distribution, it is identified which frequencies have significant attenuation in space. The gain parameters of these frequency bands are dynamically adjusted using frequency response filtering methods (such as minimum phase filters), thereby outputting a corrected frequency response curve with spatial equalization characteristics.
[0074] Next, based on the head's horizontal and vertical pitch angles, the sound source coordinates in the real-time sound field template are spatially remapped. Using the user's head angle data, the position of the sound source in three-dimensional space is adjusted to align with the user's head orientation. Specifically, based on the relative relationship between the head angle and the sound source position, a geometric transformation method is used to adjust the sound source coordinates in the template, outputting calibrated sound image coordinates. This process ensures that the sound source is aligned with the user's head orientation, allowing the user to perceive the sound's location more accurately.
[0075] Based on this, sound field widening processing is performed according to the sound image width angle. First, the sound image width angle is converted from an angular quantity into a normalized widening coefficient, for example, 60° is mapped to 0.8, 90° to 1.0, 120° to 1.2, etc. Then, the widening coefficient is used to adjust the spatial diffusion ratio of each channel in the template, for example, to increase the sound source expansion angle of the left and right channels and enhance the coverage density of the central channel, thereby outputting an expanded sound field distribution.
[0076] Finally, the corrected frequency response curve, calibrated sound image coordinates, and expanded sound field distribution are fused to reconstruct the real-time scene sound field. Through a data fusion algorithm, all the above information is integrated, combining frequency response compensation, sound source remapping, and sound field expansion to generate a dynamic sound field model that meets user needs.
[0077] Step S500: After loading the real-time scene sound field into the smart speaker, perform pre-calculation of sound field parameters based on trajectory prediction according to the spatial pose sequence obtained by the camera skeleton time-series tracking, and perform low-latency update of the real-time scene sound field according to the calculation results.
[0078] In this embodiment, after loading the real-time scene sound field onto the smart speaker device, the user's spatial pose sequence is obtained based on camera skeleton temporal tracking technology. This sequence includes the three-dimensional position and orientation information of key points such as the head and torso within consecutive time frames. Subsequently, a pose prediction process is performed, that is, by analyzing the motion trend of the pose sequence, the future pose point that the user is about to reach is calculated using trajectory prediction methods (such as LSTM).
[0079] Based on this, sound field parameters are pre-calculated for each future pose point. This includes predicting sound field attenuation compensation parameters (such as frequency response energy compensation and distance attenuation compensation) and sound image localization parameters (such as virtual sound source direction and azimuth modulation value) for that location based on the user's position relative to the speaker array. When the tracking frame returned by the subsequent binocular vision system matches any predicted future pose point, the corresponding pre-calculated parameters are invoked to quickly complete the low-latency update of the real-time scene sound field.
[0080] Furthermore, the method provided in the application embodiments also includes:
[0081] Based on the spatial pose sequence, pose prediction is performed, and the future pose point is output. The sound field attenuation compensation parameters and sound image positioning parameters are pre-calculated according to the future pose point. When the binocular vision return frame satisfies the future pose point, the sound field attenuation compensation parameters and sound image positioning parameters are triggered to update the real-time scene sound field with low latency.
[0082] In this embodiment, pose prediction is first performed based on a spatial pose sequence. The spatial pose sequence is a multi-dimensional sequence containing timestamps, formed by continuously acquiring the user's head position and orientation data in three-dimensional space through skeleton temporal tracking using binocular cameras. During pose prediction, a Long Short-Term Memory (LSTM) neural network model is invoked to model this spatial pose sequence, utilizing its ability to remember and predict temporal features to output a set of predicted spatial pose points for future moments. These predicted pose points include the three-dimensional position coordinates of the user's head and the corresponding orientation angle, thus obtaining a future pose point sequence for subsequent sound field parameter prediction calculations.
[0083] Subsequently, pre-calculation of sound field parameters is performed based on the output future pose points. First, according to the spatial coordinate relationship between each future pose point and each sound source unit in the speaker array, the distance of sound wave propagation from each sound source unit to the future pose point is calculated using a spherical wave propagation model, and the spatial attenuation intensity of the sound pressure level is estimated accordingly, outputting sound field attenuation compensation parameters. Next, based on the orientation angle of the future pose point, the binaural hearing parameters corresponding to that angle are retrieved from the Head-Related Transfer Function (HRTF) database to obtain sound image localization information, including inter-aural time difference and inter-aural sound intensity difference, and the corresponding sound image localization parameters are generated. After completing this step, the sound field attenuation compensation parameters and sound image localization parameters corresponding to each future pose point are obtained.
[0084] Upon receiving image feedback frames from the binocular vision system, the system extracts the current user's actual spatial pose information and compares it with the sequence of future pose points. The comparison process calculates the Euclidean distance and orientation angle difference between the current pose and each predicted pose point in 3D space to determine if a preset matching threshold is met. When a match is detected between the current pose and any future pose point, the system immediately loads the corresponding sound field attenuation compensation parameters and acoustic image localization parameters, using them for real-time scene sound field update control, thereby completing low-latency sound field parameter updates based on prediction matching.
[0085] In summary, the embodiments of this application have at least the following technical effects:
[0086] This application captures a real-time audio stream through a preset audio window, performs time-frequency analysis on the real-time audio stream, and matches it to obtain a real-time sound field template. It then locates the user's real-time spatial pose coordinates based on binocular visual sensing data. Using the real-time spatial pose as a reference, it performs sound wave propagation attenuation analysis in the listening area, outputting a four-dimensional sound image space vector. The four-dimensional sound image space vector is used to perform spatial sound field mapping calculations on the real-time sound field template to obtain the real-time scene sound field. After loading the real-time scene sound field into the smart speaker, it performs pre-calculation of sound field parameters based on trajectory prediction according to the spatial pose sequence obtained from the camera skeleton time-series tracking, and updates the real-time scene sound field with low latency based on the calculation results. This invention solves the technical problem in the prior art that it cannot adjust the sound field in real time according to changes in user position and environment. By combining real-time audio stream analysis, spatial pose positioning, and four-dimensional sound image space mapping, it achieves the technical effect of precise and dynamic adjustment of the sound field.
[0087] Example 2 is based on the same inventive concept as the intelligent speaker control method based on sound field adaptive adjustment in the foregoing examples, such as... Figure 2As shown, this application provides an intelligent audio control system based on adaptive sound field adjustment. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0088] The time-frequency analysis module 11 is used to extract the real-time audio stream based on a preset audio window, and then perform time-frequency analysis on the real-time audio stream to match and obtain a real-time sound field template; the coordinate positioning module 12 is used to locate the user's real-time spatial pose coordinates based on binocular visual sensing data; the spatial vector positioning module 13 is used to perform sound wave propagation attenuation analysis in the listening area based on the real-time spatial pose and output a four-dimensional sound image spatial vector; the mapping operation module 14 is used to perform spatial sound field mapping operation on the real-time sound field template using the four-dimensional sound image spatial vector to obtain the real-time scene sound field; the low-latency update module 15 is used to load the real-time scene sound field into the smart speaker, perform pre-calculation of sound field parameters based on trajectory prediction according to the spatial pose sequence obtained by camera skeleton time-series tracking, and perform low-latency update of the real-time scene sound field according to the calculation results.
[0089] Furthermore, the system is also used to implement the following functions:
[0090] Real-time head coordinates are extracted from the real-time spatial pose; the room acoustic parameter set and speaker array coordinates are locally invoked; spatial sound propagation attenuation analysis is performed based on the real-time head coordinates and room acoustic parameter set, and the sound pressure level gradient distribution and sound image width angle are output; after constructing the direct sound vector based on the speaker array coordinates and real-time head coordinates, the head horizontal direction angle and head vertical pitch angle are calculated and output based on the direct sound vector; the sound pressure level gradient distribution, head horizontal direction angle, head vertical pitch angle and sound image width angle are integrated to output the four-dimensional sound image space vector.
[0091] Furthermore, the system is also used to implement the following functions:
[0092] Based on the room acoustic parameter set, the reflected sound path is fitted, and the reflected sound pressure level attenuation is output. After calculating the output direct sound path length according to the coordinates of the speaker array and the real-time head coordinates, the direct sound pressure level attenuation of the direct sound path length is calculated using a spherical wave diffusion attenuation model. The direct sound pressure level attenuation and the reflected sound pressure level attenuation are fused to generate the sound pressure level gradient distribution. After calculating the proportion of reflected sound pressure level to the direct sound pressure level attenuation and the reflected sound pressure level attenuation, the proportion of reflected sound pressure level is mapped to the sound image width angle.
[0093] Furthermore, the system is also used to implement the following functions:
[0094] The pre-stored room acoustic parameter set is extracted from the local parameter library. This set includes a wall absorption coefficient table, an air absorption coefficient function, and a room geometric model. The coordinates of the speaker array and the real-time head coordinates are projected onto the room geometric model for collaborative reflection region analysis to locate the dominant reflection region, which is located on the primary reflection surface. Based on the primary reflection surface, the speaker array coordinates are mirrored into an equivalent virtual sound source array. The weighted center coordinates of the equivalent virtual sound source array are then calculated as the virtual sound source center. The wall absorption coefficient table and air absorption coefficient function are then used to analyze the room acoustic parameters. After the sound absorption interference is loaded into the room geometry model as a sound absorption interference term, the average reflection path length is calculated based on the virtual sound source center; the spherical wave diffusion attenuation is calculated based on the average reflection path length; the wall sound absorption coefficient table is indexed using the regional material of the dominant reflection area to obtain the main sound absorption coefficient, and the additional wall sound absorption attenuation is calculated based on the main sound absorption coefficient; the air absorption attenuation is calculated based on the average reflection path length and the air absorption coefficient function; and the superposition value of the spherical wave diffusion attenuation, the air absorption attenuation, and the additional wall sound absorption attenuation is used as the reflected sound pressure level attenuation.
[0095] Furthermore, the system is also used to implement the following functions:
[0096] Based on the inverse square distance ratio array of the sound array coordinates and the real-time head coordinates, a weighted acoustic focusing center is calculated; a virtual reference ray is constructed pointing from the acoustic focusing center to the real-time head coordinates; using the virtual reference ray as the guiding direction, a joint ray cluster projection is performed on the room geometry model according to the sound array coordinates to generate a ray cluster array; point cloud collision spatial clustering analysis is performed on the ray cluster array to locate the dominant reflection region, wherein the dominant reflection region is located on the main reflection surface.
[0097] Furthermore, the system is also used to implement the following functions:
[0098] The sound pressure level gradient distribution, head horizontal angle, head vertical pitch angle, and sound image width angle are extracted from the four-dimensional sound image space vector. The sound pressure level gradient distribution is used to dynamically compensate the frequency response curve of the real-time sound field template, and a corrected frequency response curve is output. Spatial remapping is performed on the sound source coordinates of the real-time sound field template according to the head horizontal angle and head vertical pitch angle, and calibrated sound image coordinates are output. The sound field widening parameter widening coefficient of the real-time sound field template is mapped according to the sound image width angle, and an extended sound field distribution is output. The corrected frequency response curve, calibrated sound image coordinates, and extended sound field distribution are fused to reconstruct the real-time scene sound field.
[0099] Furthermore, the system is also used to implement the following functions:
[0100] Based on the spatial pose sequence, pose prediction is performed, and the future pose point is output. The sound field attenuation compensation parameters and sound image positioning parameters are pre-calculated according to the future pose point. When the binocular vision return frame satisfies the future pose point, the sound field attenuation compensation parameters and sound image positioning parameters are triggered to update the real-time scene sound field with low latency.
[0101] Furthermore, the system is also used to implement the following functions:
[0102] Using 1 / K of the preset audio window as a sliding segmentation window; constrained by a preset sliding step size, the real-time audio stream is segmented into multiple segmented audio streams using the sliding segmentation window; multiple multi-dimensional segmentation feature vectors of the multiple segmented audio streams are extracted; the multiple multi-dimensional segmentation feature vectors are compared with a music type vector library to obtain multiple sets of music type confidence values; based on the temporal relationship of the multiple segmented audio streams, a music type temporal fitting analysis is performed on the multiple sets of music type confidence values to output the real-time music type; a standard sound field template bound to the real-time music type is called as the real-time sound field template.
[0103] Furthermore, the system is also used to implement the following functions:
[0104] A pre-constructed parallel feature processing framework is provided, comprising a temporal dynamic feature extraction unit, a spectral structure feature extraction unit, a harmonic structure feature extraction unit, and a rhythmic texture feature extraction unit. A first segmented audio stream is input into the parallel feature processing framework. After the temporal dynamic feature extraction unit, spectral structure feature extraction unit, harmonic structure feature extraction unit, and rhythmic texture feature extraction unit perform feature extraction operations in parallel, the feature extraction results are normalized to generate a first multidimensional segmented feature vector. This process is repeated to extract the multiple multidimensional segmented feature vectors of the multiple segmented audio streams using the parallel feature processing framework.
[0105] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0106] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0107] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. An intelligent sound control method based on sound field adaptive adjustment, characterized in that, The method comprises: After intercepting the real-time audio stream based on a preset audio window, the real-time sound field template is matched by performing time-frequency analysis on the real-time audio stream; The real-time spatial pose coordinates of the user are located according to the binocular vision sensing data; Based on the real-time spatial pose, the sound wave propagation attenuation analysis of the listening area is performed, and the four-dimensional sound image space vector is output; The real-time sound field template is subjected to spatial sound field mapping operation by using the four-dimensional sound image space vector, and the real-time scene sound field is obtained; After loading the real-time scene sound field to the intelligent sound, the space pose sequence obtained according to the camera skeleton time sequence tracking is used to perform sound field parameter precalculation based on trajectory prediction, and the low-delay update of the real-time scene sound field is performed according to the calculation result; Wherein, based on the real-time spatial pose, the sound wave propagation attenuation analysis of the listening area is performed, and the four-dimensional sound image space vector is output, the method comprises: The real-time head coordinates are extracted from the real-time spatial pose; The room acoustic parameter set and the sound array coordinates are called locally; According to the real-time head coordinates and the room acoustic parameter set, the spatial sound propagation attenuation analysis is performed, and the sound pressure level gradient distribution and the sound image width angle are output; After constructing the direct sound vector according to the sound array coordinates and the real-time head coordinates, the direct sound vector is taken as the reference to solve and output the head horizontal direction angle and the head vertical pitch angle; The sound pressure level gradient distribution, the head horizontal direction angle, the head vertical pitch angle and the sound image width angle are integrated to output the four-dimensional sound image space vector; Wherein, the real-time sound field template is subjected to spatial sound field mapping operation by using the four-dimensional sound image space vector, and the real-time scene sound field is obtained, the method comprises: The sound pressure level gradient distribution, the head horizontal direction angle, the head vertical pitch angle and the sound image width angle are extracted from the four-dimensional sound image space vector; The real-time sound field template is subjected to frequency response curve dynamic compensation by using the sound pressure level gradient distribution, and the corrected frequency response curve is output; According to the head horizontal direction angle and the head vertical pitch angle, the sound source coordinates of the real-time sound field template are subjected to spatial remapping, and the calibrated sound image coordinates are output; According to the sound image width angle, the sound field widening parameters of the real-time sound field template are subjected to widening coefficient mapping, and the expanded sound field distribution is output; The real-time scene sound field is reconstructed by fusing the corrected frequency response curve, the calibrated sound image coordinates and the expanded sound field distribution.
2. The intelligent sound control method based on sound field self-adaptive adjustment according to claim 1, wherein, According to the real-time head coordinates and the room acoustic parameter set, the spatial sound propagation attenuation analysis is performed, and the sound pressure level gradient distribution and the sound image width angle are output, the method comprises: Based on the room acoustic parameter set, the reflected sound path fitting is performed, and the reflected sound pressure level attenuation is output; After calculating and outputting the direct sound path length according to the sound array coordinates and the real-time head coordinates, the direct sound pressure level attenuation of the direct sound path length is calculated by applying the spherical wave diffusion attenuation model; The direct sound pressure level attenuation and the reflected sound pressure level attenuation are fused to generate the sound pressure level gradient distribution; After calculating the reflected sound pressure level proportion of the direct sound pressure level attenuation and the reflected sound pressure level attenuation, the reflected sound pressure level proportion is mapped to the sound image width angle. 3.The smart sound control method based on sound field self-adaptive adjustment according to claim 2, wherein, The method comprises: extracting a pre-stored set of room acoustic parameters from a local parameter library, wherein the set of room acoustic parameters comprises a wall sound absorption coefficient table, an air absorption coefficient function, and a room geometry model; projecting the sound array coordinates and real-time head coordinates to the room geometry model for collaborative reflection area analysis to locate a dominant reflection area, wherein the dominant reflection area is on a main reflection surface; mirroring the sound array coordinates to an equivalent virtual sound source array based on the main reflection surface, and then calculating the weighted center coordinates of the equivalent virtual sound source array as a virtual sound source center; loading the wall sound absorption coefficient table and the air absorption coefficient function as sound absorption interference terms to the room geometry model, and then calculating the average reflection path length based on the virtual sound source center; calculating the spherical wave diffusion attenuation based on the average reflection path length; obtaining the main sound absorption coefficient by indexing the wall sound absorption coefficient table with the area material of the dominant reflection area, and then calculating the wall sound absorption additional attenuation based on the main sound absorption coefficient; calculating the air absorption attenuation based on the average reflection path length and the air absorption coefficient function; superimposing the spherical wave diffusion attenuation, the air absorption attenuation, and the wall sound absorption additional attenuation to obtain the reflected sound pressure level attenuation.
4. The intelligent sound control method based on sound field self-adaptive adjustment according to claim 3, characterized in that, The method comprises: calculating the acoustic focus center based on the inverse square distance array of the sound array coordinates and real-time head coordinates; constructing a virtual reference ray pointing to the real-time head coordinates from the acoustic focus center; projecting the sound array coordinates on the room geometry model to generate a ray cluster array based on the virtual reference ray as a guide direction; performing point cloud collision space clustering analysis on the ray cluster array to locate the dominant reflection area, wherein the dominant reflection area is on a main reflection surface.
5. The intelligent sound control method based on sound field adaptive adjustment according to claim 1, characterized in that, The method further comprises: performing pose pre-judgment based on the spatial pose sequence to output a future pose point; pre-calculating sound field attenuation compensation parameters and sound image positioning parameters based on the future pose point; triggering low-delay updates of the sound field attenuation compensation parameters and sound image positioning parameters to the real-time scene sound field when the binocular vision feedback frame meets the future pose point.
6. The intelligent sound control method based on sound field self-adaptive adjustment according to claim 1, characterized in that, The method comprises: taking a real-time audio stream based on a preset audio window, and then matching a real-time sound field template by performing time-frequency analysis on the real-time audio stream, the method comprising: taking 1 / K of the preset audio window as a sliding segmentation window; segmenting the real-time audio stream into multiple segmented audio streams using the sliding segmentation window with a preset sliding step as a constraint; extracting multiple multi-dimensional segmented feature vectors of the multiple segmented audio streams; comparing the multiple multi-dimensional segmented feature vectors with a music type vector library to obtain multiple groups of music type confidence values; performing music type time sequence fitting analysis on the multiple groups of music type confidence values based on the time sequence relationship of the multiple segmented audio streams to output a real-time music type. A standard sound field template bound to the real-time music type is called as the real-time sound field template.
7. The intelligent sound control method based on sound field adaptive adjustment according to claim 6, characterized in that, A plurality of multi-dimensional segmented feature vectors of the plurality of segmented audio streams are extracted, and the method comprises: A pre-constructed parallel feature processing framework is constructed, wherein the parallel feature processing framework comprises a time domain dynamic feature extraction unit, a spectral structure feature extraction unit, a harmonic structure feature extraction unit, and a rhythm texture feature extraction unit; A first segmented audio stream is input into the parallel feature processing framework, and after performing feature extraction operations in parallel via the time domain dynamic feature extraction unit, the spectral structure feature extraction unit, the harmonic structure feature extraction unit, and the rhythm texture feature extraction unit, unitization processing of the feature extraction results is performed to generate a first multi-dimensional segmented feature vector; By analogy, the plurality of multi-dimensional segmented feature vectors of the plurality of segmented audio streams are extracted through the parallel feature processing framework.
8. An intelligent sound system control system based on sound field adaptive adjustment, characterized in that, The system is used to execute an intelligent sound control method based on sound field adaptive adjustment according to any one of claims 1-7, and the system comprises: A time-frequency analysis module is configured to, after intercepting a real-time audio stream based on a preset audio window, match a real-time sound field template by performing time-frequency analysis on the real-time audio stream; A coordinate positioning module is configured to position a real-time spatial pose coordinate of a user based on binocular vision sensing data; A spatial vector positioning module is configured to, based on the real-time spatial pose, perform listening area sound wave propagation attenuation analysis and output a four-dimensional sound image spatial vector; A mapping operation module is configured to perform spatial sound field mapping operation on the real-time sound field template using the four-dimensional sound image spatial vector to obtain a real-time scene sound field; A low-delay update module is configured to, after loading the real-time scene sound field to an intelligent sound, perform sound field parameter pre-computation based on trajectory prediction according to a spatial pose sequence obtained by camera skeleton time sequence tracking, and perform low-delay update of the real-time scene sound field according to the calculation result.
Citation Information
Patent Citations
Intelligent sound equipment sound field adjusting method and intelligent sound equipment system
CN119136140A
Sound effect optimization method and device based on coaxial loudspeaker echo wall sound box
CN119485107A