Sound effect configuration method and device, and electronic device
Patent Information
- Application Number
- CN202511075274.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-07-31
AI Technical Summary
然而,在复杂声学场景下,难以实现精准的空间声场重构,导致音效配置效果欠佳
Smart Images

Figure CN120929041B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, specifically to a sound effect configuration method, apparatus, and electronic device. Background Technology
[0002] Current spatial sound effects are primarily configured based on sound content characteristics or target object location information. However, in complex acoustic scenarios, it is difficult to achieve accurate spatial sound field reconstruction, resulting in suboptimal sound effect configuration. Summary of the Invention
[0003] In view of the above problems, this disclosure provides a sound effect configuration method, apparatus and electronic device.
[0004] According to a first aspect of this disclosure, a sound effect configuration method is provided, comprising obtaining feature information of a first part and a second part of a target object; determining the relative positional relationship between the first part and the second part based on the feature information of the first part and the second part; and determining sound effect parameters of a sound source based on the relative positional relationship between the first part and the second part, wherein the sound effect parameters of the sound source are used to configure the spatial sound effect signal of the sound source relative to the target object.
[0005] According to an embodiment of this disclosure, determining the sound effect parameters of a sound source based on the relative positional relationship between the first part and the second part includes: in response to the relative positional relationship between the first part and the second part satisfying a first directional relationship, determining the first directional data corresponding to the sound source based on the original directional data of the sound source; and determining the sound effect parameters based on the first directional data.
[0006] According to an embodiment of this disclosure, determining the sound effect parameters of a sound source based on the relative positional relationship between the first part and the second part further includes: in response to the relative positional relationship between the first part and the second part satisfying a second orientation relationship, calculating the rotation information of the first part based on feature information; adjusting the original orientation data of the sound source based on the rotation information to obtain the second orientation data corresponding to the sound source, and determining the sound effect parameters based on the second orientation data.
[0007] According to embodiments of this disclosure, determining the sound effect parameters of a sound source based on the relative positional relationship between the first part and the second part further includes: in response to the relative positional relationship between the first part and the second part satisfying a third positional relationship, determining the geometric relationship parameters of the first part and the second part based on feature information; adjusting the original orientation data of the sound source based on the geometric relationship parameters to obtain the third positional data corresponding to the sound source, and determining the sound effect parameters based on the third positional data.
[0008] According to embodiments of this disclosure, determining the sound effect parameters of a sound source based on the relative positional relationship between the first and second parts includes: in response to a target object moving from the first position to the second position, updating the origin of the target coordinate system based on feature information, wherein the target coordinate system is defined based on the target object; determining the fourth azimuth data of the sound source in the updated target coordinate system based on the original azimuth data of the sound source, and determining the sound effect parameters based on the fourth azimuth data.
[0009] According to embodiments of this disclosure, obtaining feature information of a first part and a second part of a target object includes: extracting first feature information of the target object using a first algorithm; determining second feature information of the target object using a second algorithm based on the first feature information; determining a three-dimensional model of the target object based on the first feature information and the second feature information; and determining feature information of the first part and the second part based on the three-dimensional model.
[0010] According to embodiments of this disclosure, determining feature information of a first part and feature information of a second part based on a three-dimensional model includes: determining the spatial coordinates of the center point of the first part and the second part of the target object based on the first feature information; obtaining the distance between each point on the three-dimensional model and the spatial coordinates of the center point, and dividing points with a distance less than a first threshold into multiple part models; and obtaining feature information of the first part and feature information of the second part based on the multiple part models.
[0011] According to embodiments of this disclosure, the method further includes: when the number of objects within a preset spatial range is greater than a second threshold, determining a target object from multiple objects based on a first algorithm.
[0012] A second aspect of this disclosure provides a sound effect configuration device, comprising: an acquisition module for acquiring feature information of a first part and a second part of a target object; a determination module for determining the relative positional relationship between the first part and the second part based on the feature information of the first part and the second part; and a configuration module for determining sound effect parameters of a sound source according to the relative positional relationship between the first part and the second part, wherein the sound effect parameters of the sound source are used to configure the spatial sound effect signal of the sound source relative to the target object.
[0013] A third aspect of this disclosure provides an electronic device, comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the following method: obtaining feature information of a first part and feature information of a second part of a target object; determining a relative positional relationship between the first part and the second part based on the feature information of the first part and the second part; and determining sound effect parameters of a sound source based on the relative positional relationship of the first part and the second part, the sound effect parameters of the sound source being used to configure a spatial sound effect signal of the sound source relative to the target object.
[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0015] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0016] Figure 1 This illustration schematically depicts an application scenario of a sound effect configuration method, apparatus, and electronic device according to embodiments of the present disclosure.
[0017] Figure 2 A flowchart illustrating a sound effect configuration method according to an embodiment of the present disclosure is shown schematically;
[0018] Figure 3 A flowchart illustrating the preparation and deployment of a 3D model according to an embodiment of the present disclosure is shown schematically;
[0019] Figure 4 This schematically illustrates a flowchart for determining the sound effect parameters of a sound source when a second directional relationship is satisfied according to an embodiment of the present disclosure;
[0020] Figure 5 This schematically illustrates a flowchart for determining the sound effect parameters of a sound source when a third-party bit relationship is satisfied according to an embodiment of the present disclosure;
[0021] Figure 6 A schematic diagram illustrating the structure of a sound effect configuration device according to an embodiment of the present disclosure is shown; and
[0022] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a sound effect configuration method according to an embodiment of the present disclosure. Detailed Implementation
[0023] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0026] It should be noted that the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information in this disclosed technical solution comply with relevant laws and regulations, necessary confidentiality measures have been taken, and it does not violate public order and good morals. In this disclosed technical solution, user authorization or consent has been obtained before acquiring or collecting user personal information.
[0027] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0028] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0029] Currently, most spatial sound effect configuration methods are static, typically improving spatial sound effects based on sound content rather than adjusting to the real-time posture of the target object. While some methods utilize location information to achieve dynamic spatial sound effect placement, these generally rely on the overall posture of the target object, not specific body part information. Furthermore, existing methods for extracting facial features and body key points can only obtain relatively sparse and coarse facial spatial position information, lacking details of specific body parts. The 3D Gaussian sputtering method can optimize the originally sparse spatial information at low cost and with high precision, thereby accurately acquiring the real-time shape and spatial position of the human head. However, the 3D Gaussian sputtering method also has certain limitations; the accurate model obtained through inference does not contain body part information, making it impossible to directly determine the specific body part of the head corresponding to each region of the model.
[0030] In view of the above, this disclosure provides a sound effect configuration method, apparatus, and electronic device, which will be described below with reference to the accompanying drawings.
[0031] Figure 1 Figure 100 schematically illustrates an application scenario of a sound effect configuration method, apparatus, and electronic device according to embodiments of the present disclosure.
[0032] It is important to note that Figure 1 The examples shown are merely examples of scenarios in which the embodiments of this disclosure can be applied, to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0033] like Figure 1 As shown, the application scenario 100 according to this embodiment may include multiple virtual sound sources 101. The virtual sound sources 101 may be distributed in different positions in the three-dimensional virtual space to simulate sound sources from various directions in the real environment, such as the sounding positions of different instruments in a virtual concert scene, or the origin points of various sound effects in a virtual game scene.
[0034] Application scenario 100 may also include audio rendering device 102, which includes, but is not limited to, high-performance computers, professional audio workstations, and mobile terminals with powerful audio processing capabilities. Audio rendering device 102 can precisely configure and play spatial sound effects through its complex internal audio signal processing architecture to create realistic surround sound effects and simulate the reflection and attenuation of sound in different spatial environments.
[0035] The audio rendering device 102 can acquire audio data sources corresponding to the virtual sound source 101, such as music files, game sound effect files, and video audio tracks. Simultaneously, the audio rendering device 102 can decode these audio data sources and, in conjunction with the generated spatial sound effect configuration information, perform spatial sound effect rendering operations. Then, the audio rendering device 102 can convert digital audio signals into analog audio signals and drive the connected speaker system to emit sound. The speaker system can include various types, such as multi-channel surround sound, two-channel stereo sound, and head-mounted virtual reality headphones (with built-in multi-speaker units). Ultimately, it can present audio content with a strong sense of space and realism to the target object, making the target object feel as if they are in a real scene constructed by the virtual sound source.
[0036] It should be understood that Figure 1 The number of virtual sound sources and audio rendering devices shown is merely illustrative. Depending on implementation needs, any number of virtual sound sources and audio rendering devices can be used.
[0037] It should be noted that the sound effect configuration method provided in this disclosure can generally be applied to an audio rendering device, and correspondingly, the sound effect configuration device provided in this disclosure can be set in the audio rendering device. The sound effect configuration method provided in this disclosure can also be executed by a server or server cluster that is different from the audio rendering device but capable of communicating with it. Correspondingly, the sound effect configuration device provided in this disclosure can also be set in a server or server cluster that is different from the audio rendering device but capable of communicating with it. It should be understood that... Figure 1 The number of audio rendering devices shown is merely illustrative. Depending on implementation needs, any number of audio rendering devices can be used.
[0038] The following will be based on Figure 1 The described scene, through Figures 2-5 The sound effect configuration method of the present disclosure will be described in detail.
[0039] like Figure 2 As shown, the sound effect configuration method of this embodiment may include S210~S230.
[0040] In S210, the feature information of the first part and the feature information of the second part of the target object are obtained.
[0041] The first part of the target object can be the head, and the second part can be the shoulders. The head can perform various movements, such as nodding up and down, turning left and right, and rotating horizontally, which can change its spatial orientation and angle. The shoulders can serve as a reference for the head, jointly defining the natural orientation of the target object. For example, when the head turns, the position of the shoulders can help determine the axis of rotation (such as the center of the neck). Simultaneously, the shoulders can reflect low-frequency sound waves, altering the spectral characteristics of the sound. This causes sounds from different directions to produce different frequency response changes after being reflected by the shoulders. For example, sounds coming from below may have their high-frequency components attenuated and low-frequency components relatively enhanced after being reflected by the shoulders, thus allowing for adjustment of the vertical positioning of the sound effect.
[0042] Feature information can include geometric features, positional features, and attitude features. Geometric features can include the size, shape, and contour of a part, reflecting its physical form and providing a basis for determining the relative positional relationships between parts. Positional features can be used to determine the coordinate position of a part in three-dimensional space, facilitating subsequent analysis of their relative positions. Attitude features can describe the rotation, tilt, and other attitude information of directional parts. For example, the attitude of the head can be represented as rotation angles (pitch, yaw, and roll) around three coordinate axes. By analyzing attitude information, the relative positional relationship between the first and second parts can be determined more accurately.
[0043] In some exemplary embodiments, the feature extraction of the first part (such as the head) can be further refined. In addition to the overall head movements and postures, feature information of specific organs on the head, such as the ears, can also be obtained, including but not limited to the direction, shape and size of the ears, in order to configure spatial sound effects more accurately.
[0044] In S220, the relative positional relationship between the first part and the second part is determined based on the feature information of the first part and the feature information of the second part.
[0045] For example, if the feature information of the first and second parts contains their position information in their respective local coordinate systems, they can be transformed into the same global coordinate system through coordinate transformation in order to calculate their relative positional relationship.
[0046] For example, based on the geometric features of the first and second parts, their geometric relationships, such as inclusion, adjacency, or symmetry, can be analyzed to determine their relative positions. For instance, when analyzing human body parts, knowing that the head connects to the neck, and the neck is located above the shoulder, the position of the head relative to the shoulder can be inferred from this geometric relationship. Furthermore, combining positional and posture features can more accurately describe the relative positional relationship between the head and shoulder.
[0047] The relative positional relationship between the first and second parts can include not only the straight-line distance between the two parts, but also rich information such as their angle, direction, and possible movement trends. For example: the head remains in a normal position without significant tilting or turning; the head rotates around a vertical axis (which can be defined based on the shoulders), changing the facing direction; the head tilts to one side, changing the relative position of the ear and shoulder, etc.
[0048] In S230, the sound effect parameters of the sound source are determined based on the relative positional relationship between the first part and the second part. The sound effect parameters of the sound source are used to configure the spatial sound effect signal of the sound source relative to the target object.
[0049] Based on the relative positional relationship between the first and second parts of the target object, the posture and orientation of the target object in three-dimensional space can be determined, and thus the sound effect parameters of the sound source can be determined, including but not limited to azimuth parameters (azimuth and elevation), distance, channel delay, channel gain, and gain / attenuation values of each frequency band.
[0050] For example, by calculating the orientation and angle of a first part (such as the head) relative to a second part (such as the shoulder), the specific location of the sound source in the perception of the target object can be determined. For instance, if the head is turned to the left at a certain angle relative to the shoulder, the left and right channel balance of the sound source can be adjusted accordingly, making the sound seem to come from the left, thereby simulating a realistic sound source localization effect.
[0051] For example, the attenuation characteristics of sound with distance can be simulated based on the relative distance between the first and second parts. As the distance between the target object and the sound source increases, the sound intensity gradually weakens, and high-frequency components attenuate more quickly. Based on these physical laws, the volume and high-frequency gain of the sound source can be dynamically adjusted to create a more realistic sense of distance.
[0052] To further enhance the realism of spatial sound effects, the Head-Related Transfer Function (HRTF) can be used to simulate the propagation characteristics of sound around the head and ears. HRTF describes the spectral changes caused by structures such as the head and auricle when sound reaches the ear from different directions. By selecting or synthesizing appropriate HRTF parameters based on the geometric relationship between the first and second parts (such as auricle shape, head occlusion, and shoulder reflection), the filtering characteristics of the sound (such as high-frequency attenuation and time difference) can be adjusted to make the virtual sound source sound like it's coming from the target direction.
[0053] For example, head (which may include ear features) and shoulder data with labeled relative positions can first be input into a near-field HRTF model for simulation. This model can default to a spherical uniform sampling method, which systematically discretizes the direction of the sound source and accurately simulates the propagation characteristics of sound reaching the ear from different directions through uniformly distributed sampling points. Based on this, the model can be fitted and adjusted according to the head contour with a weight of 0.1, making the model more closely resemble the actual situation of the target object, thus obtaining an HRTF model for the target object. Then, based on this HRTF model, the original input audio signal (such as music, speech, environmental sound effects, etc.) can be filtered specifically, and the sound effect parameters of the sound source can be determined by combining the original spatial information of the sound source (i.e., the original orientation data). For example, when the sound source is located to the left front of the target object, the model can simulate the subtle differences in sound as it travels around the head, is reflected by the shoulder, and reaches both ears based on the accurate relative position. These differences include time difference, sound level difference, and spectral variations. This allows for corresponding delay, gain adjustment, and spectral modification of the sound signal, simulating the process of sound propagating from a specific direction to the ear in a real environment and determining the sound source's audio parameters. Finally, based on these determined parameters, the spatial audio signal of the sound source relative to the target object can be configured, including audio decoding, mixing, and output. This ensures that the final output sound is perceived by the listener as originating from a specific virtual location, achieving precise localization of the virtual sound source. Finally, the processed audio signal can be mapped to the corresponding speaker for output.
[0054] Understandably, sound source parameters determined by the relative positional relationship between the first and second parts of the target object can more accurately simulate the propagation path and characteristics such as reflection and diffraction of sound in a real environment, thus enabling more precise configuration of spatial sound effects. Furthermore, by configuring sound effects based on individual characteristic information, a unique auditory experience can be provided to the target object.
[0055] In embodiments of this disclosure, obtaining feature information of a first part and feature information of a second part of a target object includes: extracting first feature information of the target object using a first algorithm; determining second feature information of the target object using a second algorithm based on the first feature information; determining a three-dimensional model of the target object based on the first feature information and the second feature information; and determining feature information of the first part and feature information of the second part based on the three-dimensional model.
[0056] Figure 3 A flowchart illustrating the preparation and deployment of a three-dimensional model according to an embodiment of the present disclosure is shown.
[0057] like Figure 3As shown, the first step is to prepare for head modeling of the target object. During this process, facial data of the target object can be acquired. This can be done by capturing video or multiple photographs from different angles of the target object facing an image acquisition device (such as a computer webcam) to capture image details of the face and shoulders in the forward direction. After acquisition, a first algorithm (such as 3D Morphable Model Fitting, 3DMM Fitting) can be used on the image data of the target object to extract the first feature information, namely facial features and body key points, and to initially generate sparse head and shoulder information. Next, a second algorithm, such as 3D Gaussian Splatting (3DGS), can be trained based on this sparse information. This algorithm iteratively optimizes the coarse head and shoulder information, extending the head modeling to the entire upper body, obtaining the relative position and movement of the shoulders and head, and finally obtaining the second feature information of the target object, namely the static model of the head and shoulders and the deformation equations that change with the target object's movements.
[0058] In the deployment phase of user head modeling, facial data can be collected first. Then, 3DMM Fitting can be used to extract real-time facial features and body key points (i.e., first feature information) from the real-time captured photos. Furthermore, using the trained static model of the user's head and shoulders, as well as the deformation equation (i.e., second feature information), algorithmic inference can be performed on real-time head and shoulder movements to obtain a 3D model that provides real-time, detailed modeling of the user's expressions and head movements.
[0059] In the spatial sound effect distance calculation and optimization stage, based on the three-dimensional model, the feature information of specific parts (such as the first part and the second part) in the three-dimensional model can be further determined, so as to configure the spatial sound effect signal based on the original audio signal and form an accurate personalized spatial sound effect.
[0060] It should be noted that the 3D Gaussian sputtering method can also be replaced by the more advanced 2D Gaussian sputtering method. The 2D Gaussian sputtering method can fit the head and shoulder shape of the target object with a large number of planar ellipsoids, ensuring the accuracy of spatial geometry, and each ellipsoid has a precise spatial normal vector, which can accelerate the calculation of spatial position and distance.
[0061] Understandably, using part information from real-time modeled 3D models allows for more precise dynamic spatial sound effect placement, without relying on external devices—a purely visual solution that adapts to various scenario requirements. Furthermore, by using a second algorithm to optimize the originally sparse spatial information with low computational cost and high accuracy, precise part feature information can ultimately be obtained.
[0062] Based on the above embodiments, in this embodiment, determining the feature information of the first part and the feature information of the second part according to the three-dimensional model may include: determining the spatial coordinates of the center point of the first part and the second part of the target object respectively based on the first feature information; obtaining the distance between each point on the three-dimensional model and the spatial coordinates of the center point, and dividing the points with a distance less than a first threshold into multiple part models; and obtaining the feature information of the first part and the feature information of the second part based on the multiple part models.
[0063] In one embodiment, firstly, a 3D coordinate system for the target object can be established with the midpoint between the two shoulders as the origin, so that the positional information of all relevant points can be described and analyzed in a unified and accurate coordinate system. Then, continue to refer to... Figure 3 The spatial coordinates E of the center points of the left and right ears in the head (i.e., the first part) can be obtained from the output of facial feature extraction. l E r Then calculate the distance from each point on the 3D model to E. l E r The distance is used to segment points whose distance is less than a first threshold into ear models. Based on this ear model, the precise spatial location of the left and right ears, as well as the specific shape and orientation of the auricle, can be obtained, thus obtaining the specific feature information of the part. Simultaneously, the spatial coordinates S of the center points of the left and right shoulders (i.e., the second part) can be obtained from the output of the body keypoint extraction. l S r Using the same method as for obtaining the ear model—that is, calculating the distances between each point on the 3D model and the center point of the shoulder and then segmenting the model—the precise spatial positions of the left and right shoulders can be obtained. Furthermore, the specific head orientation, height, and shape of the target object can be obtained through a real-time 3D head model.
[0064] It should be noted that the preset distance threshold can be determined through multiple experiments and verifications. This not only ensures that the segmented part model includes the main structure of the part, but also avoids mistakenly including irrelevant surrounding points in the model.
[0065] Understandably, by combining the facial features and body key points extracted by the first algorithm with a precise 3D model, more accurate part feature information can be obtained, thereby making the configuration of spatial sound effects more accurate.
[0066] Based on the above embodiments, in this embodiment, the sound effect parameters of the sound source are determined according to the relative positional relationship between the first part and the second part, including: in response to the relative positional relationship between the first part and the second part satisfying a first directional relationship, determining the first directional data corresponding to the sound source based on the original directional data of the sound source; and determining the sound effect parameters based on the first directional data.
[0067] When the relative positions of the first part (head) and the second part (shoulder) satisfy the first orientation relationship, that is, the head does not show obvious tilting or turning relative to the shoulder, specifically, the head does not rotate or the deflection angle is less than a preset threshold. In this case, the first orientation data corresponding to the sound source can be determined based on the original orientation data of the sound source. In this process, the original orientation data of the sound source can first be obtained, that is, the position coordinates in the world coordinate system. It can determine the position coordinates of the first part (such as the head) and the second part (such as the shoulder) of the target object in the world coordinate system based on feature information, and then determine the position coordinates of the left and right shoulders of the target object in the world coordinate system. and Calculate the coordinates of the midpoints of the two shoulders. Then, the coordinates of the sound source in the world coordinate system Convert to the origin of the target coordinate system relative coordinates for reference Finally, the sound source coordinates in the target coordinate system can be used as a reference. Calculate the first orientation data of the sound source in the target coordinate system.
[0068] The first location data can include the azimuth angle θ and elevation angle φ of the sound source, as well as the distance r to the head (accurate to the left / right ear). The azimuth angle θ represents the horizontal angle of the sound source relative to the front of the target object. For example, when the sound source is directly in front of the target object, the azimuth angle θ is 0°; when the sound source is at a 45° angle to the right of the target object, the azimuth angle θ is 45°. The elevation angle φ represents the vertical angle of the sound source relative to the horizontal plane of the target object. For example, when the sound source is directly above the target object's head, the elevation angle φ is 90°; when it is horizontally positioned directly in front of the target object, the elevation angle φ is 0°. The distance r to the head (left / right ear) represents the spatial distance between the sound source and the left / right ears of the target object.
[0069] For example, suppose in a virtual reality game, the player (target object) stands at (0,0,0) at the origin of the world coordinate system, with the midpoints of both shoulders also at (0,0,0), and the target object's head is not rotated (i.e., satisfying the first orientation relationship). At this time, a sound source is located at coordinates (3,4,2) in the world coordinate system. Since the origin of the target coordinate system coincides with the origin of the world coordinate system, the relative coordinates... The original coordinates can be (3,4,2). Then, based on the relative coordinates (3,4,2), the distance r from the sound source to the origin of the target coordinate system can be calculated. The azimuth angle θ = arctan2(4,3) ≈ 53.13°, and the elevation angle... .
[0070] For example, the first azimuth data (sound source azimuth angle θ, elevation angle φ, and distance r to the left / right ear) can be directly input into the target object's personalized HRTF model, directly using the parameter set of the pre-computed standard HRTF model (based on head posture modeling from the front). Since sound direction (azimuth angle θ, elevation angle φ) is usually discretely sampled, for example, HRTF data is measured every 5° in one direction, but the actual sound may come from any angle, such as 32.5°. To obtain accurate spatial sound effects, interpolation can be performed. Interpolation can use a spherical uniform interpolation method similar to the spherical uniform sampling method. Simultaneously, it can be fitted and adjusted according to the head contour with a weight of 0.1 to better adapt to the individual characteristics of the target object. Ultimately, accurate spatial sound effect parameters that conform to the individual characteristics of the target object can be obtained.
[0071] It is understandable that when the user's head is in a first position relative to the shoulder, the first position data can be calculated based on the original position data of the sound source, and then the sound effect parameters can be determined. This can simulate the effect of sound attenuation with distance, making the sound effect more realistic.
[0072] Based on the above embodiments, in this embodiment, determining the sound effect parameters of the sound source according to the relative positional relationship between the first part and the second part may further include: responding to the fact that the relative positional relationship between the first part and the second part satisfies a second orientation relationship, calculating the rotation information of the first part according to the feature information; adjusting the original orientation data of the sound source based on the rotation information to obtain the second orientation data corresponding to the sound source, and determining the sound effect parameters according to the second orientation data.
[0073] Figure 4 The flowchart illustrating the determination of sound effect parameters of a sound source when a second orientation relationship is satisfied according to an embodiment of the present disclosure is shown.
[0074] like Figure 4 As shown, when the relative positional relationship between the first part (head) and the second part (shoulder) satisfies the second orientation relationship, that is, the head rotates, and the rotation angle is greater than or equal to a preset threshold. At this time, the rotation information of the head can be calculated based on the feature information, so as to adjust the original orientation data of the sound source using the rotation information to obtain the second orientation data. Finally, the sound effect parameters can be determined based on the second orientation data.
[0075] For example, first, a target coordinate system can be established (with the midpoint between the two shoulders as the origin). Then, the head orientation vector can be calculated using key points on the head (such as the nose and the top of the head). For instance, the coordinates of the nose and the top of the head in the world coordinate system can be selected, and vector operations can be used to obtain a vector representing the head orientation. Furthermore, the direction of the rotation axis can be determined using the cross product of vectors, and the rotation angle can be calculated using the dot product of vectors combined with the magnitude of the vectors. Based on the calculated rotation axis and angle, a rotation matrix R (i.e., rotation information) can be generated. This matrix can characterize the transformation of the target object's head from its initial state to its current posture. In addition, to avoid abrupt changes in the sound field during head rotation, damping can be applied to the rotation matrix. For example, the transformation matrix R′ = 0.75R can be calculated so that the rotation of the spherical spatial sound effect is slightly less than the actual head rotation, providing a spatial relative feel for the target object. Since the input sound source can contain original orientation data, such as coordinates (x, y, z) in the world coordinate system, the coordinates of the sound source can be transferred from the world coordinate system to the currently rotated target coordinate system using the calculated rotation matrix R or R′ (rotation matrix or transformation matrix). Mathematically, the original orientation data can be multiplied by a rotation matrix to obtain the second orientation data in the target coordinate system.
[0076] In some exemplary embodiments, the time delay can be recalculated based on the new relative position of the sound source and the ears. This is achieved by calculating the distance difference between the sound source and the left and right ears, combined with the speed of sound in air. Furthermore, when the head rotates, the orientation of the sound source relative to the ears changes, and the frequency components of the sound received by the left and right ears can change accordingly. This allows for updating the frequency response difference between the left and right ears based on the azimuth angle after rotation. The scattering and reflection of sound by the auricle can affect the spectral characteristics of the sound; sound sources from different orientations will exhibit different frequency component changes after being processed by the auricle. The filtering frequency band and attenuation level can be adjusted based on the geometric relationship between the auricle and the new orientation of the sound source. For example, for lateral sound sources, attenuation is enhanced above 8kHz.
[0077] In other embodiments, based on the original orientation data of the sound source, its first orientation data in a target coordinate system aligned with the initial head posture (where the head has not shown significant tilting or turning) can be determined. In response to head rotation, i.e., when the relative positional relationship between the first and second parts satisfies a second orientation relationship, the second orientation data can be generated by correcting the first orientation data. That is, when the target object's head rotates, the first orientation data can be adjusted according to rotation information (such as a rotation matrix) to record the absolute positional change of the sound source, thus obtaining the corrected second orientation data of the sound source in the target coordinate system.
[0078] Understandably, determining the orientation data based on the rotation of the target object's head ensures that the sound source's sound effects can accurately reflect its new position relative to the target object when the head rotates, thus providing a more realistic and immersive auditory experience.
[0079] Based on the above embodiments, in this embodiment, determining the sound effect parameters of the sound source according to the relative positional relationship between the first part and the second part further includes: in response to the relative positional relationship between the first part and the second part satisfying a third positional relationship, determining the geometric relationship parameters of the first part and the second part according to the feature information; adjusting the original orientation data of the sound source based on the geometric relationship parameters to obtain the third positional data corresponding to the sound source, and determining the sound effect parameters according to the third positional data.
[0080] Figure 5 This schematically illustrates a flowchart of determining the sound effect parameters of a sound source when a third-party bit relationship is satisfied according to an embodiment of the present disclosure.
[0081] like Figure 5 As shown, when the relative positional relationship between the first part (head) and the second part (shoulder) satisfies a third positional relationship, for example, the head tilts to one side and the tilt angle is greater than or equal to a preset threshold, geometric relationship parameters can be determined based on feature information, including but not limited to deflection angle, auricle angle, and projected area.
[0082] Based on the established geometric parameters, the original azimuth data of the sound source can be adjusted through coordinate transformation and trigonometric function operations. This compensates for changes in acoustic characteristics caused by head tilt, resulting in a new sound propagation direction vector and distance (i.e., third-party positional data). Then, sound effect parameters can be determined based on this third-party positional data. In one embodiment, this can be implemented using an HRTF model. HRTF determines the third-party positional data of the sound source relative to the head based on the geometric parameters of the head and shoulders, thereby recalculating the acoustic path (such as path length, propagation angle, and information on reflected surfaces and diffraction points). Based on this path information, sound effect parameters such as the filtering coefficient and propagation delay coefficient in the model can be adjusted to accurately reflect the sound propagation characteristics under the new geometric relationship.
[0083] In the HRTF model, azimuth and elevation are crucial input parameters. By mapping the calculated deflection angle to changes in azimuth and elevation, the filter coefficients and propagation path parameters in the model can be adjusted. Changes in the auricle angle affect sound reflection and diffraction around the auricle, thus altering the sound's spectral characteristics. The HRTF model can simulate the auricle's modulation effect on sound by establishing a mapping relationship between the auricle angle and a spectral filter, adjusting the filter parameters based on real-time auricle angle data. Changes in the projected area reflect the degree of head obstruction of the sound source. When the head tilts, the projected area on one side may increase, leading to increased attenuation during sound propagation. In the HRTF model, the sound intensity can be dynamically adjusted by introducing an attenuation coefficient related to the projected area.
[0084] For example, in a game scene, when the player (i.e., the target object) tilts their head to the right, geometric parameters such as the head's deflection angle, auricle angle, and projected area can be obtained in real time based on the feature information of the first and second parts of the head. Then, the HRTF model can recalculate the acoustic path based on these parameters. For a sound source on the right, the path length of the sound to the right ear can be shortened, reducing propagation delay, while the intensity of high-frequency components can be increased, thus simulating the effect of the sound being "closer". For a sound source on the left, the model can increase the path length of the sound to the left ear, increasing propagation delay, and can introduce more reverberation and low-frequency attenuation to simulate the effect of sound propagating from a distance and being blocked by the head, ultimately creating a more realistic spatial sound effect.
[0085] Understandably, adjusting the original location data of the sound source based on geometric parameters can compensate for changes in acoustic characteristics caused by head tilt, thereby obtaining accurate and reliable third-party location data and providing data support for creating a realistic sound experience.
[0086] Based on the above embodiments, in this embodiment, the sound effect parameters of the sound source are determined according to the relative positional relationship between the first part and the second part, including: in response to the target object moving from the first position to the second position, updating the origin of the target coordinate system according to the feature information, wherein the target coordinate system is defined based on the target object; determining the fourth position data of the sound source in the updated target coordinate system based on the original position data of the sound source, and determining the sound effect parameters according to the fourth position data.
[0087] First, a target coordinate system based on the target object can be defined. The origin of this coordinate system can be set as the midpoint between the two shoulders of the target object. The azimuth data of all sound sources can be calculated with reference to this origin.
[0088] When the target object moves from a first position to a second position, the system can respond quickly to this change. Based on previously acquired feature information of the target object's body parts (including data that can accurately locate the positions of the two shoulders), the origin of the target coordinate system is updated synchronously. For example, in a virtual reality game scene, the player (i.e., the target object) moves freely in the virtual environment, moving from one end of the room to the other. The system can continuously track changes in the position of the player's two shoulders. Once a movement is detected, the origin of the coordinate system is immediately updated based on the new midpoint position of the two shoulders, so that the entire coordinate system dynamically adjusts as the player moves.
[0089] After updating the origin of the target coordinate system, the fourth azimuth data of the sound source in the updated target coordinate system can be determined based on the original azimuth data of the sound source. The original azimuth data can be the position information of the sound source in the initial coordinate system (such as the world coordinate system or other fixed coordinate systems), which can be transformed into the updated target coordinate system using a coordinate transformation algorithm to obtain the fourth azimuth data.
[0090] It should be noted that the fourth-position data is the new baseline data after the target object has moved, and it is at the same level as the first-position data. The first-position data can be the position information of the sound source relative to the target object when the target object is in its initial position, while the fourth-position data is the position information of the sound source relative to the updated target object after the target object has moved. They only reflect the translational positional relationship of the sound source relative to the target object in space.
[0091] Based on fourth-position data, the sound effect parameters of the sound source can be further determined. For example, when the fourth-position data shows that the sound source is close to the target object, the volume of the sound can be appropriately increased to give the user a strong sense of impact; when the sound source is located to the side of the target object, the stereo effect can be adjusted to create a realistic feeling that the sound is coming from the side.
[0092] The movement of the target object can dynamically update the origin of the target coordinate system. In other words, the coordinate axes defined with the midpoint between the two shoulders as the origin will move along with the target object. Furthermore, when the target object's head rotates, a rotation matrix can be superimposed synchronously.
[0093] For example, in a virtual reality scene, players not only move within the environment but also rotate their heads to observe their surroundings. When a player's head rotates to the left by a certain angle, a corresponding rotation matrix can be generated based on the angle and direction of rotation and superimposed onto the updated target coordinate system. In this way, the location data of the sound source can be adjusted in real time according to the player's head rotation, ensuring that the sound always enters the player's ears from the correct direction. For instance, a sound source originally located directly in front of the player can have its location relative to the player's current head orientation recalculated by superimposing the rotation matrix after the player rotates their head to the left. This ensures that the sound effect heard by the player matches the actual visual scene, further enhancing the sense of immersion.
[0094] Understandably, by updating the sound source's audio parameters in real time and accurately based on the movement of the target object, users can be provided with a highly realistic and personalized spatial audio experience.
[0095] In embodiments of this disclosure, the method further includes: when the number of objects within a preset spatial range is greater than a second threshold, determining a target object from multiple objects based on a first algorithm.
[0096] In practical applications, spatial sound effect configuration typically faces complex and ever-changing environments, where the number of objects can significantly impact the accuracy and effectiveness of sound effect processing. Since the sound effect configuration method provided in this embodiment primarily focuses on single-person sound effect scenarios, it aims to provide users with a highly personalized and accurate spatial sound effect experience. Therefore, when the number of objects within a preset spatial range is 1, i.e., not exceeding a second threshold (e.g., 2), the target object can be clearly identified. However, when the number of objects within the preset spatial range is greater than 2, it indicates the presence of a multi-person scenario. In this case, a specific filtering mechanism can be set, utilizing a first algorithm to analyze and filter multiple objects. This involves comprehensively considering factors such as the object's salience, motion state, distance from the camera, and position within the frame, filtering out non-salience background figures that do not conform to the characteristics of the target object.
[0097] For example, in a party setting, multiple people are active within a pre-defined space. A first algorithm can extract and analyze the characteristics of each person. If someone is far from the camera, occupies a small proportion of the frame, and has minimal movement, that person can be identified as a non-prominent background figure and filtered out. Conversely, individuals who are closer to the camera, stand out more in the frame, and whose movement matches the characteristics of the target object can be identified as the target object. This allows for focused configuration of precise spatial sound effects for the target object, avoiding deviations or errors in sound effect configuration due to interference from multiple people.
[0098] Understandably, by using the first algorithm to filter out inconspicuous background figures, it is possible to ensure that the target object can still be identified relatively accurately in multi-person scenes, providing a reliable foundation for subsequent sound effect configuration.
[0099] Figure 6 A block diagram of a sound effect configuration device according to an embodiment of the present disclosure is shown schematically.
[0100] like Figure 6 As shown, the sound effect configuration device 600 includes an acquisition module 610, a determination module 620, and a configuration module 630.
[0101] According to some embodiments of this disclosure, the sound effect configuration device 600 can be used to implement a reference. Figures 2-5 The sound effect configuration method described according to embodiments of the present disclosure.
[0102] The acquisition module 610 can perform, for example, operation S210, to obtain feature information of the first part and the second part of the target object.
[0103] The determining module 620 can perform, for example, operation S220, to determine the relative positional relationship between the first part and the second part based on the feature information of the first part and the feature information of the second part.
[0104] The configuration module 630 can perform, for example, operation S230, to determine the sound effect parameters of the sound source based on the relative positional relationship between the first part and the second part, wherein the sound effect parameters of the sound source are used to configure the spatial sound effect signal of the sound source relative to the target object.
[0105] For example, any plurality of the acquisition module 610, determination module 620, and configuration module 630 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the acquisition module 610, determination module 620, and configuration module 630 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the acquisition module 610, determination module 620, and configuration module 630 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0106] It should be understood that the sound effect configuration device in the embodiments of this disclosure corresponds to the sound effect configuration method in the embodiments of this disclosure, and their specific implementation details are the same, so they will not be repeated here.
[0107] It should be noted that the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information in this disclosed technical solution comply with relevant laws and regulations, necessary confidentiality measures have been taken, and it does not violate public order and good morals. In this disclosed technical solution, user authorization or consent has been obtained before acquiring or collecting user personal information.
[0108] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a sound effect configuration method according to an embodiment of the present disclosure.
[0109] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0110] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in one or more memories.
[0111] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the I / O interface 705: an input section 706 including target hardware, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0112] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0113] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703 described above.
[0114] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the sound effect configuration method provided in the embodiments of this disclosure.
[0115] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0116] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0117] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0118] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0120] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0121] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A method for configuring sound effects, comprising: Obtain feature information of the first and second parts of the target object; Based on the feature information of the first part and the feature information of the second part, the relative positional relationship between the first part and the second part is determined; Based on the relative positional relationship between the first part and the second part, the sound effect parameters of the sound source are determined, and the sound effect parameters of the sound source are used to configure the spatial sound effect signal of the sound source relative to the target object; The step of determining the sound effect parameters of the sound source based on the relative positional relationship between the first part and the second part includes: in response to the relative positional relationship between the first part and the second part satisfying a first directional relationship, determining the first directional data corresponding to the sound source based on the original directional data of the sound source; and determining the sound effect parameters based on the first directional data.
2. The method according to claim 1, wherein determining the sound effect parameters of the sound source based on the relative positional relationship between the first part and the second part further includes: In response to the fact that the relative positional relationship between the first part and the second part satisfies the second orientation relationship, the rotation information of the first part is calculated based on the feature information; Based on the rotation information, the original azimuth data of the sound source is adjusted to obtain the second azimuth data corresponding to the sound source, and the sound effect parameters are determined based on the second azimuth data.
3. The method according to claim 1, wherein determining the sound effect parameters of the sound source based on the relative positional relationship between the first part and the second part further includes: In response to the fact that the relative positional relationship between the first part and the second part satisfies a third positional relationship, the geometric relationship parameters of the first part and the second part are determined based on the feature information. The original orientation data of the sound source is adjusted based on the geometric relationship parameters to obtain the third-party position data corresponding to the sound source, and the sound effect parameters are determined based on the third-party position data.
4. The method according to claim 1, wherein determining the sound effect parameters of the sound source based on the relative positional relationship between the first part and the second part includes: In response to the target object moving from a first position to a second position, the origin of the target coordinate system is updated based on the feature information, wherein the target coordinate system is defined based on the target object; Based on the original location data of the sound source, the fourth location data of the sound source in the updated target coordinate system is determined, and the sound effect parameters are determined according to the fourth location data.
5. The method according to claim 1, wherein obtaining the feature information of the first part and the feature information of the second part of the target object includes: The first feature information of the target object is extracted using the first algorithm; Based on the first feature information, the second algorithm is used to determine the second feature information of the target object; Based on the first feature information and the second feature information, a three-dimensional model of the target object is determined; The feature information of the first part and the feature information of the second part are determined based on the three-dimensional model.
6. The method according to claim 5, wherein determining the feature information of the first part and the feature information of the second part based on the three-dimensional model comprises: Based on the first feature information, the spatial coordinates of the center points of the first and second parts of the target object are determined respectively. Obtain the distance between each point on the 3D model and the spatial coordinates of the center point, and divide the points whose distance is less than a first threshold into multiple part models; Based on multiple part models, feature information of the first part and feature information of the second part are obtained.
7. The method according to claim 5, further comprising: If the number of objects within a preset space is greater than a second threshold, the target object is determined from multiple objects based on the first algorithm.
8. A sound effect configuration device, comprising: The acquisition module is used to obtain feature information of the first part and the second part of the target object. The determining module is used to determine the relative positional relationship between the first part and the second part based on the feature information of the first part and the feature information of the second part; A configuration module is used to determine the sound effect parameters of a sound source based on the relative positional relationship between the first part and the second part, wherein the sound effect parameters of the sound source are used to configure the spatial sound effect signal of the sound source relative to the target object; wherein, determining the sound effect parameters of the sound source based on the relative positional relationship between the first part and the second part includes: in response to the relative positional relationship between the first part and the second part satisfying a first directional relationship, determining the first directional data corresponding to the sound source based on the original directional data of the sound source; and determining the sound effect parameters based on the first directional data.
9. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors perform the following method: The process involves obtaining feature information of a first part and a second part of a target object; determining the relative positional relationship between the first part and the second part based on their respective feature information; and determining sound effect parameters of a sound source based on this relative positional relationship. The sound effect parameters are used to configure the spatial sound effect signal of the sound source relative to the target object. Specifically, determining the sound effect parameters based on the relative positional relationship between the first part and the second part includes: determining first azimuth data corresponding to the sound source based on its original azimuth data, in response to the first azimuth relationship satisfying the relative positional relationship between the first part and the second part; and determining the sound effect parameters based on the first azimuth data.
Citation Information
Patent Citations
Audio processing device and method, augmented reality device, equipment and storage medium
CN120188497A