Audio data processing method, device, electronic device and computer storage medium

By receiving audio data and changing its spatial information, the problem of insufficient simulation of sound source spatial information in the prior art is solved, and more realistic audio data output is achieved and user experience is improved.

CN114038486BActive Publication Date: 2025-08-08BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111618120.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-08-08
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively simulate the spatial information of the sound source in audio data processing, resulting in insufficient user experience.

Method used

By receiving the target audio data and set spatial information, the spatial information of the audio data is changed using a pre-trained model or acoustic data mixing technology, so that the receiver can sense that the sound source has a set spatial location and environment.

Benefits of technology

Improves the real effect and user experience when outputting audio data, and can create preset audio effects in a variety of audio and video playback scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114038486B_ABST
    Figure CN114038486B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio data processing method, apparatus, electronic device, and computer storage medium, relating to the field of computer technology, particularly to technical fields such as speech technology. A specific implementation scheme comprises: receiving target audio data and set spatial information; processing the target audio data based on the receiving location of the target audio data and the set spatial information to change the spatial information of the target audio data, thereby obtaining playback audio data output to the receiving location. Embodiments of the present disclosure can improve the realistic effect of audio data output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to technical fields such as voice technology. Background Art

[0002] With the development of computer technology, various branch technologies related to computer data have also continued to develop. This has enabled the simulation and display of various information in people's lives through computer technology. For example, computer technology can generate images on a display screen, conveying and displaying computer-related visual information to users. Computer technology can also generate voice data. For example, by recording sound and playing it back to the user, the original sound effect can be reproduced.

[0003] With the further development of various computer-related technologies, users are constantly generating new requirements for computer simulation information or computer processing information. In order to meet the new requirements continuously generated by users, it is necessary to further improve the technology of computer processing of voice and other information. Summary of the Invention

[0004] The present disclosure provides an audio data processing method, device, electronic device, and computer storage medium.

[0005] According to one aspect of the present disclosure, there is provided an audio data method, comprising:

[0006] Receive target audio data and set spatial information;

[0007] According to the receiving position of the target audio data and the set spatial information, the target audio data is processed to change the spatial information of the target audio data, and the playback audio data output to the receiving position is obtained.

[0008] According to another aspect of the present disclosure, there is provided an audio data device, comprising:

[0009] A spatial information receiving module, configured to receive target audio data and set spatial information;

[0010] The output module is used to process the target audio data according to the receiving position of the target audio data and the set spatial information to change the spatial information of the target audio data, and obtain the playback audio data output to the receiving position.

[0011] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0012] at least one processor; and

[0013] a memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method in any embodiment of the present disclosure.

[0015] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method in any embodiment of the present disclosure.

[0016] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program / instruction, which implements the method in any embodiment of the present disclosure when the computer program / instruction is executed by a processor.

[0017] According to the technology disclosed herein, the spatial information of the target audio data can be changed, so that when the playback audio data generated by processing the target audio data is output to the receiver, the receiver can perceive that the sound source has the set spatial information. This can be applied to a variety of scenarios involving audio and video playback to create preset audio effects and improve user experience.

[0018] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0020] Figure 1 is a schematic diagram of an audio data processing method according to an embodiment of the present disclosure;

[0021] Figure 2 is a schematic diagram of a sound source position and a receiving position according to an example of the present disclosure;

[0022] Figure 3 is a schematic diagram of an audio data processing method according to another embodiment of the present disclosure;

[0023] Figure 4 is a schematic diagram of an audio data processing method according to another embodiment of the present disclosure;

[0024] Figure 5 is a schematic diagram of an audio data processing method according to an example of the present disclosure;

[0025] Figure 6 is a schematic diagram of the division of sound zones according to an example of the present disclosure;

[0026] Figure 7 is a schematic diagram of an audio data processing device according to an embodiment of the present disclosure;

[0027] Figure 8 is a schematic diagram of an audio data processing device according to another embodiment of the present disclosure;

[0028] Figure 9 is a schematic diagram of an audio data processing device according to another embodiment of the present disclosure;

[0029] Figure 10 is a schematic diagram of an audio data processing device according to another embodiment of the present disclosure;

[0030] Figure 11 It is a block diagram of an electronic device used to implement the audio data processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0032] The present disclosure provides an audio data method, such as Figure 1 Shown, including:

[0033] Step S11: receiving target audio data and set spatial information;

[0034] Step S12: processing the target audio data according to the receiving position of the target audio data and the set spatial information to change the spatial information of the target audio data, and obtaining the playback audio data output to the receiving position.

[0035] In this embodiment, the target audio data can be the original audio data to be processed, which may not have any spatial information or may have certain spatial information. For example, if the audio data generated by the original sound source is recorded at a certain position away from the original sound source as the target audio data, the recorded target audio data can reflect the position information between the recording point and the original sound source and the environmental information of the original sound source. If the audio data of the original sound source is directly played or the audio data is recorded at the position of the original sound source as the target audio data, the target audio data may not contain position information. If the environment in which the original sound source is located does not have any sound-blocking substances, the target audio data generated by the original sound source will not have environmental information.

[0036] The set spatial information can be the spatial information that the receiver should perceive, based on the receiver's settings or predetermined reception requirements. For example, when listening to audio data, the receiver can perceive the sound source of the audio data as being located to the left, right, or directly in front of the receiver, closer or farther away. Alternatively, the receiver can perceive the sound source of the audio data as being generated in the metaverse, a closed room, a hall, or an open external environment.

[0037] The Metaverse, as described above, can be a virtual world connected and created through technological means, mirroring and interacting with the real world, a digital living space with a new social system. Essentially, the Metaverse is a virtualization and digitization of the real world, requiring significant changes to content production, economic systems, user experience, and physical content. However, the development of the Metaverse is gradual, ultimately taking shape through the continuous integration and evolution of numerous tools and platforms, supported by shared infrastructure, standards, and protocols. It leverages extended reality technology to provide an immersive experience, digital twin technology to generate a mirror image of the real world, and blockchain technology to build an economic system. It seamlessly integrates the virtual and real worlds across economic, social, and identity systems, while allowing every user to create content and edit the world.

[0038] The spatial information may include position information, such as at least one of absolute coordinates, relative coordinates, a distance from a preset reference point, a relative angle from a preset reference point, and the like.

[0039] The spatial information may also include spatial perception information that is desired to be generated for the receiver, or spatial perception information that is desired to be generated by the receiver for the sound source.

[0040] For example, if the preset spatial information includes coordinates (x, y) in the world coordinate system, it means that the receiver is expected to perceive that the sound source is approximately located at coordinates (x, y) when listening to the audio data played back after processing the target audio data. For another example, if the preset spatial information includes a relative distance L, it means that the receiver is expected to perceive that the sound source is approximately located at a distance L from the receiver when listening to the audio data played back after processing the target audio data.

[0041] For example, if the preset spatial information includes the south direction in the world coordinate system, it is expected that the recipient, when listening to the audio data generated after processing the target audio data, will be able to perceive the sound source as being approximately located in the south direction. For another example, if the preset spatial information includes the direction directly in front of the recipient, it is expected that the recipient, when listening to the audio data generated after processing the target audio data, will be able to perceive the sound source as being approximately located in front of them.

[0042] For another example, if the preset spatial information includes a cube-shaped space with a volume of approximately V, it means that the receiver is expected to be able to perceive the sound source and that they are in the cube-shaped space with a volume of approximately V when listening to the audio data played after the target audio data is processed. For another example, if the preset spatial information includes a spherical space with a volume of approximately V, it means that the receiver is expected to be able to perceive the sound source and that they are in the spherical space with a volume of approximately V when listening to the audio data played after the target audio data is processed.

[0043] In this embodiment, the receiving position of the target audio can be an absolute position or a relative position. For example, the receiver's position is always relative to the origin. For another example, the receiver's position can be at a certain coordinate in a preset world coordinate system or a certain coordinate in a relative coordinate system. The receiving position can be obtained through configuration information or determined through positioning data.

[0044] Processing the target audio data to change the spatial information of the target audio data based on the receiving position and the set spatial information of the target audio may include: determining an adjustment amount corresponding to the spatial information to be adjusted or added to the target audio data based on the receiving position and the set spatial information, and processing the target audio data.

[0045] For example, when the target audio data is output at the playback position, it does not have any spatial information. The adjustment amount corresponding to the set spatial information can be added to the target audio data to add the spatial information, so that after the output playback audio data is listened to by the receiver, the listener can perceive that the sound source of the playback audio data has the set spatial information.

[0046] For another example, when the target audio data is output at the playback position, it has original spatial information. Then, the adjustment amount corresponding to the set spatial information can be added to the target audio data to change the spatial information, so that after the output playback audio data is listened to by the receiver, the listener can perceive that the sound source of the playback audio data has the set spatial information.

[0047] like Figure 2 As shown, assuming that the target audio data does not originally have spatial information, Figure 2 When playing at position A in the coordinate system shown, and the receiver is at position B in the coordinate system, the receiver can perceive the sound source as being located at position A in the coordinate system. By modifying the spatial information of the target audio data, the receiver can perceive different spatial information about the sound source. For example, the receiver can perceive the sound source as being located in a room at position C in the relative coordinate system.

[0048] Alternatively, the receiver is in Figure 2At position B in the coordinate system shown, audio data is directly output to the receiver at position B. Spatial information can be added to the target audio data or the original spatial information of the target audio data can be changed so that the receiver can perceive that the sound source is at position C when receiving the played audio data.

[0049] In a specific implementation, the target audio data is processed according to the receiving position of the target audio and the set spatial information. This can be done by using a pre-trained model to perform calculations based on the input receiving position, the set spatial information and the target audio data, and output the playback audio data.

[0050] In another specific implementation method, the target audio data is processed according to the receiving position of the target audio and the set spatial information. The sound wave data corresponding to the receiving position and the set spatial information can be pre-recorded, and the sound wave data is mixed with the target audio data to obtain the playback audio data.

[0051] In this embodiment, the spatial information of the target audio data can be changed so that when the playback audio data generated by processing the target audio data is output to the receiver, the receiver can perceive that the sound source has the set spatial information. This can be applied to a variety of scenarios involving audio and video playback to create preset audio effects and improve user experience.

[0052] In one embodiment, processing the target audio data according to the receiving position of the target audio data and the set spatial information includes:

[0053] According to the receiving position and the set spatial information, the target data is searched in the preset corresponding relationship;

[0054] The target audio data is processed according to the target data.

[0055] In this embodiment, the target data may be at least one of a processing parameter or a processing function. For example, when adding or modifying spatial information of the target audio data, if the preset operation is a filtering operation, the target data may be a parameter of the filtering operation. For another example, if the operation for adding or modifying spatial information of the target audio data is not set, the target data may be determined to include a convolution calculation and convolution calculation parameters. When processing the target audio data based on the target data, the convolution calculation may be performed on the target audio data according to the convolution calculation parameters.

[0056] In a possible implementation, the target data may include processing parameters. In this case, what processing to perform on the target data may be predetermined or defaulted.

[0057] In another possible implementation, the target data may include a processing method and processing parameters. In this case, what processing is performed on the target data may not be determined in advance.

[0058] In this embodiment, a correspondence between the receiving position, the set spatial information, and the target data for processing the target audio data can be pre-established. After determining the receiving position and the set spatial information, the preset correspondence can be searched to quickly determine how to process the target audio data.

[0059] In one embodiment, Figure 3 As shown, the set spatial information includes a set position; searching for target data in a preset correspondence according to the received position and the set spatial information includes:

[0060] Step S31: Determine the relative position based on the received position and the set position;

[0061] Step S32: determining target data according to the relative position and the preset corresponding relationship;

[0062] Step S33: Process the target audio data according to the target data.

[0063] In this embodiment, at least one of a relative distance and a relative angle between the received position and the set position in the set spatial information may be determined as the relative position.

[0064] For example, the receiving position may be a coordinate in a set coordinate system, the set position may be another coordinate in the set coordinate system, and the relative position may be the difference between the two coordinates. The set coordinate system may be a plane coordinate system or a three-dimensional coordinate system.

[0065] For another example, one of the receiving position and the setting position may be a relative origin, and the other of the receiving position and the setting position may be a coordinate relative to the relative origin.

[0066] In one implementation, the relative position within a certain angle range and distance range can be divided into at least one sound zone based on a combination of the angle and distance between the receiver and the set position, with the receiver or set position as the origin. For example, when the relative distance between the receiver and the set position is within a first range and the angle is within a second range, the receiver or set position is determined to be in the first sound zone; when the relative distance between the receiver and the set position is within a third range and the angle is within a fourth range, the receiver or set position is determined to be in the second sound zone, and so on. The corresponding relationship is recorded as a corresponding relationship between the sound zone and the target data. Therefore, after determining the receiving position and the set position, the sound zone corresponding to the relative distance between the two can be determined, and the target data can be further determined by searching the corresponding relationship.

[0067] The preset correspondence may be a correspondence between relative positions and target data. For example, the relative distance may be divided into N ranges, each corresponding to a type of target data. For another example, the relative angle may be divided into M ranges, each corresponding to a type of target data.

[0068] In another possible implementation, the set spatial information includes not only the set location but also the set environment. For example, the set environment can be set to be in a valley, an open plain, a seaside, a metaverse, etc. Each set environment can correspond to the spatial environment effect when the sound data is played, such as the presence of walls, natural mountain sides, water flow and direction, or the surrounding metaverse. Specifically, for example, a certain room can be constructed to simulate environments such as valleys, plains, seaside, and metaverse, thereby obtaining target data corresponding to the set environment.

[0069] In this embodiment, the target data is queried according to the relative position, so that the target audio data can be processed quickly.

[0070] In one embodiment, processing the target audio data according to the target data includes:

[0071] Determine the filtering operation according to the target data;

[0072] The target audio data is processed by performing a filtering operation.

[0073] In this embodiment, the filtering operation is determined according to the target data, and when the filtering operation is a default operation, the target data is used as a parameter of the filtering operation.

[0074] In this embodiment, the filtering operation is determined based on the target data. Alternatively, when there is no default operation, the specific type of the operation to be performed is determined to be a filtering operation, as well as the specific parameters of the filtering operation to be performed.

[0075] In this embodiment, the spatial information of the target audio data can be changed through the filtering operation, so that the spatial information of the playback audio data obtained after processing the target audio data can be close to the set spatial information.

[0076] In one embodiment, the relative position includes at least one of a relative distance and a relative angle.

[0077] In a specific implementation, the relative position also includes a combination of relative distance and relative angle.

[0078] The relative distance and relative angle can be the distance or angle in a plane coordinate system or the distance or angle in a three-dimensional coordinate system.

[0079] In this embodiment, the target audio data can be processed to generate distance and angle effects corresponding to the set spatial information, so that the audio playback product can produce sound effects close to the actual scene and meet more user needs.

[0080] In one embodiment, Figure 4 As shown, the audio data processing method further includes:

[0081] Step S41: Acquire first audio data received by a receiving device located at a receiving location; the first audio data is generated at the receiving location by second audio data played by a sound source located at a predetermined spatial information;

[0082] Step S42: Obtain target data according to the first audio data and the second audio data;

[0083] Step S43: Generate a preset corresponding relationship according to the target data, the preset spatial information and the receiving position.

[0084] The above steps may be performed before receiving the target audio data to predetermine the corresponding relationship.

[0085] When playing audio data in a real-world environment, the receiver's listening experience will vary depending on the playback environment. For example, if the sound source is closer to the receiver, the volume will be louder, and the receiver will perceive the sound source as being closer. If the sound source is farther away, the volume will be lower, and the receiver will perceive the sound source as being farther away. Furthermore, if the sound source is located at different angles to the receiver, the receiver will also perceive the sound source in different directions.

[0086] In this embodiment, the first audio data may be audio sound wave data received at the receiving location. The second audio data may be the original audio data played by the sound source, that is, audio data received at the sound source location without any playback environment information. Specifically, the first audio data is audio data with the spatial information of the sound source and the spatial information of the receiver added to the second audio data.

[0087] In a specific implementation, a room for generating corresponding sound effects can be manually set according to a preset environment corresponding to preset spatial information, a sound source can be set at a set position in the room, and audio data generated by the sound source can be received at a receiving position to obtain the above-mentioned first audio data.

[0088] Obtaining the target data based on the first and second audio data may involve determining a difference in spatial information between the first and second audio data to obtain the target data. Thus, after processing other audio data lacking spatial information according to the target data, the other audio data can have the difference in spatial information between the first and second audio data. When a receiver receives the processed other audio data at the receiving location, the receiver can perceive that the sound source of the other audio data has similar spatial information as the first audio data.

[0089] In this embodiment, audio data can be received while simulating real spatial information, and target data can be determined based on the difference between the received audio data and the original audio data played by the sound source. A correspondence between the target data and the set spatial information and receiving position is established, thereby ensuring that other audio data, after being processed according to the target data, can produce a spatial information effect similar to the first audio data.

[0090] In one embodiment, obtaining target data according to the first audio data and the second audio data includes:

[0091] Restoring the first audio data into second audio data to determine parameters of the restoration operation;

[0092] The target data is obtained according to the parameters of the restore operation.

[0093] Restoring the first audio data to the second audio data to determine the parameters of the restoration operation may include using a certain processing method to restore the first audio data to the second audio data, and determining the operation parameters used in the restoration process. When the target audio data is subsequently processed, the same processing method and operation parameters may be used.

[0094] In this embodiment, the restoration operation may be an operation opposite to the operation of processing the target audio data, such as a filtering operation, a deconvolution operation, or the like, which may change the spatial information of the audio data.

[0095] In this embodiment, the second audio data and the target data are obtained by restoring the first audio data, so that the correspondence between the target data and the set spatial information and the receiving position can be recorded subsequently. The correspondence can be used to determine how to process the target audio data, thereby improving the efficiency of processing the target audio data.

[0096] In one embodiment, obtaining target data according to parameters of the restoration operation includes:

[0097] Acoustic characteristic parameters are obtained according to the calibration head-related transfer function and the parameters of the restoration operation;

[0098] According to the acoustic characteristic parameters, the target data is obtained.

[0099] In this embodiment, the calibrated head-related transfer function can be obtained by calibrating the head-related transfer function. The head-related transfer function (HRTF) in this embodiment can also be called ATF (Anatomical Transfer Function), which is a sound localization algorithm.

[0100] When the restoration operation is an operation such as deconvolution, the parameters of the restoration operation may be acoustic feature parameters to be calibrated, and specifically may include at least one of a room acoustic impulse response (RIR), a spatial orientation feature, and the like.

[0101] In this embodiment, the spatial orientation feature is a feature that represents the position information of the audio data, for example, a feature that represents information such as the distance and angle of the audio source.

[0102] In this embodiment, the acoustic characteristic parameters are obtained according to the calibration head-related transfer function and the parameters of the restoration operation. The acoustic characteristic parameters may be obtained by calibrating and calculating the parameters of the restoration operation using the calibration head-related transfer function.

[0103] In this embodiment, the parameters of the restoration operation can be calibrated, thereby producing a relatively accurate effect for most user groups.

[0104] In one embodiment, obtaining target data according to acoustic characteristic parameters includes:

[0105] Performing dynamic equalization processing on the acoustic characteristic parameters to obtain processed acoustic characteristic parameters;

[0106] The processed acoustic feature parameters are used as target data.

[0107] In this embodiment, dynamic equalization (EQ) processing is performed on the acoustic characteristic parameters, and a dynamic equalizer can be used to perform dynamic equalization processing according to the frequency band of the audio data, so that the acoustic characteristic parameters of each frequency band can transition smoothly, avoiding sudden changes or jumps in the playback effect of the processed audio data.

[0108] In one embodiment, the calibrated head-related transfer function is calculated by weighted averaging a plurality of individualized head-related transfer functions.

[0109] In this embodiment, HRTF functions corresponding to different groups of people can be obtained from an open source HRTF platform, and multiple HRTF functions can be weighted averaged and normalized into an overall function, so that the calibrated head-related transfer function can be as consistent as possible with the perceptual characteristics of the general public.

[0110] In one embodiment, the set spatial information also includes a set environment.

[0111] The set environment can be, for example, a metaverse environment, a valley environment, a closed room environment, a conference room environment, a studio environment, a stadium environment, an open space, etc.

[0112] The set environment can include both conventional propagation media for audio data, such as the Earth's air layer, and unconventional propagation media, such as solids, special gases, and liquids. Therefore, by configuring the set environment, it is possible to simulate sounds produced in a variety of different environments, thereby enhancing the authenticity of the user's experience in games, movies, and simulations.

[0113] In this embodiment, the target audio data may be processed so that the obtained playback audio data has the effect of being played in a set environment.

[0114] In one example of the present disclosure, the audio data processing method includes the following steps: Figure 5 Steps shown:

[0115] Step S51: creating a virtual space direction filter database.

[0116] In this example, HRTF open source data is combined with artificial ears to model and restore the recorded data.

[0117] People have two ears, but they can locate sounds from three-dimensional space, thanks to the human ear's analysis system for sound signals. The spatial information of the signal transmitted from any point in space to the human ear (in front of the eardrum) can be described or generated by the operation of a filtering system. The original audio data of the sound source is processed by the filter to obtain the audio data received by the eardrums of the two ears at the receiving position. If this set of filters (transfer functions) that describe the spatial information is obtained, that is, a specific HRTF (HEAD RELATED TRANSFER FUNCTION), the sound signal from this direction in space can be restored. In general, HRTF is highly personalized, so in this disclosed example, multiple HRTFs in the open source HRTF dataset are used, and normalized calculations are performed according to certain weights to obtain a set of HRTFs as calibration HRTFs to achieve a situation that is suitable for most people.

[0118] When establishing a correspondence between the set spatial information, receiving location, and target data, specific audio is played through high-fidelity speakers. High-fidelity speakers can minimize the impact of the speakers themselves on the spatial information of the audio data. The set audio can include human voices, small-frequency signals (signals below a certain frequency threshold), and white noise, thereby covering various frequency bands. A self-made artificial ear is used as a receiver. Multiple sets of stereo data are collected at preset recording points (i.e., the receiving locations in the aforementioned embodiment), and modeling is performed based on the collected stereo data. The stereo data is equivalent to the first audio data in the aforementioned embodiment. Its relevant acoustic characteristics are restored through deconvolution. Acoustic characteristics may include RIR, spatial orientation characteristics, etc., and are processed by calibrating HRTFs for orientation calibration and creating a more immersive room experience. Filter coefficients for multiple orientations (sound zones or relative positions) are obtained, i.e., the acoustic characteristic parameters in the aforementioned embodiment. At the same time, based on the spectral distribution of the aforementioned specific audio, dynamic EQ adjustment is performed on the acoustic characteristic parameters to obtain the final filter coefficients, i.e., the target data. In this example, filter coefficients may be collected for each relative position within a certain spatial range, and corresponding relationships may be established to form a set of spatial directionality filter databases.

[0119] Step S52: Establish a sound zone.

[0120] In actual use, the user's position changes in real time. In this example, based on this and combined with the sensitivity of the human ear to position, multiple sound zones are divided according to the direction and distance of the sound source. Figure 6 As shown, each sector area can correspond to a sound zone, including the sector area 62 closest to the center of the circle to the sector area 61 closest to the circumference. When dividing the sound zone, the receiving position can be relatively fixed and set as the relative origin. For the entire area within a distance range near the origin, that is, Figure 6 The circular area shown is used to divide the sound zones.

[0121] Still refer to Figure 6 When dividing the sound zones, first, a circular area with a radius of a set length and centered at the origin can be divided into multiple large preliminary sector-shaped areas at 45° intervals (or other values between 1° and 360°, such as 5°, 10°, 15°, 20°, 25°, 30°, 35°, and 40°). These preliminary sector-shaped areas can then be further divided into multiple sound zones at 0.5-meter intervals (or other lengths within the range of 0.1 to 10 meters, such as 0.2 meters).

[0122] For another example, a circular area with a radius of 0.5 meters (or other intervals such as 0.2 meters) centered at the origin can be divided into a tone zone every 45 degrees, for a total of 8 tone zones. An annular area with a radius of 0.5-1.0 meters centered at the origin can be divided into a tone zone every 45 degrees, for a total of 8 tone zones. Similarly, annular areas with radii of 1.0-1.5, 1.5-2.0, 2.0-2.5, etc. can each be divided into 8 tone zones.

[0123] Step S53: performing audio zone aggregation according to the received audio stream and direction information.

[0124] In practice, the RTC (Real-Time Clock) of received audio data varies and may not be regular. Furthermore, the number and location of users are dynamically changing, and the number of audio zones and their division can be fixed by default. Therefore, in this example, audio zone aggregation is used to pre-process the received audio streams. By analyzing the spatial orientation information of each audio stream and assigning it to the corresponding audio zones according to a specific algorithm, a zone_table (a table of correspondence between audio zones and target data) is maintained to record the correspondence between the corresponding audio zones and the target data used to process the target audio data. Multiple user audio streams are aggregated according to the audio zones before proceeding to the next step of processing.

[0125] Step S54: audio data processing.

[0126] The corresponding sound zone corresponds to a set of spatial orientation filter coefficients. The sound zone divided in step S53 and the corresponding target data are used to perform filter convolution calculation or filtering operation to generate spatial sound effect audio data in the corresponding direction, that is, playback audio data.

[0127] The present disclosure also provides an audio data processing device, such as Figure 7 Shown, including:

[0128] The spatial information receiving module 71 is used to receive target audio data and set spatial information;

[0129] The output module 72 is used to process the target audio data according to the receiving position of the target audio data and the set spatial information to change the spatial information of the target audio data, and obtain the playback audio data output to the receiving position.

[0130] In one embodiment, Figure 8 As shown, the output module includes:

[0131] A search unit 81 is configured to search for target data in a preset correspondence according to the received position and the set spatial information;

[0132] The search result processing unit 82 is used to process the target audio data according to the target data.

[0133] In one embodiment, the set spatial information includes a set location; and the search unit is further configured to:

[0134] Determine the relative position based on the receiving position and the set position;

[0135] Determine target data based on relative position and preset corresponding relationship;

[0136] The target audio data is processed according to the target data.

[0137] In one embodiment, the search unit is further configured to:

[0138] Determine the filtering operation according to the target data;

[0139] The target audio data is processed by performing a filtering operation.

[0140] In one embodiment, the relative position includes at least one of a relative distance and a relative angle.

[0141] In one embodiment, Figure 9 As shown, the audio data processing device further includes:

[0142] The audio data acquisition module 91 is configured to acquire first audio data received by a receiving device located at a receiving location; the first audio data is generated at the receiving location by second audio data played by a sound source located at a predetermined spatial information;

[0143] A target data acquisition module 92 is configured to obtain target data based on the first audio data and the second audio data;

[0144] The correspondence generation module 93 is configured to generate a preset correspondence according to the target data, the preset spatial information and the receiving position.

[0145] In one embodiment, Figure 10 As shown, the target data acquisition module includes:

[0146] A parameter determination unit 101 is configured to restore the first audio data into the second audio data to determine parameters of the restoration operation;

[0147] The parameter processing unit 102 is used to obtain target data according to the parameters of the restoration operation.

[0148] In one embodiment, the parameter processing unit is further configured to:

[0149] Acoustic characteristic parameters are obtained according to the calibration head-related transfer function and the parameters of the restoration operation;

[0150] According to the acoustic characteristic parameters, the target data is obtained.

[0151] In one embodiment, the parameter processing unit is further configured to:

[0152] Performing dynamic equalization processing on the acoustic characteristic parameters to obtain processed acoustic characteristic parameters;

[0153] The processed acoustic feature parameters are used as target data.

[0154] In one embodiment, the calibrated head-related transfer function is calculated by weighted averaging a plurality of individualized head-related transfer functions.

[0155] In one embodiment, the set spatial information also includes a set environment.

[0156] The embodiments of the present disclosure can be applied to the field of computer technology, and in particular, can be applied to the field of speech processing technology.

[0157] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0158] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0159] Figure 11 A schematic block diagram of an example electronic device 110 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0160] like Figure 11As shown, the device 110 includes a computing unit 111, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 112 or a computer program loaded from a storage unit 118 into a random access memory (RAM) 113. Various programs and data required for the operation of the device 110 can also be stored in the RAM 113. The computing unit 111, the ROM 112, and the RAM 113 are connected to each other via a bus 114. An input / output (I / O) interface 115 is also connected to the bus 114.

[0161] Various components in device 110 are connected to I / O interface 115, including an input unit 116, such as a keyboard and mouse; an output unit 117, such as various types of displays and speakers; a storage unit 118, such as a magnetic disk and optical disk; and a communication unit 119, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 119 allows device 110 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0162] The computing unit 111 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 111 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 111 performs the various methods and processes described above, such as the audio data processing method. For example, in some embodiments, the audio data processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 118. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 110 via the ROM 112 and / or the communication unit 119. When the computer program is loaded into the RAM 113 and executed by the computing unit 111, one or more steps of the audio data processing method described above can be performed. Alternatively, in other embodiments, the computing unit 111 can be configured to perform the audio data processing method by any other appropriate means (e.g., by means of firmware).

[0163] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0164] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0165] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0167] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0168] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0169] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0170] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for processing audio data, comprising: Acquiring first audio data received by a receiving device disposed at a receiving position; The first audio data is formed at the receiving position by the second audio data played by a sound source set to the set spatial information; Obtaining target data according to the first audio data and the second audio data; generating a preset corresponding relationship according to the target data, the set spatial information, and the receiving location; Receive target audio data and set spatial information; processing the target audio data according to the receiving position of the target audio data and the set spatial information to change the spatial information of the target audio data, thereby obtaining playback audio data output to the receiving position; The step of obtaining target data according to the first audio data and the second audio data includes: Restoring the first audio data to the second audio data to determine parameters of the restoration operation; Obtaining target data according to the parameters of the restoration operation; The processing of the target audio data according to the receiving position of the target audio data and the set spatial information includes: Searching for target data in the preset corresponding relationship according to the receiving position and the set spatial information; The target audio data is processed according to the target data.

2. The method according to claim 1, wherein The set spatial information includes a set location; and searching for target data in a preset corresponding relationship according to the received location and the set spatial information includes: determining a relative position according to the received position and the set position; determining target data according to the relative position and the preset corresponding relationship; The target audio data is processed according to the target data.

3. The method according to claim 2, wherein: The processing of the target audio data according to the target data includes: determining a filtering operation according to the target data; The target audio data is processed by performing the filtering operation.

4. The method according to claim 2 or 3, wherein: The relative position includes at least one of a relative distance and a relative angle.

5. The method according to claim 1, wherein Obtaining target data according to the parameters of the restoration operation includes: Obtaining acoustic characteristic parameters according to the calibration head-related transfer function and the parameters of the restoration operation; Target data is obtained according to the acoustic characteristic parameters.

6. The method according to claim 5, wherein: Obtaining target data according to the acoustic characteristic parameters includes: Performing dynamic equalization processing on the acoustic characteristic parameters to obtain processed acoustic characteristic parameters; The processed acoustic feature parameters are used as target data.

7. The method according to claim 5 or 6, wherein: The calibrated head-related transfer function is obtained by weighted averaging a plurality of individualized head-related transfer functions.

8. The method according to any one of claims 1, 2, 3, 5, and 6, wherein: The set spatial information includes a set environment.

9. An audio data processing device, comprising: an audio data acquisition module, configured to acquire first audio data received by a receiving device disposed at a receiving position; The first audio data is formed at the receiving position by the second audio data played by a sound source set to the set spatial information; a target data acquisition module, configured to obtain target data according to the first audio data and the second audio data; a correspondence generation module, configured to generate a preset correspondence based on the target data, the set spatial information, and the receiving position; A spatial information receiving module, configured to receive target audio data and set spatial information; an output module, configured to process the target audio data according to the receiving position of the target audio data and the set spatial information to change the spatial information of the target audio data, and obtain playback audio data output to the receiving position; Wherein, the target data acquisition module includes: a parameter determination unit, configured to restore the first audio data to the second audio data to determine parameters of the restoration operation; a parameter processing unit, configured to obtain target data according to the parameters of the restoration operation; Wherein, the output module includes: a search unit, configured to search for target data in the preset correspondence according to the receiving position and the set spatial information; The search result processing unit is used to process the target audio data according to the target data.

10. The device according to claim 9, wherein The set spatial information includes a set location; the search unit is further configured to: determining a relative position according to the received position and the set position; determining target data according to the relative position and the preset corresponding relationship; The target audio data is processed according to the target data.

11. The device according to claim 10, wherein The search unit is further configured to: determining a filtering operation according to the target data; The target audio data is processed by performing the filtering operation.

12. The device according to claim 10 or 11, wherein The relative position includes at least one of a relative distance and a relative angle.

13. The device according to claim 9, wherein The parameter processing unit is further configured to: Obtaining acoustic characteristic parameters according to the calibration head-related transfer function and the parameters of the restoration operation; Target data is obtained according to the acoustic characteristic parameters.

14. The device according to claim 13, wherein The parameter processing unit is further configured to: Performing dynamic equalization processing on the acoustic characteristic parameters to obtain processed acoustic characteristic parameters; The processed acoustic feature parameters are used as target data.

15. The device according to claim 13 or 14, wherein The calibrated head-related transfer function is obtained by weighted averaging a plurality of individualized head-related transfer functions.

16. The device according to any one of claims 9, 10, 11, 13, and 14, wherein: The set spatial information includes a set environment.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

19. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Electronic apparatus, control method thereof, and recording medium

    US20210337340A1