Method, apparatus, and storage medium for generating object audio data
By acquiring the sound data and current location information of the sound object, the object's audio data is synthesized in real time, solving the problem that existing technologies cannot record the object's audio data in real time, and achieving accurate location information acquisition and data generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2022-05-05
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies lack methods for real-time recording of audio data of sound objects, making it impossible to accurately obtain the location information of each sound object.
By acquiring the sound data and current location information of the sound object, the audio data of the object is synthesized in real time. The location information of the recording terminal is acquired using one-way, two-way or mixed transceiver methods, and the sound data is synchronized and synthesized with the location information.
It enables real-time and accurate acquisition of the location information of each sound object and generation of object audio data, solving the problem that existing technologies cannot record object audio data in real time.
Smart Images

Figure CN117355894B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of communication, and particularly relates to a method and device for generating object audio data, an electronic device and a storage medium. BACKGROUND
[0002] MPEG (Moving Picture Experts Group) next generation audio codec standard MPEG-H 3D Audio is an ISO / IEC 23008-3 international standard, in which a brand new audio format, object audio, is used. The object audio can mark the direction of sound, so that the listener can hear the sound coming from a specific direction, no matter whether the listener uses a headphone or a sound system, and no matter how many speakers the sound system has.
[0003] In the related art, a single-channel audio is pre-recorded, and then combined with the position information of the pre-prepared single-channel audio to generate object audio data. However, this method needs to be made by a production device in the later stage, and there is still a lack of a method for real-time recording of sound objects. SUMMARY
[0004] The embodiments of the present disclosure provide a method and device for generating object audio data, an electronic device and a storage medium, which can accurately obtain the position information of each sound object in real time, and generate object audio data in real time.
[0005] In a first aspect, the embodiments of the present disclosure provide a method for generating object audio data, which comprises: obtaining sound data of at least one sound object; obtaining current position information of the at least one sound object; and synthesizing the sound data and the current position information of the at least one sound object to generate object audio data.
[0006] In the technical solution, the sound data of at least one sound object is obtained, the current position information of the at least one sound object is obtained, and the sound data and the current position information of the at least one sound object are synthesized to generate object audio data. Thus, the position information of each sound object can be accurately obtained in real time, and object audio data can be generated in real time.
[0007] In some embodiments, the obtaining of the current position information of the at least one sound object comprises: obtaining the current position information of at least one recording terminal for recording the sound data of the at least one sound object.
[0008] In some embodiments, before the synthesizing of the sound data of the at least one sound object and the current position information, the method further comprises: synchronizing the sound data of the at least one sound object and the current position information.
[0009] In some embodiments, the obtaining of the current position information of the at least one recording terminal recording the sound data of the at least one sound object comprises: obtaining the current position information of the at least one recording terminal in a one-way transceiving mode, a two-way transceiving mode or a hybrid transceiving mode.
[0010] In some embodiments, the obtaining of the position information of the at least one recording terminal in the hybrid transceiving mode comprises: obtaining first positioning reference information in the one-way transceiving mode; obtaining second positioning reference information in the two-way transceiving mode; and determining the current position information of the at least one recording terminal according to the first positioning reference information and the second positioning reference information.
[0011] In some embodiments, the first positioning reference information is one of angle information and distance information, and the second positioning reference information is the other of the angle information and the distance information.
[0012] In some embodiments, the obtaining of the current position information of the at least one recording terminal in the one-way transceiving mode comprises: receiving a first positioning signal sent by the at least one recording terminal in a broadcast mode, and generating the current position information of the at least one recording terminal according to the first positioning signal.
[0013] In some embodiments, the obtaining of the position information of the at least one recording terminal in the two-way transceiving mode comprises: receiving a positioning initiation signal sent by the at least one recording terminal in a broadcast mode; sending a response signal to the at least one recording terminal; receiving a second positioning signal sent by the at least one recording terminal, and generating the current position information of the at least one recording terminal according to the second positioning signal.
[0014] In some embodiments, each of the recording terminals corresponds to a sound object, and the position of the recording terminal moves with the sound source of the sound object.
[0015] In some embodiments, the method further comprises: obtaining initial position information of the at least one sound object.
[0016] In some embodiments, the synthesizing the sound data and the current position information of the at least one sound object to generate the object audio data comprises: obtaining audio parameters and taking the audio parameters as header information of the object audio data; at each sampling time, saving sound data of each sound object as object audio signals and saving the current position information as object audio auxiliary data to generate the object audio data.
[0017] In some embodiments, the saving the sound data and the current position information in units of frames is further included.
[0018] In a second aspect, the embodiments of the present disclosure provide an object audio data generation apparatus, which comprises: a data obtaining unit configured to obtain sound data of at least one sound object; an information obtaining unit configured to obtain current position information of the at least one sound object; and a data generation unit configured to synthesize the sound data and the current position information of the at least one sound object to generate object audio data.
[0019] In some embodiments, the information obtaining unit is specifically configured to obtain current position information of at least one recording terminal recording the sound data of the at least one sound object.
[0020] In some embodiments, the apparatus further comprises a synchronization processing unit configured to synchronize the sound data of the at least one sound object and the current position information.
[0021] In some embodiments, the information obtaining unit is specifically configured to obtain the current position information of the at least one recording terminal in a one-way transceiving mode, a two-way transceiving mode or a mixed transceiving mode.
[0022] In some embodiments, the information obtaining unit comprises: a first information obtaining module configured to obtain first positioning reference information in the one-way transceiving mode; a second information obtaining module configured to obtain second positioning reference information in the two-way transceiving mode; and a first current information obtaining module configured to determine the current position information of the at least one recording terminal according to the first positioning reference information and the second positioning reference information.
[0023] In some embodiments, the first positioning reference information is one of angle information and distance information, and the second positioning reference information is the other of the angle information and the distance information.
[0024] In some embodiments, the information obtaining unit comprises a second current information obtaining module configured to receive a first positioning signal broadcast by the at least one sound recording terminal and generate current position information of the at least one sound recording terminal according to the first positioning signal.
[0025] In some embodiments, the information obtaining unit comprises a signal receiving module configured to receive a positioning initiation signal broadcast by the at least one sound recording terminal, a signal sending module configured to send a response signal to the at least one sound recording terminal, and a third current information obtaining module configured to receive a second positioning signal sent by the at least one sound recording terminal and generate current position information of the at least one sound recording terminal according to the second positioning signal.
[0026] In some embodiments, each sound recording terminal corresponds to a sound object, and the position of the sound recording terminal moves with a sound source of the sound object.
[0027] In some embodiments, the apparatus further comprises an initial position obtaining unit configured to obtain initial position information of the at least one sound object.
[0028] In some embodiments, the data generating unit comprises a parameter obtaining module configured to obtain audio parameters and use the audio parameters as header file information of the object audio data, and an audio data generating module configured to save sound data of each sound object as object audio signals and save the current position information as object audio auxiliary data at each sampling time to generate the object audio data.
[0029] In some embodiments, the data generating unit further comprises a processing module configured to save the sound data and the current position information in units of frames.
[0030] In a third aspect, an electronic device is provided, which comprises at least one processor and a memory connected to the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.
[0031] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of the first aspect.
[0032] In a fifth aspect, the embodiments of the present disclosure provide a computer program product, comprising computer instructions, wherein the computer instructions realize the method in the first aspect when executed by a processor.
[0033] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the background art, the drawings needed to be used in the embodiments of the present disclosure or the background art will be described below.
[0035] Figure 1 is a flowchart of a method for generating object audio data provided by the embodiments of the present disclosure;
[0036] Figure 2 is a flowchart of another method for generating object audio data provided by the embodiments of the present disclosure;
[0037] Figure 3 is a flowchart of still another method for generating object audio data provided by the embodiments of the present disclosure;
[0038] Figure 4 is a flowchart of still another method for generating object audio data provided by the embodiments of the present disclosure;
[0039] Figure 5 is a flowchart of still another method for generating object audio data provided by the embodiments of the present disclosure;
[0040] Figure 6 is a flowchart of still another method for generating object audio data provided by the embodiments of the present disclosure;
[0041] Figure 7 is a structural diagram of a device for generating object audio data provided by the embodiments of the present disclosure;
[0042] Figure 8 is a structural diagram of another device for generating object audio data provided by the embodiments of the present disclosure;
[0043] Figure 9 is a structural diagram of an information acquisition unit in a device for generating object audio data provided by the embodiments of the present disclosure;
[0044] Figure 10 is a structural diagram of another information acquisition unit in a device for generating object audio data provided by the embodiments of the present disclosure;
[0045] Figure 11 is a structural diagram of still another information acquisition unit in a device for generating object audio data provided by the embodiments of the present disclosure;
[0046] Figure 12 is a structural diagram of another object audio data generation apparatus provided by an embodiment of the present disclosure;
[0047] Figure 13 is a structural diagram of a data generation unit in the object audio data generation apparatus provided by an embodiment of the present disclosure;
[0048] Figure 14 is a structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0049] In order to make the ordinary person skilled in the art better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings.
[0050] Unless otherwise required by the context, throughout the specification and claims, the term "comprising" is to be interpreted as meaning "including but not limited to". In the description of the specification, the terms "some embodiments", "an embodiment", "one embodiment" or the like are intended to mean that the particular feature, structure, material or characteristic being described is included in at least one embodiment or example of the present disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. In addition, the particular features, structures, materials or characteristics described can be included in any suitable way in one or more embodiments or examples.
[0051] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the drawings are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. The terms "first", "second" are used for description purposes only, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0052] At least one of the present disclosure can also be described as one or more, multiple can be two, three, four or more, the present disclosure does not make restrictions. In the embodiment of the present disclosure, for a technical feature, the technical features in the technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D" and the like. The "first", "second", "third", "A", "B", "C" and "D" described technical features have no order or size order.
[0053] The correspondence relationship shown in each table in the present disclosure can be configured or predefined. The values of the information in each table are only examples, and other values can be configured, and the present disclosure does not limit. When configuring the correspondence relationship between the information and each parameter, it is not necessarily required to configure all the correspondence relationships shown in each table. For example, in the table in the present disclosure, the correspondence relationship shown in some rows can also not be configured. For another example, the above table can be appropriately deformed and adjusted, for example, split, merged, etc. The parameter name shown in the title of each table above can also use other names understandable by the communication device, and the parameter value or representation method can also use other values or representation methods understandable by the communication device. Each table above can also use other data structures when implemented, for example, arrays, queues, containers, stacks, linear tables, pointers, linked lists, trees, graphs, structures, classes, heaps, hash tables, or hash tables, etc.
[0054] Those skilled in the art can appreciate that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.
[0055] For the related art, the object audio data acquisition method cannot realize the direct recording of the object audio data, and cannot obtain the real sound object position information. The embodiment of the present disclosure provides an object audio data generation method, device, electronic equipment and storage medium to obtain the position information of each sound object in real time and accurately, and to record and generate object audio data in real time, to solve the problem in the related art.
[0056] Specifically, the object audio data generation method, device, electronic equipment and storage medium provided by the embodiment of the present disclosure are specifically described below with reference to the drawings.
[0057] It should be noted that the object audio data generation method of the embodiments of the present disclosure can be executed by an object audio data generation apparatus of the embodiments of the present disclosure, which can be implemented in a software and / or hardware manner, and can be configured in an electronic device, wherein the electronic device can install and run an object audio data generation program. The electronic device can include, but is not limited to, a smart phone, a tablet computer, and other hardware devices with various operating systems.
[0058] In the embodiments of the present disclosure, the position information refers to the position information of each microphone or sound object relative to the listener (Audience) as the origin. The position information can be expressed in a rectangular coordinate system (xyz) or a spherical coordinate system (θ, γ, r). They can be converted by formula (1) as follows.
[0059]
[0060] In formula (1), xyz respectively represent the position coordinates of the microphone or sound object on the x-axis (front-back direction), y-axis (left-right direction), and z-axis (up-down direction) of the rectangular coordinate system. θ, γ, and r respectively represent the horizontal direction angle (the angle between the mapping of the line connecting the microphone or sound object and the origin on the horizontal plane and the x-axis) of the microphone or sound object on the spherical coordinate system; the vertical direction angle (the angle between the line connecting the microphone or sound object and the origin and the horizontal plane); and the straight-line distance from the microphone or sound object to the origin.
[0061] The position information described above is the position information in a three-dimensional coordinate system. If in a two-dimensional coordinate system, the position information can be expressed in a rectangular coordinate system (x, y) or a polar coordinate system (θ, r). They can be converted by formula (2) as follows.
[0062] x = rcosθ
[0063] y = sinθ (2)
[0064] In formula (2), the meanings of the variables are the same as those in formula (1).
[0065] Regardless of the coordinate system (rectangular coordinate system, spherical coordinate system, or other coordinate system) used for expression or the transformation form of changing the origin of the coordinate system, it does not affect the specific implementation of the present disclosure or the claim of the right of the present disclosure.
[0066] Therefore, in the description of the present disclosure, for the sake of simplicity, the position information will be expressed in the spherical coordinate system (polar coordinate system).
[0067] In the embodiments of the present disclosure, object audio refers to various sound formats that can describe sound objects. Point sound objects containing position information or surface sound objects that can roughly determine the center position can all be sound objects. Object audio is generally composed of two parts, sound signal itself (Audio Data) and accompanying position information (Object Audio Metadata). Among them, the sound signal itself can be regarded as a monaural audio signal, which can be in the form of PCM (Pulse-code modulation), DSD (Direct Stream Digital) and other uncompressed formats, or MP3 (MPEG-1 or MPEG-2 Audio Layer III), AAC (Advanced Audio Coding), Dolby Digital and other compressed formats. The accompanying position information is the position information shown in 1. at any time t.
[0068] If there are multiple object audios, their formats can be that the sound signals and position information of each object audio are combined separately; or that all the sound signals of the objects are combined together, the position information is combined together, and the corresponding information of the first sound signal corresponding to the first position information is added in the sound signal or the position information.
[0069] Please refer to Figure 1 , Figure 1 is a flowchart of an object audio data generation method provided by the embodiments of the present disclosure.
[0070] As Figure 1 indicated, the method can include but is not limited to the following steps:
[0071] S1, obtaining sound data of at least one sound object.
[0072] In the embodiments of the present disclosure, the sound data of at least one sound object can be obtained by recording the sound signal of the sound object through a sound collecting device, and the sound data of the sound object is obtained. The at least one sound object can include one or more sound objects. In the case of including one sound object, the sound signal of the sound object is recorded through one sound collecting device. In the case of including multiple sound objects, the sound signals of the multiple sound objects are recorded through multiple sound collecting devices.
[0073] Among them, the sound collecting device can be a microphone or other device capable of collecting sound information, and the embodiments of the present disclosure do not make specific limitations thereto.
[0074] S2, obtaining current position information of the at least one sound object.
[0075] In the embodiments of the present disclosure, the current position information of the sound object can be obtained at the same time as the sound data of the sound object is obtained, so as to obtain the sound data and the position information of the sound object in real time.
[0076] In the embodiments of the present disclosure, the sound data of the sound object can be obtained by one or more sound collection devices, and the current position information of the sound object can be obtained in the case that the relative position between the sound collection device and the sound object is fixed. The current position information of the sound object can be determined according to the relative position between the sound collection device and the sound object.
[0077] In the embodiments of the present disclosure, in the case that the relative position between the sound collection device and the sound object is fixed, the sound object moves and the sound collection device moves with the sound object, so that the position information of each sound object can be obtained in real time.
[0078] It should be noted that, in the embodiments of the present disclosure, the current position information of the sound collection device can be obtained by an ultrasonic positioning method. For example, the sound collection device is provided with an ultrasonic transceiver, and the current position information of the sound collection device can be obtained by collecting the ultrasonic signal in the sound collection device. Alternatively, the current position information of the sound collection device can also be obtained by other methods, which are not limited in the embodiments of the present disclosure.
[0079] S3, synthesizing the sound data and the current position information of the at least one sound object to generate object audio data.
[0080] In the embodiments of the present disclosure, the sound data and the current position information of the sound object are synthesized to generate the object audio data after the sound data and the current position information of the at least one sound object are obtained, so as to obtain the sound data and the position information of the sound object in real time.
[0081] In the embodiments of the present disclosure, the sound data and the current position information of the sound object can be synthesized by combining the sound data and the current position information according to time, and stored in a specific file storage format to generate the object audio data.
[0082] By implementing the embodiments of the present disclosure, the sound data of the at least one sound object is obtained, the current position information of the at least one sound object is obtained, and the sound data and the current position information of the at least one sound object are synthesized to generate the object audio data. Therefore, the position information of each sound object can be obtained in real time and accurately, and the object audio data can be recorded and generated in real time.
[0083] As Figure 2 shown, the method can include, but is not limited to, the following steps:
[0084] S21, obtaining sound data of at least one sound object.
[0085] The related description of S21 can be referred to the related description in the above embodiments, which will not be repeated here.
[0086] S22, obtaining current position information of the at least one sound object.
[0087] In the embodiments of the present disclosure, the current position information of the sound object can be obtained at the same time as the sound data of the sound object is obtained, so as to obtain the sound data and the position information of the sound object in real time.
[0088] In some embodiments, the current position information of the at least one sound object includes: obtaining the current position information of at least one recording terminal recording the sound data of the at least one sound object.
[0089] In the embodiments of the present disclosure, the sound data of the sound object can be obtained by recording the sound signal of the sound object through the recording terminal. In the case of one sound object, the sound data of the sound object can be recorded through one recording terminal. In the case of multiple sound objects, the sound data of the sound objects can be recorded through multiple recording terminals.
[0090] The recording terminal includes a microphone, and the sound data of the sound object can be recorded through the microphone in the recording terminal.
[0091] In the embodiments of the present disclosure, the current position information of the at least one sound object can be obtained, and the current position information of the recording terminal recording the sound data of the sound object can be obtained. In the case of one or more sound objects, the current position information of the recording terminal recording the sound data of the one or more sound objects can be obtained.
[0092] In some embodiments, the current position information of the at least one recording terminal recording the sound data of the at least one sound object includes: obtaining the current position information of the at least one recording terminal in a one-way transceiving mode, a two-way transceiving mode or a mixed transceiving mode.
[0093] In the embodiments of the present disclosure, the current position information of the at least one recording terminal recording the sound data of the at least one sound object includes: obtaining the current position information of the at least one recording terminal recording the sound data of the sound object in the case of one sound object, and obtaining the current position information of the at least one recording terminal recording the sound data of each sound object in the case of multiple sound objects.
[0094] The current position information of the recording terminal can be acquired by one-way transceiving, or the current position information of at least one recording terminal can be acquired by two-way transceiving, or the current position information of the recording terminal can be acquired by mixed transceiving.
[0095] The current position information of the recording terminal can be acquired by one-way transceiving, or the current position information of at least one recording terminal can be acquired by two-way transceiving, or the current position information of the recording terminal can be acquired by mixed transceiving.
[0096] In some embodiments, the position information of the at least one recording terminal is acquired by mixed transceiving, including: acquiring first positioning reference information by one-way transceiving; acquiring second positioning reference information by two-way transceiving; and determining the current position information of the at least one recording terminal according to the first positioning reference information and the second positioning reference information.
[0097] In the embodiments of the present disclosure, the position information of the recording terminal is acquired by mixed transceiving, the first positioning reference information can be acquired by one-way transceiving, and the second positioning reference information can be acquired by two-way transceiving, and the current position information of the recording terminal is determined according to the first positioning reference information and the second positioning reference information.
[0098] The first positioning reference information and the second positioning reference information are different.
[0099] In some embodiments, the first positioning reference information is one of angle information and distance information, and the second positioning reference information is the other of the angle information and the distance information.
[0100] In the embodiments of the present disclosure, the position information of the recording terminal is acquired by mixed transceiving, the angle information can be acquired by one-way transceiving, and the distance information can be acquired by two-way transceiving, and the current position information of the recording terminal is determined according to the angle information and the distance information.
[0101] Alternatively, in the embodiments of the present disclosure, the position information of the recording terminal is acquired by mixed transceiving, the distance information can be acquired by one-way transceiving, and the angle information can be acquired by two-way transceiving, and the current position information of the recording terminal is determined according to the distance information and the angle information.
[0102] In the embodiments of the present disclosure, the first positioning reference information and the second positioning reference information can be acquired by sound waves or ultrasonic waves, or can also be acquired by electromagnetic wave signals such as UWB (Ultra Wide Band), WiFi or BT.
[0103] In some embodiments, the current position information of the at least one sound recording terminal is acquired in a one-way transceiving manner, including: receiving a first positioning signal sent by the at least one sound recording terminal in a broadcast manner, and generating the current position information of the at least one sound recording terminal according to the first positioning signal.
[0104] In the embodiments of the present disclosure, the current position information of the sound recording terminal is acquired in a one-way transceiving manner, which can be acquired by receiving a first positioning signal sent by the sound recording terminal in a broadcast manner, and generating the current position information of the sound recording terminal according to the first positioning signal. The current position information of the sound recording terminal can be acquired by a TDOA (time difference of arrival) method.
[0105] The first positioning signal sent by the sound recording terminal in a broadcast manner can be a sound wave or an ultrasonic wave, or can also be an electromagnetic wave signal such as UWB (Ultra Wide Band), WiFi or BT.
[0106] In some embodiments, the position information of the at least one sound recording terminal is acquired in a two-way transceiving manner, including: receiving a positioning start signal sent by the at least one sound recording terminal in a broadcast manner; sending a response signal to the at least one sound recording terminal; receiving a second positioning signal sent by the at least one sound recording terminal, and generating the current position information of the at least one sound recording terminal according to the second positioning signal.
[0107] In the embodiments of the present disclosure, the current position information of the sound recording terminal is acquired in a two-way transceiving manner, which can be acquired by receiving a positioning start signal sent by the sound recording terminal in a broadcast manner, sending a response signal to the sound recording terminal, receiving a second positioning signal sent by the sound recording terminal, and generating the current position information of the sound recording terminal according to the second positioning signal. The position information of the at least one sound recording terminal can be acquired by a TOF (time of flight) method.
[0108] The positioning start signal sent by the sound recording terminal in a broadcast manner can be a sound wave or an ultrasonic wave, or can also be an electromagnetic wave signal such as UWB (Ultra Wide Band), WiFi or BT.
[0109] The second positioning signal sent by the sound recording terminal can be a sound wave or an ultrasonic wave, or can also be an electromagnetic wave signal such as UWB (Ultra Wide Band), WiFi or BT.
[0110] In some embodiments, each sound recording terminal corresponds to a sound object, and the position of the sound recording terminal moves with the sound source of the sound object.
[0111] In the embodiments of the present disclosure, each recording terminal corresponds to one sound object, and in the case that there is one sound object, the sound data of the sound object is recorded by one or more recording terminals corresponding to the sound object.
[0112] In the embodiments of the present disclosure, the position of the recording terminal moves along with the sound source of the sound object, and it can be understood that the current position information of the at least one sound object is obtained, including: obtaining the current position information of at least one recording terminal for recording the sound data of the at least one sound object. The recording terminal corresponds to one sound object, and the position of the recording terminal is relatively fixed with the sound source of the sound object. In the case that the sound source of the sound object moves, the recording terminal moves along with the movement of the sound source of the sound object.
[0113] In some embodiments, the initial position information of the at least one sound object is obtained.
[0114] In the embodiments of the present disclosure, the initial position information and the current position information of the sound object are obtained, and the sound data of the sound object is obtained, so that the sound data and the position information of the sound object are obtained in real time.
[0115] In the embodiments of the present disclosure, the initial position information and the current position information of the sound object are obtained, and the sound data of the sound object is obtained, so that the sound data and the position information of the sound object are obtained in real time.
[0116] S23, synchronizing the sound data and the current position information of the at least one sound object.
[0117] In the embodiments of the present disclosure, in the case that the sound data and the current position information of the sound object are obtained, the sound data and the current position information of the sound object are synchronized, and the sound data and the current position information can be synchronized according to time.
[0118] S24, synthesizing the sound data and the current position information of the at least one sound object to generate object audio data.
[0119] In the embodiments of the present disclosure, in the case that the sound data and the current position information of the at least one sound object are obtained, the sound data and the current position information of the sound object are synthesized to generate object audio data.
[0120] In some embodiments, synthesizing the sound data and the current position information of the at least one sound object to generate object audio data includes: obtaining audio parameters and taking the audio parameters as header file information of the object audio data; at each sampling time, saving the sound data of each sound object as object audio signals and saving the current position information as object audio auxiliary data to generate object audio data.
[0121] In this embodiment of the disclosure, the generated object audio data can have multiple storage formats, such as a first format for saving as a file, a second format for real-time playback, etc.
[0122] For example, in the first format: file packing mode[], the sound data of at least one sound object is combined into a single audio message, which can be saved in raw-pcm format, uncompressed wav format (in which case a single sound object is considered as a channel of a wav file), or encoded into various compressed formats. The current position information of at least one sound object is also combined and saved as object audiometadata.
[0123] For example, the second format, low delay mode, uses a frame with a certain time length. Within each frame, the audio data is saved in the same format as in file packing mode, and the audio data at that time and the current position information are concatenated to form the object audio data of that frame. At this time, the object audio data of each frame are sent to the playback device or saved in chronological order.
[0124] In this embodiment of the disclosure, the audio parameters obtained may include the sampling rate, bit depth, and the number N of sound objects. obj (Number of objects) etc., and use audio parameters as header information of object audio data. For each sampling time, save the sound data of each sound object as object audio signal, and save the current position information as object audio auxiliary data to generate object audio data.
[0125] In one possible implementation, such as Figure 3 As shown, [s51] obtains the number N of sound objects. obj And the current location information of at least one synchronized sound object. Sound data of a sound object
[0126] [s52] Determine the storage format and decide whether to save / transfer in file packing mode or low delay mode.
[0127] [s53a] The basic audio parameters, such as sampling rate, bit depth, and the number of sound objects N, are specified. obj(Number of objects) and the like are recorded as header information in the object audio file.
[0128] [s54a] When it is judged as the file packing mode, the object audio information is saved in the file packing mode, in detail as follows:
[0129] [s541a] The object audio information is saved in the raw-pcm format, in detail as follows:
[0130] For the first sampling time, the sound data of the sound object obtained from the audio data sampled at t=1 is saved in the natural order of the sound sources, and the sound data of each sound object occupies a length of wBitsPerSample bits.
[0131] At each sampling time t thereafter, the sound data of the sound object obtained from the audio data sampled at t is saved in the natural order of the sound sources, and the sound data of each sound object occupies a length of wBitsPerSample bits. At each sampling time t thereafter, the sound data of the sound object obtained from the audio data sampled at t is saved in the natural order of the sound sources, and the sound data of each sound object occupies a length of wBitsPerSample bits.
[0132] The saving format can be seen from Table 1 as follows:
[0133]
[0134] Table 1
[0135] [s542a] The current position information of the at least one sound object obtained is saved as object audio auxiliary data, and is saved in the natural order of the sound sources at the first sampling point, and the saving format can be seen from Table 2 as follows:
[0136] Table 2
[0137]
[0138] iSampleOffset: the serial number of the sampling point;
[0139] Object_index: the serial number of the currently recorded sound source;
[0140] Object_Azimuth: θ of the currently recorded sound source;
[0141] Object_Elevation: γ of the currently recorded sound source;
[0142] Object_Radius: r of the current recorded sound source.
[0143] At the following sampling point, it is determined whether there is a change in the position of at least one sound object, and if there is, the sound source with the changed position of the sampling point is saved. The saving format is shown in Table 2 above.
[0144] In this way, a certain time interval, such as N sampling points, can be specified for determination and saving once, so as to save storage space.
[0145] [s55a] In the embodiments of the present disclosure, the audio parameters of the header file information of the object audio data, and the sound data of the sound object as the object audio signal and the current position information as the object audio auxiliary data are spliced to generate complete object audio data.
[0146] In this way, the splicing method is shown in Tables 3 to 6 as follows:
[0147]
[0148] Table 3
[0149]
[0150] Table 4
[0151]
[0152] Table 5
[0153]
[0154] Table 6
[0155] In another possible implementation, as shown in Figure 3 , the number N obj of sound objects is obtained , and the current position information of at least one sound object after synchronization
[0156] [s53b] The audio parameters can be obtained, such as the sampling rate (Sampling rate), the bit depth (bit depth), the number N obj (Number of objects) of sound objects, and the audio parameters are taken as the header file information of the object audio data, [s54b] in the case of the storage format being the second format low delay mode[], the object audio data is saved in the low delay mode, and the details are as follows:
[0157] [s541b] In a frame unit, the sound data of all the sampling points contained in the current frame, for the first sampling moment, the sound data of the sound object obtained by sampling at t=1 The sound data of each sound object occupies a length of wBitsPerSample bits.
[0158] [s542b] In a frame unit, the position information of the sound object of all the sampling points contained in the current frame At each sampling moment t thereafter, the sound data of the sound object obtained by sampling at t The object audio signal of the sound object obtained at t-1 is recorded in the natural order of the sound source Thereafter, the sound data of each sound object occupies a length of wBitsPerSample bits.
[0159] The saving format is shown in Table 1.
[0160] [s543b] In the embodiments of the present disclosure, the audio parameters as the header file information of the object audio data, the sound data of the sound object as the object audio signal, and the current position information as the object audio auxiliary data are spliced to generate complete object audio data.
[0161] The splicing manner is shown in Tables 7 to 9.
[0162]
[0163] Table 7
[0164]
[0165] Table 8
[0166]
[0167] Table 9
[0168] In some embodiments, the sound data and the current position information are further saved in a frame unit.
[0169] [s55]In the embodiments of the present disclosure, first, the header file information is recorded or transmitted, and for each frame, the sound data of the sound object and the recorded object audio metadata are spliced to become the object audio information of the frame. The object audio data of each frame is spliced in time sequence and saved, or the object audio data of each frame is directly transmitted after being obtained, so as to realize low delay transmission. The combined object audio data is saved in the memory or the disk as needed, or transmitted to a playing device, or encoded into the MPEG-H 3D Audio format, the Dolby Atmos format or other encoding formats supporting object audio, and saved or transmitted.
[0170] By implementing the embodiments of the present disclosure, the current position information of each sound object can be obtained in real time and accurately by using the positioning technology, and the object audio data can be generated in real time by recording.
[0171] For the convenience of understanding, an exemplary embodiment is provided in the embodiments of the present disclosure.
[0172] As shown in Figure 4 In a possible implementation, in the embodiments of the present disclosure, the sound data of the sound object is obtained by the recording terminal, the sound data of each sound object is collected by one recording terminal, the sound data of multiple sound objects is obtained by multiple recording terminals, and the sound data of at least one sound object is sent to the recording module.
[0173] The recording terminal can send the positioning signal, which is received by a plurality of receiving ends (antennas or microphones) in the positioning module, Figure 4 In the embodiment, the sound signal is transmitted to the recording module in a wired manner, but can also be transmitted in a wireless (WiFi or BT, etc.) manner. The receiving end in the positioning module receives the positioning signal sent by the recording terminal to obtain the current position information of the sound object.
[0174] It should be noted that Figure 4 In the embodiment, only the case of obtaining the current position information of the sound object by using the one-way transceiving mode is shown. In the embodiments of the present disclosure, the current position information of the sound object can also be obtained by using the two-way transceiving mode or the mixed transceiving mode, etc. In the case of using the two-way transceiving mode, the recording terminal can send the positioning start signal in addition to the positioning signal, and can also receive the response signal returned by the positioning module. In addition to receiving the positioning signal and the positioning start signal sent by the recording terminal, the positioning module can also send the response signal.
[0175] As shown in Figure 5As shown, in the embodiment of the present disclosure, when recording the object audio data of the sound object, first, the sound signal of the corresponding sound object is recorded by each recording device, and a ranging signal is emitted. The sound information (sound data) of the sound object and the position information (current position information) of the sound object are acquired respectively; the sound information (sound data) of the sound object and the position information (current position information) are synchronized, and then the sound information (sound data) of each sound object and the position information (current position information) are combined to generate a complete object audio signal (object audio data), thereby completing the recording of the object audio data.
[0176] As shown, the process of combining the sound information (sound data) of each sound object and the position information (current position information) to generate a complete object audio signal (object audio data) can specifically include: Figure 6
[0177] [S301] Acquire the number N of sound objects, the characteristic parameters of the positioning signals emitted by each recording terminal, and the position information of the positioning module. The number N of sound objects and the characteristic parameters of the positioning signals emitted by each recording terminal can be previously agreed upon, or can be transmitted to the recording module by each recording terminal when transmitting the sound signal to the recording module, and then transmitted to the positioning module by the recording module.
[0178] [S302] According to the position information of the positioning module, determine the coordinate origin position of the position information. And assign an initial position to each sound object.
[0179] [S303] Demodulate and extract the positioning features of the positioning signals received at each receiving device (antenna or microphone) of the positioning module, for subsequent positioning of each recording terminal through the features.
[0180] [S304-S311] For each sound object to be positioned, the position information is determined respectively.
[0181] S305-S306 are to determine whether there is a positioning signal or a positioning start signal of a certain sound object from the received positioning features, if so, the information is obtained and different positioning schemes are adopted according to the positioning method. For example, when using one-way transceiving, the position information of the sound object is obtained by using the TDOA (time difference of arrival) method, and when using two-way transceiving, the position information of the sound object is obtained by using the TOF (time of flight) method. Or UWB indoor positioning scheme, etc. If two-way transceiving is used, the positioning module must perform two-way data communication with each recording terminal.
[0182] The synchronization module obtains the sound information (sound data) of the sound object from the recording module and the location information (current location information) of the sound object from the positioning module. It synchronizes the sound information (sound data) and location information (current location information) of the synchronized sound object and sends them to the combination module.
[0183] The combination module obtains the position information (current position information) of each synchronized sound object from the synchronization module. And sound information (sound data) The location information (current location information) of the sound object and the sound information (sound data) of the sound object are combined to form a complete object audio signal.
[0184] Depending on the application, there are two ways to save the audio signal: file packing mode[] for saving and low delay mode[] for real-time playback.
[0185] In file packing mode, the audio information (audio data) of each audio object is combined into a multi-object audio information, which can be saved in raw-pcm format, uncompressed wav format (in which case a single object is regarded as a channel of a wav file), or encoded into various compressed formats. The audio object position information of each object is also combined and saved as object audio metadata.
[0186] In low-delay mode, a certain time length τ is defined as a frame. Within each frame, the audio is saved in the same format as in file-packing mode, and the audio information and auxiliary audio data at that time are concatenated to form the object audio information of that frame. At this time, the audio information of each frame is sent to the playback device or saved in chronological order.
[0187] Figure 7 This is a structural diagram of an object audio data generation device provided in an embodiment of this disclosure.
[0188] like Figure 7 As shown, the object audio data generation device 1 includes: a data acquisition unit 11, an information acquisition unit 12, and a data generation unit 13.
[0189] The data acquisition unit 11 is configured to acquire the sound data of at least one sound object.
[0190] The information acquisition unit 12 is configured to acquire the current location information of at least one sound object.
[0191] The data generating unit 13 is configured to synthesize the sound data of the at least one sound object and the current position information to generate the object audio data.
[0192] In some embodiments, the information obtaining unit 12 is specifically configured to obtain the current position information of at least one recording terminal that records the sound data of the at least one sound object.
[0193] As shown in Figure 8 In some embodiments, the object audio data generating apparatus 1 further comprises a synchronization processing unit 14 configured to synchronize the sound data of the at least one sound object and the current position information.
[0194] In some embodiments, the information obtaining unit 12 is specifically configured to obtain the current position information of the at least one recording terminal in a one-way transceiving manner, a two-way transceiving manner or a mixed transceiving manner.
[0195] As shown in Figure 9 In some embodiments, the information obtaining unit 12 comprises a first information obtaining module 121, a second information obtaining module 122 and a first current information obtaining module 123.
[0196] The first information obtaining module 121 is configured to obtain the first positioning reference information in a one-way transceiving manner.
[0197] The second information obtaining module 122 is configured to obtain the second positioning reference information in a two-way transceiving manner.
[0198] The first current information obtaining module 123 is configured to determine the current position information of the at least one recording terminal according to the first positioning reference information and the second positioning reference information.
[0199] In some embodiments, the first positioning reference information is one of angle information and distance information, and the second positioning reference information is the other of angle information and distance information.
[0200] As shown in Figure 10 In some embodiments, the information obtaining unit 12 comprises a second current information obtaining module 124 configured to receive a first positioning signal broadcasted by the at least one recording terminal and generate the current position information of the at least one recording terminal according to the first positioning signal.
[0201] As shown in Figure 11 In some embodiments, the information obtaining unit 12 comprises a signal receiving module 125, a signal sending module 126 and a third current information obtaining module 127.
[0202] The signal receiving module 125 is configured to receive the positioning initiation signal sent by the at least one sound recording terminal in a broadcast manner.
[0203] The signal sending module 126 is configured to send a response signal to the at least one sound recording terminal.
[0204] The third current information obtaining module 127 is configured to receive the second positioning signal sent by the at least one sound recording terminal, and generate the current position information of the at least one sound recording terminal according to the second positioning signal.
[0205] In some embodiments, each sound recording terminal corresponds to a sound object, and the position of the sound recording terminal moves along with the sound source of the sound object.
[0206] As shown in Figure 12 In some embodiments, the object audio data generation apparatus 1 further comprises an initial position obtaining unit 15 configured to obtain the initial position information of the at least one sound object.
[0207] As shown in Figure 13 In some embodiments, the data generation unit 13 comprises a parameter obtaining module 131 and an audio data generation module 132.
[0208] The parameter obtaining module 131 is configured to obtain the audio parameter, and take the audio parameter as the header file information of the object audio data.
[0209] The audio data generation module 132 is configured to save the sound data of each sound object as the object audio signal and save the current position information as the object audio auxiliary data at each sampling time, so as to generate the object audio data.
[0210] Please continue to refer to Figure 13 In some embodiments, the data generation unit 13 further comprises a processing module 133.
[0211] The processing module 133 is configured to save the sound data and the current position information in units of frames.
[0212] As to the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the method, and will not be described in detail here.
[0213] The object audio data generation apparatus provided by the embodiments of the present disclosure can perform the object audio data generation method as described in some embodiments above, and has the same beneficial effects as the object audio data generation method described above, which will not be described here.
[0214] Figure 14FIG. 1 is a block diagram of an electronic device 100 according to an exemplary embodiment for a method of generating object audio data.
[0215] The electronic device 100 can be, for example, a mobile phone, a computer, a digital broadcasting terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0216] As shown, the electronic device 100 can include one or more of the following components: a processing component 101, a memory 102, a power supply component 103, a multimedia component 104, an audio component 105, an input / output (I / O) interface 106, a sensor component 107, and a communication component 108. Figure 14
[0217] The processing component 101 generally controls the overall operations of the electronic device 100, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 101 can include one or more processors 1011 to execute instructions to complete all or part of steps of the above-described methods. In addition, the processing component 101 can include one or more modules to facilitate the interaction between the processing component 101 and other components. For example, the processing component 101 can include a multimedia module to facilitate the interaction between the multimedia component 104 and the processing component 101.
[0218] The memory 102 is configured to store various types of data to support operations of the electronic device 100. Examples of these data include instructions for any application or method operating on the electronic device 100, contact data, phonebook data, messages, pictures, videos, etc. The memory 102 can be implemented by any type of volatile or non-volatile memory devices or a combination thereof, such as SRAM (Static Random-Access Memory), EEPROM (Electrically Erasable Programmable read only memory), EPROM (Erasable Programmable Read-Only Memory), PROM (Programmable read-only memory), ROM (Read-Only Memory), magnetic storage, flash memory, magnetic or optical disks.
[0219] The power component 103 provides power to various components of the electronic device 100. The power component 103 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 100.
[0220] The multimedia component 104 includes a touch display screen providing an output interface between the electronic device 100 and a user. In some embodiments, the touch display screen can include an LCD (Liquid Crystal Display) and a TP (Touch Panel). The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundary of a touch or swipe action, but also detect the duration and pressure associated with the touch or swipe action. In some embodiments, the multimedia component 104 includes a front-facing camera and / or a rear-facing camera. The front-facing camera and / or the rear-facing camera can receive external multimedia data when the electronic device 100 is in an operation mode, such as a shooting mode or a video mode. Each of the front-facing camera and the rear-facing camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0221] The audio component 105 is configured to output and / or input audio signals. For example, the audio component 105 includes a microphone that is configured to receive external audio signals when the electronic device 100 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 102 or transmitted via the communication component 108. In some embodiments, the audio component 105 also includes a speaker for outputting audio signals.
[0222] The I / O interface 2112 provides an interface between the processing component 101 and peripheral interface modules, which can be a keyboard, a click wheel, a button, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0223] The sensor component 107 includes one or more sensors for providing status assessments for various aspects of the electronic device 100. For example, the sensor component 107 can detect an open / closed position of the electronic device 100, relative positioning of components, such as a display and keypad of the electronic device 100, a change in position of the electronic device 100 or a component of the electronic device 100, the presence or absence of user contact with the electronic device 100, the orientation or acceleration / deceleration of the electronic device 100, and a temperature change of the electronic device 100. The sensor component 107 can include a proximity sensor configured to detect presence of a nearby object without any physical contact. The sensor component 107 can also include a light sensor, such as a CMOS (complementary metal-oxide semiconductor) or CCD (charge-coupled device) image sensor, utilized in imaging applications. In some embodiments, the sensor component 107 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0224] The communication component 108 is configured to facilitate wired or wireless communication between the electronic device 100 and other devices. The electronic device 100 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 108 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 108 can also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on RFID (radio frequency identification) technology, IrDA (infrared data association) technology, UWB (ultra-wideband) technology, Bluetooth (BT) technology, and other technologies.
[0225] In an example embodiment, the electronic device 100 can be implemented by one or more ASICs (Application Specific Integrated Circuit), DSPs (Digital Signal Processor), digital signal processing devices (DSPD), PLDs (Programmable Logic Device), FPGAs (Field Programmable Gate Array), controllers, microcontrollers, microprocessors or other electronic elements for performing the object audio data generation method described above. It should be noted that the implementation process and technical principles of the electronic device in this embodiment are referred to the above description of the object audio data generation method of the embodiments of the present disclosure, and will not be described here.
[0226] The electronic device 100 provided by the embodiments of the present disclosure can perform the object audio data generation method as described in some of the above embodiments, which has the same beneficial effects as the object audio data generation method described above, and will not be described here.
[0227] In order to achieve the above-mentioned embodiments, the present disclosure further provides a storage medium.
[0228] The instructions in the storage medium are executed by the processor of the electronic device, so that the electronic device can perform the object audio data generation method as described above. For example, the storage medium can be a ROM (Read Only Memory Image), a RAM (Random Access Memory), a CD-ROM (Compact Disc Read-Only Memory), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0229] In order to achieve the above-mentioned embodiments, the present disclosure further provides a computer program product, which is executed by the processor of the electronic device, so that the electronic device can perform the object audio data generation method as described above.
[0230] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. This application is intended to cover any variations, uses or adaptations of the present disclosure following the general principles thereof and including such departures from the present disclosure as come within known use or custom in the art. It is intended to cover the application as broadly as possible in the spirit and scope of the application. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the present disclosure is indicated by the following claims.
[0231] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, device and unit described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described here.
[0232] The above merely describes specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present disclosure, which should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for generating object audio data, characterized in that, include: Real-time acquisition of sound data from at least one sound object; Obtaining the current location information of the at least one sound object includes: obtaining the current location information of at least one recording terminal that records the sound data of the at least one sound object; determining the current location information of the at least one sound object based on the current location information of the at least one recording terminal that records the sound data of the at least one sound object; wherein each recording terminal corresponds to one sound object, and the position of the recording terminal moves along with the sound source of the sound object. The sound data and current location information of the at least one sound object are synchronized, and the sound data and current location information of the at least one sound object are synthesized to generate object audio data.
2. The method as described in claim 1, characterized in that, The current location information of at least one recording terminal for acquiring the sound data of the at least one sound object includes: The current location information of the at least one recording terminal is obtained using a one-way, two-way, or hybrid transmission and reception method.
3. The method as described in claim 2, characterized in that, The method of obtaining the location information of at least one recording terminal using a hybrid transceiver approach includes: The first positioning reference information is obtained using the unidirectional transmission and reception method described above; The second positioning reference information is obtained using the bidirectional transmission and reception method described above; The current location information of the at least one recording terminal is determined based on the first positioning reference information and the second positioning reference information.
4. The method as described in claim 3, characterized in that, The first positioning reference information is one of angle information and distance information, and the second positioning reference information is the other of the angle information and the distance information.
5. The method as described in claim 2, characterized in that, The step of obtaining the current location information of the at least one recording terminal using the one-way transceiver method includes: The system receives a first positioning signal broadcast by the at least one recording terminal and generates the current location information of the at least one recording terminal based on the first positioning signal.
6. The method as described in claim 2, characterized in that, The step of obtaining the location information of the at least one recording terminal using the bidirectional transceiver method includes: Receive the positioning start signal broadcast by the at least one recording terminal; Send a response signal to the at least one recording terminal; The system receives a second positioning signal sent by the at least one recording terminal and generates the current location information of the at least one recording terminal based on the second positioning signal.
7. The method as described in claim 1, characterized in that, Also includes: Obtain the initial position information of the at least one sound object.
8. The method according to any one of claims 1 to 7, characterized in that, The step of synthesizing the sound data of at least one sound object and the current location information to generate object audio data includes: Obtain the audio parameters and use the audio parameters as the header information of the object's audio data; At each sampling time, the sound data of each sound object is saved as the object audio signal, and the current position information of each sound object is saved as the object audio auxiliary data to generate the object audio data.
9. The method as described in claim 8, characterized in that, Also includes: The sound data and the current location information are saved in frames.
10. An apparatus for generating object audio data, characterized in that, include: The data acquisition unit is configured to acquire sound data of at least one sound object in real time. The information acquisition unit is configured to acquire the current location information of the at least one sound object, including: acquiring the current location information of at least one recording terminal that records the sound data of the at least one sound object; and determining the current location information of the at least one sound object based on the current location information of the at least one recording terminal that records the sound data of the at least one sound object, wherein each recording terminal corresponds to one sound object, and the position of the recording terminal moves along with the sound source of the sound object. The data generation unit is configured to synchronize the sound data and current location information of the at least one sound object, and to synthesize the sound data and current location information of the at least one sound object to generate object audio data.
11. The apparatus as claimed in claim 10, characterized in that, The information acquisition unit is specifically configured as follows: The current location information of the at least one recording terminal is obtained using a one-way, two-way, or hybrid transmission and reception method.
12. The apparatus as claimed in claim 11, characterized in that, The information acquisition unit includes: The first information acquisition module is configured to acquire first positioning reference information in the unidirectional transmission and reception method; The second information acquisition module is configured to acquire second positioning reference information in the bidirectional transmission and reception method; The first current information acquisition module is configured to determine the current location information of the at least one recording terminal based on the first positioning reference information and the second positioning reference information.
13. The apparatus as claimed in claim 12, characterized in that, The first positioning reference information is one of angle information and distance information, and the second positioning reference information is the other of the angle information and the distance information.
14. The apparatus as claimed in claim 11, characterized in that, The information acquisition unit includes: The second current information acquisition module is configured to receive a first positioning signal broadcast by the at least one recording terminal, and generate the current location information of the at least one recording terminal based on the first positioning signal.
15. The apparatus as claimed in claim 11, characterized in that, The information acquisition unit includes: The signal receiving module is configured to receive a positioning start signal broadcast by the at least one recording terminal; The signal transmitting module is configured to send a response signal to the at least one recording terminal; The third current information acquisition module is configured to receive a second positioning signal sent by the at least one recording terminal, and generate the current location information of the at least one recording terminal based on the second positioning signal.
16. The apparatus as claimed in claim 10, characterized in that, The device further includes: The initial position acquisition unit is configured to acquire the initial position information of the at least one sound object.
17. The apparatus as claimed in any one of claims 10 to 16, characterized in that, The data generation unit includes: The parameter acquisition module is configured to acquire audio parameters and use the audio parameters as header information of the object's audio data. The audio data generation module is configured to save the sound data of each sound object as the object audio signal at each sampling time, and save the current position information of each sound object as the object audio auxiliary data, so as to generate the object audio data.
18. The apparatus as claimed in claim 17, characterized in that, The data generation unit further includes: The processing module is configured to save the sound data and the current location information in frames.
19. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 9.
20. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 9.
21. A computer program product comprising computer instructions, characterized in that, The computer instructions, when executed by a processor, implement the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Method, device and electronic equipment for realizing recording of object audio
CN105070304A
Method and system for determining a position of a microphone
CN110320498A