Audio separation method, vocal extraction method, audio processing system, electronic device, and computer program product
By recording posture data in the audio acquisition device and dynamically adjusting the beam direction, the stability problem of the audio separation solution during posture changes is solved, and accurate separation and stable locking of the target sound source are achieved.
Patent Information
- Application Number
- CN202411954601.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing audio separation solutions based on spatial information have poor stability in separation effects when facing changes in the posture of the audio acquisition device, making it difficult to accurately separate the target sound source.
By synchronously recording the posture data in the audio acquisition device and dynamically adjusting the beam direction according to the posture data during the audio separation process, dynamic locking of the target sound source is achieved, ensuring that the beam direction always points to or is close to the target sound source.
Even if the posture of the audio acquisition device changes, the audio data of the target sound source can be accurately separated, improving the stability and accuracy of the audio separation effect.
Smart Images

Figure CN119785816B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio separation technology, and in particular to an audio separation method, a human voice extraction method, an audio processing system, an electronic device, and a computer program product. Background Art
[0002] Audio separation technology is a key branch of audio signal processing, aiming to extract independent sound source signals from mixed audio signals. Traditional audio separation methods rely primarily on voiceprint recognition, which analyzes the spectral and temporal characteristics of different sound sources to distinguish and separate them. However, this approach faces numerous challenges in practical applications, such as complex processes and poor noise immunity.
[0003] To overcome the limitations of voiceprint separation technology, researchers have proposed an audio separation solution based on spatial information. The core idea behind this solution is to leverage the spatial positioning capabilities of microphone arrays, using a predetermined direction as the beam direction toward the sound source, and then separating the audio signal based on this direction. The advantage of this spatial separation solution is its relative simplicity, eliminating the need for complex voiceprint analysis.
[0004] However, although the spatial separation scheme has certain advantages in simplifying the audio separation process, it also has an obvious defect, namely, the stability of the separation effect is poor. Summary of the Invention
[0005] Based on the above technical status, the present application proposes an audio separation method, a human voice extraction method, an audio processing system, an electronic device and a computer program product.
[0006] According to a first aspect of an embodiment of the present application, there is provided an audio separation method, the method comprising:
[0007] Acquire data to be processed from an audio acquisition device, wherein the data to be processed includes: multi-source audio data recorded by the audio acquisition device and posture data of the audio acquisition device during the recording of the multi-source audio data; separate the audio data of a target sound source from the multi-source audio data, and adjust the beam direction to point to or approach the target sound source according to the posture data during the separation process.
[0008] According to a second aspect of an embodiment of the present application, a method for extracting a human voice is provided, the method comprising:
[0009] Acquire data to be processed from an audio acquisition device, wherein the data to be processed includes: multi-source audio data recorded by the audio acquisition device and posture data of the audio acquisition device during the recording of the multi-source audio data; separate audio data of a target sound source from the multi-source audio data, and adjust a beam direction to point to or approach the target sound source according to the posture data during the separation process to obtain audio data of the target sound source; perform sound processing on the audio data of the target sound source to obtain target human voice data, wherein the sound processing includes: human voice extraction and / or human voice enhancement.
[0010] According to a third aspect of an embodiment of the present application, an audio separation method is provided, which is applied to an audio acquisition device, wherein the audio acquisition device includes: an audio acquisition unit with a sound recording function, and a posture acquisition unit with a posture detection function. The method includes:
[0011] In response to a start-up operation, multi-source audio data is recorded by the audio acquisition unit, and posture data of the audio acquisition unit during the recording of the multi-source audio data is obtained by the posture acquisition unit; the multi-source audio data and the posture data are sent to a target device, so that the target device separates the audio data of the target sound source from the multi-source audio data, and adjusts the beam direction to point to or approach the target sound source according to the posture data during the separation process.
[0012] According to a fourth aspect of an embodiment of the present application, an electronic device is provided, comprising a memory and a processor; the memory is connected to the processor for storing a program; the processor is used to implement the method described in the first aspect, the second aspect, or the third aspect by running the program in the memory.
[0013] According to a fifth aspect of the embodiments of the present application, there is provided an electronic device, an audio processing system, comprising an audio acquisition device and a target device;
[0014] The audio acquisition device is configured to perform the audio separation method as described in the third aspect;
[0015] and / or,
[0016] The target device is configured to execute the audio separation method as described in the first aspect, or execute the human voice extraction method as described in the second aspect.
[0017] According to a sixth aspect of an embodiment of the present application, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect, the second aspect, or the third aspect is implemented.
[0018] According to a seventh aspect of the embodiments of the present application, a computer program product is provided, comprising: a computer program, which implements the method described in the first aspect, the second aspect, or the third aspect when executed by a processor.
[0019] In an embodiment of the present application, the acquired data to be processed includes multi-source audio data recorded by the audio acquisition device and posture data of the audio acquisition device during the process of recording multi-source audio data. Therefore, the posture changes of the audio acquisition device during the process of recording multi-source audio data can be determined through the posture data. Then, the audio data of the target sound source is separated from the multi-source audio data, and the beam direction is adjusted to point to or approach the target sound source according to the posture data during the separation process, thereby achieving dynamic locking of the target sound source. Therefore, during the process of recording multi-source audio data, even if the posture of the audio acquisition device changes, the audio data of the target sound source can be accurately separated, thereby improving the stability of the audio separation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0021] Figure 1 This is a flowchart of an audio separation method provided in an embodiment of the present application;
[0022] Figure 2 A flowchart of a method for extracting human voice provided in an embodiment of the present application;
[0023] Figure 3 A second flowchart of an audio separation method provided in an embodiment of the present application;
[0024] Figure 4 This is one of the structural diagrams of an audio separation device provided in an embodiment of the present application;
[0025] Figure 5 A schematic diagram of the structure of a human voice extraction device provided in an embodiment of the present application;
[0026] Figure 6 This is a second structural diagram of an audio separation device provided in an embodiment of the present application;
[0027] Figure 7 A schematic diagram of the architecture of an audio processing system provided in an embodiment of the present application;
[0028] Figure 8A schematic diagram of the process of obtaining human voice through processing by a cloud device provided in an embodiment of the present application;
[0029] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0031] Overview
[0032] As described in the background, spatial information-based audio separation relies primarily on a multi-microphone array in an audio capture device. It first determines the differences in the time it takes multiple microphones to receive the same sound source. Based on these differences, it then selects the closest fixed beam direction as the direction pointing to the sound source. This fixed beam direction is consistently selected throughout the separation process to localize the sound source. Finally, audio signal separation is performed based on the localized sound source.
[0033] However, the spatial posture of the audio capture device recording audio may change during the recording process. For example, if the wearer of the audio capture device walks, sits down, or other activities during recording, the spatial posture of the audio capture device may change significantly. In this case, continuing to perform audio separation based on a fixed beam direction will inevitably affect the separation results.
[0034] In response to the above-mentioned technical status, the inventors propose that the posture data of the audio acquisition device can be synchronously recorded during the process of recording audio data, and then the posture data can be used to dynamically adjust the beam direction during the separation of the audio data to achieve dynamic locking of the sound source. The embodiment of the present application proposes an audio separation method. The acquired data to be processed includes multi-source audio data recorded by the audio acquisition device and the posture data of the audio acquisition device during the process of recording the multi-source audio data. Therefore, the posture changes of the audio acquisition device during the process of recording the multi-source audio data can be determined through the posture data. Then, the audio data of the target sound source is separated from the multi-source audio data, and the beam direction is adjusted to point to or approach the target sound source according to the posture data during the separation process to achieve dynamic locking of the target sound source. Therefore, during the process of recording multi-source audio data, even if the posture of the audio acquisition device changes, the audio data of the target sound source can be accurately separated, thereby improving the stability of the audio separation effect.
[0035] Exemplary Methods
[0036] See also Figure 1 In an exemplary embodiment, an audio separation method is provided. The method can be applied to any electronic device capable of processing audio data, such as, but not limited to, a server, a computer, or a smartphone. The electronic device can even be a dedicated recording device with audio acquisition capabilities. The method can include:
[0037] S101: Acquire data to be processed from an audio acquisition device.
[0038] In this step, the audio capture device can be any electronic device with a recording function. In some embodiments, the audio capture device can be a wearable audio capture device. For example, the audio capture device can be a wearable smart ID card. When worn by an employee, it can serve as an ID card and also record the employee's work.
[0039] The data to be processed include: multi-source audio data recorded by an audio acquisition device. Among them, the multi-source audio data is the object that needs to be audio separated, which contains sounds from multiple sound sources. The multiple sound sources include but are not limited to people and animals. For example, the multiple sound sources include multiple different people. Regarding the data format of the multi-source audio data, it can be any audio data format. For example, the multi-source audio data can be audio data formed by sub-band coding (SBC) technology. Regarding the data content or semantic content of the multi-source audio data, there is no limitation here, and it can be any semantic content.
[0040] The data to be processed also includes posture data of the audio capture device during the recording of multi-source audio data. This posture data may represent the posture of the audio capture device in space. It will be appreciated that the posture data in the data to be processed is continuous posture data within a time period. This time period is the time period during which the multi-source audio data was recorded. Therefore, the posture data in the data to be processed can be used to determine the posture of the audio capture device or changes in posture during the recording of the multi-source audio data.
[0041] S102: Separate the audio data of the target sound source from the multi-source audio data, and adjust the beam direction to point to or approach the target sound source according to the posture data during the separation process.
[0042] In this step, the target sound source can be any sound source in the multi-source audio data. In some embodiments, user input can be received, and the sound source to be separated can be determined based on the user input, and the sound source to be separated can be used as the target sound source. In some embodiments, the target sound source can be a target person, and the audio data from the separated target sound source is the voice of the separated target person.
[0043] Spatial separation techniques can be used to separate audio data. In some embodiments, a beam direction is first determined, and then this beam direction is used to indicate the direction of a target sound source, thereby localizing the target sound source. Finally, based on the localization result / beam direction of the target sound source, the audio data of the target sound source is separated from the multi-source audio data. In this embodiment, after the beam direction is determined, the audio data is not always separated using this fixed beam direction. Instead, the beam direction is dynamically adjusted using posture data so that the beam direction always accurately points toward or approaches the target sound source, thereby locking on to the target sound source. Finally, while maintaining lock on to the target sound source, the audio data of the target sound source is separated. This allows the target sound source to be locked on in a timely manner, even if the posture of the audio capture device changes significantly during the recording of multi-source audio data, rather than separating the audio based on the original position of the target sound source. For example, in scenarios where the audio capture device is used dynamically to record, the target sound source can be locked on to the target sound source, achieving better separation results.
[0044] In an embodiment of the present application, the acquired data to be processed includes multi-source audio data recorded by the audio acquisition device and posture data of the audio acquisition device during the process of recording multi-source audio data. Therefore, the posture changes of the audio acquisition device during the process of recording multi-source audio data can be determined through the posture data. Then, the audio data of the target sound source is separated from the multi-source audio data, and the beam direction is adjusted to point to or approach the target sound source according to the posture data during the separation process, thereby achieving dynamic locking of the target sound source. Therefore, during the process of recording multi-source audio data, even if the posture of the audio acquisition device changes, the audio data of the target sound source can be accurately separated, thereby improving the stability of the audio separation effect.
[0045] In order to avoid frequent adjustments to the beam direction, which would result in excessive hardware overhead or increased complexity in the separation process, in some embodiments of the present application, after obtaining the data to be processed from the audio acquisition device, the method further includes:
[0046] According to the posture data, the posture changes of the audio acquisition device in each recording period are determined; according to the posture changes in each recording period, the target recording period in which the beam direction needs to be adjusted is determined.
[0047] During the separation process, the beam direction is adjusted to point towards or approach the target sound source based on the attitude data, including:
[0048] During the separation process, the beam direction within the target recording period is adjusted to point to or approach the target sound source according to the posture data of the target recording period.
[0049] It should be noted that when recording multi-source audio data, the audio capture device's posture may change frequently. For example, when recording with a wearable audio capture device, the user may walk, sit, and so on, causing the device's posture to change frequently. Adjusting the beam direction for each change would inevitably increase the hardware burden, complicate the separation process, and even affect the separation speed.
[0050] To this end, certain screening conditions can be set to screen the recording time periods that require adjustment of the beam direction, i.e., the target recording time periods. The beam direction is then adjusted only during the target recording time periods. The number of target recording time periods is not limited and can be one or more. In some embodiments, when there are multiple target recording time periods, the beam direction between any two adjacent target recording time periods is consistent with the beam direction within the previous target recording time period. For example, the overall time period of the recording process starts from 0 minutes and 0 seconds and ends at 10 minutes and 20 seconds. If the audio acquisition device undergoes significant changes in the recording period from 1 minute and 20 seconds to 1 minute and 30 seconds and the recording period from 5 minutes and 10 seconds to 6 minutes and 20 seconds, then the target recording time periods include a first target recording time period from 1 minute and 20 seconds to 1 minute and 30 seconds and a second target recording time period from 5 minutes and 10 seconds to 6 minutes and 20 seconds. In the process of separating the audio, the beam direction between 0 minutes and 0 seconds and 1 minute and 19 seconds can be determined based on spatial separation technology. During the first target recording period, the beam direction is adjusted to the first beam direction. The beam direction from 1 minute 31 seconds to 5 minutes 9 seconds remains the first beam direction. During the second target recording period, the beam direction is adjusted to the second beam direction. The beam direction from 6 minutes 21 seconds to 10 minutes 20 seconds remains the second beam direction.
[0051] Posture changes include the magnitude and frequency of changes in the audio capture device during the recording period. In some embodiments, filtering conditions can be set based on scenarios where the audio capture device's posture changes. For example, if the wearer's sitting down causes the audio capture device's posture to change, a filtering condition can be set such that posture changes indicate a significant change in the audio capture device within a relatively short period of time.
[0052] In some embodiments, a plurality of neural network models can be pre-trained for audio separation of multi-source audio data recorded in different postures. In the separation process, when the beam direction of the target recording period is adjusted to point to or close to the target sound source according to the posture data of the target recording period, the corresponding neural network model can be selected according to the posture data of the target recording period, and then the selected neural network model is used for audio separation.
[0053] In the embodiments of the present application, in the case of frequent changes of the audio acquisition device, the target recording period can be screened for adjustment of the beam direction, thereby avoiding frequent adjustment of the beam direction, which leads to excessive hardware device overhead or increases the complexity of the separation process.
[0054] In some embodiments of the present application, the posture change condition includes: time domain features and / or frequency domain features of the posture data; and the multi-source audio data is recorded by the target person wearing the audio acquisition device.
[0055] According to the posture change condition of each recording period, the target recording period in which the beam direction needs to be adjusted is determined, including:
[0056] According to the time domain features and / or frequency domain features of the posture data of each recording period, a behavior label corresponding to the target person in each recording period is determined; the recording period corresponding to the target behavior label is determined as the target recording period; wherein the target behavior label includes a behavior label identical to any dangerous behavior label in the behavior label library.
[0057] It should be noted that, since the target person wears the audio acquisition device, the target person implements different behaviors during the recording of the multi-source audio data, and the posture of the audio acquisition device will change differently. Conversely, different posture changes of the audio acquisition device correspond to different behaviors of the target person. Therefore, the behavior of the target person is closely related to the posture of the audio acquisition device.
[0058] The embodiments utilize the time domain features and / or frequency domain features of the posture data to inversely deduce the behavior of the target person. If the separation effect is poor after the behavior is implemented and the original beam direction is used for separation, the behavior is defined as a dangerous behavior, and the recording period in which the behavior occurs is regarded as the target recording period. Here, a label library is pre-set, and the behavior label of the dangerous behavior is stored in the label library. Therefore, after the behavior of the target person is determined, the target recording period is determined based on the comparison result of the behavior label of the behavior and the label library. For example, the determined behavior includes behavior A, the label corresponding to behavior A is a, and if the label library contains the behavior label a, then behavior A is regarded as a dangerous behavior, and the recording period in which behavior A occurs is regarded as the target recording period.
[0059] In the embodiment of the present application, the behavior of the target person is deduced with the help of the time domain characteristics and frequency domain characteristics of the posture data, and then the target recording period is determined based on the comparison results of the behavior label of the behavior and the label library. The process is simple.
[0060] In some embodiments of the present application, determining a target recording period for which the beam direction needs to be adjusted based on the posture change in each recording period includes:
[0061] According to the posture changes in each recording period, the recording period with low-frequency and large-scale posture changes is determined; the recording period with low-frequency and large-scale posture changes is determined as the target recording period.
[0062] It should be noted that multi-source audio data can be recorded when the target person is wearing an audio acquisition device. Therefore, when the target person performs different behaviors during the recording of multi-source audio data, the posture of the audio acquisition device will change differently. For example, when the target person performs behaviors such as walking, going up and down stairs, the corresponding posture data will show small changes in posture at high frequency. Since such behaviors have little effect on the accuracy of audio separation, the beam direction can be left unchanged during such behaviors. When the target person changes their standing posture, sitting posture, wears the device improperly, falls, etc., the corresponding posture data will show low-frequency and large-scale changes in posture. Since such behaviors have a greater impact on the accuracy of audio separation, the beam direction needs to be adjusted during such behaviors, that is, the period of such behaviors is determined as the target recording period.
[0063] In the embodiment of the present application, the beam direction can be adjusted when the posture changes at a low frequency and a large amplitude, thereby avoiding a significant impact on the audio separation accuracy.
[0064] In some embodiments of the present application, obtaining data to be processed from an audio acquisition device includes:
[0065] Receiving an encoded file from an audio acquisition device, wherein the encoded file is a single file formed by encoding multi-source audio data recorded by the audio acquisition device and posture data during the recording of the multi-source audio data by the audio acquisition device;
[0066] Decode the encoded file to obtain multi-source audio data and posture data.
[0067] It should be noted that the device for audio separation and the audio acquisition device can be different electronic devices. In this way, the audio acquisition device is required to transmit relevant data to the device for audio separation. In some embodiments, the audio acquisition device is required to upload relevant data to the cloud, and the audio separation device downloads relevant data from the cloud. To facilitate uploading and downloading, the audio acquisition device can encode multi-source audio data and posture data into a single file, and subsequently upload and download the single file. The audio separation device can decode the single file to obtain multi-source audio data and posture data. This application does not limit the encoding method.
[0068] In the embodiment of the present application, multi-source audio data and posture data are transmitted in the form of a single file, which improves the convenience of transmission.
[0069] In some embodiments of the present application, the step of obtaining posture data in the data to be processed includes:
[0070] Obtaining the raw data output by the inertial measurement unit during the process of the audio acquisition device recording multi-source audio data;
[0071] The raw data is converted into posture data representing the posture of the audio acquisition device.
[0072] It should be noted that the inertial measurement unit can measure the three-axis attitude angle (or angular velocity) and acceleration of an object, which will not be described in detail here. The audio acquisition device is associated with an inertial measurement unit (IMU for short). For example, an IMU is provided in the audio acquisition device. In this embodiment, the process of converting the raw data output by the IMU into attitude data occurs in the audio separation device. In this way, the pressure on the audio acquisition device and the requirements for the audio acquisition device are reduced. For example, the audio acquisition device only needs to have a recording function and an IMU detection function.
[0073] In the embodiment of the present application, the audio separation device converts the raw data output by the IMU into posture data, thereby reducing the pressure on the audio acquisition device and lowering the requirements for the audio acquisition device.
[0074] To improve the quality of the posture data, in some embodiments of the present application, after converting the original data into posture data representing the posture of the audio acquisition device, the method further includes:
[0075] The posture data is filtered to remove data representing data noise and / or glitches.
[0076] It should be noted that gesture data may contain data noise and glitches caused by various external factors (such as hardware factors of the audio acquisition device or collisions with other objects during operation of the audio acquisition device). These data noise and glitches reduce the quality of the gesture data. To this end, after obtaining the gesture data, data filtering is first performed to filter out the data representing data noise and / or glitches to obtain better quality gesture data, which will be used for audio separation later.
[0077] In the embodiment of the present application, data filtering is performed on the posture data to eliminate the influence of data noise and burrs, thereby improving the quality of the posture data.
[0078] See also Figure 2 According to another aspect of the present application, in an exemplary embodiment, a method for extracting human voice is provided, the method comprising:
[0079] S201: Acquire data to be processed from an audio acquisition device, wherein the data to be processed includes: multi-source audio data recorded by the audio acquisition device and posture data during the recording process of the multi-source audio data by the audio acquisition device;
[0080] S202: Separating audio data of a target sound source from the multi-source audio data, and adjusting a beam direction to point toward or approach the target sound source according to the posture data during the separation process, thereby obtaining audio data of the target sound source;
[0081] S203: Performing sound processing on the audio data of the target sound source to obtain target human voice data, wherein the sound processing includes: human voice extraction and / or human voice enhancement.
[0082] It should be noted that for S201-S202, reference can be made to the similar content of S101-S102 in the above embodiment and will not be repeated here. In this embodiment, the target sound source is any designated person. After separating the audio data of the target sound source, the audio data typically contains interference such as noise. Therefore, through voice extraction and / or voice enhancement, this noise interference can be filtered out to obtain higher-quality target voice data, further improving the audio separation effect.
[0083] In some embodiments, after obtaining the target voice data, the identity of the target person in the target voice data can be identified, and different business processes can be performed based on the identity recognition results. For example, if the identity recognition result is a staff member, the target voice data can be used to conduct a work quality inspection to examine whether the staff member's work methods are appropriate and meet standards. If the identity recognition result is a customer, the target voice data can be used to create a persona profile, and appropriate experts can be assigned to provide services based on the persona profile.
[0084] Optionally, after obtaining the to-be-processed data originating from the audio acquisition device, the method further comprises:
[0085] According to the posture data, determine the posture change of the audio acquisition device in each recording period; according to the posture change of each recording period, determine the target recording period in which the beam direction needs to be adjusted;
[0086] In the separation process, the beam direction is adjusted to point to or approach the target sound source according to the posture data, comprising:
[0087] In the separation process, the beam direction is adjusted to point to or approach the target sound source according to the posture data of the target recording period.
[0088] Optionally, the posture change includes: time domain features and / or frequency domain features of the posture data; the multi-source audio data is recorded by the target personnel wearing the audio acquisition device;
[0089] According to the posture change of each recording period, determine the target recording period in which the beam direction needs to be adjusted, comprising:
[0090] According to the time domain features and / or frequency domain features of the posture data of each recording period, determine the behavior label corresponding to the target personnel in each recording period; the recording period corresponding to the target behavior label is determined as the target recording period; wherein, the target behavior label includes the same behavior label as any dangerous behavior label in the behavior label library.
[0091] Optionally, according to the posture change of each recording period, determine the target recording period in which the beam direction needs to be adjusted, comprising:
[0092] According to the posture change of each recording period, determine the recording period with low frequency and large amplitude of the posture; the recording period with low frequency and large amplitude of the posture is determined as the target recording period.
[0093] Optionally, obtaining the to-be-processed data originating from the audio acquisition device, comprising:
[0094] Receiving an encoded file originating from the audio acquisition device, wherein the encoded file is a single file formed by encoding the multi-source audio data recorded by the audio acquisition device and the posture data in the process of recording the multi-source audio data by the audio acquisition device; decoding the encoded file to obtain the multi-source audio data and the posture data.
[0095] Optionally, the step of obtaining the posture data in the to-be-processed data, comprising:
[0096] Obtaining the original data output by the inertial measurement unit in the process of recording the multi-source audio data by the audio acquisition device; converting the original data into posture data representing the posture of the audio acquisition device.
[0097] Optionally, after converting the original data into posture data representing the posture of the audio acquisition device, the method further includes:
[0098] The posture data is filtered to remove data representing data noise and / or glitches.
[0099] The human voice extraction method provided in this embodiment is based on the same concept as the audio separation method provided in the above embodiment of this application. For technical details not fully described in this embodiment, please refer to the specific processing content of the audio separation method provided in the above embodiment of this application, and will not be repeated here.
[0100] See also Figure 3 According to another aspect of the present application, in an exemplary embodiment, an audio separation method is provided, which is applied to an audio acquisition device, wherein the audio acquisition device includes: an audio acquisition unit with a sound recording function, and a posture acquisition unit with a posture detection function.
[0101] S301: In response to a start operation, the audio acquisition unit records and obtains multi-source audio data, and the posture acquisition unit obtains, through the posture acquisition unit, posture data during the recording process of the multi-source audio data.
[0102] S302: Sending multi-source audio data and gesture data to a target device, so that the target device separates the audio data of a target sound source from the multi-source audio data, and adjusts the beam direction to point to or approach the target sound source according to the gesture data during the separation process.
[0103] It should be noted that the audio acquisition device can be a wearable audio acquisition device. For example, the audio acquisition device can be a wearable smart ID card. When worn by an employee, it can serve as an ID card and also record the employee's work. The posture acquisition unit includes but is not limited to an IMU. For example, the posture data can be the raw data output by the IMU.
[0104] Multi-source audio data is an object that requires audio separation, and it contains sounds from multiple sources. The multiple sound sources include but are not limited to people and animals. For example, the multiple sound sources include multiple different people. Regarding the data format of the multi-source audio data, it can be any audio data format. For example, the multi-source audio data can be audio data formed by sub-band coding (SBC) technology. Regarding the data content or semantic content of the multi-source audio data, there is no limitation here, and it can be any semantic content.
[0105] The posture data may represent the posture of the audio capture device in space. It is understood that the posture data in the data to be processed is continuous posture data within a time period. This time period is the time period for recording multi-source audio data. Therefore, the posture data in the data to be processed can be used to determine the posture of the audio capture device or its posture changes during the recording of the multi-source audio data.
[0106] In some embodiments, the audio acquisition unit and the gesture acquisition unit can be controlled by the same start button, and can start and end operations synchronously. Furthermore, the audio acquisition unit and the gesture acquisition unit can use the same clock device for timing, or after the audio acquisition unit and the gesture acquisition unit have completed their operations, the data generated by the two units are time-aligned.
[0107] In an embodiment of the present application, the posture of the device during the audio recording process is taken into account, so that during the audio separation process, the beam direction is dynamically adjusted to point to or approach the sound source according to the posture data, so that the beam direction is locked to the sound source in real time during the separation process, and the separation effect is greatly enhanced.
[0108] In some embodiments of the present application, sending multi-source audio data and gesture data includes:
[0109] Encode multi-source audio data and gesture data into a single file and send the single file.
[0110] It should be noted that the target device for audio separation and the audio acquisition device are different electronic devices. In this way, the audio acquisition device is required to transmit the relevant data to the target device. In some embodiments, the audio acquisition device is required to upload the relevant data to the cloud, and the target device downloads the relevant data from the cloud. To facilitate uploading and downloading, the audio acquisition device can encode the multi-source audio data and posture data into a single file, and subsequently upload and download the single file. The target device can decode the single file to obtain the multi-source audio data and posture data. This application does not limit the encoding method.
[0111] In the embodiment of the present application, multi-source audio data and posture data are transmitted in the form of a single file, which improves the convenience of transmission.
[0112] Exemplary devices
[0113] Accordingly, the present application also provides an audio separation device, see Figure 4 As shown, the device includes:
[0114] The data acquisition module 401 is used to acquire data to be processed from an audio acquisition device, wherein the data to be processed includes: multi-source audio data recorded by the audio acquisition device and posture data during the recording process of the multi-source audio data by the audio acquisition device.
[0115] The audio separation module 402 is used to separate the audio data of the target sound source from the audio data of multiple sound sources, and adjust the beam direction to point to or approach the target sound source according to the posture data during the separation process.
[0116] In some embodiments of the present application, the device further comprises:
[0117] The recording period screening module is used to determine the posture changes of the audio acquisition device in each recording period based on the posture data; based on the posture changes in each recording period, determine the target recording period that needs to adjust the beam direction;
[0118] The audio separation module 402 is specifically configured to adjust the beam direction in the target recording period to point to or approach the target sound source according to the posture data of the target recording period during the separation process.
[0119] In some embodiments of the present application, the posture change includes: time domain features and / or frequency domain features of posture data; multi-source audio data is recorded by the target person wearing an audio acquisition device;
[0120] The recording period screening module is specifically used to determine the behavior label corresponding to the target person in each recording period based on the time domain characteristics and / or frequency domain characteristics of the posture data in each recording period; the recording period corresponding to the target behavior label is determined as the target recording period; wherein the target behavior label includes a behavior label that is the same as any dangerous behavior label in the behavior label library.
[0121] In some embodiments of the present application, the recording period screening module is specifically used to determine the recording period with low-frequency and large-scale posture changes based on the posture changes in each recording period; and determine the recording period with low-frequency and large-scale posture changes as the target recording period.
[0122] In some embodiments of the present application, the data acquisition module 401 is specifically used to receive an encoded file from an audio acquisition device, wherein the encoded file is a single file formed by encoding multi-source audio data recorded by the audio acquisition device and posture data during the process of recording the multi-source audio data by the audio acquisition device; the encoded file is decoded to obtain the multi-source audio data and posture data.
[0123] In some embodiments of the present application, the step of the data acquisition module 401 acquiring the posture data in the data to be processed includes:
[0124] The original data output by the inertial measurement unit during the process of the audio acquisition device recording multi-source audio data is obtained; the original data is converted into posture data representing the posture of the audio acquisition device.
[0125] In some embodiments of the present application, the device further comprises:
[0126] The data filtering module is used to filter the posture data to remove data representing data noise and / or glitches.
[0127] The audio separation device provided in this embodiment is based on the same concept as the audio separation method provided in the above embodiments of this application. It can execute the audio separation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects. For technical details not fully described in this embodiment, please refer to the specific processing content of the audio separation method provided in the above embodiments of this application, and will not be repeated here.
[0128] Accordingly, the present invention also provides a human voice extraction device, see Figure 5 As shown, the device includes:
[0129] The data acquisition module 401 is used to acquire data to be processed from an audio acquisition device, wherein the data to be processed includes: multi-source audio data recorded by the audio acquisition device and posture data during the recording process of the multi-source audio data by the audio acquisition device.
[0130] The audio separation module 402 is used to separate the audio data of the target sound source from the audio data of multiple sound sources, and adjust the beam direction to point to or approach the target sound source according to the posture data during the separation process.
[0131] The vocal processing module 403 is configured to perform sound processing on the audio data of the target sound source to obtain target vocal data, wherein the sound processing includes: vocal extraction and / or vocal enhancement.
[0132] In some embodiments of the present application, the device further comprises:
[0133] The recording period screening module is used to determine the posture changes of the audio acquisition device in each recording period based on the posture data; based on the posture changes in each recording period, determine the target recording period that needs to adjust the beam direction;
[0134] The audio separation module 402 is specifically configured to adjust the beam direction within the target recording period to point to or approach the target sound source according to the posture data of the target recording period during the separation process.
[0135] In some embodiments of the present application, the posture change includes: time domain features and / or frequency domain features of posture data; multi-source audio data is recorded by the target person wearing an audio acquisition device;
[0136] The recording period screening module is specifically used to determine the behavior label corresponding to the target person in each recording period based on the time domain characteristics and / or frequency domain characteristics of the posture data in each recording period; the recording period corresponding to the target behavior label is determined as the target recording period; wherein the target behavior label includes a behavior label that is the same as any dangerous behavior label in the behavior label library.
[0137] In some embodiments of the present application, the recording period screening module is specifically used to determine the recording period with low-frequency and large-scale posture changes based on the posture changes in each recording period; and determine the recording period with low-frequency and large-scale posture changes as the target recording period.
[0138] In some embodiments of the present application, the data acquisition module 401 is specifically used to receive an encoded file from an audio acquisition device, wherein the encoded file is a single file formed by encoding multi-source audio data recorded by the audio acquisition device and posture data during the process of recording the multi-source audio data by the audio acquisition device; the encoded file is decoded to obtain the multi-source audio data and posture data.
[0139] In some embodiments of the present application, the step of the data acquisition module 401 acquiring the posture data in the data to be processed includes:
[0140] The original data output by the inertial measurement unit during the process of the audio acquisition device recording multi-source audio data is obtained; the original data is converted into posture data representing the posture of the audio acquisition device.
[0141] In some embodiments of the present application, the device further comprises:
[0142] The data filtering module is used to filter the posture data to remove data representing data noise and / or glitches.
[0143] The vocal extraction device provided in this embodiment shares the same concept as the vocal extraction method provided in the aforementioned embodiments of this application. It can implement the vocal extraction method provided in any of the aforementioned embodiments of this application and possesses the corresponding functional modules and beneficial effects. Technical details not fully described in this embodiment can be found in the specific processing details of the vocal extraction method provided in the aforementioned embodiments of this application and will not be further elaborated here.
[0144] Accordingly, the embodiment of the present application further provides an audio separation device, which is applied to an audio acquisition device, wherein the audio acquisition device includes: an audio acquisition unit with a sound recording function, a posture acquisition unit with a posture detection function, see Figure 6 As shown, the device includes:
[0145] A response module 601 is configured to, in response to a start operation, obtain multi-source audio data by recording through an audio acquisition unit, and obtain, through a posture acquisition unit, posture data during the recording process of the multi-source audio data by the audio acquisition unit;
[0146] The sending module 602 is used to send multi-source audio data and posture data to the target device, so that the target device separates the audio data of the target sound source from the multi-source audio data, and adjusts the beam direction to point to or approach the target sound source according to the posture data during the separation process.
[0147] In some embodiments of the present application, the sending module 602 is specifically configured to encode the multi-source audio data and the gesture data into a single file and send the single file.
[0148] The audio separation device provided in this embodiment is based on the same concept as the audio separation method provided in the above embodiments of this application. It can execute the audio separation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects. For technical details not fully described in this embodiment, please refer to the specific processing content of the audio separation method provided in the above embodiments of this application, and will not be repeated here.
[0149] It should be understood that the modules in the above apparatus can be implemented in the form of processor calling software. For example, the apparatus includes a processor connected with a memory, the memory stores instructions, and the processor calls the instructions stored in the memory to implement any of the above methods or to realize the functions of the units of the apparatus. The processor can be a general processor, such as a CPU or a microprocessor, and the memory can be an internal memory or an external memory of the apparatus. Alternatively, the units in the apparatus can be implemented in the form of hardware circuit. The functions of some or all of the units can be realized by the design of the hardware circuit, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the units are realized by the design of the logical relationship of the elements in the circuit. For another example, in another implementation, the hardware circuit can be realized by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of some or all of the units. All the units of the above apparatus can be implemented in the form of processor calling software, or all the units can be implemented in the form of hardware circuit, or part of the units are implemented in the form of processor calling software, and the remaining part is implemented in the form of hardware circuit.
[0150] In the embodiments of the present application, the processor is a circuit with signal processing capability. In one implementation, the processor can be a circuit with instruction reading and running capability, such as CPU, microprocessor, GPU, or DSP, etc. In another implementation, the processor can realize certain functions through the logical relationship of the hardware circuit, which is fixed or can be reconfigured. For example, the processor is an ASIC or a PLD implemented hardware circuit, such as FPGA, etc. In the reconfigurable hardware circuit, the process of the processor loading the configuration document to realize the configuration of the hardware circuit can be understood as the process of the processor loading the instructions to realize the functions of some or all of the units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as NPU, TPU, DPU, etc.
[0151] It can be seen that each unit in the above apparatus can be one or more processors (or processing circuits) configured to implement the above methods, such as CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0152] In addition, the various units in the above apparatus may be fully or partially integrated together, or may be implemented independently. In one implementation, these units are integrated together and implemented in the form of a system-on-chip (SOC). The SOC may include at least one processor for implementing any of the above methods or implementing the functions of the various units of the apparatus. The at least one processor may be of different types, such as a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0153] Exemplary Systems
[0154] Accordingly, an embodiment of the present application further provides an audio processing system, the audio processing system comprising: an audio acquisition device and a target device;
[0155] The audio acquisition device is configured to execute the audio separation method applied to the audio acquisition device provided in the above embodiments;
[0156] and / or,
[0157] The target device is configured to execute the audio separation method provided in the above embodiments, or execute the human voice extraction method provided in the above embodiments.
[0158] The audio processing system provided in this embodiment is based on the same concept as the audio separation method and voice extraction method provided in the aforementioned embodiments of this application. For technical details not fully described in this embodiment, please refer to the specific processing content of the audio separation method and voice extraction method provided in the aforementioned embodiments of this application, and will not be repeated here.
[0159] For ease of understanding, the audio processing system is described below as an example. Figure 7 As shown, the audio processing system includes a terminal device 701 and a cloud device 702.
[0160] After receiving the user's start-up operation, the terminal device 701 simultaneously collects IMU data and records audio. It then encodes the IMU data and the audio data into a single file. The single file is stored locally and uploaded to the cloud device 702.
[0161] The cloud device 702 first parses a single file to obtain separate audio data and IMU data. The IMU data is solved and processed to extract behavioral features related to the wearer's behavior; the audio data is decoded by SBC to obtain decoded audio data. The behavioral data and the decoded audio data are used as two inputs, and the audio data of the target sound source is separated using a spatial separation algorithm. If the audio data recorded by the terminal device 701 contains sales audio and customer audio, the separated audio data will contain two copies, one containing only sales audio from the sales sound source, and the other containing only customer audio from the customer sound source. After the two audio data are subjected to voice extraction and enhancement processing, the sales voice and customer voice are obtained. They are then transcribed for quality inspection and portrait generation.
[0162] The cloud device 702 calculates and processes the IMU data, separates the audio data, and obtains the human voice. Figure 8 Shown, including:
[0163] S801: Start.
[0164] S802: Data filtering, converting the IMU data from raw data state to attitude data, and filtering out data noise and burrs.
[0165] S803: Behavior feature extraction: extract behavior features from the posture data obtained in S802 to determine the behavior features of each recording period. The behavior features include: behavior features of posture with small amplitude and high frequency changes, and behavior features of posture with low frequency and large amplitude changes.
[0166] S804: Dynamic beam adjustment: During the audio data separation process, a target recording period requiring dynamic beam adjustment is determined based on the behavioral characteristics, and the beam direction is adjusted to point toward or approach the target sound source based on the gesture data during the target recording period.
[0167] S805: Voice Enhancement: Voice extraction and enhancement are performed on the audio data separated in S804 to obtain voice data.
[0168] S806: End.
[0169] Exemplary electronic devices
[0170] The present application embodiment provides an electronic device, see Figure 9 As shown, the device includes:
[0171] Memory 900 and processor 910;
[0172] The memory 900 is connected to the processor 910 and is used to store programs;
[0173] The processor 910 is configured to implement the audio separation method or the human voice extraction method disclosed in any of the above embodiments by running the program stored in the memory 900 .
[0174] Specifically, the electronic device may further include: a bus, a communication interface 920 , an input device 930 and an output device 940 .
[0175] The processor 910, the memory 900, the communication interface 920, the input device 930 and the output device 940 are interconnected via a bus.
[0176] A bus may include a pathway that transfers information between components of a computer system.
[0177] Processor 910 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, or the like, or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. Alternatively, it can be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware components.
[0178] The processor 910 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0179] The memory 900 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include program code, which may include computer operating instructions. More specifically, the memory 900 may include read-only memory (ROM), other types of static storage devices that can store static information and instructions, random access memory (RAM), other types of dynamic storage devices that can store information and instructions, disk storage, flash, etc.
[0180] The input device 930 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0181] Output device 940 may include devices that allow information to be output to a user, such as a display screen, printer, speakers, etc.
[0182] The communication interface 920 may include any transceiver or similar device for communicating with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0183] The processor 910 executes the program stored in the memory 900 and calls other devices, which can be used to implement the various steps of any audio separation method or human voice extraction method provided in the above embodiments of the present application.
[0184] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in a memory through the data interface to execute the audio separation method or human voice extraction method introduced in any of the above embodiments. The specific processing process and its beneficial effects can be found in the embodiment introduction of the above-mentioned audio separation method or human voice extraction method.
[0185] Exemplary computer program products and storage media
[0186] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the audio separation method or human voice extraction method according to various embodiments of the present application described in any of the above embodiments of this specification.
[0187] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0188] In addition, an embodiment of the present application may also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to execute the steps of the audio separation method or human voice extraction method according to various embodiments of the present application described in any of the above embodiments of this specification.
[0189] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0190] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0191] The steps in the methods of each embodiment of the present application can be adjusted in sequence, merged, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0192] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be merged, divided, and deleted according to actual needs.
[0193] In the several embodiments provided in this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or submodules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.
[0194] The modules or submodules described as separate components may or may not be physically separate, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules may be selected to achieve the purpose of this embodiment according to actual needs.
[0195] In addition, each functional module or submodule in each embodiment of the present application may be integrated into a processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into a single module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or software functional modules or submodules.
[0196] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0197] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, software units executed by a processor, or a combination of the two. The software units may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0198] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0199] The above description of the disclosed embodiments will enable those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is to be construed in the widest manner consistent with the principles and novel features disclosed herein.
Claims
1. An audio separation method, characterized in that: The method comprises: Acquire data to be processed from an audio acquisition device, wherein the data to be processed includes: multi-source audio data recorded by the audio acquisition device and posture data during the process of the audio acquisition device recording the multi-source audio data; Determining, based on the posture data, a posture change of the audio acquisition device in each recording period; the posture change includes: time domain features and / or frequency domain features of the posture data; According to the posture changes in each recording period, determine the target recording period where the beam direction needs to be adjusted; Separating audio data of a target sound source from the multi-source audio data, and during the separation process, adjusting the beam direction within the target recording period to point toward or approach the target sound source according to the posture data of the target recording period; According to the posture changes in each recording period, determine the target recording period that needs to adjust the beam direction, including: Determining the behavior label corresponding to the target person in each recording period based on the time domain characteristics and / or frequency domain characteristics of the posture data in each recording period; The recording period corresponding to the target behavior label is determined as the target recording period; wherein, the target behavior label includes a behavior label that is the same as any dangerous behavior label in the behavior label library; dangerous behavior refers to a behavior that will cause the separation effect to deteriorate when separating the audio according to the current beam direction after the behavior is implemented.
2. The method according to claim 1, characterized in that According to the posture changes in each recording period, determine the target recording period that needs to adjust the beam direction, including: According to the posture changes in each recording period, determine the recording period with low frequency and large changes in posture; The recording period in which the posture changes in low frequency and large amplitude is determined as the target recording period.
3. The method according to claim 1, characterized in that Obtain data to be processed from the audio capture device, including: Receiving an encoded file from the audio acquisition device, wherein the encoded file is a single file formed by encoding multi-source audio data recorded by the audio acquisition device and gesture data during the recording of the multi-source audio data by the audio acquisition device; The encoded file is decoded to obtain the multi-source audio data and the posture data.
4. The method according to claim 1, wherein The step of obtaining the posture data in the data to be processed includes: Acquire raw data output by an inertial measurement unit during a process in which the audio acquisition device records the multi-source audio data; The original data is converted into posture data representing the posture of the audio acquisition device.
5. The method according to claim 4, characterized in that After converting the original data into posture data representing the posture of the audio acquisition device, the method further includes: The posture data is subjected to data filtering to remove data representing data noise and / or glitches.
6. A method for extracting human voice, characterized in that: The method comprises: Acquire data to be processed from an audio acquisition device, wherein the data to be processed includes: multi-source audio data recorded by the audio acquisition device and posture data during the process of the audio acquisition device recording the multi-source audio data; Determining, based on the posture data, a posture change of the audio acquisition device in each recording period; the posture change includes: time domain features and / or frequency domain features of the posture data; According to the posture changes in each recording period, determine the target recording period where the beam direction needs to be adjusted; Separating audio data of a target sound source from the multi-source audio data, and during the separation process, adjusting the beam direction within the target recording period to point toward or approach the target sound source according to the posture data of the target recording period, thereby obtaining the audio data of the target sound source; Performing sound processing on the audio data of the target sound source to obtain target human voice data, wherein the sound processing includes: human voice extraction and / or human voice enhancement; According to the posture changes in each recording period, determine the target recording period that needs to adjust the beam direction, including: Determining the behavior label corresponding to the target person in each recording period based on the time domain characteristics and / or frequency domain characteristics of the posture data in each recording period; The recording period corresponding to the target behavior label is determined as the target recording period; wherein, the target behavior label includes a behavior label that is the same as any dangerous behavior label in the behavior label library; dangerous behavior refers to a behavior that will cause the separation effect to deteriorate when separating the audio according to the current beam direction after the behavior is implemented.
7. An audio separation method, characterized in that: Applied to an audio acquisition device, wherein the audio acquisition device comprises: an audio acquisition unit with a sound recording function, and a posture acquisition unit with a posture detection function, the method comprising: In response to a start operation, the audio acquisition unit records and obtains multi-source audio data, and the posture acquisition unit obtains, through the posture acquisition unit, posture data during the recording of the multi-source audio data by the audio acquisition unit; The multi-source audio data and the posture data are sent to a target device, so that the target device determines the posture changes of the audio acquisition device in each recording period based on the posture data, determines the behavior label corresponding to the target person in each recording period based on the time domain characteristics and / or frequency domain characteristics of the posture data included in the posture changes in each recording period, determines the recording period corresponding to the target behavior label as the target recording period, separates the audio data of the target sound source from the multi-source audio data, and during the separation process, adjusts the beam direction in the target recording period to point to or approach the target sound source based on the posture data of the target recording period; wherein the target behavior label includes a behavior label that is the same as any dangerous behavior label in a behavior label library; dangerous behavior indicates a behavior that will result in a poor separation effect when the audio is separated according to the current beam direction after the behavior is implemented.
8. The method according to claim 7, characterized in that Sending the multi-source audio data and the gesture data includes: The multi-source audio data and the gesture data are encoded into a single file, and the single file is transmitted.
9. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the method according to any one of claims 1 to 8 by running the program in the memory.
10. An audio processing system, characterized in that: Including audio acquisition device and target device; Wherein, the audio acquisition device is configured to perform the audio separation method according to claim 7 or 8; and / or, The target device is configured to execute the audio separation method according to any one of claims 1 to 5, or execute the human voice extraction method according to claim 6.
11. A computer program product, characterized in that include: A computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Directional recording method and apparatus, and recording device
CN106486147A
Voice separation method, device, system, and storage medium
CN111145775A