Audio recording method, device, electronic device, system and storage medium

By separating the speech signal into different audio tracks and optimizing it in multi-speaker recording scenarios, the problem of inconsistent recording quality between recording devices is solved, and high-quality audio output is achieved.

CN116668588BActive Publication Date: 2026-03-20ANHUI IFLYREC TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In multi-speaker recording scenarios, the different recording quality and distance of each voice acquisition device lead to inconsistent voice volume, resulting in overload, popping distortion, and mutual interference in the merged voice, thus affecting the quality of the recorded audio.

Method used

Speaker separation technology separates the speech signal into different audio tracks, and each track is optimized, including equalization, noise reduction, compression, delay compensation, and adaptive gain adjustment, before being mixed.

Benefits of technology

It improves the audio quality of multi-speaker recording scenarios, avoids problems such as voice interference, noise and distortion, and enhances the overall quality of the recorded audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116668588B_ABST
    Figure CN116668588B_ABST
Patent Text Reader

Abstract

The application provides a recording method, device, electronic equipment, system and storage medium, and relates to the technical field of voice processing. The method comprises the following steps: acquiring voice signals collected by at least one voice collection device; performing speaker separation on the voice signals, and adding target voice signals of each speaker separated out into different audio tracks, wherein the target voice signals and the audio tracks are in one-to-one correspondence; respectively performing optimization processing on the target voice signals in each audio track to obtain target audio track signals of each audio track; and performing mixing processing on the target audio track signals of each audio track to obtain a recording audio file. The technical scheme provided by the application can improve the quality of a recording audio in a multi-speaker recording scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, and in particular to a recording method, device, electronic equipment, system and storage medium. BACKGROUND

[0002] Meetings occupy a large proportion in daily office work, and recording in meetings can help quickly record meeting content. Compared with other recording scenarios, there are multiple speakers in a meeting scenario, and some speakers can also be personnel remotely accessing the meeting. In the case of a large number of personnel, there will be more noise, which will affect the quality of the recorded audio.

[0003] In related technologies, the speech of meeting personnel can be collected through at least one microphone in a meeting place, and in the case of accessing a remote meeting, the system speech transmitted by a remote meeting system to the meeting place can also be collected through internal recording. Then, the same linear gain processing is performed on the speech, and the gain processed speech is merged to obtain the recording audio of the meeting. However, due to different recording qualities of the speech collection devices, different distances of the speakers from the speech collection devices, and other reasons, the sizes of the speech collected by the speech collection devices are inconsistent. After uniform linear gain processing is performed on the collected speech, the larger speech signal is prone to overload, resulting in distortion of the merged speech, and mutual interference and distortion exist when the speech collection devices simultaneously collect sound, which makes the quality of the finally recorded audio poor. SUMMARY

[0004] To solve the problems in the prior art, the present application provides a recording method, device, electronic equipment, system and storage medium to improve the quality of the recording audio in a multi-speaker recording scenario.

[0005] The present application provides a recording method, comprising:

[0006] obtaining speech signals collected by at least one speech collection device;

[0007] performing speaker separation on the speech signals, and adding target speech signals of each separated speaker to different audio tracks, the target speech signals and the audio tracks corresponding one by one;

[0008] respectively performing optimization processing on the target speech signals in each audio track to obtain target audio track signals of each audio track;

[0009] performing audio mixing processing on the target audio track signals of each audio track to obtain a recording audio file.

[0010] According to the recording method provided by the application, the target voice signals in each audio track are respectively optimized to obtain target audio track signals of each audio track, which comprises:

[0011] The target voice signals in each audio track are synchronized to obtain synchronized voice signals of each audio track.

[0012] The synchronized voice signals of each audio track are respectively optimized to obtain target audio track signals corresponding to each audio track respectively.

[0013] According to the recording method provided by the application, the target voice signals in each audio track are synchronized to obtain synchronized voice signals of each audio track, which comprises:

[0014] Audio features of the target voice signals in each audio track are respectively extracted to obtain audio feature data corresponding to each audio track respectively.

[0015] Audio feature offsets between the target voice signals in each audio track are determined based on the audio feature data.

[0016] The target voice signals in each audio track are synchronized based on the audio feature offsets to obtain the synchronized voice signals of each audio track.

[0017] According to the recording method provided by the application, the audio feature offsets comprise time offsets and / or frequency offsets; the target voice signals in each audio track are synchronized based on the audio feature offsets to obtain the synchronized voice signals of each audio track, which comprises:

[0018] The target voice signals in each audio track are adjusted to reduce the time offsets and / or the frequency offsets to obtain the synchronized voice signals of each audio track.

[0019] According to the recording method provided by the application, the optimization processing comprises at least one of equalization, noise reduction, compression, delay compensation and adaptive gain adjustment.

[0020] According to the recording method provided by the application, the at least one voice collecting device comprises a first voice collecting device and a second voice collecting device; the voice signals collected by the at least one voice collecting device are obtained, which comprises:

[0021] The first voice signals collected by the first voice collecting device and the second voice signals collected by the second voice collecting device are obtained, wherein the second voice signals are voice signals output through a loudspeaker in the second voice collecting device.

[0022] The application further provides a recording device, comprising:

[0023] a voice acquisition module, configured to acquire voice signals collected by at least one voice collection device;

[0024] a voice separation module, configured to perform speaker separation on the voice signals, and add target voice signals of each speaker separated out into different audio tracks, the target voice signals and the audio tracks corresponding one by one;

[0025] an optimization module, configured to perform optimization processing on the target voice signals in each audio track respectively, to obtain target audio track signals of each audio track;

[0026] a mix sound module, configured to perform mix sound processing on the target audio track signals of each audio track, to obtain a recording audio file.

[0027] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the recording method according to any one of the above when executing the computer program.

[0028] The application further provides a non-transitory computer readable storage medium, having a computer program stored thereon, wherein the computer program is executable by a processor to implement the recording method according to any one of the above.

[0029] The application further provides a computer program product, comprising a computer program, wherein the computer program is executable by a processor to implement the recording method according to any one of the above.

[0030] The recording method, device, electronic device, system and storage medium provided by the application can separate voice signals collected by at least one voice collection device, add target voice signals of each speaker separated out into different audio tracks, so that the target voice signals of each speaker correspond to one audio track, and can process the target voice signals of each speaker separately, improve the quality of the target voice signals in each audio track through optimization processing on the target voice signals in each audio track respectively, and improve the overall quality of the recording audio file obtained through mix sound processing on the target audio track signals in each audio track. In this way, the target voice signals of different speakers are placed in different audio tracks, and based on multi-audio track recording, the problems of voice interference, noise and distortion can be avoided, thereby improving the quality of the recording audio in a multi-speaker scenario. BRIEF DESCRIPTION OF DRAWINGS

[0031] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0032] Figure 1 is a flowchart of the recording method provided by the embodiment of the present application;

[0033] Figure 2 is a structural schematic diagram of the recording device provided by the embodiment of the present application;

[0034] Figure 3 is a structural schematic diagram of the electronic device provided by the embodiment of the present application;

[0035] Figure 4 is a structural schematic diagram of the recording system provided by the embodiment of the present application. DETAILED DESCRIPTION

[0036] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0037] It should be noted that the serial numbers of the described objects in the present application, such as "first", "second", etc., are only used to distinguish the described objects, and do not have any sequence or technical meaning.

[0038] The following will be described in combination with Figure 1 The recording method of the present application will be described. The recording method can be applied to electronic devices such as terminal devices or servers. Among them, the terminal device can include mobile phones, computers, tablet computers, wearable devices, etc.; the server can include independent servers, cluster servers or cloud servers, etc. The recording method can also be applied to the recording device provided in the electronic device such as the terminal device or the server, and the recording device can be realized by software, hardware or combination of both. The following will take the recording method applied to the electronic device as an example to describe the recording method provided by the embodiment of the present application.

[0039] Figure 1 The flowchart of the recording method provided by the embodiment of the present application is exemplarily shown, and with reference to Figure 1 The recording method can include the following steps 110-140.

[0040] Step 110: obtaining a voice signal collected by at least one voice collection device.

[0041] In a conference site, a voice signal of a conference participant can be collected by at least one voice collection device. The electronic device can obtain the collected voice signal from the at least one voice collection device, and the voice signal can include voice signals of multiple speakers.

[0042] For example, the voice collection device can be a microphone device, and at least one microphone device can be deployed in the conference site to collect voice signals of the participants in the conference site.

[0043] For example, in a remote conference scenario, the at least one voice collection device can include a first voice collection device and a second voice collection device, and obtaining the voice signal collected by the at least one voice collection device can include: obtaining a first voice signal collected by the first voice collection device and a second voice signal collected by the second voice collection device, the second voice signal being a voice signal output by a loudspeaker in the second voice collection device.

[0044] Specifically, the first voice collection device can collect a first voice signal in the conference site; the second voice collection device can include a loudspeaker, and a second voice signal transmitted by a remote end of a remote conference can be output by the loudspeaker, and the second voice collection device can collect the second voice signal output by the loudspeaker or collect the second voice signal in an internal recording manner.

[0045] Step 120: performing speaker separation on the voice signal, and adding the separated target voice signal of each speaker to a different audio track.

[0046] The target voice signal and the audio track correspond to each other, that is, the target voice signal of one speaker corresponds to one audio track.

[0047] After the electronic device obtains the voice signal collected by the voice collection device, the electronic device can perform speaker separation on the voice signal to obtain a target voice signal of each speaker, and then set an audio track for each speaker and add the target voice signal of the speaker to the corresponding audio track. For example, after performing speaker separation, three target voice signals of three speakers are obtained, which are a target voice signal A of speaker 1, a target voice signal B of speaker 2, and a target voice signal C of speaker 3. Three audio tracks can be allocated, the target voice signal A is added to audio track 1, the target voice signal B is added to audio track 2, and the target voice signal C is added to audio track 3. In this way, the target voice signals of different speakers are placed in different audio tracks, and the target voice signal in each audio track can be processed individually, that is, the target voice signal of each speaker can be processed individually.

[0048] Exemplarily, voice activity detection (VAD) can be performed on the voice signal, voiceprint features or timbre features can be extracted from the segmented voice segments, and then the voiceprint features or timbre features can be clustered to realize speaker separation of the voice signal.

[0049] Step 130: The target voice signals in each audio track are respectively optimized to obtain target audio track signals of each audio track.

[0050] After adding the target voice signals of different speakers into different audio tracks and making the target voice signals of one speaker correspond to one audio track, the target voice signals in the audio tracks can be respectively optimized to improve the quality of the target voice signals in each audio track.

[0051] The optimization processing can include at least one of equalization, noise reduction, compression, delay compensation and adaptive gain, but is not limited thereto.

[0052] The noise reduction method can include a digital filtering algorithm, an adaptive filtering algorithm, a wavelet transform algorithm or a signal amplitude-based noise reduction algorithm, and the noise interference in the synchronous voice signal can be removed by noise reduction. The distortion optimization can include a signal recovery method and a least square method. Equalization can reduce the dynamic range of the synchronous voice signal.

[0053] Step 140: The target audio track signals of each audio track are mixed to obtain a recording audio file.

[0054] After the target voice signals in each audio track are optimized, the target audio track signals of each audio track with high quality can be obtained.

[0055] It can be understood that in the embodiments of the present application, the speech signals in a multi-speaker scene can be collected based on multi-track recording by using at least one speech collection device, and the speech signals can be added to different audio tracks according to different speakers based on speaker separation technology, so that the target speech signals of different speakers can be processed separately in different audio tracks. Due to the differences in speech collection quality of different speech collection devices, and the different volumes of different speakers, the quality of the collected speech signals of different speakers also differs, and echo is likely to occur when each speech collection device collects speech signals at the same time. These will bring about phenomena such as uncoordinated sound, distortion and noise, and if the speech signals are directly mixed and combined, it will lead to problems such as audio quality degradation and speech confusion, and reduce the quality of the final recorded audio.

[0056] The recording method provided by the embodiments of the present application separates the speech signals collected by at least one speech collection device, and adds the target speech signals of each speaker separated to different audio tracks, so that the target speech signals of each speaker correspond to an audio track. In this way, the target speech signals of each speaker can be processed separately, and by optimizing the target speech signals in each audio track respectively, the quality of the target speech signals in each audio track can be improved, and when the target audio signals in each audio track are mixed, the overall quality of the recorded audio file obtained by mixing can be improved. In this way, by placing the target speech signals of different speakers in different audio tracks, based on multi-track recording, problems such as speech interference, noise and distortion can be avoided, thereby improving the quality of the recorded audio in a multi-speaker scene.

[0057] Based on Figure 1 The recording method of the corresponding embodiments, in an example embodiment, after placing the target speech signals of different speakers in different audio tracks, the target speech signals of each audio track can be individually optimized and mixed; or the target speech signals between the audio tracks can be synchronized, and the synchronized speech signals of each audio track can be individually optimized, and then the optimized speech signals of each audio track can be mixed. By synchronizing, the echo effect, noise and distortion between the target speech signals of different speakers can be eliminated, and the quality of the recorded audio file can be further improved.

[0058] Specifically, the target speech signals in the audio tracks are respectively optimized to obtain the target audio track signals of the audio tracks, which can include: synchronizing the target speech signals in the audio tracks to obtain the synchronized speech signals of the audio tracks; and respectively optimizing the synchronized speech signals of the audio tracks to obtain the target audio track signals respectively corresponding to the audio tracks.

[0059] Specifically, the target speech signals in the audio tracks can be compared, and the target speech signals are adjusted according to the comparison result to synchronize the target speech signals. For example, the target speech signal of one of the audio tracks can be taken as a reference, the differences between the target speech signals of the other audio tracks and the target speech signal taken as the reference are determined, and the target speech signals of the other audio tracks are adjusted based on the differences to reduce the differences until the differences are less than a set threshold or 0, thereby synchronizing the target speech signals. Alternatively, the audio tracks can be grouped, the target speech signals of the audio tracks in the same group are synchronized first, and then the synchronization between the groups is performed to synchronize the target speech signals.

[0060] For example, the target speech signals of the audio tracks can be respectively optimized first, and then the synchronized speech signals of the audio tracks are obtained to obtain the target audio track signals respectively corresponding to the audio tracks.

[0061] For example, at least one of the time, frequency and spectral width of the target speech signals can be synchronized, but the application is not limited thereto.

[0062] Based on Figure 1 In an example embodiment, the recording method of the corresponding embodiment can include: extracting the audio features of the target speech signals in the audio tracks to obtain the audio feature data respectively corresponding to the audio tracks; determining the audio feature offsets between the target speech signals in the audio tracks based on the audio feature data; and synchronizing the target speech signals in the audio tracks based on the audio feature offsets to obtain the synchronized speech signals of the audio tracks.

[0063] For example, the audio feature data can include at least one of the time information, frequency and spectral width, but the application is not limited thereto.

[0064] For example, the audio analysis tools such as Adobe Audition or Audacity can be used to analyze the sampling rate and bit depth of each audio track, and the Fast Fourier Transform (FFT) is used to obtain the audio feature data respectively corresponding to the audio tracks.

[0065] For example, after obtaining the audio feature data corresponding to each audio track, the audio feature data of each audio track can be compared to determine the audio feature offset, such as the time offset and / or the frequency offset, between the target speech signals in each audio track. Then, the target speech signals in each audio track are corrected based on the audio feature offset so that the audio feature offset is less than a preset offset threshold, achieving the purpose of synchronizing the target speech signals in each audio track. In this way, when mixing the synchronized speech signals of each audio track, the distortion and noise in the mixing process can be reduced, ensuring the quality of the recorded audio.

[0066] In an example embodiment, the audio feature offset can include a time offset and / or a frequency offset. Accordingly, synchronizing the target speech signals in each audio track based on the audio feature offset to obtain the synchronized speech signals of each audio track can include:

[0067] Adjusting the target speech signals in each audio track to reduce the time offset and / or the frequency offset to obtain the synchronized speech signals of each audio track.

[0068] For example, three audio tracks are established by speaker separation, the target speech signal A of speaker 1 is added to audio track 1, the target speech signal B of speaker 2 is added to audio track 2, and the target speech signal C of speaker 3 is added to audio track 3. The time information and frequency of the target speech signals corresponding to audio track 1, audio track 2, and audio track 3 are compared to determine the time offset and frequency offset between the target speech signals corresponding to each audio track. The target speech signals corresponding to audio track 1, audio track 2, and audio track 3 are corrected using the time offset and frequency offset so that the frequencies and times of all target speech signals are substantially consistent, thereby eliminating echo effects, distortion, and noise in the subsequent mixing process, and ensuring the quality of the recorded audio.

[0069] For example, the target speech signal A in audio track 1 can be taken as a reference to calculate the time offset a1 and the frequency offset b1 of the target speech signal B in audio track 2 relative to the target speech signal A. The target speech signal B is adjusted according to the time offset a1 and the frequency offset b1 so that the time offset a1 and the frequency offset b1 are less than a preset offset threshold or equal to 0, ensuring that the time information and frequency of the target speech signal B are substantially consistent with those of the target speech signal A. The target speech signal C in audio track 3 is processed in the same way as the target speech signal B, and the target speech signals A, B, and C are synchronized so that their frequency and time information are substantially consistent.

[0070] The recording method provided by the embodiment of the present application can be based on a multi-track recording method, and the voice signals collected by each voice collection device are separated by speaker separation technology, and then the target voice signals of each speaker separated are added to different audio tracks, so that the target voice signals of one speaker correspond to one audio track. In this way, the voices of different speakers can be more accurately placed in different audio tracks, facilitating the processing of the voices of each speaker. Then, for the target voice signals of each audio track, the sampling rate and bit depth parameters can be analyzed by using an audio analysis tool, the time offset and frequency offset between the target voice signals in each audio track are determined, the target voice signals of each audio track are corrected by the time offset and frequency offset, so that the time and frequency of the target voice signals in all audio tracks are consistent, and the distortion and noise in the subsequent mixing process are effectively reduced. Then, at least one optimization processing of equalization, noise reduction, compression, delay compensation and adaptive gain is performed on each synchronized voice signal after audio correction, and the target audio track signals obtained after the optimization processing of each audio track are merged, and the quality of the final obtained recording audio file is improved.

[0071] The recording method provided by the embodiment of the present application is based on multi-track recording, which can avoid voice confusion and distortion and improve the quality of the recording audio.

[0072] The recording device provided by the present application is described below, and the recording device described below can be correspondingly referred to the recording method described above.

[0073] Figure 2 An exemplary structure diagram of the recording device provided by the embodiment of the present application is shown, and the recording device 200 can include a voice acquisition module 210, a voice separation module 220, an optimization module 230 and a mixing module 240. Figure 2 As shown in the figure, the recording device 200 can include a voice acquisition module 210, a voice separation module 220, an optimization module 230 and a mixing module 240.

[0074] In an example embodiment, the optimization module 230 includes a synchronization unit and an optimization unit. The synchronization unit is configured to synchronize the target voice signals in each audio track to obtain synchronized voice signals of each audio track. The optimization unit is configured to perform optimization processing on the synchronized voice signals of each audio track to obtain the target audio track signals corresponding to each audio track, respectively.

[0075] In an example embodiment, the synchronizing unit comprises: a feature extraction subunit configured to extract audio features of the target speech signals in each audio track respectively, to obtain audio feature data corresponding to each audio track respectively; a determination subunit configured to determine audio feature offsets between the target speech signals in each audio track based on the audio feature data; and a synchronization subunit configured to synchronize the target speech signals in each audio track based on the audio feature offsets, to obtain synchronized speech signals of each audio track.

[0076] In an example embodiment, the audio feature offsets comprise time offsets and / or frequency offsets. Correspondingly, the synchronization subunit is specifically configured to adjust the target speech signals in each audio track to reduce the time offsets and / or the frequency offsets, to obtain the synchronized speech signals of each audio track.

[0077] In an example embodiment, the optimization unit is specifically configured to perform at least one optimization processing of equalization, noise reduction, compression, delay compensation and adaptive gain adjustment on the synchronized speech signals of each audio track.

[0078] In an example embodiment, the optimization processing comprises at least one of equalization, noise reduction, compression, delay compensation and adaptive gain.

[0079] In an example embodiment, the at least one speech collecting device comprises a first speech collecting device and a second speech collecting device. Specifically, the speech obtaining module 210 is specifically configured to obtain a first speech signal collected by the first speech collecting device and a second speech signal collected by the second speech collecting device, wherein the second speech signal is a speech signal output through a loudspeaker in the second speech collecting device.

[0080] Figure 3 An example structure diagram of an electronic device is shown in Figure 3 The electronic device can comprise a processor 310, a communication interface 320, a memory 330 and a communication bus 340, wherein the processor 310, the communication interface 320 and the memory 330 complete mutual communication through the communication bus 340. The processor 310 can invoke a logical instruction in the memory 330 to execute a recording method provided by any of the above method embodiments, which can comprise: obtaining speech signals collected by at least one speech collecting device; performing speaker separation on the speech signals, and adding target speech signals of each separated speaker into different audio tracks, wherein the target speech signals and the audio tracks are in one-to-one correspondence; performing optimization processing on the target speech signals in each audio track respectively, to obtain target audio track signals of each audio track; and performing mixing processing on the target audio track signals of each audio track, to obtain a recording audio file.

[0081] Moreover, the logic instructions in the memory 330 described above can be implemented in the form of software function units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0082] Figure 4 An exemplary structure diagram of the recording system provided by the embodiment of the present application is shown. Referring to Figure 4 The recording system can include an electronic device 410 and at least one voice collection device, such as a microphone device 420. Figure 4 Two voice collection devices, i.e., a first voice collection device 421 and a second voice collection device 422, are exemplarily shown in the figure. Each voice collection device can be in communication connection with the electronic device 410. The electronic device 410 can be, for example, a smart speaker. Figure 3 The electronic device corresponding to the embodiment.

[0083] It should be noted that Figure 4 Two voice collection devices are exemplarily shown in the figure, but this does not constitute a unique limitation on the number of voice collection devices in the recording system of the embodiment of the present application. The number of voice collection devices in the recording system is not limited in the embodiment of the present application.

[0084] Exemplarily, the voice collection device can be a microphone device. The recording system can include at least one microphone device in communication connection with the electronic device 410. These microphone devices can be deployed in different areas in a multi-speaker scenario to collect the voice signals of the speakers in different areas. For example, in a conference scenario, a microphone device can be configured for each participant.

[0085] Exemplarily, in the conference scenario, the remote conference can also be accessed, the voice signal in the local conference place can include the voice signal of the local participant and the voice signal of the remote participant output by the loudspeaker, at least one microphone device and one recording device can be deployed in the local conference place, the recording device can output the voice signal of the remote participant through the loudspeaker thereof, the voice signal collected by each microphone device and the voice signal of the remote participant recorded or collected by the loudspeaker of the recording device can be sent to the electronic device 410, and the electronic device 410 can process the received voice signal based on the recording method provided in the embodiments of the present application.

[0086] Exemplarily, one recording device can be deployed in the local conference place, the recording device can be in communication connection with at least one microphone device deployed in the local conference place, and the first voice signal of the participant in the local conference place can be collected through the microphone device. The recording device accesses the remote conference system, can receive the second voice signal of the remote participant sent by the remote conference system, and output the second voice signal through the loudspeaker thereof. The recording device can record the first voice signal collected by the microphone device and record the second voice signal of the remote participant, and then send the first voice signal and the second voice signal to the electronic device 410, the electronic device 410 can perform speaker separation on the first voice signal and the second voice signal respectively, add the voice signals of different speakers to different audio tracks according to different timbres, then correct the time and frequency of the voice signals in all audio tracks to be consistent, and perform optimization processing such as equalization, noise reduction, compression, delay compensation and adaptive gain on the voice signals in each audio track respectively, and mix the optimized audio tracks to obtain the final recording audio file.

[0087] In this way, when the internal recording and the external recording exist at the same time, that is, when the microphone device and the loudspeaker of the recording device work at the same time, the multi-audio track recording can be performed based on the speaker separation, and the voice signals of different speakers can be optimized and corrected respectively, which effectively avoids the problems such as echo influence, noise, volume imbalance and mutual interference between voice signals, and improves the quality of the recording audio.

[0088] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer-readable storage medium, and the computer program, when executed by a processor, enables a computer to perform the recording method provided by any of the above method embodiments, which may, for example, include: obtaining voice signals collected by at least one voice collection device; performing speaker separation on the voice signals, and adding target voice signals of each speaker separated out into different audio tracks, wherein the target voice signals and the audio tracks correspond one-to-one; respectively performing optimization processing on the target voice signals in each audio track to obtain target audio track signals of each audio track; and performing mixing processing on the target audio track signals of each audio track to obtain a recording audio file.

[0089] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, enables a computer to perform the recording method provided by any of the above method embodiments, which may, for example, include: obtaining voice signals collected by at least one voice collection device; performing speaker separation on the voice signals, and adding target voice signals of each speaker separated out into different audio tracks, wherein the target voice signals and the audio tracks correspond one-to-one; respectively performing optimization processing on the target voice signals in each audio track to obtain target audio track signals of each audio track; and performing mixing processing on the target audio track signals of each audio track to obtain a recording audio file.

[0090] The device embodiments described above are merely illustrative, wherein the units shown as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed on multiple network units. Some or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0091] From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in terms of the contribution to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0092] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application, but not to limit the same; and although the present application has been described in detail with reference to the foregoing embodiments, it should be appreciated by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A recording method, characterized in that, include: Acquire multiple voice signals collected by multiple voice acquisition devices; The multiple speech signals are segmented by speech activity detection to obtain multiple speech signal segments. The timbre features of each speech signal segment are extracted, and based on each timbre feature, the speech signal segments are clustered to obtain the target speech signal of each speaker. The separated target speech signals of each speaker are added to different audio tracks, and the target speech signal and the audio track correspond one-to-one. The target speech signal in each of the audio tracks is optimized to obtain the target audio track signal for each of the audio tracks. The optimization process includes: synchronizing the target speech signals in each of the audio tracks; The target audio track signals of each of the aforementioned audio tracks are mixed to obtain a recorded audio file.

2. The recording method according to claim 1, characterized in that, The step of optimizing the target speech signal in each of the audio tracks to obtain the target audio track signal for each of the audio tracks includes: The target speech signals in each of the audio tracks are synchronized to obtain synchronized speech signals for each of the audio tracks; The synchronous speech signals of each audio track are optimized to obtain the target audio track signals corresponding to each audio track.

3. The recording method according to claim 2, characterized in that, The step of synchronizing the target speech signals in each of the audio tracks to obtain synchronized speech signals for each of the audio tracks includes: The audio features of the target speech signal in each of the audio tracks are extracted to obtain the audio feature data corresponding to each of the audio tracks. Based on the audio feature data, determine the audio feature offset between the target speech signals in each of the audio tracks; The target speech signal in each of the audio tracks is synchronized based on the audio feature offset to obtain the synchronized speech signal for each of the audio tracks.

4. The recording method according to claim 3, characterized in that, The audio feature offset includes a time offset and / or a frequency offset; the step of synchronizing the target speech signal in each of the audio tracks based on the audio feature offset to obtain the synchronized speech signal for each of the audio tracks includes: With the goal of reducing the time offset and / or the frequency offset, the target speech signal in each of the audio tracks is adjusted to obtain the synchronized speech signal of each of the audio tracks.

5. The recording method according to any one of claims 1 to 4, characterized in that, The optimization process includes at least one of equalization, noise reduction, compression, delay compensation, and adaptive gain.

6. The recording method according to any one of claims 1 to 4, characterized in that, The plurality of voice acquisition devices includes a first voice acquisition device and a second voice acquisition device; The acquisition of voice signals collected by multiple voice acquisition devices includes: The system acquires a first voice signal acquired by the first voice acquisition device and a second voice signal acquired by the second voice acquisition device, wherein the second voice signal is a voice signal output through a speaker in the second voice acquisition device.

7. A recording device, characterized in that, include: The voice acquisition module is used to acquire multiple voice signals collected by multiple voice acquisition devices; The speech separation module is used to perform speech activity detection and segmentation on the multiple speech signals to obtain multiple speech signal segments, extract the timbre features of each speech signal segment, and cluster the speech signal segments based on each timbre feature to obtain the target speech signal of each speaker, and add the separated target speech signals of each speaker to different audio tracks, wherein the target speech signal and the audio track correspond one-to-one. The optimization module is used to optimize the target speech signal in each of the audio tracks to obtain the target audio track signal for each of the audio tracks. The optimization process includes: synchronizing the target speech signals in each of the audio tracks; The mixing module is used to mix the target audio track signals of each of the audio tracks to obtain a recorded audio file.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the recording method as described in any one of claims 1 to 6.

9. A recording system, characterized in that, It includes multiple voice acquisition devices and the electronic device as described in claim 8, wherein the voice acquisition devices are communicatively connected to the electronic device.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the recording method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Media file processing method, server and mobile terminal

    CN108174236A

  • Multi-source data processing method and device and readable storage medium

    CN111833898A

  • Recording data processing method and system, electronic device and storage medium

    CN112562712A