Conference audio processing method, device and storage medium thereof
By obtaining the current location information of the conference audio signal and matching it with a preset mapping table, the problem of traditional voiceprint recognition being cumbersome and costly is solved, achieving rapid and accurate identification of the speaker's identity and reducing costs.
Patent Information
- Application Number
- CN202211351357.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-10-31
AI Technical Summary
Identifying speakers through voiceprint recognition in traditional meeting recordings is a cumbersome and costly process.
The system acquires current location information from multiple sets of conference audio signals collected from different acquisition channels and matches it with a preset location identity mapping table. If the match is successful, the speaker's identity is determined; otherwise, a new identity is created in the mapping table.
It enables rapid and accurate identification of speakers, simplifies processing steps, reduces technical costs, and improves the efficiency and accuracy of conference audio processing.
Smart Images

Figure CN115914905B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio technology, in particular to a conference audio processing method and device and a storage medium thereof. BACKGROUND
[0002] With the continuous development of speech recognition technology, more and more industries begin to use speech recognition technology, among which, conference recording is an important application scenario of speech recognition technology.
[0003] In the traditional method, the conference content can be recorded by recording, and the speaker can be identified based on the voiceprint recognition method. However, the identification process of the voiceprint recognition method is relatively cumbersome, and identification errors are prone to occur. In addition, the use of the voiceprint recognition method usually requires the purchase of a specific voiceprint recognition algorithm, which increases the cost of the recording device.
[0004] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY
[0005] The main purpose of the present application is to provide a conference audio processing method, device and storage medium, which aims to solve the technical problems of cumbersome process of identifying speakers by voiceprint recognition in conventional conference recording and high implementation cost.
[0006] To achieve the above purpose, the present application provides a conference audio processing method, which comprises:
[0007] Obtaining current orientation information in a plurality of groups of conference audio signals collected by different collection channels;
[0008] Matching the current orientation information with a preset orientation-identity mapping table;
[0009] If the matching is successful, the target identity corresponding to the current orientation information in the orientation-identity mapping table is taken as the speaker identity of the conference audio signal.
[0010] Optionally, after the step of matching the orientation information with the preset orientation information, the method further comprises:
[0011] If the matching is not successful, a new identity is established in the orientation-identity mapping table according to the current orientation information;
[0012] The new identity is taken as the speaker identity of the conference audio signal.
[0013] Optionally, the plurality of groups of conference audio signals comprise a first sound signal and a second sound signal, and the step of obtaining current orientation information in a plurality of groups of conference audio signals collected by different collection channels comprises:
[0014] acquire a first sound signal and a second sound signal of a current speaker based on a preset paired earphone;
[0015] determine a first arrival time of the first sound signal and a second arrival time of the second sound signal;
[0016] determine a volume difference between the first sound signal and the second sound signal;
[0017] determine current orientation information of the current speaker relative to the preset paired earphone according to the first arrival time, the second arrival time and the volume difference.
[0018] Optionally, the orientation information includes angle information and distance information, and the step of determining the current orientation information of the current speaker relative to the preset paired earphone according to the first arrival time, the second arrival time and the volume difference includes:
[0019] determining angle information of the current speaker relative to the preset paired earphone according to the volume difference and a preset angle formula;
[0020] determining distance information of the current speaker relative to the preset paired earphone according to the first arrival time, the second arrival time and a preset distance formula.
[0021] Optionally, before the step of matching the current orientation information with a preset orientation identity mapping table, it includes:
[0022] acquiring compensation information of the multiple sets of conference audio signals, and correcting current orientation information in the multiple sets of conference audio signals according to the compensation information.
[0023] Optionally, the multiple sets of conference audio signals include a first sound signal and a second sound signal, and the step of acquiring compensation information of the multiple sets of conference audio signals includes:
[0024] acquiring a real-time position and a starting position of a preset paired earphone when the first sound signal or the second sound signal is received;
[0025] comparing the real-time position and the starting position to obtain an offset position of the preset paired earphone as the compensation information.
[0026] Optionally, the step of acquiring a starting position of a preset paired earphone includes:
[0027] when the preset paired earphone is started, acquiring a current position of the preset paired earphone as the starting position.
[0028] Optionally, the multiple sets of conference audio signals comprise bone conduction sound signals, and before the step of acquiring current position information in the multiple sets of conference audio signals collected by different collection channels, the method further comprises:
[0029] determining whether the volume value of the bone conduction sound signal is greater than a preset threshold value;
[0030] If the volume value is greater than the preset threshold value, a preset main speaker identity is taken as the speaker identity of the bone conduction sound signal.
[0031] The application further provides a conference audio processing device, which comprises a memory, a processor and a conference audio processing program stored in the memory and executable on the processor, and the conference audio processing program is configured to implement the steps of the conference audio processing method.
[0032] The application further provides a storage medium, which is a computer readable storage medium, and the conference audio processing program is stored on the computer readable storage medium and is executed by a processor to implement the steps of the conference audio processing method.
[0033] The application discloses a conference audio processing method, device and storage medium, and the conference audio processing method is applied to earphones. The current position information in the multiple sets of conference audio signals collected by different collection channels is acquired, and then the current position information is matched with a preset position identity mapping table. The speaker corresponding to the audio signal is determined according to the matching result of the current position information. If the matching is successful, the target identity corresponding to the current position information in the position identity mapping table is taken as the speaker identity of the conference audio signal. In the conference process, the conference is recorded by using the earphones that are convenient to carry, and then the acquired conference audio is processed to obtain the current position information in the multiple sets of conference audio signals. Then, the fast and accurate identification of the speaker identity is realized according to the current position information. The cumbersome process of identifying the speaker by using the voiceprint recognition mode is avoided. The fixed distance between the recording function of the earphones and the double-recording microphone of the earphones further reduces the processing steps of the conference audio, reduces the technical cost, and further improves the efficiency of the conference audio processing and the accuracy of the identity identification. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 is a structural schematic diagram of a conference audio processing device related to a hardware running environment of an embodiment scheme of the application;
[0035] Figure 2 is a scene schematic diagram related to a conference audio processing method related to an embodiment scheme of the application;
[0036] Figure 3 This is a flowchart illustrating a more complete embodiment of the conference audio processing method according to the present invention.
[0037] Figure 4 This is a flowchart illustrating the first embodiment of the conference audio processing method according to the present invention.
[0038] Figure 5 for Figure 4 A detailed flowchart of an embodiment of step S10;
[0039] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0040] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0041] Reference Figure 1 , Figure 1 This is a schematic diagram of the hardware operating environment of the conference audio processing device involved in the embodiments of the present invention.
[0042] like Figure 1 As shown, the conference audio processing device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0043] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the conference audio processing equipment and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0044] like Figure 1As shown, the memory 1005 as a storage medium can include an operating system, a data storage module, a network communication module, a user interface module, and a conference audio processing program.
[0045] In Figure 1 In the conference audio processing device shown, the network interface 1004 is mainly used for data communication with other devices; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the conference audio processing device of the application can be arranged in the conference audio processing device, and the conference audio processing device calls the conference audio processing program stored in the memory 1005 through the processor 1001 and performs the following operations:
[0046] Obtain the current orientation information in the multiple groups of conference audio signals collected by different collection channels;
[0047] Match the current orientation information with a preset orientation-identity mapping table;
[0048] If the matching is successful, the target identity corresponding to the current orientation information in the orientation-identity mapping table is taken as the speaker identity of the conference audio signal.
[0049] Further, the processor 1001 can call the conference audio processing program stored in the memory 1005 and further perform the following operations:
[0050] After the step of matching the current orientation information with the preset orientation-identity mapping table, the following steps are further included:
[0051] If the matching is not successful, a new identity is established in the orientation-identity mapping table according to the current orientation information;
[0052] The new identity is taken as the speaker identity of the conference audio signal.
[0053] Further, the multiple groups of conference audio signals include a first sound signal and a second sound signal, and the step of obtaining the current orientation information in the multiple groups of conference audio signals collected by different collection channels includes:
[0054] Collecting the first sound signal and the second sound signal of the current speaker based on the preset paired earphones respectively;
[0055] Determining the first arrival time of the first sound signal and the second arrival time of the second sound signal;
[0056] Determining the volume difference between the first sound signal and the second sound signal;
[0057] According to the first arrival time, the second arrival time and the volume difference, current position information of the current speaker relative to the preset paired earphone is determined.
[0058] Further, the position information includes angle information and distance information, and the step of determining the current position information of the current speaker relative to the preset paired earphone according to the first arrival time, the second arrival time and the volume difference includes:
[0059] According to the volume difference and a preset angle formula, angle information of the current speaker relative to the preset paired earphone is determined.
[0060] According to the first arrival time, the second arrival time and a preset distance formula, distance information of the current speaker relative to the preset paired earphone is determined.
[0061] Further, the processor 1001 can invoke the conference audio processing program stored in the memory 1005, and further perform the following operations:
[0062] Before the step of matching the current position information with the preset position identity mapping table, the following operation is included:
[0063] Compensation information of the plurality of groups of conference audio signals is obtained, and the current position information in the plurality of groups of conference audio signals is corrected according to the compensation information.
[0064] Further, the plurality of groups of conference audio signals include a first sound signal and a second sound signal, and the step of obtaining the compensation information of the plurality of groups of conference audio signals includes:
[0065] When the first sound signal or the second sound signal is received, a real-time position and a starting position of the preset paired earphone are obtained.
[0066] The real-time position and the starting position are compared to obtain an offset position of the preset paired earphone as the compensation information.
[0067] Further, the step of obtaining the starting position of the preset paired earphone includes:
[0068] When the preset paired earphone is started, a current position of the preset paired earphone is obtained as the starting position.
[0069] Further, the processor 1001 can invoke the conference audio processing program stored in the memory 1005, and further perform the following operations:
[0070] The multiple sets of conference audio signals include bone conduction sound signals, and before the step of acquiring current position information in the multiple sets of conference audio signals collected by different collection channels, the method further includes:
[0071] determining whether the volume value of the bone conduction sound signal is greater than a preset threshold value;
[0072] If the volume value is greater than the preset threshold value, a preset main speaker identity is taken as the speaker identity of the bone conduction sound signal.
[0073] The application provides a conference audio processing method, and an example is described below with the aid of a specific conference audio processing method scene schematic diagram, referring to Figure 2 The application is applied to a conference scene, and from the top of the head of a conference participant, the conference participant is seated around a conference table, a main speaker wears an earphone on the left side of the conference table, and after the conference starts, each speaker speaks one by one; the main speaker, as a conference content recorder, records the conference by using a preset paired earphone that can be carried, the preset paired earphone is a tws (true wireless stereo) Bluetooth earphone (hereinafter referred to as a Bluetooth earphone), the Bluetooth earphone includes two talk microphones of a left ear and a right ear, a first sound signal and a second sound signal are collected by the left and right ear talk microphones (a first talk microphone and a second talk microphone) respectively, and then are sent to a conference audio processing device; a first conference participant on the left side of the main speaker speaks as a current speaker, and a dashed arrow represents a propagation path of a generated audio signal of the current speaker to the left ear talk microphone and the right ear talk microphone. Figure 2 Since the current speaker is located on the left side of the main speaker, the position of the current speaker is closer to the left ear talk microphone, the propagation speed of sound in air is constant, therefore, the time when the first sound signal is collected by the left ear talk microphone is earlier than the time when the second sound signal is collected by the right ear talk microphone, that is, the first arrival time is earlier than the second arrival time; because the left ear talk microphone and the right ear talk microphone are separated by a head, the head blocks sound according to the head shadow effect, therefore, the volume (intensity) of the first sound signal collected by the left ear talk microphone closer to the current speaker is greater than that of the second sound signal, and then the volume difference between the two can be obtained by comparison; the first arrival time, the second arrival time and the volume difference in the multiple sets of conference audio signals collected are used to obtain the position information of the current speaker relative to the center point of the Bluetooth earphone (the main speaker), and then the identity of the current speaker is determined according to the position information.
[0074] The application provides a conference audio processing method, and an example is described below with the aid of a specific conference audio processing method scene schematic diagram, referring to Figure 3The preset paired earphone is a kind of tws Bluetooth earphone, wherein the earphone supports LE Audio (new generation Bluetooth audio technology), that is, stereo recording can be performed through left and right ear microphones, and the earphone further comprises a bone conduction microphone, that is, an associated bone conduction recording device, and the bone conduction microphone can more clearly identify the voice of the wearer than the call microphone. The earphone transmits all received sound signals to an associated mobile terminal, which is a kind of conference audio processing device, and a conference audio processing program is run on the mobile terminal, which can receive the sound signals transmitted by the earphone, convert the received sound signals into audio signals, that is, conference recordings, and determine the speaker corresponding to the audio signals according to the position information in the audio signals. When the wearer wears and starts the earphone, the left and right ear microphones of the earphone receive first and second sound signals respectively, and the bone conduction microphone receives a bone conduction sound signal. If the mobile terminal determines that the volume value of the received bone conduction sound signal is greater than a preset threshold, it can be determined that the speaker of the sound signal received by the earphone at this time is the wearer, that is, the wearer, the bone conduction sound signal is matched with the identity of the wearer, it is determined that the speaker of the bone conduction sound signal is the wearer, and the bone conduction sound signal is saved as an audio signal. If the mobile terminal determines that the volume value of the received bone conduction sound signal is less than or equal to the preset threshold, it indicates that the speaker at this time is not the wearer, then the gyro in the earphone is used to determine whether the wearer has a head rotation behavior. If the wearer has a head rotation, in order to avoid the deviation of the position information in the audio signal caused by the head rotation, the position information in the audio signal is compensated according to the position offset of the gyro. If the wearer does not have a head rotation, the speaker corresponding to the audio signal is directly determined according to the position information in the audio signal. Further, after the conference ends, all received audio signals are combined to generate a conference recording, wherein the conference recording can combine all speech contents of each speaker together, or display the identity information and speech content of the speaker in the conference recording audio with the same color, or highlight the speech audio segment of the speaker after clicking the identity information of the speaker in the conference recording. Further, the conference recording can be converted into a text form according to the requirement, and the conference minutes are generated by associating with the identity information of the speaker, so that the reader of the conference minutes can more clearly know the speech content of each speaker.
[0075] The embodiment of the present application provides a conference audio processing method, referring to Figure 4 In the first embodiment of the conference audio processing method, the conference audio processing method comprises:
[0076] Step S10, acquiring current position information in a plurality of groups of conference audio signals collected by different collection channels.
[0077] In some embodiments, it is necessary to note that the conference audio processing method is executed by a conference audio processing device, which is a terminal that can be connected to a network, such as a mobile phone, a notebook computer, a desktop computer, a smart bracelet, a Bluetooth headset, etc. The conference audio processing device is connected to the preset paired headset through, but not limited to, the Internet, a local area network, or a physical connection, thereby obtaining multiple groups of conference sound signals collected by different talk microphones of the preset paired headset, and generating an audio signal according to the sound signals, wherein the sound signals include a first sound signal, a second sound signal, and a bone conduction sound signal, and the sound signals are collected and transmitted by different microphones in the preset paired headset.
[0078] The preset paired headset refers to a sound pickup device that can convert the energy of sound into an audio (sound frequency) signal. For example, the preset paired headset is a headset that supports LE Audio, i.e., can perform stereo recording through left and right talk microphones (first talk microphone and second talk microphone), and also includes a bone conduction microphone. The bone conduction microphone refers to a microphone that collects sound signals caused by slight vibrations of the head and neck during speech and converts them into electrical signals. Compared with a talk microphone, the bone conduction microphone can more clearly identify the speaker's own voice.
[0079] For example, the conference audio processing device is a mobile phone terminal, and the preset paired headset is a Bluetooth headset that includes a first talk microphone, a second talk microphone, and a bone conduction microphone. When the conference starts, the speaker speaks, and the Bluetooth headset simultaneously collects sound signals in the conference through the three built-in microphones and sends them to the mobile phone terminal. Then, the mobile phone terminal generates an audio signal based on the received three sound signals and extracts the current orientation information in the audio signal to determine the corresponding speaker of the audio signal.
[0080] The speaker is a person object who speaks in the conference, i.e., the source of the audio signal. The audio signal can select the optimal sound signal as the audio signal according to the amplitude, frequency, tone, etc. of the received three sound signals, or can combine the received three sound signals after noise reduction to generate an audio signal. This embodiment does not limit this.
[0081] Step S20, match the current orientation information with a preset orientation identity mapping table.
[0082] In some embodiments, the conference audio processing device is a mobile terminal, which extracts the current orientation information from the received audio signal and matches the extracted current orientation information with a preset orientation-identity mapping table. The orientation-identity mapping table is a static table storing orientation information and identity information, each of which corresponds to the other, and a query value of one of the information can be mapped to a specified information. The current orientation information refers to the position information of the current speaker or the source (sound source) of the audio signal relative to the center point of the preset paired earphone, including angle information and distance information. The angle information refers to the angle of the current speaker relative to the center point of the preset paired earphone, and the distance information refers to the distance of the current speaker relative to the center point of the preset paired earphone.
[0083] For example, the mobile terminal extracts the current orientation information (+45°, 1m) from the audio signal, matches the (+45°, 1m) with the preset orientation-identity mapping table, determines whether the preset orientation information stores the (+45°, 1m) orientation information, and if the preset orientation-identity mapping table stores the (+45°, 1m), the matching is successful. If the preset orientation-identity mapping table does not store the (+45°, 1m), the matching is unsuccessful.
[0084] Step S21, if the matching is successful, the target identity corresponding to the current orientation information in the orientation-identity mapping table is taken as the speaker identity of the conference audio signal.
[0085] In some embodiments, the conference audio processing device is a mobile terminal, which matches the current orientation information obtained from the audio signal with a preset orientation-identity mapping table. If the matching is successful, it indicates that the speaker corresponding to the audio signal stores a relevant orientation-identity mapping relationship in the mobile terminal. Then, the mobile terminal takes the target identity corresponding to the current orientation information in the orientation-identity mapping table as the speaker identity of the conference audio signal, further determines the speaker corresponding to the audio signal according to the speaker identity, and generates a corresponding conference recording file.
[0086] The conference recording file refers to an audio file that can be played or processed in a terminal such as a computer. The conference recording file generated by the conference audio processing device can be in the form of a complete audio file of the current conference, in which the identity information of each speaker is recorded together with the speech content, or a segmented audio file of the current conference, which can be segmented according to the identity information of the speaker or the content of the conference. This embodiment does not limit this.
[0087] Step S22, if the matching is not successful, a new identity is established in the orientation-identity mapping table according to the current orientation information.
[0088] Step S23, the new identity is taken as the speaker identity of the conference audio signal.
[0089] In some embodiments, exemplarily, the conference audio processing device is a mobile phone terminal. The mobile phone terminal matches the current orientation information obtained from the audio signal with a preset orientation-identity mapping table. If the matching is not successful, it indicates that the speaker corresponding to the audio signal is a new speaker. Then the mobile phone terminal establishes a new identity in the orientation-identity mapping table according to the orientation information, and takes the new identity as the speaker identity of the conference audio signal. Meanwhile, the new identity and the current orientation information establish a new mapping relationship in the orientation-identity mapping table, so that when a new audio signal is received subsequently, the speaker corresponding to the new audio signal can be determined, and a corresponding conference recording file can be generated. Thus, the step of confirming the identity of the collected multiple sets of conference audio signals is simplified, and the operation cost is reduced.
[0090] The identity refers to various information capable of identifying the speaker identity alone or in combination with other information, including but not limited to the name, company, position, department, employee number, orientation information, etc. of the speaker. The user of the mobile phone terminal can manually input or modify the identity information, or the mobile phone terminal can automatically identify the content and generate the identity information.
[0091] In the embodiment, the conference audio processing method is applied to an earphone. The current orientation information in a plurality of groups of conference audio signals collected by different collection channels is acquired, and the current orientation information is matched with a preset orientation-identity mapping table. The speaker corresponding to the audio signal is determined according to the matching result of the current orientation information. If the matching is successful, the target identity corresponding to the current orientation information in the orientation-identity mapping table is taken as the speaker identity of the conference audio signal. If the matching is not successful, a new identity is established in the orientation-identity mapping table according to the current orientation information. In the conference process, the earphone is used to collect a plurality of groups of conference audio signals in multiple channels. The use of the earphone makes the collection of conference audio signals more convenient, reduces the equipment cost of conference recording, and the distance between the two ear microphones is constant when the earphone is worn, avoiding the multiple correction of parameters in the recording process, so that the acquisition of the current orientation information in the conference audio signal is more simple and fast. Then, the acquired conference audio is processed to obtain the current orientation information in a plurality of groups of conference audio signals, and then the current orientation information is used to realize the rapid and accurate identification of the speaker identity, avoiding the cumbersome process of identifying the speaker by using the voiceprint recognition method. The fixed distance between the recording function of the earphone and the earphone double-recording microphone further reduces the processing steps of the conference audio, reduces the technical cost, and further improves the efficiency of conference audio processing and the accuracy of identity identification.
[0092] In another embodiment, before the step S10 of acquiring the current orientation information in a plurality of groups of conference audio signals collected by different collection channels, the method further comprises:
[0093] Step A, determining whether the volume value of the bone conduction sound signal is greater than a preset threshold.
[0094] Step B, if the volume value is greater than the preset threshold, taking the preset main speaker identity as the speaker identity of the bone conduction sound signal.
[0095] In some embodiments, an exemplary, there will be a main speaker or host during the meeting, and the speech of the main speaker needs to be recorded clearly and in detail. At this time, the main speaker only needs to wear the bone conduction microphone of the pre-paired earphone. The bone conduction microphone can collect the bone conduction sound signal completely by using the slight vibration of the head and neck bone caused by the main speaker's speech. Then, the volume value of the bone conduction sound signal collected by the bone conduction microphone is compared with the preset threshold value. Although the bone conduction microphone will still collect the sound signal of other speakers when the wearer does not speak, since the working principle of the bone conduction microphone is to collect sound by using the slight vibration of the head and neck bone caused by the speaker's speech, only when the wearer speaks, the sound signal collected by the bone conduction microphone will be clear and have a larger volume value. Then, by judging whether the volume value of the bone conduction sound signal is greater than the preset threshold value, it can be determined whether the source of the sound signal is the wearer (the main speaker). If the volume value is greater than the preset threshold value, it indicates that the sound signal collected by the bone conduction microphone at this time is the sound signal emitted by the wearer (the main speaker), and the preset main speaker identity is taken as the speaker identity of the bone conduction sound signal. Compared with the communication microphone, the bone conduction sound signal collected by the bone conduction microphone in the present embodiment can record the voice of the wearer more clearly, and by judging the volume value, the corresponding speaker of the audio signal can be quickly determined, further improving the accuracy and speed of identifying the speaker identity during the conference recording process.
[0096] Further, in another embodiment of the conference audio processing method of the present application, with reference to Figure 5 , the step of acquiring the orientation information in the audio signal comprises:
[0097] Step S11, acquiring a first sound signal and a second sound signal of a current speaker based on a pre-paired earphone.
[0098] Step S12, determining a first arrival time of the first sound signal and a second arrival time of the second sound signal.
[0099] Step S13, determining the volume difference between the first sound signal and the second sound signal.
[0100] Step S14, determining the current orientation information of the current speaker relative to the pre-paired earphone according to the first arrival time, the second arrival time and the volume difference.
[0101] In some embodiments, the conference audio processing device collects a first sound signal and a second sound signal of a current speaker based on a first microphone and a second microphone of a preset paired headset, the first microphone collects the first sound signal, and the second microphone collects the second sound signal. Due to the difference (or no difference) between the positions of the speaker relative to the first microphone and the second microphone, the time and volume value of the same sound reaching the two microphones are different when the sound is transmitted in three-dimensional space, and then the position information of the sound relative to the preset paired headset is determined based on the time of arrival and the volume value of the sound and a preset position information library. The position information library contains a preset angle formula and a distance formula, which can determine the distance and angle of the sound based on the information in the sound signals obtained through different collection channels of the same sound, wherein the information in the sound signals includes the time of arrival, the volume value, etc.
[0102] Optionally, the step S14 of determining the current position information of the current speaker relative to the preset paired headset based on the first time of arrival, the second time of arrival, and the volume difference includes:
[0103] The step S141 of determining the angle information of the current speaker relative to the preset paired headset based on the volume difference and a preset angle formula.
[0104] In some embodiments, the position information includes angle information and distance information. Since the distance of the speaker (sound source) relative to the first microphone and the second microphone of the preset paired headset is different (or no difference), the volume values of the first sound signal and the second sound signal collected by different microphones are different, so the volume difference between the first sound signal and the second sound signal can be obtained by analyzing (for example, spectral analysis) the first sound signal and the second sound signal, and then the angle information can be obtained by combining a preset angle formula. For example, the preset angle formula is:
[0105] θ = inv_ILD
[0106] Wherein, θ is the angle information, and ILD is the volume difference between the first sound signal and the second sound signal.
[0107] The preset angle formula refers to a formula that can obtain the angle of the sound source relative to the center position of the microphone based on the volume values of the same sound transmitted to different microphones, and the present embodiment does not limit this.
[0108] The step S142 of determining the distance information of the current speaker relative to the preset paired headset based on the first time of arrival, the second time of arrival, and a preset distance formula.
[0109] In some embodiments, due to the difference in distance of the speaker (sound source) relative to the first and second talk-through microphones of the preset paired earphone (the difference can also not exist), the time of the first and second sound signals collected by different talk-through microphones is different, so that the first arrival time of the first sound signal and the second arrival time of the second sound signal can be obtained by analyzing (for example, spectral analysis) the first and second sound signals, and the distance information can be obtained in combination with a preset distance formula and angle information. Exemplarily, the preset distance formula is:
[0110] L = a x sin θ + t1 x c
[0111] wherein L is the distance information, θ is the angle information, t1 is the first arrival time, t2 is the second arrival time, and t1 is less than t2, a is the distance of the center point of the preset paired earphone relative to the first or second talk-through microphone, and c is the propagation speed of sound.
[0112] The preset distance formula refers to a formula capable of obtaining the distance of the sound source relative to the center position of the talk-through microphone through the arrival time of the same sound transmitted to different talk-through microphones, and the present embodiment does not limit this.
[0113] In an implementable manner, the angle information of the current speaker relative to the preset paired earphone can also be determined through the arrival time difference of the first and second sound signals. Exemplarily, the arrival time difference of the first and second sound signals is obtained by analyzing (for example, spectral analysis) the first and second sound signals, and the angle information is obtained in combination with a preset second angle formula. The preset second angle formula is:
[0114]
[0115] wherein ITD is the arrival time difference, a is the distance of the center point of the preset paired earphone relative to the first or second talk-through microphone, c is the propagation speed of sound, and θ is the angle information.
[0116] The angle information of the current speaker relative to the preset paired earphone is determined through the arrival time difference of the first and second sound signals, and the angle information obtained through the volume difference of the first and second sound signals is compared to verify the accuracy of the angle information result.
[0117] In the embodiment, the conference audio processing device collects the first sound signal and the second sound signal of the current speaker through the preset paired earphone respectively, and then determines the first arrival time of the first sound signal, the second arrival time of the second sound signal, and the volume difference between the first sound signal and the second sound signal, and then determines the current orientation information of the current speaker relative to the preset paired earphone, i.e. the angle information and the distance information, according to the first arrival time, the second arrival time and the volume difference. Through the binaural microphone arranged in the earphone, the time and the volume value of the same sound transmitted to the binaural microphone are obtained, and then the position source of the audio signal can be accurately identified according to the arrival time and the volume value, so that the corresponding speaker is determined based on the orientation information. The conference audio processing method collects the multi-channel multi-group conference audio signals by using the convenient and portable earphone in the conference. The use of the earphone makes the collection of the conference audio signals more convenient, reduces the equipment cost of the conference recording, and the distance between the binaural microphones is constant when the earphone is worn, i.e. the distance between the center point of the preset paired earphone and the first or second microphone is constant, which avoids the multiple corrections of the parameters in the recording process, makes the acquisition of the current orientation information in the conference audio signal more simple and fast, avoids the cumbersome identification process of using the voiceprint recognition method, reduces the cost of the conference audio processing device, realizes the accurate and fast identification of the speaker in the conference recording, and further improves the user experience.
[0118] Optionally, before the step of matching the current orientation information with the preset orientation identity mapping table in the step S20, the method further comprises:
[0119] In step S15, the compensation information of the multi-group conference audio signals is obtained, and the current orientation information in the multi-group conference audio signals is corrected according to the compensation information.
[0120] In some embodiments, after the conference audio processing device acquires the current orientation information in the audio signal, in order to avoid the misjudgment of the current orientation information of the identified sound signal caused by the rotation or movement of the head of the earphone wearer, it is necessary to further judge whether the orientation information needs to be corrected or calibrated. Then, the conference audio processing device obtains the compensation information of the audio signal. If the compensation information of the audio signal is obtained, it indicates that the preset paired earphone has deviated from the starting position when receiving the first sound signal and / or the second sound signal. In order to avoid the determination error of the orientation information caused by the deviation of the preset paired earphone, and to avoid the judgment error of the speaker, the conference audio processing device corrects the orientation information according to the obtained compensation information to obtain the corrected orientation information, and then determines the speaker corresponding to the audio signal according to the corrected orientation information.
[0121] For example, the conference audio processing device obtains the orientation information of the audio signal as 0°, i.e., the speaker is in front of the preset paired earphone, and then obtains compensation information of the audio signal. If the compensation information is 45° east, it indicates that the preset paired earphone has deviated 45° east compared with the initial position, and then the conference audio processing device corrects the orientation information 0° according to the compensation information 45° east to obtain the corrected orientation information 45° east, i.e., the actual speaker is in the direction 45° east of the preset paired earphone, and then performs the step S20 of matching with the preset orientation information according to the corrected orientation information 45° east.
[0122] In this embodiment, the conference audio processing device judges whether the preset paired earphone moves in position when collecting the sound signal after obtaining the orientation information of the audio signal, so as to avoid the judgment error of the orientation information of the audio signal caused by the position movement of the preset paired earphone, and then to obtain the accurate orientation information by obtaining the compensation information of the audio signal and correcting the orientation information of the audio signal according to the compensation information, so as to realize the accurate identification of the speaker corresponding to the audio signal and further improve the accuracy of the speaker identification of the audio signal.
[0123] Optionally, the step S15 of obtaining the compensation information of the plurality of groups of conference audio signals comprises:
[0124] In step S151, the real-time position and the initial position of the preset paired earphone are obtained when the first sound signal or the second sound signal is received.
[0125] In some embodiments, the conference audio processing device receives the first sound signal or the second sound signal, which indicates that the speaker is speaking at this time. In order to avoid the judgment error of the orientation information of the audio signal caused by the rotation or movement of the head of the earphone wearer, the conference audio processing device obtains the real-time position and the initial position of the preset paired earphone. The real-time position is the real-time position information of the angular motion detection device in the preset paired earphone, and the initial position is the position of the angular motion detection device in the preset paired earphone when it first works. The position is taken as the reference position of the conference audio, and the orientation information is determined. The conference audio processing device determines whether the preset paired earphone moves or deviates in angle by monitoring the real-time position of the angular motion detection device in the preset paired earphone in real time. The angular motion detection device refers to an angular motion detection device using the momentum moment sensitive shell of a high-speed rotating body relative to the inertial space to rotate around one or two axes perpendicular to the rotation axis, such as a gyroscope.
[0126] Optionally, before the step S151 of acquiring the starting position of the preset paired earphone, the method comprises the steps of:
[0127] Step C: when the preset paired earphone is started, the current position of the preset paired earphone is acquired as the starting position.
[0128] In some embodiments, the user can correct the position of the preset paired earphone according to the needs before recording the conference, and after determining the recording position, the preset paired earphone is started. At this time, the conference audio processing device acquires the current position of the preset paired earphone and stores it as the starting position, and then determines whether the preset paired earphone moves in the subsequent recording process, and then determines the correct azimuth information of the audio signal; the starting position can be re-set or a new current position can be acquired as a new starting position during the recording process according to the user's needs. By acquiring the current position of the preset paired earphone as the starting position when the preset paired earphone is started, the accurate judgment of whether the preset paired earphone moves is realized, and then the compensation information of the audio signal is determined, the azimuth information of the audio signal is compensated and corrected, and then the accurate azimuth information of the audio signal is determined.
[0129] Step S152: comparing the real-time position and the starting position to obtain the offset position of the preset paired earphone as the compensation information.
[0130] In some embodiments, the conference audio processing device compares the acquired real-time position and starting position to obtain the offset position of the preset paired earphone, and takes the offset position as compensation information for compensating and correcting the azimuth information of the audio signal.
[0131] For example, the preset paired earphone is an earphone, and the earphone has a gyroscope for determining whether the earphone moves; the main speaker wears the earphone to record the conference, and when recording the second speaker, the head of the main speaker rotates 30° to the east. At this time, the real-time position of the earphone is 30° to the east, and the starting position of the earphone acquired by the conference audio processing device is 0°. Then, the conference audio processing device compares the real-time position 30° to the east with the starting position 0° to obtain the offset position 30° to the east, which is the compensation information. The conference audio processing device corrects the azimuth information of the acquired audio signal according to the compensation information, and then obtains the corrected accurate azimuth information.
[0132] In the embodiment, the conference audio processing device realizes accurate judgment on whether the preset paired earphone moves by acquiring the real-time position and the starting position of the preset paired earphone when the first sound signal or the second sound signal is received, and then comparing the real-time position and the starting position, and then obtains the offset position of the preset paired earphone as the compensation information, and then corrects the orientation information of the audio signal according to the compensation information, realizes accurate judgment on the orientation information of the audio signal, and further improves the accuracy of audio signal speaker recognition.
[0133] It should be noted that in this document, the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article, or system. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or system that includes the element.
[0134] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of software products, which are stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and include a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the methods described in various embodiments of the present application.
[0135] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A conference audio processing method, characterized by, The conference audio processing method is applied to a headset, and the conference audio processing method comprises the following steps: Obtaining current orientation information in a plurality of groups of conference audio signals collected by different collection channels, wherein the current orientation information comprises angle information and distance information; Matching the current orientation information with a preset orientation-identity mapping table, wherein the orientation-identity mapping table is a static table in which orientation information and identity information are stored in one-to-one correspondence; If the matching is successful, taking a target identity corresponding to the current orientation information in the orientation-identity mapping table as a speaker identity of the conference audio signal; The plurality of groups of conference audio signals comprise a first sound signal and a second sound signal, and the step of obtaining the current orientation information in the plurality of groups of conference audio signals collected by different collection channels comprises the following steps: Collecting the first sound signal and the second sound signal of a current speaker based on a preset paired headset; Determining a first arrival time of the first sound signal and a second arrival time of the second sound signal; Determining a volume difference between the first sound signal and the second sound signal; According to the volume difference and a preset angle formula, determining angle information of the current speaker relative to the preset paired headset; According to the first arrival time, the second arrival time and a preset distance formula, determining distance information of the current speaker relative to the preset paired headset.
2. The conference audio processing method of claim 1, wherein, After the step of matching the current orientation information with the preset orientation-identity mapping table, the method further comprises the following steps: If the matching is not successful, establishing a new identity in the orientation-identity mapping table according to the current orientation information; Taking the new identity as the speaker identity of the conference audio signal.
3. The conference audio processing method of claim 1, wherein, Before the step of matching the current orientation information with the preset orientation-identity mapping table, the method comprises the following step: Obtaining compensation information of the plurality of groups of conference audio signals, and correcting the current orientation information in the plurality of groups of conference audio signals according to the compensation information.
4. The conference audio processing method of claim 3, wherein, The plurality of groups of conference audio signals comprise a first sound signal and a second sound signal, and the step of obtaining the compensation information of the plurality of groups of conference audio signals comprises the following steps: When the first sound signal or the second sound signal is received, obtaining a real-time position and a starting position of a preset paired headset; Comparing the real-time position and the starting position to obtain an offset position of the preset paired headset as the compensation information.
5. The conference audio processing method of claim 4, wherein, The step of obtaining the starting position of the preset paired headset comprises the following step: When the preset paired headset is started, obtaining a current position of the preset paired headset as the starting position.
6. The conference audio processing method of any one of claims 1 to 5, wherein, The plurality of groups of conference audio signals comprise a bone conduction sound signal, and before the step of obtaining the current orientation information in the plurality of groups of conference audio signals collected by different collection channels, the method further comprises the following steps: Determining whether a volume value of the bone conduction sound signal is greater than a preset threshold value; If the volume value is greater than the preset threshold value, taking a preset main speaker identity as a speaker identity of the bone conduction sound signal.
7. A conference audio processing device, characterized by The device comprises a memory, a processor, and a conference audio processing program stored on the memory and executable on the processor, the conference audio processing program being configured to implement the steps of the conference audio processing method according to any one of claims 1 to 6.
8. A storage medium, characterized by The storage medium has stored thereon a conference audio processing program, the conference audio processing program being executable by the processor to implement the steps of the conference audio processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice acquisition method and device for intelligent wearable equipment, and related components
CN109412544A
Conference summary transcription method and device and storage medium
CN112037791A
Spokesman role information processing method and device
CN114003192A