A conference data processing method and related device
By using a combined identification method involving conference terminals and information processing equipment, and leveraging facial and voiceprint feature databases, the problem of insufficient accuracy in identifying speakers in multi-location joint meetings has been solved, achieving higher accuracy and privacy protection.
Patent Information
- Application Number
- CN201980102782.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-31
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2039-12-31
AI Technical Summary
In multi-location joint meetings, how to accurately determine the correspondence between speeches and speakers during the meeting process, especially the problem of insufficient recognition accuracy of individual meeting terminals in traditional solutions.
By using a joint identification method involving conference terminals and conference information processing equipment, and utilizing facial feature databases and voiceprint feature databases, the speaker's identity is determined by analyzing the location of the sound source in an audio segment using both facial and voiceprint recognition methods. This process generates additional information and, combined with the sound source location information, prevents database fusion and protects privacy and security.
This improved the accuracy of speaker identification, protected the privacy and security of attendees, and prevented information leaks.
Smart Images

Figure CN114762039B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a conference data processing method and related equipment. Background Art
[0002] Whether it is a remote meeting or a local meeting, the speech information during the meeting often needs to be sorted. When sorting, each speech (which can be voice content, text content, or speech time period, etc.) needs to be matched with the corresponding speaker. Taking the speech as a speech text as an example, the following example illustrates a form of correspondence between the speech text and the speaker.
[0003] Zhang San: This month's sales are not ideal. Let's analyze the reasons for the problem together.
[0004] Li Si: Now is the traditional off-season, and at the same time, our competitors’ promotion efforts are greater than ours, resulting in our market share being taken away by our competitors.
[0005] Wang Wu: I think the competitiveness of the product has declined, and several market problems have arisen, resulting in customers not buying it.
[0006] Zhang San: I also interviewed several customers and several salespeople.
[0007] It can be seen that in the process of organizing speech information, it is necessary to determine the speakers corresponding to each part of the conference audio. Currently, there is a solution for a single conference terminal to identify the speakers of each part of the audio. However, with the development of communication technology, there are also situations where joint meetings are held in multiple locations. In this case, how to accurately determine the correspondence between the speeches and speakers during the meeting is a technical problem that technical personnel in this field are studying. Summary of the Invention
[0008] The embodiment of the present invention discloses a conference data processing method and related equipment, which can more accurately determine the correspondence between speeches and speakers during a conference.
[0009] In a first aspect, an embodiment of the present application provides a conference data processing method, which is applied to a conference system, wherein the conference system includes a conference terminal and a conference information processing device, and the method includes:
[0010] The conference terminal collects audio clips of the first conference venue according to the sound source orientation during the conference, wherein the sound source orientations of two audio clips that are adjacent in time sequence are different; in specific implementation, the conference terminal can continuously detect (for example, it can detect when there is sound, or it can detect according to a preset frequency) the sound source orientation during the conference, and when the sound source orientation changes, it starts collecting the next audio clip. For example, the audio orientation is orientation 1 from the 0th minute to the 6th minute of the conference, the audio orientation is 2 from the 6th minute to the 10th minute of the conference, and the audio orientation is 3 from the 10th minute to the 15th minute of the conference; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio clip, collect the audio from the 6th minute to the 10th minute as an audio clip, and collect the audio from the 10th minute to the 15th minute as an audio clip. It can be understood that multiple audio clips can be collected in this way.
[0011] The conference terminal generates first additional information corresponding to each of the multiple collected audio clips, wherein the first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio clip includes the start time and end time of the audio clip; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio clip identified by the identification information belongs to.
[0012] The conference terminal sends the conference audio recorded during the conference and the first additional information corresponding to the multiple audio segments to the conference information processing device. The conference audio is segmented by the conference information processing device into multiple audio segments (for example, segmented according to the direction of the sound source, so that the sound source directions of two audio segments adjacent in time sequence obtained by segmentation are different) and attached with corresponding second additional information, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment; the first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants and speeches in the first venue. It may be a correspondence between all participants who have spoken and their speeches, or it may be a correspondence between participants who have spoken more and their speeches, and of course it may be other situations.
[0013] Using the above method, the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0014] In conjunction with the first aspect, in a first possible implementation of the first aspect, the conference system further includes a facial feature library, the facial feature library includes facial features, and the conference terminal generates first additional information corresponding to each of the multiple collected audio clips, including:
[0015] performing facial recognition on a target image according to the facial feature library, wherein the target image is a facial image at the location of the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and the processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0016] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the facial recognition result and identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, where different face IDs identify different facial features; optionally, it may also include a facial identity identifier for identifying a specific person. It should be noted that facial recognition may be performed but no facial features are identified. In this case, the recognition result may be empty or filled in according to preset rules.
[0017] In this implementation, the facial image at the sound source location of the audio clip is identified based on the facial feature library to preliminarily obtain information for determining the speaker's identity, which can improve the accuracy of subsequent determination of the speaker's identity.
[0018] In combination with the first aspect, or any possible implementation of the first aspect, in a second possible implementation of the first aspect, the conference system further includes a first voiceprint feature library, the first voiceprint feature library includes voiceprint features, and the first additional information corresponding to each of the multiple collected audio clips generated by the conference terminal includes:
[0019] Determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; other audio segments in the multiple audio segments may be processed in a manner similar to the processing manner of the first audio segment;
[0020] Generate first additional information corresponding to the first audio clip, where the first additional information corresponding to the first audio clip includes the matching result of the voiceprint matching and the identification information of the first audio clip. The matching result here can be a voiceprint ID used to identify the voiceprint feature, and different voiceprint IDs identify different voiceprint features. It should be noted that there may be cases where voiceprint recognition (or matching) is performed but no voiceprint feature is identified. In this case, the matching result may be empty or filled in according to preset rules.
[0021] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0022] In combination with the first aspect, or any possible implementation of the first aspect, in a third possible implementation of the first aspect, the conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, the first voiceprint feature library includes voiceprint features, and the conference terminal generates first additional information corresponding to each of the multiple collected audio clips, including:
[0023] performing facial recognition on a target image according to the facial feature library, wherein the target image is an image at the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0024] Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library;
[0025] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the face recognition, the matching result of the voiceprint matching and the identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, and different face IDs may identify different facial features; optionally, it may also include a face identity identifier for marking a specific person; it should be noted that there may be cases where face recognition is performed but facial features are not identified. In this case, the recognition result may be empty or filled with content according to preset rules. The matching result here may be a voiceprint ID for identifying voiceprint features, and different voiceprint IDs may identify different voiceprint features; it should be noted that there may be cases where voiceprint recognition (or matching) is performed but voiceprint features are not identified. In this case, the matching result may be empty or filled with content according to preset rules.
[0026] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library, and the facial image of the sound source direction of the audio clip is identified based on the facial feature library, so as to preliminarily obtain information for determining the identity of the speaker. The combination of these two aspects of information can improve the accuracy of subsequent determination of the speaker's identity.
[0027] In combination with the first aspect, or any possible implementation of the first aspect, in a fourth possible implementation of the first aspect, the method further includes:
[0028] If the voiceprint features of the first audio clip cannot be matched in the first voiceprint feature library, the voiceprint features of the first audio clip are saved in the first voiceprint feature library. Optionally, when the number of audio clips at a certain sound source location accumulates to a certain level, a voiceprint feature that can be distinguished from the voiceprint features of other users (i.e., the voiceprint feature corresponding to the sound source location) can be determined based on the accumulated audio clips. In this way, the first voiceprint feature library can be continuously enriched and improved, making subsequent voiceprint recognition based on the first voiceprint feature library more accurate.
[0029] In combination with the first aspect, or any possible implementation of the first aspect, in a fifth possible implementation of the first aspect, the first audio segment is multi-channel audio, and determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library includes:
[0030] Performing sound source separation on the first audio segment to obtain multiple mono audio channels, each mono audio channel being audio of a speaker, and then determining a voiceprint feature of each mono audio channel;
[0031] The conference terminal matches the voiceprint features of the multiple mono audios respectively from the first voiceprint feature library.
[0032] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0033] In combination with the first aspect, or any possible implementation of the first aspect, a fifth possible implementation of the first aspect further includes:
[0034] The conference information processing device determines the speaker identity of the first audio segment based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment (for example, including the face ID) and the information for determining the speaker identity in the corresponding second additional information (for example, including the voiceprint ID); optionally, when the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment, further confirmation can be made in combination with an audio segment similar to the first audio segment, for example, the voiceprint ID and the face ID are identified for the second audio segment, but the voiceprint ID is identified for the first audio segment, but the face ID is not identified, and the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment. The voiceprint ID identified by the segment is the same as the voiceprint ID identified for the second audio segment, so it can be considered that the second audio segment and the first audio segment are similar audio segments, and therefore it is considered that the face ID corresponding to the first audio segment and the face ID corresponding to the second audio segment are the same, and therefore the face identity identifiers corresponding to the first audio segment and the second audio segment are also the same; optionally, when the conference information processing device also obtains the sound source orientation information, the role of the sound source orientation information can also refer to the role of the voiceprint ID here; wherein, the first audio segment is one of the multiple audio segments; the processing method of other audio segments in the multiple audio segments can refer to the processing method of the first audio segment;
[0035] The conference information processing device generates a conference record, wherein the conference record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
[0036] In this implementation, the speaker's identity is determined jointly by the conference terminal and the conference information processing device, which can make the determined speaker's identity more accurate; and, since both the conference terminal and the conference information processing device can perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each possess without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0037] In a second aspect, an embodiment of the present application provides a conference data processing method, which is applied to a conference system, wherein the conference system includes a conference terminal and a conference information processing device, and the method includes:
[0038] The conference terminal collects the conference audio of the first conference site during the conference;
[0039] The conference terminal performs voice segmentation on the conference audio according to the sound source orientation in the conference audio to obtain multiple audio segments, wherein the sound source orientations of two audio segments adjacent in time sequence are different; in specific implementation, the conference terminal can continuously detect (for example, detect when someone is speaking, or detect at a preset frequency) the sound source orientation during the conference, and start collecting the next audio segment when the sound source orientation changes. For example, the audio orientation is orientation 1 from the 0th minute to the 6th minute of the conference, the audio orientation is 2 from the 6th minute to the 10th minute of the conference, and the audio orientation is 3 from the 10th minute to the 15th minute of the conference; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio segment, the audio from the 6th minute to the 10th minute as an audio segment, and the audio from the 10th minute to the 15th minute as an audio segment. It can be understood that multiple audio segments can be collected in this way.
[0040] The conference terminal generates first additional information corresponding to each of the multiple audio clips, wherein the first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio clip includes the start time and end time of the audio clip; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio clip identified by the identification information belongs to.
[0041] The conference terminal sends the conference audio and first additional information corresponding to the multiple audio segments to the conference information processing device. The conference audio is segmented into multiple audio segments by the conference information processing device and accompanied by corresponding second additional information. The second additional information corresponding to each audio segment includes information for identifying the speaker of the audio segment and identification information of the corresponding audio segment. The conference information processing device uses the first additional information and the second additional information to generate a correspondence between the participants and speeches at the first conference site. This may be a correspondence between all participants who have spoken and their speeches, or a correspondence between participants who have spoken more and their speeches, or other situations are also possible.
[0042] Using the above method, the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0043] In conjunction with the second aspect, in a first possible implementation of the second aspect, the conference system further includes a facial feature library, the facial feature library includes facial features, and the conference terminal generates the first additional information corresponding to each of the multiple audio clips, including:
[0044] performing facial recognition on a target image according to the facial feature library, wherein the target image is a facial image at the location of the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and the processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0045] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the facial recognition result and identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, where different face IDs identify different facial features; optionally, it may also include a facial identity identifier for identifying a specific person. It should be noted that facial recognition may be performed but no facial features are identified. In this case, the recognition result may be empty or filled in according to preset rules.
[0046] In this implementation, the facial image at the sound source location of the audio clip is identified based on the facial feature library to preliminarily obtain information for determining the speaker's identity, which can improve the accuracy of subsequent determination of the speaker's identity.
[0047] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a second possible implementation of the second aspect, the conference system further includes a first voiceprint feature library, the first voiceprint feature library includes voiceprint features, and the conference terminal generates the first additional information corresponding to each of the multiple audio clips, including:
[0048] Determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; other audio segments in the multiple audio segments may be processed in a manner similar to the processing manner of the first audio segment;
[0049] Generate first additional information corresponding to the first audio clip, where the first additional information corresponding to the first audio clip includes the matching result of the voiceprint matching and the identification information of the first audio clip. The matching result here can be a voiceprint ID used to identify the voiceprint feature, and different voiceprint IDs identify different voiceprint features. It should be noted that there may be cases where voiceprint recognition (or matching) is performed but no voiceprint feature is identified. In this case, the matching result may be empty or filled in according to preset rules.
[0050] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0051] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a third possible implementation of the second aspect, the conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, the first voiceprint feature library includes voiceprint features, and the conference terminal generates the first additional information corresponding to each of the multiple audio clips, including:
[0052] performing facial recognition on a target image according to the facial feature library, wherein the target image is an image at the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0053] Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library;
[0054] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the face recognition, the matching result of the voiceprint matching and the identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, and different face IDs may identify different facial features; optionally, it may also include a face identity identifier for marking a specific person; it should be noted that there may be cases where face recognition is performed but facial features are not identified. In this case, the recognition result may be empty or filled with content according to preset rules. The matching result here may be a voiceprint ID for identifying voiceprint features, and different voiceprint IDs may identify different voiceprint features; it should be noted that there may be cases where voiceprint recognition (or matching) is performed but voiceprint features are not identified. In this case, the matching result may be empty or filled with content according to preset rules.
[0055] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library, and the facial image of the sound source direction of the audio clip is identified based on the facial feature library, so as to preliminarily obtain information for determining the identity of the speaker. The combination of these two aspects of information can improve the accuracy of subsequent determination of the speaker's identity.
[0056] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a fourth possible implementation of the second aspect, the method further includes:
[0057] If the voiceprint features of the first audio clip cannot be matched in the first voiceprint feature library, the voiceprint features of the first audio clip are saved in the first voiceprint feature library. Optionally, when the number of audio clips at a certain sound source location accumulates to a certain level, a voiceprint feature that can be distinguished from the voiceprint features of other users (i.e., the voiceprint feature corresponding to the sound source location) can be determined based on the accumulated audio clips. In this way, the first voiceprint feature library can be continuously enriched and improved, making subsequent voiceprint recognition based on the first voiceprint feature library more accurate.
[0058] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a fifth possible implementation of the second aspect, the first audio segment is multi-channel audio, and determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library includes:
[0059] Performing sound source separation on the first audio segment to obtain multiple mono audio channels, each mono audio channel being audio of a speaker, and then determining a voiceprint feature of each mono audio channel;
[0060] The conference terminal matches the voiceprint features of the multiple mono audios respectively from the first voiceprint feature library.
[0061] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0062] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, a sixth possible implementation of the second aspect further includes:
[0063] The conference information processing device determines the speaker identity of the first audio segment based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment (for example, including the face ID) and the information for determining the speaker identity in the corresponding second additional information (for example, including the voiceprint ID); optionally, when the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment, further confirmation can be made in combination with an audio segment similar to the first audio segment, for example, the voiceprint ID and the face ID are identified for the second audio segment, but the voiceprint ID is identified for the first audio segment, but the face ID is not identified, and the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment. The voiceprint ID identified by the segment is the same as the voiceprint ID identified for the second audio segment, so it can be considered that the second audio segment and the first audio segment are similar audio segments, and therefore it is considered that the face ID corresponding to the first audio segment and the face ID corresponding to the second audio segment are the same, and therefore the face identity identifiers corresponding to the first audio segment and the second audio segment are also the same; optionally, when the conference information processing device also obtains the sound source orientation information, the role of the sound source orientation information can also refer to the role of the voiceprint ID here; wherein, the first audio segment is one of the multiple audio segments; the processing method of other audio segments in the multiple audio segments can refer to the processing method of the first audio segment;
[0064] The conference information processing device generates a conference record, wherein the conference record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
[0065] In this implementation, the speaker's identity is determined jointly by the conference terminal and the conference information processing device, which can make the determined speaker's identity more accurate; and, since both the conference terminal and the conference information processing device can perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each possess without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0066] In a third aspect, an embodiment of the present application provides a conference information processing method, which is applied to a conference system, the conference system including a conference terminal and a conference information processing device, the method comprising:
[0067] The conference information processing device receives conference audio and first additional information corresponding to multiple audio clips sent by a conference terminal at a first conference site, wherein the conference audio is recorded by the conference terminal during the conference, and the multiple audio clips are obtained by voice segmentation of the conference audio or collected at the first conference site based on the sound source orientation. The first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip, wherein the sound source orientations of two audio clips that are adjacent in time sequence are different; for example, the audio orientation from the 0th minute to the 6th minute of the conference audio is orientation 1, the audio orientation from the 6th minute to the 10th minute of the conference audio is orientation 2, and the audio orientation from the 10th minute to the 15th minute of the conference audio is orientation 3; then, the conference information processing device will segment the audio from the 0th minute to the 6th minute into one audio clip, segment the audio from the 6th minute to the 10th minute into one audio clip, and segment the audio from the 10th minute to the 15th minute into one audio clip. It can be understood that multiple audio clips can be segmented in this way.
[0068] The conference information processing device performs voice segmentation on the conference audio to obtain multiple audio segments;
[0069] The conference information processing device performs voiceprint recognition on the multiple audio segments obtained by voice segmentation to obtain the second additional information corresponding to each of the multiple audio segments, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment. Optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio segment includes the start time and end time of the audio segment; optionally, the identification information can also be generated according to a preset rule. It should be noted that the audio segments segmented by the conference information processing device may be exactly the same as or not exactly the same as the audio segments segmented by the conference terminal. For example, the audio segments segmented by the conference terminal may be S1, S2, S3, and S4, while the audio segments segmented by the conference information processing device may be S1, S2, S3, S4, and S5. The embodiment of this application will focus on the same parts of the audio segments segmented by the two (i.e., the above-mentioned multiple audio segments) for description, and the processing of other audio segments is not limited here.
[0070] The conference information processing device generates a correspondence between the participants and speeches in the first venue based on the first additional information and the second additional information. The correspondence may be a correspondence between all participants who have spoken and their speeches, or a correspondence between participants who have spoken more and their speeches. Of course, other situations are also possible.
[0071] Using the above method, the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0072] In conjunction with the third aspect, in a first possible implementation of the third aspect, the conference information processing device generates a correspondence between each participant in the first venue and the speech based on the first additional information and the second additional information, including:
[0073] The conference information processing device determines the speaker identity of the first audio segment based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment (for example, including the face ID) and the information for determining the speaker identity in the corresponding second additional information (for example, including the voiceprint ID); optionally, when the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment, further confirmation can be made in combination with an audio segment similar to the first audio segment, for example, the voiceprint ID and the face ID are identified for the second audio segment, but the voiceprint ID is identified for the first audio segment, but the face ID is not identified, and the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment. The voiceprint ID identified by the segment is the same as the voiceprint ID identified for the second audio segment, so it can be considered that the second audio segment and the first audio segment are similar audio segments, and therefore it is considered that the face ID corresponding to the first audio segment and the face ID corresponding to the second audio segment are the same, and therefore the face identity identifiers corresponding to the first audio segment and the second audio segment are also the same; optionally, when the conference information processing device also obtains the sound source orientation information, the role of the sound source orientation information can also refer to the role of the voiceprint ID here; wherein, the first audio segment is one of the multiple audio segments; the processing method of other audio segments in the multiple audio segments can refer to the processing method of the first audio segment;
[0074] The conference information processing device generates a conference record, wherein the conference record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
[0075] In combination with the third aspect, or any of the foregoing possible implementations of the third aspect, in a second possible implementation of the third aspect, the conference system further includes a second voiceprint feature library, the second voiceprint feature library including voiceprint features, and the conference information processing device performs voiceprint recognition on multiple audio segments obtained by speech segmentation to obtain second additional information corresponding to each of the multiple audio segments, including:
[0076] The conference information processing device determines the voiceprint features of the first audio segment and matches the voiceprint features of the first audio segment from the second voiceprint feature library. The second additional information includes the matching result of the voiceprint matching and the identification information of the first audio segment. The first audio segment is one of the multiple audio segments.
[0077] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0078] In combination with the third aspect, or any of the foregoing possible implementations of the third aspect, in a third possible implementation of the third aspect, the first audio segment is a multi-source segment, and the conference information processing device determines the voiceprint feature of the first audio segment and matches the voiceprint feature of the first audio segment from the second voiceprint feature library, including:
[0079] The conference information processing device performs sound source separation on the first audio segment to obtain a plurality of mono audios;
[0080] The conference information processing device determines the voiceprint features of each mono audio, and matches the voiceprint features of the multiple mono audios from the second voiceprint feature library respectively.
[0081] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0082] In combination with the third aspect, or any one of the above-mentioned possible implementations of the third aspect, in a fourth possible implementation of the third aspect, the first additional information corresponding to the first audio clip includes the face recognition result and / or voiceprint recognition result of the first audio clip and the identification information of the first audio clip, and the first audio clip is one of the multiple audio clips.
[0083] In a fourth aspect, an embodiment of the present application provides a conference terminal, which is applied to a conference system. The conference terminal includes a processor, a memory, and a communication interface, wherein the memory is used to store a computer program, and the processor calls the computer program to perform the following operations:
[0084] During the meeting, audio clips of the first conference venue are collected based on the sound source orientation, wherein the sound source orientations of two audio clips that are adjacent in time sequence are different; in specific implementation, the conference terminal can continuously detect (for example, it can detect according to a preset frequency) the sound source orientation during the meeting, and when the sound source orientation changes, it starts collecting the next audio clip. For example, the audio orientation is orientation 1 from the 0th minute to the 6th minute of the meeting, the audio orientation is 2 from the 6th minute to the 10th minute of the meeting, and the audio orientation is 3 from the 10th minute to the 15th minute of the meeting; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio clip, collect the audio from the 6th minute to the 10th minute as an audio clip, and collect the audio from the 10th minute to the 15th minute as an audio clip. It can be understood that multiple audio clips can be collected in this way.
[0085] Generate first additional information corresponding to each of the multiple collected audio clips, wherein the first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio clip includes the start time and end time of the audio clip; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio clip identified by the identification information belongs to.
[0086] The conference audio recorded during the conference and the first additional information corresponding to the multiple audio segments are sent to the conference information processing device through the communication interface. The conference audio is divided by the conference information processing device into multiple audio segments (for example, divided according to the direction of the sound source, so that the sound source directions of two audio segments adjacent in time sequence obtained by division are different) and attached with corresponding second additional information, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment; the first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants and speeches in the first venue. It may be a correspondence between all participants who have spoken and their speeches, or it may be a correspondence between participants who have spoken more and their speeches, and of course it may be other situations.
[0087] It can be seen that the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0088] In conjunction with the fourth aspect, in a first possible implementation of the fourth aspect, the conference system further includes a facial feature library, the facial feature library including facial features, and in generating the first additional information corresponding to each of the multiple collected audio segments, the processor is specifically configured to:
[0089] performing facial recognition on a target image according to the facial feature library, wherein the target image is a facial image at the location of the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and the processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0090] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the facial recognition result and identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, where different face IDs identify different facial features; optionally, it may also include a facial identity identifier for identifying a specific person. It should be noted that facial recognition may be performed but no facial features are identified. In this case, the recognition result may be empty or filled in according to preset rules.
[0091] In this implementation, the facial image at the sound source location of the audio clip is identified based on the facial feature library to preliminarily obtain information for determining the speaker's identity, which can improve the accuracy of subsequent determination of the speaker's identity.
[0092] In combination with the fourth aspect, or any possible implementation of the fourth aspect, in a second possible implementation of the fourth aspect, the conferencing system further includes a first voiceprint feature library, the first voiceprint feature library including voiceprint features, and in terms of generating the first additional information corresponding to each of the multiple collected audio clips, the processor is specifically configured to:
[0093] Determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; other audio segments in the multiple audio segments may be processed in a manner similar to the processing manner of the first audio segment;
[0094] Generate first additional information corresponding to the first audio clip, where the first additional information corresponding to the first audio clip includes the matching result of the voiceprint matching and the identification information of the first audio clip. The matching result here can be a voiceprint ID used to identify the voiceprint feature, and different voiceprint IDs identify different voiceprint features. It should be noted that there may be cases where voiceprint recognition (or matching) is performed but no voiceprint feature is identified. In this case, the matching result may be empty or filled in according to preset rules.
[0095] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0096] In combination with the fourth aspect, or any possible implementation of the fourth aspect, in a third possible implementation of the fourth aspect, the conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, and the first voiceprint feature library includes voiceprint features. In terms of generating the first additional information corresponding to each of the multiple collected audio clips, the processor is specifically configured to:
[0097] performing facial recognition on a target image according to the facial feature library, wherein the target image is an image at the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0098] Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library;
[0099] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the face recognition, the matching result of the voiceprint matching and the identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, and different face IDs may identify different facial features; optionally, it may also include a face identity identifier for marking a specific person; it should be noted that there may be cases where face recognition is performed but facial features are not identified. In this case, the recognition result may be empty or filled with content according to preset rules. The matching result here may be a voiceprint ID for identifying voiceprint features, and different voiceprint IDs may identify different voiceprint features; it should be noted that there may be cases where voiceprint recognition (or matching) is performed but voiceprint features are not identified. In this case, the matching result may be empty or filled with content according to preset rules.
[0100] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library, and the facial image of the sound source direction of the audio clip is identified based on the facial feature library, so as to preliminarily obtain information for determining the identity of the speaker. The combination of these two aspects of information can improve the accuracy of subsequent determination of the speaker's identity.
[0101] In combination with the fourth aspect, or any possible implementation of the fourth aspect, in a fourth possible implementation of the fourth aspect, the processor is further configured to:
[0102] If the voiceprint features of the first audio clip cannot be matched in the first voiceprint feature library, the voiceprint features of the first audio clip are saved in the first voiceprint feature library. Optionally, when the number of audio clips at a certain sound source location accumulates to a certain level, a voiceprint feature that can be distinguished from the voiceprint features of other users (i.e., the voiceprint feature corresponding to the sound source location) can be determined based on the accumulated audio clips. In this way, the first voiceprint feature library can be continuously enriched and improved, making subsequent voiceprint recognition based on the first voiceprint feature library more accurate.
[0103] In combination with the fourth aspect, or any possible implementation of the fourth aspect, in a fifth possible implementation of the fourth aspect, the first audio segment is multi-channel audio, and in determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, the processor is specifically configured to:
[0104] Performing sound source separation on the first audio segment to obtain multiple mono audio channels, each mono audio channel being audio of a speaker, and then determining a voiceprint feature of each mono audio channel;
[0105] The voiceprint features of the multiple mono audios are matched respectively from the first voiceprint feature library.
[0106] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0107] In a fifth aspect, an embodiment of the present application provides a conference terminal, which is applied to a conference system. The conference terminal includes a processor, a memory, and a communication interface, wherein the memory is used to store a computer program, and the processor calls the computer program to perform the following operations:
[0108] Collect the conference audio of the first conference venue during the conference;
[0109] The conference audio is voice-segmented according to the sound source orientation in the conference audio to obtain multiple audio segments, wherein the sound source orientations of two audio segments adjacent in time sequence are different; in specific implementation, the conference terminal can continuously detect (for example, it can detect according to a preset frequency) the sound source orientation during the conference, and when the sound source orientation changes, it starts to collect the next audio segment. For example, the audio orientation is orientation 1 from the 0th minute to the 6th minute of the conference, the audio orientation is 2 from the 6th minute to the 10th minute of the conference, and the audio orientation is 3 from the 10th minute to the 15th minute of the conference; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio segment, the audio from the 6th minute to the 10th minute as an audio segment, and the audio from the 10th minute to the 15th minute as an audio segment. It can be understood that multiple audio segments can be collected in this way.
[0110] Generate first additional information corresponding to each of the multiple audio clips, wherein the first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio clip includes the start time and end time of the audio clip; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio clip identified by the identification information belongs to.
[0111] The conference audio and first additional information corresponding to the multiple audio segments are sent to a conference information processing device via the communication interface. The conference audio is segmented into multiple audio segments by the conference information processing device and accompanied by corresponding second additional information. The second additional information corresponding to each audio segment includes information for identifying the speaker of the audio segment and identification information of the corresponding audio segment. The first additional information and the second additional information are used by the conference information processing device to generate a correspondence between participants and speeches at the first conference site. This may be a correspondence between all participants who have spoken and their speeches, or a correspondence between participants who have spoken more and their speeches, or other situations are also possible.
[0112] It can be seen that the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0113] In conjunction with the fifth aspect, in a first possible implementation of the fifth aspect, the conference system further includes a facial feature library, the facial feature library including facial features, and in generating the first additional information corresponding to each of the multiple audio segments, the processor is specifically configured to:
[0114] performing facial recognition on a target image according to the facial feature library, wherein the target image is a facial image at the location of the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and the processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0115] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the facial recognition result and identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, where different face IDs identify different facial features; optionally, it may also include a facial identity identifier for identifying a specific person. It should be noted that facial recognition may be performed but no facial features are identified. In this case, the recognition result may be empty or filled in according to preset rules.
[0116] In this implementation, the facial image at the sound source location of the audio clip is identified based on the facial feature library to preliminarily obtain information for determining the speaker's identity, which can improve the accuracy of subsequent determination of the speaker's identity.
[0117] In combination with the fifth aspect, or any of the foregoing possible implementations of the fifth aspect, in a second possible implementation of the fifth aspect, the conference system further includes a first voiceprint feature library, the first voiceprint feature library including voiceprint features, and in generating the first additional information corresponding to each of the multiple audio clips, the processor is specifically configured to:
[0118] Determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; other audio segments in the multiple audio segments may be processed in a manner similar to the processing manner of the first audio segment;
[0119] Generate first additional information corresponding to the first audio clip, where the first additional information corresponding to the first audio clip includes the matching result of the voiceprint matching and the identification information of the first audio clip. The matching result here can be a voiceprint ID used to identify the voiceprint feature, and different voiceprint IDs identify different voiceprint features. It should be noted that there may be cases where voiceprint recognition (or matching) is performed but no voiceprint feature is identified. In this case, the matching result may be empty or filled in according to preset rules.
[0120] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0121] In combination with the fifth aspect, or any of the foregoing possible implementations of the fifth aspect, in a third possible implementation of the fifth aspect, the conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, and the first voiceprint feature library includes voiceprint features. In terms of generating the first additional information corresponding to each of the multiple audio clips, the processor is specifically configured to:
[0122] performing facial recognition on a target image according to the facial feature library, wherein the target image is an image at the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0123] Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library;
[0124] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the face recognition, the matching result of the voiceprint matching and the identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, and different face IDs may identify different facial features; optionally, it may also include a face identity identifier for marking a specific person; it should be noted that there may be cases where face recognition is performed but facial features are not identified. In this case, the recognition result may be empty or filled with content according to preset rules. The matching result here may be a voiceprint ID for identifying voiceprint features, and different voiceprint IDs may identify different voiceprint features; it should be noted that there may be cases where voiceprint recognition (or matching) is performed but voiceprint features are not identified. In this case, the matching result may be empty or filled with content according to preset rules.
[0125] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library, and the facial image of the sound source direction of the audio clip is identified based on the facial feature library, so as to preliminarily obtain information for determining the identity of the speaker. The combination of these two aspects of information can improve the accuracy of subsequent determination of the speaker's identity.
[0126] In combination with the fifth aspect, or any of the foregoing possible implementations of the fifth aspect, in a fourth possible implementation of the fifth aspect, the processor is further configured to:
[0127] If the voiceprint features of the first audio clip cannot be matched in the first voiceprint feature library, the voiceprint features of the first audio clip are saved in the first voiceprint feature library. Optionally, when the number of audio clips at a certain sound source location accumulates to a certain level, a voiceprint feature that can be distinguished from the voiceprint features of other users (i.e., the voiceprint feature corresponding to the sound source location) can be determined based on the accumulated audio clips. In this way, the first voiceprint feature library can be continuously enriched and improved, making subsequent voiceprint recognition based on the first voiceprint feature library more accurate.
[0128] In combination with the fifth aspect, or any of the foregoing possible implementations of the fifth aspect, in a fifth possible implementation of the fifth aspect, the first audio segment is multi-channel audio, and in determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, the processor is specifically configured to:
[0129] Performing sound source separation on the first audio segment to obtain multiple mono audio channels, each mono audio channel being audio of a speaker, and then determining a voiceprint feature of each mono audio channel;
[0130] The voiceprint features of the multiple mono audios are matched respectively from the first voiceprint feature library.
[0131] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0132] In a sixth aspect, an embodiment of the present application provides a conference information processing device, the device comprising a processor, a memory, and a communication interface, wherein the memory is used to store a computer program, and the processor calls the computer program to perform the following operations:
[0133] The communication interface receives conference audio and first additional information corresponding to multiple audio clips sent by a conference terminal at the first conference site, wherein the conference audio is recorded by the conference terminal during the conference, and the multiple audio clips are obtained by voice segmentation of the conference audio or collected at the first conference site based on the sound source orientation. The first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip, wherein the sound source orientations of two audio clips that are adjacent in time sequence are different; for example, the audio orientation from the 0th minute to the 6th minute of the conference audio is orientation 1, the audio orientation from the 6th minute to the 10th minute of the conference audio is orientation 2, and the audio orientation from the 10th minute to the 15th minute of the conference audio is orientation 3; then, the conference information processing device will segment the audio from the 0th minute to the 6th minute into one audio clip, the audio from the 6th minute to the 10th minute into one audio clip, and the audio from the 10th minute to the 15th minute into one audio clip. It can be understood that multiple audio clips can be segmented in this way.
[0134] Performing voice segmentation on the conference audio to obtain multiple audio segments;
[0135] Voiceprint recognition is performed on the multiple audio segments obtained by speech segmentation to obtain the second additional information corresponding to each of the multiple audio segments, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment. Optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio segment includes the start time and end time of the audio segment; optionally, the identification information can also be generated according to a preset rule. It should be noted that the audio segments segmented by the conference information processing device may be exactly the same as or not exactly the same as the audio segments segmented by the conference terminal. For example, the audio segments segmented by the conference terminal may be S1, S2, S3, and S4, while the audio segments segmented by the conference information processing device may be S1, S2, S3, S4, and S5. The embodiment of this application will focus on the same parts of the audio segments segmented by the two (i.e., the multiple audio segments mentioned above), and the processing of other audio segments is not limited here.
[0136] The correspondence between the participants and speeches in the first venue is generated based on the first additional information and the second additional information. It may be the correspondence between all participants who have spoken and their speeches, or the correspondence between participants who have spoken more and their speeches, or of course other situations.
[0137] It can be seen that the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0138] In conjunction with the sixth aspect, in a first possible implementation of the sixth aspect, in generating a correspondence between each participant in the first venue and the speech based on the first additional information and the second additional information, the processor is specifically configured to:
[0139] The speaker identity of the first audio segment is determined based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment (for example, including face ID) and the information for determining the speaker identity in the corresponding second additional information (for example, including voiceprint ID); optionally, when the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment, further confirmation can be made in combination with audio segments similar to the first audio segment, for example, the voiceprint ID and face ID are identified for the second audio segment, but the voiceprint ID is identified for the first audio segment, but the face ID is not identified, and the face ID is identified for the first audio segment. The voiceprint ID is the same as the voiceprint ID identified for the second audio segment, so the second audio segment and the first audio segment can be considered to be similar audio segments, so the face ID corresponding to the first audio segment and the face ID corresponding to the second audio segment are considered to be the same, and therefore the face identity identifiers corresponding to the first audio segment and the second audio segment are also the same; optionally, when the conference information processing device also obtains the sound source orientation information, the role of the sound source orientation information can also refer to the role of the voiceprint ID here; wherein, the first audio segment is one of the multiple audio segments; the processing method of other audio segments in the multiple audio segments can refer to the processing method of the first audio segment;
[0140] A meeting record is generated, wherein the meeting record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
[0141] In combination with the sixth aspect, or any of the foregoing possible implementations of the sixth aspect, in a second possible implementation of the sixth aspect, the conferencing system further includes a second voiceprint feature library, the second voiceprint feature library including voiceprint features, and in performing voiceprint recognition on multiple audio segments obtained by speech segmentation to obtain second additional information corresponding to each of the multiple audio segments, the processor is specifically configured to:
[0142] Determine the voiceprint features of the first audio segment, and match the voiceprint features of the first audio segment from the second voiceprint feature library, wherein the second additional information includes a matching result of the voiceprint matching and identification information of the first audio segment, and the first audio segment is one of the multiple audio segments.
[0143] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0144] In combination with the sixth aspect, or any of the foregoing possible implementations of the sixth aspect, in a third possible implementation of the sixth aspect, the first audio segment is a multi-source segment. In determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the second voiceprint feature library, the processor is specifically configured to:
[0145] Performing sound source separation on the first audio clip to obtain a plurality of mono audios;
[0146] Determine the voiceprint feature of each mono audio, and match the voiceprint features of the multiple mono audios from the second voiceprint feature library.
[0147] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0148] In combination with the sixth aspect, or any one of the above-mentioned possible implementations of the sixth aspect, in the fourth possible implementation of the sixth aspect, the first additional information corresponding to the first audio clip includes the face recognition result and / or voiceprint recognition result of the first audio clip and the identification information of the first audio clip, and the first audio clip is one of the multiple audio clips.
[0149] In combination with any of the above aspects, or any possible implementation scheme of any of the above aspects, in a possible implementation manner, the speech includes at least one of a speech text and a speech time period.
[0150] In combination with any of the above aspects, or any possible implementation scheme of any of the above aspects, in a possible implementation manner, the first venue is any one of multiple venues of the conference.
[0151] In a seventh aspect, an embodiment of the present application provides a terminal comprising a functional unit for implementing the method described in the first aspect, or any possible implementation of the first aspect, or the second aspect, or any possible implementation of the second aspect.
[0152] In an eighth aspect, an embodiment of the present application provides a conference information processing device, which includes a functional unit for implementing the method described in the third aspect, or any possible implementation manner of the third aspect.
[0153] In a ninth aspect, an embodiment of the present application provides a conference system, including a conference terminal and a conference information processing device:
[0154] The conference terminal is the conference terminal described in the fourth aspect, or any possible implementation of the fourth aspect, the fifth aspect, or any possible implementation of the fifth aspect, or the seventh aspect;
[0155] The conference information processing device is the conference information processing device described in the sixth aspect, or any possible implementation manner of the sixth aspect.
[0156] In the tenth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, it implements the method described in the first aspect, or any possible implementation of the first aspect, or the second aspect, or any possible implementation of the second aspect, or the third aspect, or any possible implementation of the third aspect.
[0157] In the eleventh aspect, an embodiment of the present application provides a computer program product, which, when running on a processor, implements the method described in the first aspect, or any possible implementation of the first aspect, or the second aspect, or any possible implementation of the second aspect, or the third aspect, or any possible implementation of the third aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0158] Figure 1 This is a schematic diagram of the architecture of a conference system provided by an embodiment of the present invention;
[0159] Figure 2 This is a flow chart of a conference information processing method provided by an embodiment of the present invention;
[0160] Figure 3 This is a flow chart of a conference information processing method provided by an embodiment of the present invention;
[0161] Figure 4 This is a flow chart of a conference information processing method provided by an embodiment of the present invention;
[0162] Figure 5A This is a schematic diagram of a voiceprint registration process provided by an embodiment of the present invention;
[0163] Figure 5B This is a schematic diagram of a display interface provided by an embodiment of the present invention;
[0164] Figure 6 This is a schematic diagram of a processing flow of a conference information processing device provided by an embodiment of the present invention;
[0165] Figure 7 This is a schematic diagram of a process for processing conference audio and first additional information provided by an embodiment of the present invention;
[0166] Figure 8 This is a flow chart of performing voiceprint recognition and speaker identification according to an embodiment of the present invention;
[0167] Figure 9 This is a flow chart of voiceprint recognition provided by an embodiment of the present invention;
[0168] Figure 10 This is a schematic diagram of an ideal conference architecture provided by an embodiment of the present invention;
[0169] Figure 11 This is a structural diagram of a conference information processing device provided by an embodiment of the present invention;
[0170] Figure 12 This is a structural diagram of a conference information processing device provided by an embodiment of the present invention;
[0171] Figure 13 This is a structural diagram of a conference information processing device provided by an embodiment of the present invention;
[0172] Figure 14 This is a structural diagram of a conference terminal provided by an embodiment of the present invention;
[0173] Figure 15 This is a structural diagram of a conference terminal provided by an embodiment of the present invention;
[0174] Figure 16 It is a structural diagram of a conference information processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0175] The embodiments of the present invention are described below with reference to the accompanying drawings.
[0176] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of a conference system provided by an embodiment of the present invention. The conference system includes a conference terminal 101, a multipoint control unit 102, and a conference information processing device 103; see Figure 2 , Figure 2 This is a flow chart of a conference information processing method provided by an embodiment of the present application. The method comprises Figure 1 The conference terminal 101, the multipoint control unit 102, and the conference information processing device 103 shown in the figure work together to complete the following process:
[0177] This conference terminal 101 belongs to the equipment with computing power, and its specific form is not limited here. The number of this conference terminal 101 can be one or more. When there are multiple conference terminals 101, these multiple conference terminals 101 can be deployed in different venues respectively. For example, a certain company has business departments in Beijing, Shenzhen and Wuhan. Then the business departments in Beijing, Shenzhen and Wuhan can be deployed with conference terminals 101 respectively. When the company is to hold a business meeting, these three business departments are equivalent to three venues. The conference terminals 101 of these three venues can be started to collect the speeches of the participants in the three venues during the meeting. Optionally, this conference terminal 101 can be a terminal with corresponding acquisition devices (such as cameras, microphones, etc.) installed, so information collection can be performed. Optionally, conference terminal 101 can also be performed by means of external acquisition devices. For example, a director camera is configured in the venue where conference terminal 101 is located. The director camera includes an array microphone for sound source location and recording the voice of the speaker. It also includes one or more cameras for video data in the venue, and video data for face recognition is obtained.
[0178] Conference terminal 101 ( Figure 2The conference terminal 101 collects audio clips or segments the recorded conference audio (such as segmentation based on sound source orientation information) to obtain audio clips, and then extracts the identity information corresponding to the audio clips. The conference audio collected locally and the extracted related information (which may be called additional information) are then packaged and sent to the multipoint control unit 102 (for example, it may be a microcontroller unit (MCU), and the subsequent description will focus on the MCU as an example). Optionally, the conference terminal 101 may utilize sound source positioning, face recognition, voiceprint recognition, voice switching detection and other technologies when extracting relevant information. The extracted relevant information may be first stored locally in the conference terminal 101 and then sent to the multipoint control unit 102 after the meeting.
[0179] The multipoint control unit 102 receives conference audio and additional information from conference terminals 101 at each venue, adds a corresponding venue ID and timestamp to the data from each venue, and then sends it to the conference information processing device 103. Optionally, the multipoint control unit 102 can also generate a multi-screen conference scene for multiple venues. Furthermore, the conference information processing device 103 can be a single device with data processing capabilities, or a cluster of multiple devices, such as a server cluster.
[0180] Conference information processing device 103 receives conference audio and associated supplementary information from multipoint control unit 102 and stores it by venue, ensuring that data from different venues is not mixed. After the meeting, it performs speech segmentation based on the conference audio and supplementary information, and matches speeches in the conference audio with corresponding speakers. Speeches can be text, audio, or duration, among other things.
[0181] It should be noted that the multipoint control unit 102 can be deployed independently of the conference information processing device 103 or within the conference information processing device 103. When deployed within the conference information processing device 103, certain transmission and reception operations mentioned above between the conference information processing device 103 and the multipoint control unit 102 may be implemented through internal signals or instructions within the device. In the subsequent description, it will be mentioned that the conference terminal 101 sends information to the conference information processing device 103. When the multipoint control unit 102 is deployed independently of the conference information processing device 103, the information is first sent by the conference terminal 101 to the multipoint control unit 102, which then forwards it to the conference information processing device (this process may involve some processing of the information). Of course, the information may also require other nodes to participate in the transfer during the transmission process. However, when the multipoint control unit 102 is deployed within the conference information processing device 103, the information can be sent directly from the conference terminal 101 to the conference information processing device 103, or it can be sent to the conference information processing device 103 after being transferred by other nodes.
[0182] See Figure 3 , Figure 3 This is a flow chart of a conference information processing method provided by an embodiment of the present invention. This method can be regarded as a method for processing conference information. Figure 2 The method includes but is not limited to the following steps:
[0183] Step S301: During a conference, the conference terminal collects audio clips from the first conference site according to the direction of the sound source.
[0184] Specifically, the meeting may be a meeting jointly held at multiple venues or a meeting at a single venue. When it is a meeting jointly held at multiple venues, the first venue here is one of the multiple venues, and the information processing method of other venues can refer to the description of the first venue; when it is a meeting at a single venue, the first venue here is the single venue.
[0185] In an embodiment of the present application, the sound source orientations of two audio clips that are adjacent in time sequence are different. In specific implementation, the conference terminal can continuously detect (for example, it can detect according to a preset frequency) the sound source orientation during the conference. When the sound source orientation changes, it starts to collect the next audio clip. For example, the audio orientation is orientation 1 from the 0th minute to the 6th minute of the conference, the audio orientation is 2 from the 6th minute to the 10th minute of the conference, and the audio orientation is 3 from the 10th minute to the 15th minute of the conference; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio clip, collect the audio from the 6th minute to the 10th minute as an audio clip, and collect the audio from the 10th minute to the 15th minute as an audio clip. It can be understood that multiple audio clips can be collected in this way.
[0186] It should be noted that it is also possible not to collect audio clips in real time, but after recording the conference audio, the conference audio can be segmented according to the sound source in the conference audio to obtain multiple audio clips, and the sound source directions of two audio clips adjacent in time sequence in the multiple audio clips are different.
[0187] It should be noted that facial images are also collected during the process of collecting audio clips and recording conference audio. For example, during a meeting, whenever someone speaks, a facial image is captured in the direction of the sound source.
[0188] Step S302: the conference terminal generates first additional information corresponding to each of the multiple collected audio segments.
[0189] Specifically, the first additional information corresponding to each audio segment can be generated uniformly after multiple audio segments are collected (or segmented), or can be generated each time an audio segment is collected (or simultaneously with the collection of the audio segment). Other generation timings can also be configured. The first additional information corresponding to each audio segment includes information for identifying the speaker of the audio segment and identification information of the corresponding audio segment.
[0190] In an embodiment of the present application, different audio clips correspond to different identification information. Optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio clip includes the start time and end time of the audio clip; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio clip identified by the identification information belongs to.
[0191] In an embodiment of the present application, the information included in the first additional information and used to determine the identity of the speaker of the audio clip may be a face recognition result, or a recognition result of other biometric features.
[0192] In the first method, the conference system also includes a facial feature library, which includes facial features. The conference terminal generates first additional information corresponding to each of the multiple audio clips collected (or segmented), including: performing facial recognition on a target image based on the facial feature library, wherein the target image is a facial image at the sound source location of the first audio clip taken during the recording of the first audio clip, and the first audio clip is one of the multiple audio clips. That is, during the collection of each audio clip, a facial image at the sound source location is collected by a director camera (or camera, or camera module). Thereafter, first additional information corresponding to the first audio clip is generated, wherein the first additional information corresponding to the first audio clip includes the recognition result of the facial recognition and identification information (such as a timestamp) of the first audio clip. The recognition result of the facial recognition here includes at least one of a facial identifier (ID) and a facial identity identifier corresponding to the facial identifier, wherein the facial identifier is used to identify facial features, and different facial features result in different facial identifiers. It is understandable that when a certain face identifier (ID) is different from other face identifiers (ID), it indicates that the face identity identifier corresponding to the certain face identifier (ID) is different from the face identity identifiers corresponding to the other face identifiers, but it is still unclear what the specific face identity identifier corresponding to the certain face identifier is; optionally, the face identity identifier can be a name such as Zhang San or Li Si. Of course, it is also possible that no face identifier is recognized based on the face image captured during the collection of the first audio clip. In this case, the recognition result can be empty or filled with other information.
[0193] It is understood that for the sake of convenience, the first audio segment is used as an example for explanation. The method for generating the first additional information corresponding to other audio segments can be the same as the method for generating the first additional information corresponding to the first audio segment. The following table 1 illustrates an example of a representation of the first additional information:
[0194] Table 1
[0195]
[0196] In a second embodiment, the conference system further includes a first voiceprint feature library containing voiceprint features. The conference terminal generates first additional information corresponding to each of the multiple collected audio clips, including: determining a voiceprint feature of a first audio clip; and matching the voiceprint feature of the first audio clip from the first voiceprint feature library based on the determined voiceprint feature of the first audio clip, where the first audio clip is one of the multiple audio clips; and generating first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes a matching result of the voiceprint matching and identification information (e.g., a timestamp) of the first audio clip. The matching result of the voiceprint matching includes a voiceprint ID. Of course, it is possible that no voiceprint ID is matched for the first audio clip. In this case, the recognition result may be empty or filled with other information. The voiceprint ID is used to identify the voiceprint feature. It is understood that different voiceprint features generally correspond to different voiceprint identities. Accordingly, different voiceprint IDs correspond to different voiceprint identities. It is understood that for the sake of convenience, the first audio segment is used as an example for description. The method for generating the first additional information corresponding to other audio segments can be the same as the method for generating the first additional information corresponding to the first audio segment. The following table 2 illustrates an example of a representation of the first additional information:
[0197] Table 2
[0198]
[0199]
[0200] In a third embodiment, the conferencing system further includes a facial feature library and a first voiceprint feature library, the facial feature library including facial features, and the first voiceprint feature library including voiceprint features. The conferencing terminal generates first additional information corresponding to each of the multiple collected audio clips, including: performing facial recognition on a target image based on the facial feature library, wherein the target image is an image taken at the location of the sound source of the first audio clip during the recording of the first audio clip, the first audio clip being one of the multiple audio clips; determining voiceprint features of the first audio clip, and matching the voiceprint features of the first audio clip with the first voiceprint feature library; and generating first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the facial recognition, the matching result of the voiceprint matching, and the identification information of the first audio clip. The principles of the third embodiment can be referred to above in the description of the first and second embodiments. Therefore, the first additional information in the third embodiment includes the information in Table 1 and Table 2 above, and is not separately illustrated here. For example, the first additional information can be the entire set of information in Table 1 and Table 2 above.
[0201] Based on the above description, the following Figure 4 This section describes the detailed operations of the conference terminal, including sound source localization, speech segmentation, voiceprint recognition, and face recognition, with examples.
[0202] In the first part, when a speaker (i.e., the speaker) speaks, the direction of the sound source is determined (i.e., the speaker's direction is determined by sound source localization), and then based on the determined sound source direction (i.e., the result of sound source localization), face recognition technology is used to identify the face information (such as face ID) at the sound source direction from the face feature library, thereby obtaining the corresponding face identity.
[0203] In the second part, the determined sound source orientation is also used as the basis for speech segmentation (which can also be described as speech segmentation) (it can also be combined with other detection methods to detect speaker switching for speech segmentation). If the audio segment currently segmented belongs to a new sound source (that is, the sound source orientation of this audio segment has no registered voiceprint information, and the process of identifying whether the sound source orientation of this audio segment has registered voiceprint information is the process of matching the voiceprint features of the first audio segment from the first voiceprint feature library. If the sound source orientation has registered voiceprint information, the voiceprint features can be matched, and the matched voiceprint features are the voiceprint information registered in the first voiceprint feature library), then the voiceprint needs to be registered, so the audio segment is saved in the first voiceprint feature library for subsequent training to obtain the voiceprint features corresponding to the sound source orientation of the audio segment. If voiceprint information has been registered for the current sound source, the voiceprint of the currently segmented audio clip is identified from the first voiceprint feature library (local real-time voiceprint library). The identified voiceprint is used in combination with the face recognition results to comprehensively determine the speaker's identity. For example, when two face IDs are recognized, the voiceprint ID identified can be further used to select one of the identities corresponding to the two face IDs. Because the first voiceprint feature library here is trained and updated by adding audio clips during the meeting, the recognition accuracy is higher.
[0204] Figure 5A The figure illustrates a process for registering voiceprints. If the sound source location to which the currently segmented audio segment belongs has no registered voiceprint information, the currently segmented audio segment is saved, and it is determined whether the sound source location has accumulated a preset number of audio segments. If the preset number of audio segments has been accumulated, the voiceprint features of the sound source location are extracted based on the accumulated audio segments, and the extracted voiceprint features of the sound source location are registered in the first voiceprint feature library (the newly registered voiceprint features correspond to a new voiceprint ID); if the preset number of audio segments has not been accumulated, the accumulation continues until the preset number of audio segments is accumulated and the voiceprint extraction is performed.
[0205] Optionally, the first voiceprint feature library includes voiceprint features pre-trained based on existing audio files (which may not be audio recorded in this meeting). If the confidence level of a certain voiceprint feature in the first voiceprint feature library for the current segmented audio segment exceeds a certain threshold, the voiceprint of the current segmented audio segment can be considered to be an existing voiceprint, i.e., it has been registered in the first voiceprint feature library. If the confidence level of a certain voiceprint feature in the first voiceprint feature library for the current audio segment does not exist for a certain voiceprint feature, and the deviation between the sound source orientation of the current segmented audio segment and the sound source orientation for which the voiceprint feature has been registered is not less than a preset reference value, the current segmented audio segment is considered to be a new sound source, and therefore a voiceprint needs to be registered for the sound source orientation of the current segmented audio segment.
[0206] In the third part, while a single person usually speaks in a conference, multiple people may speak simultaneously or interrupt. In this case, the mixing of multiple voices can lead to inaccurate speech recognition results (e.g., recognition of corresponding text). Blind source separation technology can be used to separate the speech and perform speech recognition on each separated mono audio channel. For example, a conference terminal sends the conference audio for blind source separation to a conference information processing device (which may be forwarded via an MCU). The conference information processing device performs blind source separation and performs speech recognition on each separated mono audio channel. When performing blind source separation, the conference information processing device first needs to determine whether multiple speakers are speaking simultaneously in the audio clip to be processed. The conference information processing device can process the audio clip to determine whether multiple speakers are speaking simultaneously, or the conference terminal can process the audio clip to determine whether multiple speakers are speaking simultaneously. If multiple speakers are speaking simultaneously, the conference terminal records the identification of the multiple speakers and obtains the number of simultaneous speakers. It then records mono data or mixed mono data for at least the number of speakers and sends it to the conference information processing device. Optionally, whether there are multiple people speaking and the number of speakers can be determined based on the number of sound sources obtained by sound source localization. For example, if there are three sound sources in an audio clip, it can be determined that there are multiple people speaking, and the number of speakers can be determined to be no less than 3.
[0207] In the fourth part, the conference terminal combines the relevant results of operations such as sound source localization, voice segmentation, voiceprint recognition, and face recognition to obtain first additional information corresponding to each audio segment, so as to subsequently send it to the conference information processing device. The first additional information may include, but is not limited to, one or more of the following information:
[0208] The voiceprint ID of an audio clip is used to identify the voiceprint characteristics of the audio clip.
[0209] The voiceprint identity of the audio clip is used to identify the identity of the person represented by the voiceprint feature.
[0210] The face ID of the audio clip is used to identify the facial features of the audio clip.
[0211] The face identifier of an audio clip is used to identify the person represented by the facial features. The identification information of the audio clip is used to identify the audio clip. Different audio clips have different audio clip identification information. For example, this identification information can be the timestamp of the audio clip, including the start and end time of the audio clip.
[0212] The venue ID corresponding to the audio clip is used to identify the venue where the audio clip was recorded, that is, to indicate which venue the corresponding audio clip comes from. Optionally, the venue ID can also be added to the first additional information after the MCU identifies the venue.
[0213] The sound source direction information of the audio clip is used to identify the sound source direction of the audio clip.
[0214] The reference identifier of the audio segment is used to indicate whether the corresponding audio segment is a speech of multiple people. For example, if it is a speech of multiple people, the reference identifier may be 1, and if it is a speech of a single person, the reference identifier may be 0.
[0215] The number of speakers in the audio clip.
[0216] The number of mono audio channels contained in the audio segment is used for more accurate speech-to-text conversion. An optional representation of the first additional information mentioned above can be shown in Table 3.
[0217] Table 3
[0218]
[0219]
[0220] It should be noted that some of the information listed above may not be included in the first additional information, such as the sound source direction of each audio clip, the reference identifier of each audio clip, the number of speakers in each audio clip, the number of mono audio contained in each audio clip, etc. This information can be obtained by the conference information processing device itself later.
[0221] Optionally, the information that needs to be sent to the conference information processing device may include a first voiceprint feature library that has been established in real time locally at the conference terminal to improve the voiceprint recognition accuracy in the conference information processing device. For example, if the conference audio received by the conference information processing device does not fully record the complete audio data of each venue, the conference information processing device may not be able to determine that all speakers have enough audio clips to create the local voiceprint feature library of the conference information processing device when performing voiceprint recognition, and therefore cannot perform accurate voiceprint recognition. When the conference information processing device receives the above-mentioned first voiceprint feature library, it has more voiceprint feature libraries. The subsequent conference information processing device can directly use the first voiceprint feature library, or it can update it based on the first voiceprint feature library to obtain a voiceprint feature library with richer voiceprint features, such as the second voiceprint feature library. The conference information processing device can perform voiceprint recognition more accurately based on the first voiceprint feature library or the second voiceprint feature library.
[0222] Step S303: the conference terminal sends the conference audio recorded during the conference and first additional information corresponding to the multiple audio segments to the conference information processing device.
[0223] Specifically, the conference audio is a longer audio segment, for example, it can be the entire audio of the first conference venue recorded by the conference terminal during the conference, which is equivalent to a longer audio segment spliced together from multiple audio segments, or longer than a longer audio segment spliced together from multiple audio segments.
[0224] Optionally, the conference terminal can first send the conference audio and first additional information corresponding to each audio clip to the MCU, for example, via the Transmission Control Protocol (TCP). The MCU then forwards this information to the conference information processing device. Optionally, the MCU can also mark the conference venue from which the conference audio originated and the corresponding timestamp and send it to the conference information processing device.
[0225] Step S304: the conference information processing device receives the conference audio and first additional information corresponding to the plurality of audio segments sent by the conference terminal of the first conference site.
[0226] Specifically, the conference information processing device can distinguish the received multiple first additional information according to the identification information in the first additional information corresponding to the multiple audio segments, that is, know which first additional information corresponds to which audio segment.
[0227] In addition, when there are multiple venues, the conference information processing device can store data from each venue according to venue classification, and then process each venue separately. The subsequent processing will be explained by taking the first venue as an example.
[0228] Step S305: The conference information processing device performs voice segmentation on the conference audio to obtain multiple audio segments.
[0229] Specifically, the conference information processing device can perform voice segmentation on the conference audio according to the sound source orientation in the audio segment. The principle is similar to that of the conference terminal collecting audio segments according to the sound source orientation or the conference terminal segmenting the conference audio according to the sound source orientation. Both are based on the sound source orientation, so that the sound source orientations of the two audio segments that are adjacent in time sequence are different. Since the conference information processing device and the conference terminal both obtain audio segments based on the sound source orientation, the conference information processing device can also segment the multiple audio segments that can be obtained on the conference terminal side. In addition, the conference information processing device may also segment other audio segments; that is, in addition to segmenting the above-mentioned multiple audio segments, the conference information processing device may also segment other audio segments. Because different devices may have different processing capabilities and accuracy when processing the same conference audio, the audio segments finally segmented are not exactly the same. The following explanation will focus on these multiple audio segments.
[0230] Optionally, if the first additional information sent by the conference terminal includes the timestamp of each audio segment segmented by the conference terminal, the conference terminal may also segment the audio segments according to the timestamp. For example, if the timestamp information is "Start 00:000 End 00:200", the conference terminal may segment the audio segment from 00:000 to 00:200 in the conference audio into one audio segment.
[0231] Step S306: the conference information processing device performs voiceprint recognition on the multiple audio segments obtained by voice segmentation to obtain second additional information corresponding to each of the multiple audio segments.
[0232] Specifically, the second additional information corresponding to each audio segment includes information for determining the speaker identity of the audio segment and identification information of the corresponding audio segment. For example, the conference system also includes a second voiceprint feature library containing voiceprint features. The conference information processing device performs voiceprint recognition on multiple audio segments obtained through speech segmentation to obtain the second additional information corresponding to each of the multiple audio segments. This includes: the conference information processing device determines the voiceprint features of a first audio segment and matches the voiceprint features of the first audio segment with those of the second voiceprint feature library. The second additional information includes a matching result of the voiceprint matching and identification information (e.g., a timestamp) of the first audio segment, where the first audio segment is one of the multiple audio segments. Optionally, the second voiceprint feature library can be a library obtained by improving the first voiceprint feature library. The matching result of the voiceprint matching here includes a voiceprint identification (ID). Of course, it is also possible that no voiceprint identification (ID) is matched for the first audio segment. In this case, the recognition result can be empty or filled with other information. The voiceprint ID is used to identify the voiceprint characteristics. It is understood that different voiceprint characteristics usually correspond to different identities of people. Correspondingly, different voiceprint IDs correspond to different identities of people. It is understood that for the convenience of description, the first audio clip is used as an example here. The method of generating the second additional information corresponding to other audio clips can be the same as the method of generating the second additional information corresponding to the first audio clip. The following table 4 illustrates an example of a representation of the second additional information:
[0233] Table 4
[0234]
[0235]
[0236] Optionally, when the first audio clip is a multi-source clip, the conference information processing device determines the voiceprint features of the first audio clip and matches the voiceprint features of the first audio clip from the second voiceprint feature library. Specifically, the conference information processing device performs sound source separation on the first audio clip to obtain multiple mono audio channels; the conference information processing device determines the voiceprint features of each mono audio channel and matches the voiceprint features of the multiple mono audio channels from the second voiceprint feature library to determine the speaker identity corresponding to each mono audio channel, thereby obtaining multiple speaker identities in the first audio clip. In other words, when multiple people speak in a certain audio clip, voiceprint features are matched for each of the multiple speakers in the audio clip, rather than just matching a single voiceprint feature.
[0237] Optionally, the conference terminal can send information such as whether multiple people are speaking, the number of speakers, and the number of audio channels (as shown in Table 4) to the conference information processing device, and the conference information processing device determines how many mono audio channels the first audio segment includes based on this information. Of course, the conference information processing device can also obtain information such as whether multiple people are speaking, the number of speakers, and the number of audio channels, and then determine how many mono audio channels the first audio segment includes based on this information.
[0238] Step S307: the conference information processing device generates a correspondence between the participants in the first conference site and the speeches according to the first additional information and the second additional information.
[0239] First, the conference information processing device determines the identity of the speaker of the first audio clip based on the information for determining the identity of the speaker in the first additional information corresponding to the first audio clip and the information for determining the identity of the speaker in the corresponding second additional information. Specifically, the identity of the speaker can be identified based on the face identity identifier and the voiceprint identity identifier included in the collection of the first additional information and the second additional information. Alternatively, the identity of the speaker can be identified based on the voiceprint ID and the voiceprint identity identifier included in the collection of the first additional information and the second additional information. Since the cost of determining the identity corresponding to the voiceprint is relatively high, the voiceprint identity identifier corresponding to the voiceprint may not be registered in advance in a specific implementation. In this case, the face ID, face identity identifier and voiceprint ID included in the collection of the first additional information and the second additional information can be used to identify the identity of the speaker.
[0240] To improve recognition accuracy or the likelihood of successful identification, face ID or voiceprint ID can also be introduced in addition to facial and voiceprint identification to identify the speaker. For example, when the voiceprint features of two voiceprints are similar, a single person's voice may match both voiceprint features. In this case, face ID can be used to further identify the speaker's precise identity. Similarly, if facial recognition obtains the facial features of two people, but only one person is actually speaking, voiceprint ID is needed to further identify the speaker's precise identity. It is understood that in this case, the combination of the first and second additional information should include the face ID or voiceprint ID.
[0241] Here are some implementation examples:
[0242] In Case 1, when the first additional information and the second additional information are combined as shown in Table 3 (it is understood that the second additional information can enhance and improve the first additional information, making the content of Table 3 more complete), if the first audio segment is audio segment S1, two faces are recognized for audio segment S1, and the corresponding facial identities include Zhang San and Liu Liu, and the direction of the sound source is Dir_1. However, for audio segment S3, the recognized facial identity is Zhang San, and the direction of the sound source is also Dir_1. Therefore, it can be determined that the speaker identity corresponding to audio segment S1 is Zhang San. Similarly, if the first audio segment is audio segment S4, the direction of the sound source is Dir_4, the voiceprint information ID corresponding to audio segment S4 is VP_4, and there is no facial information. However, for audio segment S6, the recognized voiceprint ID is VP_4, and the corresponding facial identity is Wang Wu. Therefore, it can be confirmed that the speaker identity corresponding to audio segment S4 is Wang Wu.
[0243] In Case 2, the first additional information and the second additional information are combined to obtain the voiceprint recognition result and face recognition result of each of the multiple audio clips. When further determining the speaker's identity based on the voiceprint recognition result and face recognition result, there are several different situations:
[0244] In case a, when the voiceprint recognition result contains a voiceprint ID (i.e., the recognition result is not empty), the voiceprint ID corresponds to the speaker identity. When the face recognition result contains a face ID (i.e., the face recognition result is not empty), the face ID also corresponds to the face identity. If the first additional information is as shown in Table 1, and the first audio segment is audio segment S1, it can be seen that the first additional information of the first audio segment (i.e., S1) in Table 1 contains face IDs, including F_ID_1 and F_ID_3. Among them, the face identity corresponding to F_ID_1 is Zhang San, and the face identity corresponding to F_ID_3 is Liu Liu. Therefore, the speaker identity of the first audio segment cannot be uniquely determined. In this case, the second additional information can be combined to further determine the speaker identity. If the second additional information is as shown in Table 5, it can be seen that the second additional information of the first audio segment (S1) in Table 5 contains voiceprint IDs, including VP_1. Among them, the voiceprint identity corresponding to VP_1 is Zhang San. Therefore, the speaker identity of the first audio segment (i.e., audio segment S1) can be finally determined to be Zhang San.
[0245] Table 5
[0246]
[0247] In case b, when there is a voiceprint ID in the voiceprint recognition result (i.e. the recognition result is not empty), the voiceprint ID does not specifically correspond to a speaker identity, but the voiceprint it represents can be clearly distinguished from the voiceprints represented by other voiceprint IDs; when there is a face ID in the face recognition result (i.e. the face recognition result is not empty), the face ID also corresponds to a speaker identity. In this case, if the first additional information is as shown in Table 1, the second additional information is as shown in Table 4, the second audio segment is audio segment S6, and the first audio segment is audio segment S4; it can be seen that in Table 1, the first additional information corresponding to the second audio segment (i.e., S6) contains a face ID, including F_ID_5, where the face identity identifier corresponding to F_ID_5 is Wang Wu; in Table 1, the first additional information corresponding to the first audio segment (i.e., S4) does not contain a face ID (i.e., the face recognition result is empty); in this case, it is necessary to combine the information in Table 4 to further confirm the identity of the speaker of the first audio segment. In Table 4, the second additional information corresponding to the first audio segment contains a voiceprint ID, including VP_4; the second additional information corresponding to the second audio segment also contains a voiceprint ID, including VP_4; therefore, based on the information in Table 4, it can be determined that the speaker identity of the first audio segment is the same as the speaker identity of the second audio segment, and through the information in Table 1, it can be determined that the face identity identifier corresponding to the second audio segment is Wang Wu, so the speaker identity of the first audio segment can be determined to be Wang Wu.
[0248] In Case 3, the first additional information is shown in Table 2, and the second additional information is shown in Table 4. That is, the first additional information contains a voiceprint recognition result, and the second additional information also contains a voiceprint recognition result. Each voiceprint ID corresponds to a speaker identity. Therefore, if the voiceprint recognition result of one of the first and second additional information is empty, the speaker identity can be determined based on the voiceprint ID of the other. When both the first and second additional information have voiceprint IDs, the speaker identity can be determined based on the combined voiceprint IDs of both, and the result is more accurate.
[0249] In case 4, the first additional information is shown in Table 3, and the second additional information is shown in Table 4. The voiceprint recognition results in Table 3 and the voiceprint recognition results in Table 4 can be regarded as mutually complementary to obtain a more accurate voiceprint recognition result. The principle of determining the identity of the speaker by combining the more accurate voiceprint recognition result and the face recognition result in Table 3 can refer to the method of determining the identity of the speaker based on the voiceprint recognition results and the face recognition results in the previous case 1, which will not be repeated here.
[0250] Furthermore, the above-mentioned sound source direction can be determined by the conference terminal and sent to the conference information processing device; of course, both the conference terminal and the conference information processing device can obtain it, and then the conference information processing device can integrate the sound source directions obtained by both for use in confirming the identity of the speaker.
[0251] After generating the speaker identity of each audio segment in the multiple audio segments, the conference information processing device generates a meeting record, wherein the meeting record includes the speech of the first audio segment and the speaker identity corresponding to the first audio segment. The first audio segment is one of the multiple audio segments, and the processing method of the other audio segments in the multiple audio segments can refer to the description of the first audio segment here. Therefore, the meeting record finally generated includes the speech of each audio segment in the multiple audio segments and the corresponding speaker identity. Optionally, the speech here includes one or more of the speech text, speech time period and speech voice, which are respectively illustrated below with examples.
[0252] For example, if the speech is a speech text, then the above method can be used to know the text content of each person's speech during the meeting. The following is an example:
[0253] Zhang San: This month's sales are not ideal. Let's analyze the reasons for the problem together.
[0254] Li Si: Now is the traditional off-season, and at the same time, our competitors’ promotion efforts are greater than ours, resulting in our market share being taken away by our competitors.
[0255] Wang Wu: I think the competitiveness of the product has declined, and several market problems have arisen, resulting in customers not buying it.
[0256] Zhang San: I also interviewed several customers and several salespeople.
[0257] This method is equivalent to marking the speaker's identity for each audio clip. This method includes a speech recognition step, which is to transcribe the audio into text.
[0258] For example, if the speech is a speaking time period, then the above method can be used to know the timing of each person speaking during the meeting. The following example is provided (the time format is seconds: milliseconds):
[0259] Zhang San: Speak between 00:00 and 10:00
[0260] Li Si: Speaks between 10:00 and 10:220
[0261] Wang Wu: Speaks between 10:220 and 13:110
[0262] Zhang San: Speaks between 13:110 and 16:010
[0263] In this way, the speaking time of each audio clip is marked with a speaker ID.
[0264] Optionally, the timing of the speech or the duration of the speech can be further used to infer the importance of the speaker (such as whether he or she is a leader).
[0265] For another example, if the speech is an audio speech, then the above method can be used to obtain the audio clips of each person speaking during the meeting. The following example is provided (the time format is seconds: milliseconds):
[0266] Zhang San: Audio clip 00:000--10:000
[0267] Li Si: Audio clip 10:000--10:220
[0268] Wang Wu: Audio clip 10:220--13:110
[0269] Zhang San: Audio clip 13:110--16:010
[0270] It can be understood that if the first audio clip is the one mentioned above that first performs sound source separation and then identifies the voiceprint features of each mono audio separately, then in the process of generating the meeting minutes, the speech of each mono audio in the first audio clip and the speaker identity corresponding to each mono audio will be specifically generated.
[0271] It should be noted that the conference in the embodiment of the present application can be jointly held in multiple venues. In this case, the above-mentioned first venue is one of the multiple venues. The conference information processing device processes the conference audio of each venue in the same way as the conference audio of the first venue. It can be processed one venue at a time in a round-robin manner. After all venues have been processed, the correspondence between the participants and speeches at each venue can be obtained, and the order of speeches can also be reflected. Optionally, a complete meeting record can be generated for each venue to reflect the correspondence between speeches and corresponding roles (such as speakers). Optionally, the conference information processing device can also display the obtained correspondence between participants and speeches (such as through a display screen), or send it to other devices for display. The display interface can be as follows. Figure 5B shown.
[0272] exist Figure 3In the described method, the conference terminal obtains conference audio and audio clips during the conference, identifies information for determining the speaker identity of each audio clip based on each audio clip, and then sends the conference audio and the identified information for determining the speaker identity of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and identifies information for determining the speaker identity of each audio clip, and then determines the speaker corresponding to each clip using the information it has identified for determining the speaker identity of each audio clip and the received information for determining the speaker identity of each audio clip. By jointly determining the speaker identity through the conference terminal and the conference information processing device, the determined speaker identity can be made more accurate; and because both the conference terminal and the conference information processing device perform identification, the conference terminal and the conference information processing device can perform identification based on their respective information bases without the need to merge the information bases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0273] According to the above introduction, after receiving the conference audio and the first additional information from the conference terminal, the conference information processing device involves a series of processing operations, such as speech segmentation, voiceprint recognition, speech recognition, sound source separation, etc. The operations of each link are illustrated below with examples in conjunction with the accompanying drawings to better understand the embodiment of the present application.
[0274] See Figure 6 , is a schematic diagram of a processing flow of a conference information processing device provided in an embodiment of the present application, including:
[0275] 601: The conference information processing device receives conference audio and first additional information of each venue, and stores them by category according to the venue. For example, the first venue is stored in one file, the second venue is stored in another file, and so on. Of course, classification can also be performed in other ways.
[0276] Optionally, the conference information processing device may record the voice data of each venue in chronological order into the same file. When segmenting, the audio of each venue is first separated independently, and then the voice of the venue is segmented and identified in combination with the first additional information of the corresponding venue, and the results are matched to the corresponding voice segments in the recorded voice file.
[0277] 602: After the conference ends, the conference information processing device extracts the conference audio of the first conference site and the first additional information from the first conference site from the received data.
[0278] 603: The conference information processing device processes the conference audio from the first venue and the first additional information from the first venue, including performing speech segmentation, voiceprint recognition, and converting each audio segment into corresponding text. This process includes generating the aforementioned second additional information. Subsequent processing includes identifying the speaker of each audio segment from the first venue based on the first and second additional information, and identifying the text content of each audio segment.
[0279] 604: The conference information processing device tags the text corresponding to each audio clip in the first conference site with the speaker's identity, that is, marks which text segment is spoken by which person, which is the aforementioned generation of the correspondence between the participants in the first conference site and the speech.
[0280] 605: The conference information processing device determines whether all conference sites have been processed in the same manner as the first conference site. If not, the process returns to steps 602 to 604 to process the remaining conference sites. If so, the process proceeds to step S606.
[0281] 606: The conference information processing device summarizes the correspondence between the participants and speeches in each conference venue to form a general meeting record.
[0282] See Figure 7 , is a flowchart of processing conference audio and first additional information provided by an embodiment of the present application, specifically a further extension of the previous step 603, including the following steps:
[0283] 701: The conference information processing device performs voice segmentation on the conference audio of the first conference site to obtain multiple audio segments, and performs voiceprint recognition on each audio segment to obtain second additional information.
[0284] 702: The conference information processing device determines whether multiple people are speaking for each audio clip. The conference terminal may identify whether multiple people are speaking and then inform the conference information processing device.
[0285] 703: If a certain audio segment is spoken by a single person, determine a speaker identifier corresponding to the audio segment according to the second additional information corresponding to the audio segment and the first additional information corresponding to the audio segment.
[0286] 704: If a certain audio segment is spoken by multiple people, perform sound source separation on the audio segment to obtain multiple mono audios, and then determine the speaker identifier corresponding to each mono audio based on the second additional information corresponding to each mono audio and the first additional information corresponding to each mono audio.
[0287] 705: The conference information processing device performs speech recognition on each audio segment to obtain text corresponding to each audio segment.
[0288] 706: The conference information processing device marks the text corresponding to each audio segment (or mono audio for multiple speakers) with the speaker identification corresponding to each audio segment (or mono audio for multiple speakers).
[0289] See Figure 8 , is a flow chart of a method for performing voiceprint recognition and determining a speaker, provided in an embodiment of the present application, including the following steps:
[0290] 801: The conference information processing device obtains the sound source direction of each segmented audio segment.
[0291] 802: The conference information processing device clusters the segmented audio segments according to the sound source orientation, and determines whether a voiceprint is registered for each sound source orientation. If a sound source orientation does not have a registered voiceprint, a voiceprint is registered for the sound source orientation when the number of audio segments at the sound source orientation meets a certain amount. The registered voiceprint is stored in the second voiceprint feature library for subsequent matching.
[0292] 803: The conference information processing device performs voiceprint recognition on each audio segment based on the second voiceprint feature library to obtain a voiceprint ID corresponding to each audio segment (i.e., obtains the second additional information); and performs speech recognition on each audio segment to obtain a text corresponding to each audio segment.
[0293] 804: The conference information processing device searches for the face ID and face identity corresponding to each audio segment from the received multiple pieces of first additional information (such as Table 1, Table 2, or Table 3) based on the voiceprint ID corresponding to each audio segment.
[0294] 805: The conference information processing device performs statistics on each audio clip to obtain the facial identification corresponding to each audio clip. For a certain audio clip, the facial identification that is searched the most times is used as the speaker identification corresponding to the audio clip.
[0295] 806: For the case where the corresponding voiceprint ID is not identified for some audio clips, the conference information processing device can search for the facial identity corresponding to the same sound source direction from the multiple received first additional information based on the sound source direction of the audio clip. The facial identity searched the most times can be used as the facial identity corresponding to the audio clip.
[0296] See Figure 9 , is a flow chart of a voiceprint recognition process provided by an embodiment of the present application, comprising the following steps:
[0297] 901: The conference information processing device obtains the sound source direction in the conference audio.
[0298] 902: The conference information processing device performs voice segmentation on the conference audio according to the sound source orientation in the conference audio to obtain multiple audio segments.
[0299] 903: The conference information processing device performs voiceprint recognition on each segmented audio segment based on the second voiceprint feature library (the second voiceprint feature library is obtained based on the first voiceprint feature library), obtains a voiceprint ID corresponding to the corresponding audio segment, and updates the voiceprint ID corresponding to the audio segment. Of course, if an audio segment is a new sound source, the voiceprint is extracted and registered by combining some audio segments in the same direction. This registered voiceprint ID can be used to identify whether the audio segment and other audio segments are from the same person.
[0300] In an optional solution of the present application, each conference terminal (such as conference terminal 1 and conference terminal 2) sends the facial recognition results to the MCU, and accordingly, the MCU forwards this information to the conference information processing device. In this way, the conference information processing device does not need to perform facial recognition itself when determining the identity of the speaker of the audio clip, which will improve privacy security. Figure 10 The diagram shows an ideal conference architecture. In some cases, the contents of facial feature library 1, facial feature library 2 and facial feature library 3 are the same, and conference terminal 1, conference terminal 2 and conference information processing equipment can all perform face recognition based on the corresponding facial feature library respectively; but in some cases, conference terminal 1 is company A, conference terminal is company B, and conference information processing equipment is service provider C. At this time, people from company A and company B register faces in their own facial feature library 1 and facial feature library 2 respectively, and these faces are not registered in facial feature library 3. Therefore, the identity of the speaker can only be identified within company A and company B, and the conference information processing equipment cannot identify the identity of the speaker. An optional solution of the present application proposes that the conference terminal perform face recognition, and the conference information processing equipment does not need to perform face recognition, thereby avoiding the problem that the conference information processing equipment cannot accurately obtain the identity of the speaker in the above special scenario.
[0301] The above describes in detail the method according to the embodiment of the present invention. The following provides an apparatus according to the embodiment of the present invention.
[0302] See Figure 11 , Figure 111 is a structural diagram of a conference information processing device 110 provided in an embodiment of the present invention. The conference information processing device 110 may be the conference terminal described above or a device in the conference terminal. The conference information processing device 110 is applied to a conference system and may include a collection unit 1101, a generation unit 1102, and a sending unit 1103. Each unit is described in detail as follows.
[0303] The collection unit 1101 is used to collect audio clips of the first conference venue according to the sound source orientation during the conference, wherein the sound source orientations of two audio clips that are adjacent in time sequence are different; in specific implementation, the conference terminal can continuously detect (for example, it can detect according to a preset frequency) the sound source orientation during the conference, and when the sound source orientation changes, it starts to collect the next audio clip. For example, the audio orientation from the 0th minute to the 6th minute of the conference is orientation 1, the audio orientation from the 6th minute to the 10th minute of the conference is audio orientation 2, and the audio orientation from the 10th minute to the 15th minute of the conference is audio orientation 3; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio clip, collect the audio from the 6th minute to the 10th minute as an audio clip, and collect the audio from the 10th minute to the 15th minute as an audio clip. It can be understood that multiple audio clips can be collected in this way.
[0304] The generation unit 1102 is used to generate first additional information corresponding to each of the multiple collected audio clips, wherein the first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio clip includes the start time and end time of the audio clip; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio clip identified by the identification information belongs to.
[0305] The sending unit 1103 is used to send the conference audio recorded during the conference and the first additional information corresponding to the multiple audio segments to the conference information processing device. The conference audio is divided by the conference information processing device into multiple audio segments (for example, divided according to the direction of the sound source, so that the sound source directions of two audio segments adjacent in time sequence obtained by the division are different) and are accompanied by corresponding second additional information, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment; the first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants and speeches in the first conference venue. It may be a correspondence between all participants who have spoken and their speeches, or it may be a correspondence between participants who have spoken more and their speeches, and of course it may be other situations.
[0306] Using the above method, the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0307] In a possible implementation, the conference system further includes a facial feature library including facial features. In terms of generating the first additional information corresponding to each of the plurality of collected audio segments, the generating unit 1102 is specifically configured to:
[0308] performing facial recognition on a target image according to the facial feature library, wherein the target image is a facial image at the location of the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and the processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0309] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the facial recognition result and identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, where different face IDs identify different facial features; optionally, it may also include a facial identity identifier for identifying a specific person. It should be noted that facial recognition may be performed but no facial features are identified. In this case, the recognition result may be empty or filled in according to preset rules.
[0310] In this implementation, the facial image at the sound source location of the audio clip is identified based on the facial feature library to preliminarily obtain information for determining the speaker's identity, which can improve the accuracy of subsequent determination of the speaker's identity.
[0311] In one possible implementation, the conference system further includes a first voiceprint feature library, which includes voiceprint features. In terms of generating the first additional information corresponding to each of the multiple collected audio segments, the generating unit 1102 is specifically configured to:
[0312] Determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; other audio segments in the multiple audio segments may be processed in a manner similar to the processing manner of the first audio segment;
[0313] Generate first additional information corresponding to the first audio clip, where the first additional information corresponding to the first audio clip includes the matching result of the voiceprint matching and the identification information of the first audio clip. The matching result here can be a voiceprint ID used to identify the voiceprint feature, and different voiceprint IDs identify different voiceprint features. It should be noted that there may be cases where voiceprint recognition (or matching) is performed but no voiceprint feature is identified. In this case, the matching result may be empty or filled in according to preset rules.
[0314] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0315] In one possible implementation, the conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, and the first voiceprint feature library includes voiceprint features. In terms of generating the first additional information corresponding to each of the multiple collected audio clips, the generating unit 1102 is specifically configured to:
[0316] performing facial recognition on a target image according to the facial feature library, wherein the target image is an image at the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0317] Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library;
[0318] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the face recognition, the matching result of the voiceprint matching and the identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, and different face IDs may identify different facial features; optionally, it may also include a face identity identifier for marking a specific person; it should be noted that there may be cases where face recognition is performed but facial features are not identified. In this case, the recognition result may be empty or filled with content according to preset rules. The matching result here may be a voiceprint ID for identifying voiceprint features, and different voiceprint IDs may identify different voiceprint features; it should be noted that there may be cases where voiceprint recognition (or matching) is performed but voiceprint features are not identified. In this case, the matching result may be empty or filled with content according to preset rules.
[0319] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library, and the facial image of the sound source direction of the audio clip is identified based on the facial feature library, so as to preliminarily obtain information for determining the identity of the speaker. The combination of these two aspects of information can improve the accuracy of subsequent determination of the speaker's identity.
[0320] In a possible implementation, the apparatus further includes:
[0321] The storage unit is configured to store the voiceprint features of the first audio segment in the first voiceprint feature library if no voiceprint features of the first audio segment can be matched in the first voiceprint feature library. Optionally, when the number of audio segments at a certain sound source location accumulates to a certain level, a voiceprint feature that can be distinguished from the voiceprint features of other users (i.e., the voiceprint feature corresponding to the sound source location) can be determined based on the accumulated audio segments. In this way, the first voiceprint feature library can be continuously enriched and improved, thereby increasing the accuracy of subsequent voiceprint recognition based on the first voiceprint feature library.
[0322] In a possible implementation, the first audio segment is multi-channel audio, and in terms of determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, the generating unit is specifically configured to:
[0323] Performing sound source separation on the first audio segment to obtain multiple mono audio channels, each mono audio channel being audio of a speaker, and then determining a voiceprint feature of each mono audio channel;
[0324] The voiceprint features of the multiple mono audios are matched respectively from the first voiceprint feature library.
[0325] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0326] It should be noted that the implementation and beneficial effects of each unit can also refer to Figure 3-Figure 10 The corresponding description of the method embodiment shown.
[0327] See Figure 12 , Figure 12 This is a structural diagram of a conference information processing device 120 provided in an embodiment of the present invention. The conference information processing device 120 may be the conference terminal described above or a device in the conference terminal. The conference information processing device 120 is applied to a conference system and may include a collection unit 1201, a segmentation unit 1202, a generation unit 1203, and a sending unit 1204. A detailed description of each unit is as follows.
[0328] The collection unit 1201 is configured to collect conference audio from the first conference site during the conference;
[0329] The segmentation unit 1202 is used to perform voice segmentation on the conference audio according to the sound source orientation in the conference audio to obtain multiple audio segments, wherein the sound source orientations of two audio segments adjacent in time sequence are different; in specific implementation, the conference terminal can continuously detect (for example, it can detect according to a preset frequency) the sound source orientation during the conference, and when the sound source orientation changes, it starts to collect the next audio segment. For example, the audio orientation from the 0th minute to the 6th minute of the conference is orientation 1, the audio orientation from the 6th minute to the 10th minute of the conference is audio orientation 2, and the audio orientation from the 10th minute to the 15th minute of the conference is audio orientation 3; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio segment, collect the audio from the 6th minute to the 10th minute as an audio segment, and collect the audio from the 10th minute to the 15th minute as an audio segment. It can be understood that multiple audio segments can be collected in this way.
[0330] The generation unit 1203 is used to generate the first additional information corresponding to each of the multiple audio segments, wherein the first additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment; optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio segment includes the start time and end time of the audio segment; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio segment identified by the identification information belongs to.
[0331] Sending unit 1204 is configured to send the conference audio and first additional information corresponding to the multiple audio segments to a conference information processing device. The conference audio is segmented into multiple audio segments by the conference information processing device and accompanied by corresponding second additional information. The second additional information corresponding to each audio segment includes information for identifying the speaker of the audio segment and identification information of the corresponding audio segment. The conference information processing device uses the first additional information and the second additional information to generate a correspondence between participants and speeches at the first conference site. This correspondence may be between all participants who have spoken and their speeches, or between participants who have spoken more frequently and their speeches, or other situations are also possible.
[0332] Using the above method, the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0333] In a possible implementation, the conference system further includes a facial feature library including facial features. In terms of generating the first additional information corresponding to each of the plurality of audio segments, the generating unit 1203 is specifically configured to:
[0334] performing facial recognition on a target image according to the facial feature library, wherein the target image is a facial image at the location of the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and the processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0335] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the facial recognition result and identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, where different face IDs identify different facial features; optionally, it may also include a facial identity identifier for identifying a specific person. It should be noted that facial recognition may be performed but no facial features are identified. In this case, the recognition result may be empty or filled in according to preset rules.
[0336] In this implementation, the facial image at the sound source location of the audio clip is identified based on the facial feature library to preliminarily obtain information for determining the speaker's identity, which can improve the accuracy of subsequent determination of the speaker's identity.
[0337] In a possible implementation, the conference system further includes a first voiceprint feature library, which includes voiceprint features. In terms of generating the first additional information corresponding to each of the multiple audio segments, the generating unit 1203 is specifically configured to:
[0338] Determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; other audio segments in the multiple audio segments may be processed in a manner similar to the processing manner of the first audio segment;
[0339] Generate first additional information corresponding to the first audio clip, where the first additional information corresponding to the first audio clip includes the matching result of the voiceprint matching and the identification information of the first audio clip. The matching result here can be a voiceprint ID used to identify the voiceprint feature, and different voiceprint IDs identify different voiceprint features. It should be noted that there may be cases where voiceprint recognition (or matching) is performed but no voiceprint feature is identified. In this case, the matching result may be empty or filled in according to preset rules.
[0340] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0341] In one possible implementation, the conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, and the first voiceprint feature library includes voiceprint features. In generating the first additional information corresponding to each of the multiple audio segments, the generating unit 1203 is specifically configured to:
[0342] performing facial recognition on a target image according to the facial feature library, wherein the target image is an image at the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0343] Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library;
[0344] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the face recognition, the matching result of the voiceprint matching and the identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, and different face IDs may identify different facial features; optionally, it may also include a face identity identifier for marking a specific person; it should be noted that there may be cases where face recognition is performed but facial features are not identified. In this case, the recognition result may be empty or filled with content according to preset rules. The matching result here may be a voiceprint ID for identifying voiceprint features, and different voiceprint IDs may identify different voiceprint features; it should be noted that there may be cases where voiceprint recognition (or matching) is performed but voiceprint features are not identified. In this case, the matching result may be empty or filled with content according to preset rules.
[0345] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library, and the facial image of the sound source direction of the audio clip is identified based on the facial feature library, so as to preliminarily obtain information for determining the identity of the speaker. The combination of these two aspects of information can improve the accuracy of subsequent determination of the speaker's identity.
[0346] In a possible implementation, the apparatus further includes:
[0347] The storage unit is configured to store the voiceprint features of the first audio segment in the first voiceprint feature library if no voiceprint features of the first audio segment can be matched in the first voiceprint feature library. Optionally, when the number of audio segments at a certain sound source location accumulates to a certain level, a voiceprint feature that can be distinguished from the voiceprint features of other users (i.e., the voiceprint feature corresponding to the sound source location) can be determined based on the accumulated audio segments. In this way, the first voiceprint feature library can be continuously enriched and improved, thereby increasing the accuracy of subsequent voiceprint recognition based on the first voiceprint feature library.
[0348] In a possible implementation, the first audio segment is multi-channel audio. In determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, the generating unit 1203 is specifically configured to:
[0349] Performing sound source separation on the first audio segment to obtain multiple mono audio channels, each mono audio channel being audio of a speaker, and then determining a voiceprint feature of each mono audio channel;
[0350] The voiceprint features of the multiple mono audios are matched respectively from the first voiceprint feature library.
[0351] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0352] It should be noted that the implementation and beneficial effects of each unit can also refer to Figure 3-Figure 10 The corresponding description of the method embodiment shown.
[0353] See Figure 13 , Figure 13 This is a structural diagram of a conference information processing device 130 provided in an embodiment of the present invention. The conference information processing device 130 may be the conference information processing device described above or a device in the conference information processing device. The conference information processing device 130 is applied to a conference system and may include a receiving unit 1301, a segmentation unit 1302, an identification unit 1303, and a generation unit 1304. A detailed description of each unit is as follows.
[0354] Receiving unit 1301 is configured to receive conference audio and first additional information corresponding to multiple audio segments sent by a conference terminal at a first conference site, wherein the conference audio is recorded by the conference terminal during the conference, and the multiple audio segments are obtained by voice segmentation of the conference audio or collected at the first conference site based on the sound source orientation. The first additional information corresponding to each audio segment includes information for identifying the speaker identity of the audio segment and identification information of the corresponding audio segment, wherein the sound source orientations of two temporally adjacent audio segments are different. For example, the audio orientation from the 0th minute to the 6th minute of the conference audio is orientation 1, the audio orientation from the 6th minute to the 10th minute of the conference audio is orientation 2, and the audio orientation from the 10th minute to the 15th minute of the conference audio is orientation 3. Then, the conference information processing device will segment the audio from the 0th minute to the 6th minute into one audio segment, the audio from the 6th minute to the 10th minute into one audio segment, and the audio from the 10th minute to the 15th minute into one audio segment. It can be understood that multiple audio segments can be segmented in this manner.
[0355] A segmentation unit 1302 is configured to perform voice segmentation on the conference audio to obtain multiple audio segments;
[0356] The recognition unit 1303 is used to perform voiceprint recognition on the multiple audio segments obtained by speech segmentation to obtain the second additional information corresponding to each of the multiple audio segments, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and the identification information of the corresponding audio segment. Optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio segment includes the start time and end time of the audio segment; optionally, the identification information can also be generated according to a preset rule. It should be noted that the audio segments segmented by the conference information processing device may be exactly the same as or not exactly the same as the audio segments segmented by the conference terminal. For example, the audio segments segmented by the conference terminal may be S1, S2, S3, and S4, while the audio segments segmented by the conference information processing device may be S1, S2, S3, S4, and S5. The embodiment of this application will focus on the same parts of the audio segments segmented by the two (i.e., the multiple audio segments mentioned above), and the processing of other audio segments is not limited here.
[0357] The generation unit 1304 is used to generate a correspondence between the participants and speeches in the first venue based on the first additional information and the second additional information. It may be a correspondence between all participants who have spoken and their speeches, or it may be a correspondence between participants who have spoken more and their speeches, and of course it may be other situations.
[0358] Using the above method, the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information database they each have, without the need to merge the information databases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0359] In a possible implementation, in terms of generating a correspondence between each participant in the first venue and a speech according to the first additional information and the second additional information, the generating unit 1304 is specifically configured to:
[0360] The speaker identity of the first audio segment is determined based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment (for example, including face ID) and the information for determining the speaker identity in the corresponding second additional information (for example, including voiceprint ID); optionally, when the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment, further confirmation can be made in combination with audio segments similar to the first audio segment, for example, the voiceprint ID and face ID are identified for the second audio segment, but the voiceprint ID is identified for the first audio segment, but the face ID is not identified, and the face ID is identified for the first audio segment. The voiceprint ID is the same as the voiceprint ID identified for the second audio segment, so the second audio segment and the first audio segment can be considered to be similar audio segments, so the face ID corresponding to the first audio segment and the face ID corresponding to the second audio segment are considered to be the same, and therefore the face identity identifiers corresponding to the first audio segment and the second audio segment are also the same; optionally, when the conference information processing device also obtains the sound source orientation information, the role of the sound source orientation information can also refer to the role of the voiceprint ID here; wherein, the first audio segment is one of the multiple audio segments; the processing method of other audio segments in the multiple audio segments can refer to the processing method of the first audio segment;
[0361] A meeting record is generated, wherein the meeting record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
[0362] In one possible implementation, the conference system further includes a second voiceprint feature library, the second voiceprint feature library including voiceprint features, and in performing voiceprint recognition on multiple audio segments obtained by speech segmentation to obtain the second additional information corresponding to each of the multiple audio segments, the recognition unit 1303 is specifically configured to:
[0363] Determine the voiceprint features of the first audio segment, and match the voiceprint features of the first audio segment from the second voiceprint feature library, wherein the second additional information includes a matching result of the voiceprint matching and identification information of the first audio segment, and the first audio segment is one of the multiple audio segments.
[0364] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0365] In a possible implementation, if the first audio segment is a multi-source segment, then in determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the second voiceprint feature library, the recognition unit 1303 is specifically configured to:
[0366] Performing sound source separation on the first audio clip to obtain a plurality of mono audios;
[0367] Determine the voiceprint feature of each mono audio, and match the voiceprint features of the multiple mono audios from the second voiceprint feature library.
[0368] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0369] In a possible implementation, the first additional information corresponding to the first audio segment includes a face recognition result and / or a voiceprint recognition result of the first audio segment and identification information of the first audio segment, and the first audio segment is one of the multiple audio segments.
[0370] It should be noted that the implementation and beneficial effects of each unit can also refer to Figure 3-Figure 10 The corresponding description of the method embodiment shown.
[0371] See Figure 14 , Figure 14 A conference terminal 140 is provided in an embodiment of the present invention. The conference terminal is applied to a conference system. The conference terminal 140 includes a processor 1401, a memory 1402 and a communication interface 1403. The processor 1401, the memory 1402, the communication interface 1403, the camera 1404 and the microphone 1405 are interconnected through a bus; of course, the camera 1404 and the microphone 1405 can also be connected to the conference terminal 140 in an external manner.
[0372] The camera 1404 is used to collect facial images or other image information during the meeting. The camera 1404 can also be a camera module, wherein the camera can also be called a camera head or a camera device.
[0373] The microphone 1405 is used to collect audio information during the conference, such as the aforementioned conference audio, audio clips, etc. The microphone 1405 can also be an array microphone, and the microphone 1405 can also be called a recording device, a recording equipment, etc.
[0374] Memory 1402 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM). Memory 1402 is used for related instructions and data. Communication interface 1403 is used to receive and send data.
[0375] The processor 1401 may be one or more central processing units (CPUs). When the processor 1401 is a CPU, the CPU may be a single-core CPU or a multi-core CPU.
[0376] The processor 1401 in the conference terminal 140 is configured to read the program code stored in the memory 1402 and perform the following operations:
[0377] During the meeting, audio clips of the first conference venue are collected based on the sound source orientation, wherein the sound source orientations of two audio clips that are adjacent in time sequence are different; in specific implementation, the conference terminal can continuously detect (for example, it can detect according to a preset frequency) the sound source orientation during the meeting, and when the sound source orientation changes, it starts collecting the next audio clip. For example, the audio orientation is orientation 1 from the 0th minute to the 6th minute of the meeting, the audio orientation is 2 from the 6th minute to the 10th minute of the meeting, and the audio orientation is 3 from the 10th minute to the 15th minute of the meeting; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio clip, collect the audio from the 6th minute to the 10th minute as an audio clip, and collect the audio from the 10th minute to the 15th minute as an audio clip. It can be understood that multiple audio clips can be collected in this way.
[0378] Generate first additional information corresponding to each of the multiple collected audio clips, wherein the first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio clip includes the start time and end time of the audio clip; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio clip identified by the identification information belongs to.
[0379] The conference audio recorded during the conference and the first additional information corresponding to the multiple audio segments are sent to the conference information processing device through the communication interface 1403. The conference audio is divided into multiple audio segments by the conference information processing device (for example, divided according to the direction of the sound source, so that the sound source directions of two audio segments adjacent in time sequence obtained by the division are different) and are accompanied by corresponding second additional information, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment; the first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants and speeches in the first venue. It may be a correspondence between all participants who have spoken and their speeches, or it may be a correspondence between participants who have spoken more and their speeches, and of course it may be other situations.
[0380] It can be seen that the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information bases they each have, without the need to merge the information bases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0381] In a possible implementation, the conference system further includes a facial feature library including facial features. In generating the first additional information corresponding to each of the plurality of collected audio segments, the processor is specifically configured to:
[0382] performing facial recognition on a target image according to the facial feature library, wherein the target image is a facial image at the location of the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and the processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0383] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the facial recognition result and identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, where different face IDs identify different facial features; optionally, it may also include a facial identity identifier for identifying a specific person. It should be noted that facial recognition may be performed but no facial features are identified. In this case, the recognition result may be empty or filled in according to preset rules.
[0384] In this implementation, the facial image at the sound source location of the audio clip is identified based on the facial feature library to preliminarily obtain information for determining the speaker's identity, which can improve the accuracy of subsequent determination of the speaker's identity.
[0385] In a possible implementation, the conference system further includes a first voiceprint feature library, the first voiceprint feature library including voiceprint features, and in generating the first additional information corresponding to each of the plurality of collected audio segments, the processor is specifically configured to:
[0386] Determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; other audio segments in the multiple audio segments may be processed in a manner similar to the processing manner of the first audio segment;
[0387] Generate first additional information corresponding to the first audio clip, where the first additional information corresponding to the first audio clip includes the matching result of the voiceprint matching and the identification information of the first audio clip. The matching result here can be a voiceprint ID used to identify the voiceprint feature, and different voiceprint IDs identify different voiceprint features. It should be noted that there may be cases where voiceprint recognition (or matching) is performed but no voiceprint feature is identified. In this case, the matching result may be empty or filled in according to preset rules.
[0388] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0389] In one possible implementation, the conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, and the first voiceprint feature library includes voiceprint features. In terms of generating the first additional information corresponding to each of the plurality of collected audio clips, the processor is specifically configured to:
[0390] performing facial recognition on a target image according to the facial feature library, wherein the target image is an image at the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0391] Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library;
[0392] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the face recognition, the matching result of the voiceprint matching and the identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, and different face IDs may identify different facial features; optionally, it may also include a face identity identifier for marking a specific person; it should be noted that there may be cases where face recognition is performed but facial features are not identified. In this case, the recognition result may be empty or filled with content according to preset rules. The matching result here may be a voiceprint ID for identifying voiceprint features, and different voiceprint IDs may identify different voiceprint features; it should be noted that there may be cases where voiceprint recognition (or matching) is performed but voiceprint features are not identified. In this case, the matching result may be empty or filled with content according to preset rules.
[0393] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library, and the facial image of the sound source direction of the audio clip is identified based on the facial feature library, so as to preliminarily obtain information for determining the identity of the speaker. The combination of these two aspects of information can improve the accuracy of subsequent determination of the speaker's identity.
[0394] In a possible implementation, the processor is further configured to:
[0395] If the voiceprint features of the first audio clip cannot be matched in the first voiceprint feature library, the voiceprint features of the first audio clip are saved in the first voiceprint feature library. Optionally, when the number of audio clips at a certain sound source location accumulates to a certain level, a voiceprint feature that can be distinguished from the voiceprint features of other users (i.e., the voiceprint feature corresponding to the sound source location) can be determined based on the accumulated audio clips. In this way, the first voiceprint feature library can be continuously enriched and improved, making subsequent voiceprint recognition based on the first voiceprint feature library more accurate.
[0396] In a possible implementation, the first audio segment is multi-channel audio. In determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, the processor is specifically configured to:
[0397] Performing sound source separation on the first audio segment to obtain multiple mono audio channels, each mono audio channel being audio of a speaker, and then determining a voiceprint feature of each mono audio channel;
[0398] The voiceprint features of the multiple mono audios are matched respectively from the first voiceprint feature library.
[0399] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0400] It should be noted that the implementation and beneficial effects of each operation can also refer to Figure 3-Figure 10 The corresponding description of the method embodiment shown.
[0401] See Figure 15 , Figure 15 A conference terminal 150 is provided in an embodiment of the present invention. The conference terminal is applied to a conference system. The conference terminal 150 includes a processor 1501, a memory 1502 and a communication interface 1503. The processor 1501, the memory 1502, the communication interface 1503, the camera 1504 and the microphone 1505 are interconnected through a bus; of course, the camera 1504 and the microphone 1505 can also be connected to the conference terminal 150 in an external manner.
[0402] The camera 1504 is used to collect facial images or other image information during the meeting. The camera 1504 can also be a camera module, wherein the camera can also be called a camera head or a camera device.
[0403] The microphone 1505 is used to collect audio information during the conference, such as the aforementioned conference audio, audio clips, etc. The microphone 1505 can also be an array microphone, and the microphone 1505 can also be called a recording device, recording equipment, etc.
[0404] Memory 1502 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM). Memory 1502 is used for related instructions and data. Communication interface 1503 is used to receive and send data.
[0405] The processor 1501 may be one or more central processing units (CPUs). When the processor 1501 is a CPU, the CPU may be a single-core CPU or a multi-core CPU.
[0406] The processor 1501 in the conference terminal 150 is configured to read the program code stored in the memory 1502 and perform the following operations:
[0407] Collect the conference audio of the first conference venue during the conference;
[0408] The conference audio is voice-segmented according to the sound source orientation in the conference audio to obtain multiple audio segments, wherein the sound source orientations of two audio segments adjacent in time sequence are different; in specific implementation, the conference terminal can continuously detect (for example, it can detect according to a preset frequency) the sound source orientation during the conference, and when the sound source orientation changes, it starts to collect the next audio segment. For example, the audio orientation is orientation 1 from the 0th minute to the 6th minute of the conference, the audio orientation is 2 from the 6th minute to the 10th minute of the conference, and the audio orientation is 3 from the 10th minute to the 15th minute of the conference; then, the conference terminal will collect the audio from the 0th minute to the 6th minute as an audio segment, the audio from the 6th minute to the 10th minute as an audio segment, and the audio from the 10th minute to the 15th minute as an audio segment. It can be understood that multiple audio segments can be collected in this way.
[0409] Generate first additional information corresponding to each of the multiple audio clips, wherein the first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio clip includes the start time and end time of the audio clip; optionally, the identification information can also be generated according to preset rules, so the conference information processing device can also process the corresponding identification according to the preset rules to determine which section of the conference the audio clip identified by the identification information belongs to.
[0410] The conference audio and first additional information corresponding to the multiple audio segments are sent to the conference information processing device via the communication interface 1503. The conference audio is segmented into multiple audio segments by the conference information processing device and accompanied by corresponding second additional information. The second additional information corresponding to each audio segment includes information for identifying the speaker of the audio segment and identification information of the corresponding audio segment. The first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants and speeches at the first conference site. This may be a correspondence between all participants who have spoken and their speeches, or a correspondence between participants who have spoken more and their speeches, or other situations are also possible.
[0411] It can be seen that the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information bases they each have, without the need to merge the information bases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0412] In a possible implementation, the conference system further includes a facial feature library including facial features. In terms of generating the first additional information corresponding to each of the plurality of audio segments, the processor is specifically configured to:
[0413] performing facial recognition on a target image according to the facial feature library, wherein the target image is a facial image at the location of the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and the processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0414] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the facial recognition result and identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, where different face IDs identify different facial features; optionally, it may also include a facial identity identifier for identifying a specific person. It should be noted that facial recognition may be performed but no facial features are identified. In this case, the recognition result may be empty or filled in according to preset rules.
[0415] In this implementation, the facial image at the sound source location of the audio clip is identified based on the facial feature library to preliminarily obtain information for determining the speaker's identity, which can improve the accuracy of subsequent determination of the speaker's identity.
[0416] In a possible implementation, the conference system further includes a first voiceprint feature library, the first voiceprint feature library including voiceprint features, and in generating the first additional information corresponding to each of the plurality of audio segments, the processor is specifically configured to:
[0417] Determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; other audio segments in the multiple audio segments may be processed in a manner similar to the processing manner of the first audio segment;
[0418] Generate first additional information corresponding to the first audio clip, where the first additional information corresponding to the first audio clip includes the matching result of the voiceprint matching and the identification information of the first audio clip. The matching result here can be a voiceprint ID used to identify the voiceprint feature, and different voiceprint IDs identify different voiceprint features. It should be noted that there may be cases where voiceprint recognition (or matching) is performed but no voiceprint feature is identified. In this case, the matching result may be empty or filled in according to preset rules.
[0419] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0420] In one possible implementation, the conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, and the first voiceprint feature library includes voiceprint features. In generating the first additional information corresponding to each of the multiple audio segments, the processor is specifically configured to:
[0421] performing facial recognition on a target image according to the facial feature library, wherein the target image is an image at the sound source of the first audio segment captured during recording of the first audio segment (e.g., captured by a director camera, a camera, or a camera module), the first audio segment being one of the plurality of audio segments, and processing of the other audio segments in the plurality of audio segments may refer to the processing of the first audio segment;
[0422] Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library;
[0423] Generate first additional information corresponding to the first audio clip, wherein the first additional information corresponding to the first audio clip includes the recognition result of the face recognition, the matching result of the voiceprint matching and the identification information of the first audio clip. The recognition result here may include a face ID for identifying facial features, and different face IDs may identify different facial features; optionally, it may also include a face identity identifier for marking a specific person; it should be noted that there may be cases where face recognition is performed but facial features are not identified. In this case, the recognition result may be empty or filled with content according to preset rules. The matching result here may be a voiceprint ID for identifying voiceprint features, and different voiceprint IDs may identify different voiceprint features; it should be noted that there may be cases where voiceprint recognition (or matching) is performed but voiceprint features are not identified. In this case, the matching result may be empty or filled with content according to preset rules.
[0424] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library, and the facial image of the sound source direction of the audio clip is identified based on the facial feature library, so as to preliminarily obtain information for determining the identity of the speaker. The combination of these two aspects of information can improve the accuracy of subsequent determination of the speaker's identity.
[0425] In a possible implementation, the processor is further configured to:
[0426] If the voiceprint features of the first audio clip cannot be matched in the first voiceprint feature library, the voiceprint features of the first audio clip are saved in the first voiceprint feature library. Optionally, when the number of audio clips at a certain sound source location accumulates to a certain level, a voiceprint feature that can be distinguished from the voiceprint features of other users (i.e., the voiceprint feature corresponding to the sound source location) can be determined based on the accumulated audio clips. In this way, the first voiceprint feature library can be continuously enriched and improved, making subsequent voiceprint recognition based on the first voiceprint feature library more accurate.
[0427] In a possible implementation, the first audio segment is multi-channel audio. In determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, the processor is specifically configured to:
[0428] Performing sound source separation on the first audio segment to obtain multiple mono audio channels, each mono audio channel being audio of a speaker, and then determining a voiceprint feature of each mono audio channel;
[0429] The voiceprint features of the multiple mono audios are matched respectively from the first voiceprint feature library.
[0430] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0431] It should be noted that the implementation and beneficial effects of each operation can also refer to Figure 3-Figure 10 The corresponding description of the method embodiment shown.
[0432] See Figure 16 , Figure 16 A conference information processing device 160 provided by an embodiment of the present invention includes a processor 1601, a memory 1602, and a communication interface 1603. The processor 1601, the memory 1602, and the communication interface 1603 are interconnected via a bus.
[0433] Memory 1602 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM). Memory 1602 is used for related instructions and data. Communication interface 1603 is used to receive and send data.
[0434] The processor 1601 may be one or more central processing units (CPUs). When the processor 1601 is a CPU, the CPU may be a single-core CPU or a multi-core CPU.
[0435] The processor 1601 in the conference information processing device 160 is configured to read the program code stored in the memory 1602 and perform the following operations:
[0436] The communication interface 1603 receives the conference audio and first additional information corresponding to multiple audio clips sent by the conference terminal of the first conference site, wherein the conference audio is recorded by the conference terminal during the conference, and the multiple audio clips are obtained by voice segmentation of the conference audio or collected at the first conference site according to the sound source direction. The first additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip, wherein the sound source directions of two audio clips that are adjacent in time sequence are different; for example, the audio direction from the 0th minute to the 6th minute of the conference audio is direction 1, the audio direction from the 6th minute to the 10th minute of the conference audio is direction 2, and the audio direction from the 10th minute to the 15th minute of the conference audio is direction 3; then, the conference information processing device will segment the audio from the 0th minute to the 6th minute into one audio clip, the audio from the 6th minute to the 10th minute into one audio clip, and the audio from the 10th minute to the 15th minute into one audio clip. It can be understood that multiple audio clips can be segmented in this way.
[0437] Performing voice segmentation on the conference audio to obtain multiple audio segments;
[0438] Voiceprint recognition is performed on the multiple audio segments obtained by speech segmentation to obtain the second additional information corresponding to each of the multiple audio segments, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment. Optionally, the identification information can be a timestamp, and the timestamp corresponding to each audio segment includes the start time and end time of the audio segment; optionally, the identification information can also be generated according to a preset rule. It should be noted that the audio segments segmented by the conference information processing device may be exactly the same as or not exactly the same as the audio segments segmented by the conference terminal. For example, the audio segments segmented by the conference terminal may be S1, S2, S3, and S4, while the audio segments segmented by the conference information processing device may be S1, S2, S3, S4, and S5. The embodiment of this application will focus on the same parts of the audio segments segmented by the two (i.e., the multiple audio segments mentioned above), and the processing of other audio segments is not limited here.
[0439] The correspondence between the participants and speeches in the first venue is generated based on the first additional information and the second additional information. It may be the correspondence between all participants who have spoken and their speeches, or the correspondence between participants who have spoken more and their speeches, or of course other situations.
[0440] It can be seen that the conference terminal obtains the conference audio and audio clips during the conference, identifies the information used to determine the identity of the speaker of each audio clip based on each audio clip, and then sends the conference audio and the identified information used to determine the identity of the speaker of the audio clip to the conference information processing device; the conference information processing device also divides the conference audio into multiple audio clips, and then identifies the information used to determine the identity of the speaker of each audio clip, and then determines the speaker corresponding to each clip by combining the information it has identified for determining the identity of the speaker of each audio clip and the received information for determining the identity of the speaker of each audio clip. By jointly determining the identity of the speaker by the conference terminal and the conference information processing device, the determined identity of the speaker can be made more accurate; and since both the conference terminal and the conference information processing device will perform identification, the conference terminal and the conference information processing device can perform identification based on the information bases they each have, without the need to merge the information bases of both parties, thereby preventing the leakage of each other's information and effectively protecting the privacy and security of the participants.
[0441] In a possible implementation, in terms of generating a correspondence between each participant in the first venue and a speech according to the first additional information and the second additional information, the processor is specifically configured to:
[0442] The speaker identity of the first audio segment is determined based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment (for example, including face ID) and the information for determining the speaker identity in the corresponding second additional information (for example, including voiceprint ID); optionally, when the speaker identity cannot be uniquely determined based on the first additional information and the second additional information corresponding to the first audio segment, further confirmation can be made in combination with audio segments similar to the first audio segment, for example, the voiceprint ID and face ID are identified for the second audio segment, but the voiceprint ID is identified for the first audio segment, but the face ID is not identified, and the face ID is identified for the first audio segment. The voiceprint ID is the same as the voiceprint ID identified for the second audio segment, so the second audio segment and the first audio segment can be considered to be similar audio segments, so the face ID corresponding to the first audio segment and the face ID corresponding to the second audio segment are considered to be the same, and therefore the face identity identifiers corresponding to the first audio segment and the second audio segment are also the same; optionally, when the conference information processing device also obtains the sound source orientation information, the role of the sound source orientation information can also refer to the role of the voiceprint ID here; wherein, the first audio segment is one of the multiple audio segments; the processing method of other audio segments in the multiple audio segments can refer to the processing method of the first audio segment;
[0443] A meeting record is generated, wherein the meeting record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
[0444] In one possible implementation, the conference system further includes a second voiceprint feature library, the second voiceprint feature library including voiceprint features, and in performing voiceprint recognition on multiple audio segments obtained by speech segmentation to obtain second additional information corresponding to each of the multiple audio segments, the processor is specifically configured to:
[0445] Determine the voiceprint features of the first audio segment, and match the voiceprint features of the first audio segment from the second voiceprint feature library, wherein the second additional information includes a matching result of the voiceprint matching and identification information of the first audio segment, and the first audio segment is one of the multiple audio segments.
[0446] In this implementation, the voiceprint of the audio clip is identified based on the voiceprint feature library to preliminarily obtain information for determining the identity of the speaker, which can improve the accuracy of subsequent determination of the speaker's identity.
[0447] In a possible implementation, if the first audio segment is a multi-source segment, then in determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the second voiceprint feature library, the processor is specifically configured to:
[0448] Performing sound source separation on the first audio clip to obtain a plurality of mono audios;
[0449] Determine the voiceprint feature of each mono audio, and match the voiceprint features of the multiple mono audios from the second voiceprint feature library.
[0450] It can be understood that when the first audio segment is multi-channel audio, performing sound source separation on it and performing voiceprint recognition on each separated mono audio can more accurately determine the correspondence between the speech and the speaker in the meeting.
[0451] In a possible implementation, the first additional information corresponding to the first audio segment includes a face recognition result and / or a voiceprint recognition result of the first audio segment and identification information of the first audio segment, and the first audio segment is one of the multiple audio segments.
[0452] It should be noted that the implementation and beneficial effects of each operation can also refer to Figure 3-Figure 10 The corresponding description of the method embodiment shown.
[0453] An embodiment of the present invention further provides a chip system, the chip system comprising at least one processor, a memory and a communication interface, the memory, the communication interface and the at least one processor being interconnected via a line, the at least one memory storing instructions; when the instructions are executed by the processor, Figure 3-Figure 10 The method flow shown is realized.
[0454] The embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a processor, implement Figure 3-Figure 10 The method flow shown is realized.
[0455] The embodiment of the present invention further provides a computer program product, which, when executed on a processor, implements Figure 3-Figure 10 The method flow shown is realized.
[0456] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A conference data processing method, characterized in that: Applied to a conference system comprising a conference terminal and a conference information processing device, the method comprises: The conference terminal collects audio clips of the first conference site according to the sound source orientation during the conference, wherein the sound source orientations of two audio clips adjacent in time sequence are different; The conference terminal generates first additional information corresponding to each of the multiple collected audio clips, wherein the first additional information corresponding to each audio clip includes a face recognition result and / or a voiceprint recognition result corresponding to the audio clip, and corresponding audio clip identification information; The conference terminal sends the conference audio recorded during the conference and the first additional information corresponding to the multiple audio clips to the conference information processing device. The conference audio is divided into multiple audio clips by the conference information processing device. The multiple audio clips are used by the conference information processing device to obtain the second additional information corresponding to each of the multiple audio clips through voiceprint recognition, wherein the second additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; the first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants in the first conference venue and the speeches.
2. The method according to claim 1, characterized in that The conference system further includes a facial feature library, the facial feature library including facial features, and the conference terminal generates first additional information corresponding to each of the plurality of collected audio clips, including: performing face recognition on a target image according to the face feature library, wherein the target image is a face image at a location of a sound source of a first audio segment captured during recording of the first audio segment, and the first audio segment is one of the plurality of audio segments; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a recognition result of the face recognition and identification information of the first audio segment.
3. The method according to claim 1, characterized in that The conference system further includes a first voiceprint feature library, the first voiceprint feature library including voiceprint features, and the first additional information corresponding to each of the plurality of collected audio clips generated by the conference terminal includes: determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a matching result of the voiceprint matching and identification information of the first audio segment.
4. The method according to claim 1, wherein The conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, the first voiceprint feature library includes voiceprint features, and the first additional information corresponding to each of the multiple collected audio clips generated by the conference terminal includes: performing face recognition on a target image according to the face feature library, wherein the target image is an image taken during recording of a first audio segment at a location of a sound source of the first audio segment, and the first audio segment is one of the plurality of audio segments; Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a recognition result of the face recognition, a matching result of the voiceprint matching, and identification information of the first audio segment.
5. The method according to claim 3 or 4, characterized in that The method further comprises: If the voiceprint feature of the first audio segment cannot be matched from the first voiceprint feature library, the voiceprint feature of the first audio segment is saved in the first voiceprint feature library.
6. The method according to claim 3 or 4, characterized in that The first audio segment is multi-channel audio, and determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library includes: Performing sound source separation on the first audio segment to obtain multiple mono audio channels, and determining a voiceprint feature of each mono audio channel; The conference terminal matches the voiceprint features of the multiple mono audios respectively from the first voiceprint feature library.
7. The method according to any one of claims 1 to 4, characterized in that Also includes: The conference information processing device determines the speaker identity of the first audio segment based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment and the information for determining the speaker identity in the corresponding second additional information; wherein the first audio segment is one of the multiple audio segments; The conference information processing device generates a conference record, wherein the conference record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
8. The method according to any one of claims 1 to 4, characterized in that The speech includes at least one of speech text, speech voice and speech time period.
9. A conference data processing method, characterized in that: Applied to a conference system comprising a conference terminal and a conference information processing device, the method comprises: The conference terminal collects the conference audio of the first conference site during the conference; The conference terminal performs voice segmentation on the conference audio according to the sound source orientation in the conference audio to obtain a plurality of audio segments, wherein the sound source orientations of two audio segments adjacent in time sequence are different; The conference terminal generates first additional information corresponding to each of the multiple audio clips, wherein the first additional information corresponding to each audio clip includes a face recognition result and / or a voiceprint recognition result corresponding to the audio clip, and identification information of the corresponding audio clip; The conference terminal sends the conference audio and the first additional information corresponding to the multiple audio clips to the conference information processing device. The conference audio is divided into multiple audio clips by the conference information processing device. The multiple audio clips are used by the conference information processing device to obtain the second additional information corresponding to each of the multiple audio clips through voiceprint recognition, wherein the second additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; the first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants in the first conference venue and the speeches.
10. The method according to claim 9, characterized in that The conference system further includes a facial feature library including facial features, and the conference terminal generates first additional information corresponding to each of the plurality of audio clips, including: performing face recognition on a target image according to the face feature library, wherein the target image is a face image at a location of a sound source of a first audio segment captured during recording of the first audio segment, and the first audio segment is one of the plurality of audio segments; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a recognition result of the face recognition and identification information of the first audio segment.
11. The method according to claim 9, characterized in that The conference system further includes a first voiceprint feature library, the first voiceprint feature library including voiceprint features, and the conference terminal generates first additional information corresponding to each of the plurality of audio clips including: determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a matching result of the voiceprint matching and identification information of the first audio segment.
12. The method according to claim 9, wherein The conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library includes facial features, the first voiceprint feature library includes voiceprint features, and the conference terminal generates the first additional information corresponding to each of the multiple audio clips, including: performing face recognition on a target image according to the face feature library, wherein the target image is an image taken during recording of a first audio segment at a location of a sound source of the first audio segment, and the first audio segment is one of the plurality of audio segments; Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a recognition result of the face recognition, a matching result of the voiceprint matching, and identification information of the first audio segment.
13. The method according to claim 11 or 12, characterized in that The method further comprises: If the voiceprint feature of the first audio segment cannot be matched from the first voiceprint feature library, the voiceprint feature of the first audio segment is saved in the first voiceprint feature library.
14. The method according to claim 11 or 12, characterized in that The first audio segment is multi-channel audio, and determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library includes: Performing sound source separation on the first audio segment to obtain multiple mono audio channels, and determining a voiceprint feature of each mono audio channel; The conference terminal matches the voiceprint features of the multiple mono audios respectively from the first voiceprint feature library.
15. The method according to any one of claims 9 to 12, characterized in that: Also includes: The conference information processing device determines the speaker identity of the first audio segment based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment and the information for determining the speaker identity in the corresponding second additional information; wherein the first audio segment is one of the multiple audio segments; The conference information processing device generates a conference record, wherein the conference record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
16. The method according to any one of claims 9 to 12, characterized in that: The speech includes at least one of speech text, speech voice and speech time period.
17. A conference information processing method, characterized in that: Applied to a conference system comprising a conference terminal and a conference information processing device, the method comprises: The conference information processing device receives conference audio and first additional information corresponding to multiple audio clips sent by a conference terminal at a first conference site, wherein the conference audio is recorded by the conference terminal during a conference, and the multiple audio clips are obtained by performing voice segmentation on the conference audio or are collected at the first conference site based on sound source orientations, and the first additional information corresponding to each audio clip includes a face recognition result and / or a voiceprint recognition result corresponding to the audio clip, as well as identification information of the corresponding audio clip, wherein the sound source orientations of two audio clips that are adjacent in time sequence are different; The conference information processing device performs voice segmentation on the conference audio to obtain multiple audio segments; The conference information processing device performs voiceprint recognition on the multiple audio segments obtained by voice segmentation to obtain second additional information corresponding to each of the multiple audio segments, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment; The conference information processing device generates a correspondence between participants in the first conference site and speeches based on the first additional information and the second additional information.
18. The method according to claim 17, characterized in that The conference information processing device generates a correspondence between each participant in the first conference site and the speech according to the first additional information and the second additional information, including: The conference information processing device determines the speaker identity of the first audio segment based on the information for determining the speaker identity in the first additional information corresponding to the first audio segment and the information for determining the speaker identity in the corresponding second additional information; wherein the first audio segment is one of the multiple audio segments; The conference information processing device generates a conference record, wherein the conference record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
19. The method according to claim 17 or 18, characterized in that The conference system further includes a second voiceprint feature library, the second voiceprint feature library including voiceprint features, and the conference information processing device performs voiceprint recognition on multiple audio segments obtained by speech segmentation to obtain second additional information corresponding to each of the multiple audio segments, including: The conference information processing device determines the voiceprint features of the first audio segment and matches the voiceprint features of the first audio segment from the second voiceprint feature library. The second additional information includes the matching result of the voiceprint matching and the identification information of the first audio segment. The first audio segment is one of the multiple audio segments.
20. The method according to claim 19, characterized in that If the first audio segment is a multi-source segment, the conference information processing device determines the voiceprint feature of the first audio segment, and matches the voiceprint feature of the first audio segment from the second voiceprint feature library, including: The conference information processing device performs sound source separation on the first audio segment to obtain a plurality of mono audios; The conference information processing device determines the voiceprint features of each mono audio, and matches the voiceprint features of the multiple mono audios from the second voiceprint feature library respectively.
21. The method according to claim 17 or 18, characterized in that The first additional information corresponding to the first audio segment includes a face recognition result and / or a voiceprint recognition result of the first audio segment and identification information of the first audio segment, where the first audio segment is one of the multiple audio segments.
22. The method according to claim 17 or 18, characterized in that The speech includes at least one of a speech text and a speech time period.
23. A conference terminal, characterized in that: The conference terminal is applied to a conference system and includes a processor, a memory, and a communication interface, wherein the memory is used to store a computer program, and the processor calls the computer program to perform the following operations: During the conference, audio clips of the first conference site are collected according to the directions of the sound sources, wherein the directions of the sound sources of two audio clips adjacent in time sequence are different; Generate first additional information corresponding to each of the multiple collected audio clips, wherein the first additional information corresponding to each audio clip includes a face recognition result and / or a voiceprint recognition result corresponding to the audio clip, and identification information of the corresponding audio clip; The conference audio recorded during the conference and the first additional information corresponding to the multiple audio clips are sent to the conference information processing device through the communication interface. The conference audio is divided into multiple audio clips by the conference information processing device. The multiple audio clips are used by the conference information processing device to obtain the second additional information corresponding to each of the multiple audio clips through voiceprint recognition, wherein the second additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; the first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants in the first venue and the speeches.
24. The conference terminal according to claim 23, characterized in that: The conference system further includes a facial feature library including facial features. In terms of generating first additional information corresponding to each of the plurality of collected audio segments, the processor is specifically configured to: performing face recognition on a target image according to the face feature library, wherein the target image is a face image at a location of a sound source of a first audio segment captured during recording of the first audio segment, and the first audio segment is one of the plurality of audio segments; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a recognition result of the face recognition and identification information of the first audio segment.
25. The conference terminal according to claim 23, characterized in that The conference system further includes a first voiceprint feature library including voiceprint features. In terms of generating first additional information corresponding to each of the plurality of collected audio segments, the processor is specifically configured to: determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a matching result of the voiceprint matching and identification information of the first audio segment.
26. The conference terminal according to claim 23, wherein: The conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library including facial features, and the first voiceprint feature library including voiceprint features. In terms of generating first additional information corresponding to each of the plurality of collected audio clips, the processor is specifically configured to: performing face recognition on a target image according to the face feature library, wherein the target image is an image taken during recording of a first audio segment at a location of a sound source of the first audio segment, and the first audio segment is one of the plurality of audio segments; Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a recognition result of the face recognition, a matching result of the voiceprint matching, and identification information of the first audio segment.
27. The conference terminal according to claim 25 or 26, characterized in that: The processor is further configured to: If the voiceprint feature of the first audio segment cannot be matched from the first voiceprint feature library, the voiceprint feature of the first audio segment is saved in the first voiceprint feature library.
28. The conference terminal according to claim 25 or 26, characterized in that: The first audio segment is multi-channel audio. In determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, the processor is specifically configured to: Performing sound source separation on the first audio segment to obtain multiple mono audio channels, and determining a voiceprint feature of each mono audio channel; The voiceprint features of the multiple mono audios are matched respectively from the first voiceprint feature library.
29. The conference terminal according to any one of claims 23 to 26, characterized in that: The speech includes at least one of speech text, speech voice and speech time period.
30. A conference terminal, characterized in that: The conference terminal is applied to a conference system and includes a processor, a memory, and a communication interface, wherein the memory is used to store a computer program, and the processor calls the computer program to perform the following operations: Collect the conference audio of the first conference venue during the conference; Segmenting the conference audio according to the sound source orientation in the conference audio to obtain a plurality of audio segments, wherein the sound source orientations of two audio segments adjacent in time sequence are different; generating first additional information corresponding to each of the plurality of audio segments, wherein the first additional information corresponding to each audio segment includes a face recognition result and / or a voiceprint recognition result corresponding to the audio segment, and identification information of the corresponding audio segment; The conference audio and the first additional information corresponding to the multiple audio clips are sent to the conference information processing device through the communication interface. The conference audio is divided into multiple audio clips by the conference information processing device. The multiple audio clips are used by the conference information processing device to obtain the second additional information corresponding to each of the multiple audio clips through voiceprint recognition, wherein the second additional information corresponding to each audio clip includes information for determining the identity of the speaker of the audio clip and identification information of the corresponding audio clip; the first additional information and the second additional information are used by the conference information processing device to generate a correspondence between the participants in the first venue and the speeches.
31. The conference terminal according to claim 30, characterized in that: The conference system further includes a facial feature library including facial features. In terms of generating the first additional information corresponding to each of the plurality of audio segments, the processor is specifically configured to: performing face recognition on a target image according to the face feature library, wherein the target image is a face image at a location of a sound source of a first audio segment captured during recording of the first audio segment, and the first audio segment is one of the plurality of audio segments; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a recognition result of the face recognition and identification information of the first audio segment.
32. The conference terminal according to claim 30, characterized in that The conference system further includes a first voiceprint feature library including voiceprint features. In terms of generating the first additional information corresponding to each of the plurality of audio segments, the processor is specifically configured to: determining a voiceprint feature of a first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, where the first audio segment is one of the multiple audio segments; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a matching result of the voiceprint matching and identification information of the first audio segment.
33. The conference terminal according to claim 30, wherein: The conference system further includes a facial feature library and a first voiceprint feature library, the facial feature library including facial features, and the first voiceprint feature library including voiceprint features. In terms of generating the first additional information corresponding to each of the plurality of audio segments, the processor is specifically configured to: performing face recognition on a target image according to the face feature library, wherein the target image is an image taken during recording of a first audio segment at a location of a sound source of the first audio segment, and the first audio segment is one of the plurality of audio segments; Determining a voiceprint feature of the first audio segment, and matching the voiceprint feature of the first audio segment from the first voiceprint feature library; First additional information corresponding to the first audio segment is generated, wherein the first additional information corresponding to the first audio segment includes a recognition result of the face recognition, a matching result of the voiceprint matching, and identification information of the first audio segment.
34. The conference terminal according to claim 32 or 33, characterized in that: The processor is further configured to: If the voiceprint feature of the first audio segment cannot be matched from the first voiceprint feature library, the voiceprint feature of the first audio segment is saved in the first voiceprint feature library.
35. The conference terminal according to claim 32 or 33, characterized in that: The first audio segment is multi-channel audio. In determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the first voiceprint feature library, the processor is specifically configured to: Performing sound source separation on the first audio segment to obtain multiple mono audio channels, and determining a voiceprint feature of each mono audio channel; The voiceprint features of the multiple mono audios are matched respectively from the first voiceprint feature library.
36. The conference terminal according to any one of claims 30 to 33, characterized in that: The speech includes at least one of speech text, speech voice and speech time period.
37. A conference information processing device, characterized in that: The system comprises a processor, a memory, and a communication interface, wherein the memory is used to store a computer program, and the processor calls the computer program to perform the following operations: Receiving, through the communication interface, conference audio and first additional information corresponding to multiple audio clips sent by a conference terminal at a first conference site, wherein the conference audio is recorded by the conference terminal during a conference, and the multiple audio clips are obtained by performing voice segmentation on the conference audio or are collected at the first conference site based on sound source orientations, and the first additional information corresponding to each audio clip includes a face recognition result and / or a voiceprint recognition result corresponding to the audio clip, as well as identification information of the corresponding audio clip, wherein the sound source orientations of two audio clips that are adjacent in time sequence are different; Performing voice segmentation on the conference audio to obtain multiple audio segments; Performing voiceprint recognition on the multiple audio segments obtained by speech segmentation to obtain second additional information corresponding to each of the multiple audio segments, wherein the second additional information corresponding to each audio segment includes information for determining the identity of the speaker of the audio segment and identification information of the corresponding audio segment; A correspondence between participants in the first conference site and speeches is generated according to the first additional information and the second additional information.
38. The apparatus according to claim 37, wherein In terms of generating a correspondence between each participant in the first venue and a speech according to the first additional information and the second additional information, the processor is specifically configured to: Determining the speaker identity of the first audio segment based on information for determining the speaker identity in the first additional information corresponding to the first audio segment and information for determining the speaker identity in the corresponding second additional information; wherein the first audio segment is one of the multiple audio segments; A meeting record is generated, wherein the meeting record includes the speech of the first audio segment and the identity of the speaker corresponding to the first audio segment.
39. The apparatus according to claim 37 or 38, characterized in that The conference system further includes a second voiceprint feature library including voiceprint features. In performing voiceprint recognition on multiple audio segments obtained by speech segmentation to obtain second additional information corresponding to each of the multiple audio segments, the processor is specifically configured to: Determine the voiceprint features of the first audio segment, and match the voiceprint features of the first audio segment from the second voiceprint feature library, wherein the second additional information includes a matching result of the voiceprint matching and identification information of the first audio segment, and the first audio segment is one of the multiple audio segments.
40. The apparatus according to claim 39, wherein If the first audio segment is a multi-source segment, then in determining the voiceprint feature of the first audio segment and matching the voiceprint feature of the first audio segment from the second voiceprint feature library, the processor is specifically configured to: Performing sound source separation on the first audio clip to obtain a plurality of mono audios; Determine the voiceprint feature of each mono audio, and match the voiceprint features of the multiple mono audios from the second voiceprint feature library.
41. The apparatus according to claim 37 or 38, characterized in that The first additional information corresponding to the first audio segment includes a face recognition result and / or a voiceprint recognition result of the first audio segment and identification information of the first audio segment, where the first audio segment is one of the multiple audio segments.
42. The apparatus according to claim 37 or 38, wherein The speech includes at least one of a speech text and a speech time period.
43. A conference system, characterized in that: It includes conference terminals and conference information processing equipment, including: The conference terminal is the conference terminal according to any one of claims 23 to 36; The conference information processing device is the conference information processing device described in any one of claims 37-42.
44. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program runs on a processor, the method according to any one of claims 1 to 22 is implemented.
45. A computer program product, characterized in that When the computer program product runs on a processor, the method according to any one of claims 1 to 22 is implemented.
Citation Information
Patent Citations
Method and device for generating conference record and conference terminal
CN110232925A
Computerized intelligent assistant for conferences
US20190341050A1