Sound emitting object determination method, apparatus, computing device, and medium

By combining voiceprint recognition and sound source localization technologies, the problem of low accuracy in identifying the speaker has been solved, enabling precise voice role separation in multi-person scenarios. This technology is suitable for voice analysis in scenarios such as meetings, interrogations, interviews, and classrooms.

CN114005451BActive Publication Date: 2026-05-12ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA GROUP HOLDING LTD
Filing Date
2020-07-28
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in identifying the voice actor, making it difficult to achieve accurate voice role separation.

Method used

结合声纹识别技术和声源定位技术,通过获取目标音频帧的位置信息和声纹特征,确定发声对象的身份,实现精确的语音角色分离。

Benefits of technology

It improves the accuracy of identifying the speaker, enabling more accurate recognition of speakers in multi-person scenarios, and is suitable for speech analysis in scenarios such as meetings, interrogations, interviews, and classrooms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114005451B_ABST
    Figure CN114005451B_ABST
Patent Text Reader

Abstract

A sound emitting object determination method, device, computing equipment and medium are disclosed. The method comprises: obtaining a target audio frame emitted by a second sound emitting object and target position information of the second sound emitting object; determining that first position information corresponding to the first target audio segment does not match the target position information, and extracting a target voiceprint feature of a second target audio segment, wherein the first target audio segment comprises the first N audio frames of the target audio frame; the sound emitting object of the audio frame in the first target audio segment is the same as the sound emitting object of the audio frame in the second target audio segment; the first target audio segment is at least part of the second target audio segment; N is an integer greater than or equal to 1; and determining the target sound emitting object of the second target audio segment according to the target voiceprint feature. The accuracy of determining the sound emitting object can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to a method, apparatus, computing device, and medium for determining a sound-emitting object. Background Technology

[0002] Speech is present in many scenarios, including daily life, meetings, and telephone conversations. In practical applications, to analyze speech signals more accurately, it's necessary not only to perform speech recognition but also to perform speech role separation to determine the speaker of each part of the speech. Once the speaker is identified, a wider range of applications emerges. For example, in a large conference room setting, by separating the speech roles, meeting minutes can be quickly taken, recording the content spoken by each speaker.

[0003] Currently, most methods for identifying the speaker rely on voiceprint recognition technology, but the accuracy is relatively low. Therefore, there is an urgent need to provide a more accurate method for identifying the speaker. Summary of the Invention

[0004] This invention provides a method, apparatus, computing device, and medium for determining a sound-emitting object, which can solve the problem of low accuracy in determining the sound-emitting object.

[0005] According to a first aspect of the present invention, a method for determining a sound-producing object is provided, comprising:

[0006] Obtain the target audio frame emitted by the second sound-emitting object and the target location information of the second sound-emitting object;

[0007] If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target voiceprint features of the second target audio segment are extracted, wherein the first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frames in the first target audio segment is the same as the voice-producing object of the audio frames in the second target audio segment; the first target audio segment is at least a part of the second target audio segment; N is an integer greater than or equal to 1;

[0008] The target speaker of the second target audio segment is determined based on the target voiceprint characteristics.

[0009] According to a second aspect of the present invention, a method for determining a sound-producing object is provided, comprising:

[0010] Obtain the target audio frame emitted by the second sound-emitting object and the target location information of the second sound-emitting object;

[0011] If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target voiceprint features of the first target audio segment are extracted. The first target audio segment includes all continuous audio frames emitted by the first voice object, and the first voice object is the voice object that emitted the audio frame preceding the target audio frame. The endpoint of the first target audio segment is the audio frame preceding the target audio frame.

[0012] The source of the first target audio segment is determined based on the target voiceprint characteristics.

[0013] According to a third aspect of the present invention, a method for determining the starting point of vocal content is provided, comprising:

[0014] Obtain the target audio frame emitted by the second sound-emitting object and the target location information of the second sound-emitting object;

[0015] If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target audio frame is determined as the starting point of the sound content of the second sound-emitting object;

[0016] Wherein, the first target audio segment includes all continuous audio frames emitted by the first sound-emitting object, the first sound-emitting object being the sound-emitting object of the previous audio frame emitting the target audio frame; the end point of the first target audio segment is the previous audio frame of the target audio frame.

[0017] According to a fourth aspect of the present invention, a method for changing the identifier of a sound-emitting object is provided, comprising:

[0018] Obtain the target audio frame emitted by the second sound-emitting object and the target location information of the second sound-emitting object;

[0019] If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target voiceprint features of the second target audio segment are extracted, wherein the first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frame in the first target audio segment is the same as the sound-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment;

[0020] The target speaker of the second target audio segment is determined based on the target voiceprint characteristics.

[0021] Change the identifier of the second vocal object and the identifier of the target vocal object, wherein the identifier is used to characterize the vocal state of the vocal object.

[0022] According to a fifth aspect of the present invention, a method for generating session records is provided, comprising:

[0023] Obtain the target audio frame emitted by the second speaker in the audio session data and the target location information of the second speaker;

[0024] If it is determined that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, then the target voiceprint features of the second target audio segment in the audio session data are extracted, wherein the first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frame in the first target audio segment is the same as the voice-producing object of the audio frame in the second target audio segment; and the first target audio segment is at least a part of the second target audio segment.

[0025] The target speaker of the second target audio segment is determined based on the target voiceprint characteristics.

[0026] The target voice object is associated with the text content corresponding to the second target audio segment to obtain the conversation record of the target voice object.

[0027] According to a sixth aspect of the present invention, a sound-emitting object determination device is provided, comprising:

[0028] The acquisition module is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object;

[0029] An extraction module is used to determine if the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the second target audio segment. The first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frames in the first target audio segment is the same as the voice-producing object of the audio frames in the second target audio segment; the first target audio segment is at least a part of the second target audio segment; and N is an integer greater than or equal to 1.

[0030] The first determining module is used to determine the target sound source of the second target audio segment based on the target voiceprint features.

[0031] According to a seventh aspect of the present invention, a sound-emitting object determination device is provided, comprising:

[0032] The acquisition module is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object;

[0033] An extraction module is used to determine that the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the first target audio segment. The first target audio segment includes all continuous audio frames emitted by the first voice object, and the first voice object is the voice object that emitted the audio frame preceding the target audio frame. The end point of the first target audio segment is the audio frame preceding the target audio frame.

[0034] The first determining module is used to determine the source of the first target audio segment based on the target voiceprint features.

[0035] According to an eighth aspect of the present invention, a device for determining the starting point of vocal content is provided, comprising:

[0036] The acquisition module is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object;

[0037] The first determining module is used to determine that if the target location information does not match the first location information corresponding to the first target audio segment, then the target audio frame is determined as the starting point of the sound content of the second sound-emitting object.

[0038] Wherein, the first target audio segment includes all continuous audio frames emitted by the first sound-emitting object, the first sound-emitting object being the sound-emitting object of the previous audio frame emitting the target audio frame; the end point of the first target audio segment is the previous audio frame of the target audio frame.

[0039] According to a ninth aspect of the present invention, a device for changing the identifier of a sound-emitting object is provided, comprising:

[0040] The acquisition module is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object;

[0041] An extraction module is used to determine if the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the second target audio segment, wherein the first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frames in the first target audio segment is the same as the sound-producing object of the audio frames in the second target audio segment; and the first target audio segment is at least a part of the second target audio segment.

[0042] The first determining module is used to determine the target sound-producing object of the second target audio segment based on the target voiceprint features;

[0043] The modification module is used to modify the identifier of the second sound-emitting object and the identifier of the target sound-emitting object, wherein the identifier is used to characterize the sound-emitting state of the sound-emitting object.

[0044] According to a tenth aspect of the present invention, a session record generation apparatus is provided, comprising:

[0045] The acquisition module is used to acquire the target audio frame emitted by the second speaker in the audio session data and the target location information of the second speaker.

[0046] An extraction module is used to determine that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, and then extract the target voiceprint features of the second target audio segment in the audio session data, wherein the first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frames in the first target audio segment is the same as the voice-producing object of the audio frames in the second target audio segment; and the first target audio segment is at least a part of the second target audio segment.

[0047] The first determining module is used to determine the target sound-producing object of the second target audio segment based on the target voiceprint features;

[0048] The association module is used to associate the target voice object with the text content corresponding to the second target audio segment to obtain the conversation record of the target voice object.

[0049] According to an eleventh aspect of the present invention, a computing device is provided, comprising: a processor and a memory storing computer program instructions;

[0050] When a processor executes computer program instructions, it implements the methods provided in the first, second, third, fourth, or fifth aspects described above.

[0051] According to a twelfth aspect of the present invention, a computer storage medium is provided, on which computer program instructions are stored, which, when executed by a processor, implement the methods provided in the first, second, third, fourth, or fifth aspects described above.

[0052] According to an embodiment of the present invention, if the target location information of the second voice-emitting object emitting the target audio frame does not match the first location information corresponding to the first target audio segment, it can be determined that the first voice-emitting object emitting the first target audio segment and the second voice-emitting object of the target audio frame are different. Next, the target voiceprint features of the second target audio segment emitted by the first voice-emitting object are extracted, and the target voice-emitting object of the second target audio segment is determined based on the target voiceprint features, i.e., the identity of the first voice-emitting object is determined, thereby achieving role separation. By combining voiceprint recognition technology and sound source localization technology, the accuracy of voice-emitting object determination can be improved, thereby increasing the accuracy of voice-emitting object determination. Attached Figure Description

[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 A schematic diagram illustrating an application scenario for the sound-generating object determination method provided in the first aspect of this application;

[0055] Figure 2 A flowchart illustrating an embodiment of the sound-generating object determination method provided in the first aspect of this application;

[0056] Figure 3 A flowchart illustrating the voiceprint matching process provided in this application;

[0057] Figure 4 A flowchart illustrating the method for determining the sound-emitting object provided in the second aspect of this application;

[0058] Figure 5 A flowchart illustrating an embodiment of the method for determining the starting point of vocal content provided in the third aspect of this application;

[0059] Figure 6 A flowchart illustrating an embodiment of the method for changing the identifier of a sound-emitting object provided in the fourth aspect of this application;

[0060] Figure 7 A flowchart illustrating an embodiment of the session record generation method provided in the fifth aspect of this application;

[0061] Figure 8 A schematic diagram of the structure of an embodiment of the sound-emitting object determination device provided in the sixth aspect of this application;

[0062] Figure 9 A schematic diagram of the structure of an embodiment of the sound-emitting object determination device provided in the seventh aspect of this application;

[0063] Figure 10 A schematic diagram of the structure of an embodiment of the sound content starting point determination device provided in the eighth aspect of this application;

[0064] Figure 11 A schematic diagram of the structure of an embodiment of the sound-emitting object identification changing device provided in the ninth aspect of this application;

[0065] Figure 12 A schematic diagram of the structure of an embodiment of the session record generation apparatus provided in the tenth aspect of this application;

[0066] Figure 13 This is a schematic diagram of an embodiment of the hardware structure of the computing device provided in this application. Detailed Implementation

[0067] The features and exemplary embodiments of various aspects of the present invention will now be described in detail. To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be practiced without some of these specific details. The following description of the embodiments is merely intended to provide a better understanding of the present invention by illustrating examples of the invention.

[0068] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0069] Figure 1 This is a schematic diagram illustrating an application scenario of the sound-generating object determination method provided in this application. For example, Figure 1 The image shows a meeting room scene. The meeting room contains multiple participants. Figure 1 The diagram only shows three participants: participant A, participant B, and participant C. The total number of participants is not limited. To facilitate rapid extraction of audio content during the meeting, the voices of multiple participants can be role-based throughout the meeting.

[0070] During the process of participants emitting audio signals, the audio acquisition device preset in the conference room can collect these audio signals. After the audio acquisition device acquires the audio signals emitted by the participants, it can process the audio signals emitted by the participants into frames according to the preset duration, thereby obtaining the audio frames emitted by the participants.

[0071] In the embodiments of this application, the audio frame currently acquired by the audio collector is used as the target audio frame.

[0072] It should be noted that, in the embodiments of this application, in order to accurately determine the sound-emitting object, it is also necessary to obtain the location information of the sound-emitting object that emits the target audio frame. For each target audio frame, the location information of the sound-emitting object that emits the target audio frame can be determined by sound source localization technology. For example, an audio acquisition device pre-installed in a conference room can be used to obtain the location information of the sound-emitting object that emits the target audio frame.

[0073] Sound source localization technology refers to the technology of acquiring the location information of a sound source. In some embodiments, the sound source location information can be the relative position information between the sound-emitting object and a preset audio acquisition device. For example, the sound source location information can include the angle between the sound-emitting object and the preset audio acquisition device.

[0074] In some embodiments, the preset audio acquisition device can be a microphone array, and the sound source localization technology can be used to locate the sound source using the microphone array. The microphone array consists of several to thousands of microphones arranged according to certain rules. After receiving the audio signal through the microphone array, a time delay estimation method is used to locate the sound source. Specifically, the microphone array receives the audio signal, calculates the time delay of the audio signal received by each microphone relative to the audio signal received at a reference point, and completes the sound source localization based on the calculated time delay.

[0075] For example, during the meeting's time period 0 to t1, participant B speaks. During the meeting's time period t1 to t2, participant A speaks, and during the meeting's time period t2 to t3, participant C speaks. Time t3 is later than time t2, and time t2 is later than time t1. The audio acquisition device in the meeting room will collect the audio frames emitted by the participants in real time.

[0076] In the embodiments of this application, after the audio acquisition device in the conference room obtains the first audio frame and the location information D1 of the sound-emitting object that emitted the first audio frame, the first audio frame is taken as the starting point of the sound content emitted by the first sound-emitting object.

[0077] Next, the audio acquisition unit continues to acquire the second audio frame and the location information D2 of the object emitting the second audio frame. When the target audio frame is the second audio frame, the first target audio segment includes the first audio frame. Then, the location information D2 of the second audio frame is matched with the location information D1 of the first audio frame. If the location information D2 matches the location information D1, it can be determined that the object emitting both the first and second audio frames is the first object, and then the acquisition continues with the third audio frame and the location information D3 of the object emitting the third audio frame.

[0078] When the third audio frame is the target audio frame, for example, the first target audio segment may include the first audio frame and the second audio frame. Since the first audio frame and the second audio frame have the same sound source, the second target audio segment may include the first audio frame and the second audio frame.

[0079] Next, the position information D3 of the third audio frame is matched with the first position information corresponding to the first target audio segment. For example, the first position information corresponding to the first target audio segment can be the average D' of the position information of the sound source emitting the first audio frame and the position information of the sound source emitting the second audio frame. If the position information D3 of the third audio frame matches the position information D', it can be determined that the sound source of the first to third audio frames is the first sound source.

[0080] Similarly, following the above method, we can determine which audio frames have the first sound source. Assuming that the sound source of audio frames 1 through M1 has been determined to be the first sound source using the above method, we continue to acquire the (M1+1)th audio frame, i.e., the target audio frame, and the target location information D of the sound source emitting this target audio frame. M1+1 .

[0081] When the target audio frame is the (M1+1)th audio frame, the first target audio segment can include the N audio frames preceding the target audio frame, where N is a positive integer greater than or equal to 1. Therefore, the first target audio segment can include audio frames from the (M1-N)th audio frame to the M1th audio frame, where M1 is a positive integer. It should be noted that if the number of audio frames acquired before the target audio frame is less than N, then the first target audio segment includes all audio frames preceding the target audio frame.

[0082] When the target audio frame is the (M1+1)th audio frame, the second target audio segment can be all or part of the audio frames emitted by the sound source of the first target audio segment. Furthermore, the second target audio segment includes the first target audio segment. For example, the second target audio segment includes all audio frames between the first and M1th audio frames.

[0083] As an example, M1 = 1000, N = 500. If the 1001st audio frame is the target audio frame, the first target audio segment includes the audio frames between the 500th and 1000th audio frames, and the second target audio segment includes the audio frames between the 1st and 1000th audio frames.

[0084] For example, the first location information corresponding to the first target audio segment can be the average value D of the location information of the speaker in each audio frame between the 500th and 1000th audio frames. Then, the location information D of the participant emitting the 1001st audio frame will be...1001 Match it with location information D". If location information D 1001 If the location information "D" does not match, it can be determined that the speaker emitting the 1001st audio frame is different from the first speaker emitting the first through 1000th audio frames, indicating a change in the speakers in the conference room. Therefore, the 1000th audio frame can be taken as the end point of the first speaker's speech, and the 1001st audio frame as the starting point of the second speaker's speech.

[0085] Next, the audio frames from the start to the end of the first speaker's speech content can be extracted, i.e., the first audio frame to the 1000th audio frame, which is the second target audio segment. Then, the target voiceprint features of the second target audio segment are extracted, and based on these features, the target speaker corresponding to the first audio frame to the 1000th audio frame is determined, i.e., the first speaker. Since participant B spoke first, based on the target voiceprint features of the first audio frame to the 1000th audio frame, the speaker corresponding to the first audio frame to the 1000th audio frame can be identified as participant B.

[0086] Among these, identifying the speaker using voiceprint features utilizes voiceprint recognition technology. Voiceprint recognition technology refers to the technique of identifying the identity of a speaker by analyzing their voiceprint characteristics. A voiceprint is the sound wave spectrum that carries speech information.

[0087] Next, the next target audio frame is acquired. Following the method described above, the endpoint of the second speaker can be determined. For example, if the speakers in audio frames 1001 to 2000 are all the second speaker, and the target location information of the speaker in audio frame 2001 does not match the average location information of the speaker in each audio frame from 1500 to 2000, then audio frame 2000 is the endpoint of the second speaker's speech content. By extracting the target voiceprint features from audio frames 1001 to 2000, the target speaker in audio frames 1001 to 2000 can be determined based on these features. Since participant A is the second speaker, the speaker in audio frames 1001 to 2000 is participant A. Similarly, the audio frames emitted by participant C can also be determined.

[0088] In the embodiments of this application, by combining voiceprint recognition technology and sound source localization technology, it is possible to accurately determine the voice object corresponding to the voice content.

[0089] It should be noted that the method for determining the sound source provided in the embodiments of this specification can be applied not only to the scenario of determining the sound source in the conference room mentioned above, but also to other scenarios, such as interrogation scenarios, interview scenarios, classroom scenarios, and other different scenarios. Here, we will only use the above-mentioned conference room scenario as an example for explanation.

[0090] Based on the application scenarios mentioned above, the following will combine... Figure 2 The method for determining the sound-producing object provided in the embodiments of this application will be described in detail.

[0091] Figure 2 This is a flowchart illustrating the method for determining the sound-emitting object provided in the first aspect of this application.

[0092] like Figure 2 As shown, the sound-generating object determination method 200 provided in this application embodiment includes:

[0093] Step 210: Obtain the target audio frame emitted by the second sound source and the target location information of the second sound source;

[0094] Step 220: If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then extract the target voiceprint features of the second target audio segment. The first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frame in the first target audio segment is the same as the sound-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment; N is an integer greater than or equal to 1.

[0095] Step 230: Determine the target speaker of the second target audio segment based on the target voiceprint characteristics.

[0096] First, we will introduce the specific implementation method of step 210.

[0097] In embodiments of this application, the voice data stream includes a time-sequential series of sample point values. These sample point values ​​are obtained by sampling the original analog sound signal at a specific audio sampling rate. A series of sample point values ​​can describe the sound. The audio sampling rate is the number of sample points collected per second, measured in Hertz (Hz). A higher audio sampling rate describes a higher frequency of sound waves. An audio frame includes a time-sequential, fixed number of sample point values.

[0098] Once the sound source emits an audio signal, a pre-defined audio acquisition device can capture that signal. After acquiring the audio signal from the sound source, the acquired target audio signal can be segmented into frames according to a preset duration, dividing the target audio signal into several frames to obtain each audio frame emitted by the sound source. In the embodiments of this application, the current audio frame acquired by the audio acquisition device is determined as the target audio frame.

[0099] In the embodiments of this application, sound source localization technology can be used to obtain the location information of the sound-emitting object that emits the target audio frame. The description of sound source localization technology is as described above and will not be repeated here.

[0100] As an example, the audio acquisition device is a microphone array. When a sound-emitting object emits an audio signal, the microphones in the array can pick it up. To process the audio signal in real time, after acquiring the audio signals from each microphone array, the target audio signals can be segmented into frames according to a preset duration. The currently obtained audio frame is then used as the target audio frame. Since the distance between each microphone in the array and the sound-emitting object is generally different, the time it takes for each microphone to receive the target audio frame is also different. Therefore, based on the time difference between the reception of the corresponding target audio frames by each microphone, the target location information of the sound-emitting object can be calculated.

[0101] The specific implementation method of step 220 is described below.

[0102] In the embodiments of this application, if the sound-emitting object is switched from sound-emitting object A to sound-emitting object B, since the positions of sound-emitting object A and sound-emitting object B are different, the position information of sound-emitting object A is different from that of sound-emitting object B. In order to accurately determine whether the sound-emitting object has switched, it is necessary to determine whether the target position information of the second sound-emitting object that emits the target audio frame matches the position information of the first sound-emitting object in the previous audio frame that emits the target audio frame.

[0103] Since the audio frame emitted by the first sound-emitting object in the preceding audio frame may include multiple frames, in order to improve the accuracy of determining whether the sound-emitting object has switched, the first position information corresponding to the first target audio segment, which includes the first N audio frames of the target audio frame, can be matched with the target position information.

[0104] The first target audio segment comprises the N audio frames preceding the target audio frame, and the sound-producing object is the same for each audio frame in the first target audio segment. In other words, the sound-producing object for each audio frame in the first target audio segment is the first sound-producing object of the preceding audio frame.

[0105] It should be noted that if the number of audio frames emitted by the first sound-emitting object obtained before the target audio frame is less than N, then all consecutive audio frames emitted by the first sound-emitting object will be taken as the first target audio segment, and the end point of the first target audio segment will be the frame preceding the target audio frame.

[0106] In some embodiments, the first location information corresponding to the first target audio segment is determined based on the location information of the sound-producing object corresponding to each audio frame in the first target audio segment. For example, the first location information is determined based on the average of the location information of the sound-producing object corresponding to each audio frame in the first target audio segment.

[0107] For example, the position information of the sound-emitting object is the angle between the sound-emitting object and the preset microphone array. Then, the first position information is the first angle corresponding to the first target audio segment, and the first angle is the average value of the angle between the sound-emitting object and the preset microphone array corresponding to each audio frame in the first target audio segment.

[0108] In the embodiments of this application, whether the target location information matches the first location information can be determined by whether the difference between the target location information and the first location information is within a preset numerical range. If the difference is within the preset numerical range, it means that the target location information matches the first location information; if the difference exceeds the preset numerical range, it means that the target location information does not match the first location information.

[0109] The matching degree between the target location information and the first location information can be characterized by the difference between the target location information and the first location information. That is, the matching degree between the target location information and the first location information is the difference between the target location information and the first location information.

[0110] In some other embodiments of this application, the matching degree between the target location information and the first location information can be characterized by the ratio of the difference between the target location information and the first location information to the first location information. If the ratio of the difference between the target location information and the first location information to the first location information is within a preset ratio range, it means that the target location information and the first location information match. If the ratio of the difference between the target location information and the first location information to the first location information is not within the preset ratio range, it means that the target location information and the first location information do not match.

[0111] No restrictions are placed on the specific implementation method for determining whether the target location information matches the first location information.

[0112] In the embodiments of this application, if the target location information does not match the first location information, it means that the second voice-emitting object that emits the target audio frame is different from the voice-emitting object of the first target audio segment, which means that the voice-emitting object may have switched. Then, voiceprint recognition technology can be further used to determine the voice-emitting object of the first target audio segment.

[0113] Since the first target audio segment may only be a part of the audio signal emitted by the first speaker (i.e., the speaker of the audio frame preceding the target audio frame), when performing voice stream role separation, it is necessary to perform role separation on the second target audio segment emitted by the first speaker.

[0114] As an example, the second target audio segment includes the first target audio segment, and the sound-producing object of the audio frames in the second target audio segment is the same as the sound-producing object of the audio frames in the first target audio segment. That is, the sound-producing object of each audio frame in the second target audio segment is the same as the sound-producing object of each audio frame in the first target audio segment, i.e., the first sound-producing object.

[0115] In some embodiments, by checking whether the target location information matches the first location information, it can be determined whether the sound-producing object of the target audio frame and the sound-producing object of the first target audio segment might be the same. If the target location information does not match the first location information, the target audio frame can be considered as the starting point of the sound content of the second sound-producing object, i.e., the first audio frame emitted, and the audio frame preceding the target audio frame is the sound-producing object of the first target audio segment, i.e., the ending point of the sound content of the first sound-producing object, i.e., the last audio frame emitted. Therefore, the sound-producing object of the first target audio segment, i.e., the starting point of the sound content of the first sound-producing object, i.e., the first audio frame emitted by the first sound-producing object, can also be obtained in advance using a similar method.

[0116] If the target location information does not match the first location information, then at least a portion of the consecutive audio frames between the first audio frame emitted by the first sound-emitting object and the preceding audio frame of the target audio frame can be determined as the second target audio segment.

[0117] For example, all audio frames between the first audio frame emitted by the sound-emitting object of the first target audio segment and the audio frame preceding the target audio frame can be identified as the second target audio segment. In other words, the start and end points of the identified second target audio segment are the points where the positional information mismatches between the two consecutive segments. If the positional information of the sound-emitting object is the angle between the sound-emitting object and the microphone array, then the start and end points of the identified second target audio segment are the points where the angle changes between the two consecutive segments.

[0118] It should be noted that both the first target audio segment and the second target audio segment consist of multiple consecutive audio frames.

[0119] When a change in location information is detected, the audio segment between the two points where the location information changes is determined as the second target audio segment, and then the target voiceprint features of the second target audio segment are extracted.

[0120] In the embodiments of this application, if the target location information matches the first location information, it means that the second sound-emitting object is the sound-emitting object of the first target audio segment, that is, the sound-emitting object corresponding to the first target audio segment and the sound-emitting object corresponding to the target audio frame are the same. Then, the next audio frame is obtained again, and the next audio frame is used as the target audio frame. The target location information of the second sound-emitting object that emits the target audio frame is obtained, that is, the process returns to step 210.

[0121] It should be noted that after the target audio frame is reacquired, the first target audio segment in step 220 will be updated, that is, the first target audio segment will include the previous target audio frame, and the first position information corresponding to the first target audio segment will also be updated accordingly.

[0122] The specific implementation method of step 230 is described below.

[0123] In the embodiments of this application, a sound database is pre-established, which includes the correspondence between the sound-producing object and the voiceprint feature, as well as the correspondence between the sound-producing object and the audio signal.

[0124] After obtaining the target voiceprint features of the second target audio segment, in order to determine the speaker of the second target audio segment, it is necessary to match the target voiceprint features with each voiceprint feature in the preset sound database, and determine the speaker corresponding to the voiceprint feature that matches the target voiceprint feature in the preset sound database as the target speaker of the second target audio segment. Furthermore, the second target audio segment can be determined as the audio signal corresponding to the target speaker, that is, the second target audio segment can be added as the audio signal corresponding to the target speaker.

[0125] In some embodiments of this application, step 230 includes: if a first voiceprint feature exists in the preset sound database and its matching degree with the target voiceprint feature satisfies a first preset matching condition, determining the voice-emitting object corresponding to the first voiceprint feature as the target voice-emitting object corresponding to the second target audio segment; if the matching degree between the first voiceprint feature and the target voiceprint feature satisfies a second preset matching condition, updating the voiceprint feature corresponding to the target voice-emitting object in the preset sound database using the target voiceprint feature.

[0126] Among them, the matching degree required to be met by the second preset matching condition is greater than the matching degree required to be met by the first preset matching condition.

[0127] In the embodiments of this application, the target voiceprint feature is used to update the voiceprint feature corresponding to the target voice-emitting object in the preset sound database only when the matching degree between the first voiceprint feature and the target voiceprint feature meets the second preset matching condition. This can improve the richness and accuracy of the voiceprint feature corresponding to the target voice-emitting object, thereby improving the accuracy of voiceprint recognition.

[0128] As an example, the first preset matching condition is that the matching degree between the voiceprint features in the preset voice database and the target voiceprint features is greater than 80%, and the second preset matching condition is that the matching degree between the voiceprint features in the preset voice database and the target voiceprint features is greater than 90%.

[0129] In the embodiments of this application, when the matching degree between the first voiceprint feature and the target voiceprint feature does not meet the second preset matching condition, the target voiceprint feature is not used to update the corresponding voiceprint feature of the target voice-emitting object in the preset sound database, and only the second target audio segment is determined as the audio signal emitted by the target voice-emitting object.

[0130] In the embodiments of this application, if there are multiple first voiceprint features, the voice-producing object corresponding to the first voiceprint feature with the highest matching degree with the target voiceprint feature is determined as the voice-producing object corresponding to the second target audio segment.

[0131] In some embodiments of this application, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold but greater than a second preset location matching degree threshold, the first preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the first preset voiceprint matching degree threshold; the second preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the second preset voiceprint matching degree threshold, wherein the second preset voiceprint matching degree threshold is greater than the first preset voiceprint matching degree threshold.

[0132] When the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the first preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the third preset voiceprint matching degree threshold; the second preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the fourth preset voiceprint matching degree threshold.

[0133] Among them, the fourth preset voiceprint matching threshold is greater than the third preset voiceprint matching threshold, the first preset voiceprint matching threshold is less than the second preset voiceprint matching threshold, the second preset voiceprint matching threshold is less than the fourth preset voiceprint matching threshold, and the first preset voiceprint matching threshold is less than the third preset voiceprint matching threshold.

[0134] In the embodiments of this application, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold but greater than a second preset location matching degree threshold, it represents that the target location information and the first location information do not match, but are relatively similar. When the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, it represents that the target location information and the first location information do not match, and the target location information and the first location information differ significantly.

[0135] When the target location information is similar to the first location information, it means that the second voice source may be the same as the voice source of the first target audio segment. In other words, the voice source of the target audio frame and the voice source of the second target audio segment may be the same person. Therefore, the voiceprint matching threshold can be set slightly lower. When the target location information differs significantly from the first location information, it means that the second voice source may be different from the voice source of the first target audio segment. In other words, the voice source of the target audio frame and the voice source of the second target audio segment may not be the same person. Therefore, the voiceprint matching threshold can be set slightly higher. That is, the first preset voiceprint matching threshold is lower than the third preset voiceprint matching threshold, and the second preset voiceprint matching threshold is lower than the fourth preset voiceprint matching threshold. By setting it this way, the accuracy of voice source identification can be improved.

[0136] In some embodiments of this application, the method for determining the sound source provided in this application further includes: when the matching degree between each voiceprint feature in the preset sound database and the target voiceprint feature satisfies a third preset matching condition, the correspondence between the target voiceprint feature and the sound source corresponding to the target voiceprint feature is stored in the preset sound database, and the sound source corresponding to the target voiceprint feature is determined as the target sound source of the second target audio segment; wherein, the third preset matching condition is used to characterize the mismatch between the voiceprint feature in the preset sound database and the target voiceprint feature.

[0137] In the embodiments of this application, when each voiceprint feature in the preset sound database does not match the target voiceprint feature, it means that the voice-producing object corresponding to each voiceprint feature in the preset sound database is not the voice-producing object corresponding to the second target audio segment.

[0138] In some embodiments, the voice-generating object corresponding to the target voiceprint feature can be determined from a pre-established correspondence between voiceprint features and voice-generating objects. Then, the correspondence between the target voiceprint feature and its corresponding voice-generating object is updated to a preset sound database. In other words, the voice-generating object corresponding to the target voiceprint feature is registered in the preset sound database.

[0139] In some other embodiments of this application, the method for determining the sound source provided in this application further includes: discarding the second target audio segment when the matching degree between each voiceprint feature in the preset sound database and the target voiceprint feature satisfies the fourth preset matching condition but not the third preset matching condition.

[0140] The fourth preset matching condition is also used to characterize the mismatch between the voiceprint features in the preset sound database and the target voiceprint features. However, the degree of mismatch that the fourth preset matching condition needs to satisfy is less than the degree of mismatch that the third preset matching condition needs to satisfy.

[0141] In other words, if the matching degree between each voiceprint feature in the preset sound database and the target voiceprint feature meets the fourth preset matching condition but does not meet the third preset matching condition, then in order to improve the accuracy of subsequent determination of the voice source, the target voiceprint feature of the second target audio segment and the voice source corresponding to the target voiceprint feature will not be registered.

[0142] In some embodiments of this application, when the matching degree between the target location information and the first location information is less than the first preset location matching degree threshold but greater than the second preset location matching degree threshold, the third preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is less than the fifth preset voiceprint matching degree threshold.

[0143] If the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the third preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is less than the sixth preset voiceprint matching degree threshold.

[0144] Among them, the fifth preset voiceprint matching threshold is less than the sixth preset voiceprint matching threshold.

[0145] In some embodiments of this application, when the matching degree between the target location information and the first location information is less than the first preset location matching degree threshold but greater than the second preset location matching degree threshold, the fourth preset matching condition is that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is less than the seventh preset voiceprint matching degree threshold.

[0146] If the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the fourth preset matching condition is that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is less than the eighth preset voiceprint matching degree threshold.

[0147] Among them, the seventh preset voiceprint matching threshold is less than the eighth preset voiceprint matching threshold.

[0148] When the target location information is similar to the first location information, it means that the second voice source may be the same as the voice source corresponding to the first target audio segment. That is, the voice source of the target audio frame and the voice source of the second target audio segment may be the same person. Therefore, the voiceprint mismatch threshold can be set slightly lower to improve the accuracy of voice source identification. When the target location information differs significantly from the first location information, it means that the second voice source may be different from the voice source corresponding to the first target audio segment. That is, the voice source of the target audio frame and the voice source of the second target audio segment may not be the same person. Therefore, the voiceprint mismatch threshold can be set slightly higher. In other words, the fifth preset voiceprint matching threshold is lower than the sixth preset voiceprint matching threshold, and the seventh preset voiceprint matching threshold is lower than the eighth preset voiceprint matching threshold. This setting improves the accuracy of voice source identification.

[0149] In other words, when the target location does not match the first location information, a mismatch between the target location information and the first location information can be categorized into two degrees: a mismatch less than a first preset location matching threshold but greater than a second preset location matching threshold, and a mismatch less than a second preset location matching threshold. Specifically, a mismatch less than the first preset location matching threshold but greater than the second preset location matching threshold indicates a slightly lower degree of mismatch, while a mismatch less than the second preset location matching threshold indicates a slightly higher degree of mismatch.

[0150] For both slightly low and slightly high mismatches between the target location information and the first location information, four preset voiceprint matching thresholds are used for voiceprint matching. When the mismatch is slightly low, indicating a close approximation, the four thresholds are: the first preset voiceprint matching threshold, the second preset voiceprint matching threshold, the fifth preset voiceprint matching threshold, and the seventh preset voiceprint matching threshold. When the mismatch is slightly high, indicating a significant difference, the four preset voiceprint matching thresholds are: the third preset voiceprint matching threshold, the fourth preset voiceprint matching threshold, the sixth preset voiceprint matching threshold, and the eighth preset voiceprint matching threshold. Furthermore, the four preset matching thresholds for slightly low mismatches are all lower than the four preset matching thresholds for the respective low mismatches. In other words, the first preset voiceprint matching threshold is less than the third preset voiceprint matching threshold, the second preset voiceprint matching threshold is less than the fourth preset voiceprint matching threshold, the fifth preset voiceprint matching threshold is less than the sixth preset voiceprint matching threshold, and the seventh preset voiceprint matching threshold is less than the eighth preset voiceprint matching threshold.

[0151] It should be noted that the seventh preset voiceprint matching threshold is less than the first preset voiceprint matching threshold, and the eighth preset voiceprint matching threshold is less than the third preset voiceprint matching threshold.

[0152] Figure 3 This diagram illustrates the voiceprint matching process provided in an embodiment of this application. For example, if the position information of the voice-emitting object is the angle between it and the microphone array, then the target position information is the target angle between the second voice-emitting object and the microphone array. The first position information is the first angle corresponding to the first target audio segment.

[0153] See Figure 3 When the matching degree between the target angle and the first angle is less than the first preset position matching degree threshold and greater than the second preset position matching degree threshold, that is, the target angle is similar to the first angle, a lower preset voiceprint matching degree threshold is used for voiceprint comparison.

[0154] In other words, if there is a first voiceprint feature in the preset voice database that matches the target voiceprint feature with a degree greater than the second preset voiceprint matching threshold, then the probability that the voice-producing object corresponding to the first voiceprint feature is the same as the voice-producing object of the second target audio segment is very high. Therefore, the voiceprint feature corresponding to the target voice-producing object is updated using the target voiceprint feature.

[0155] If a first voiceprint feature exists in the preset voice database with a matching degree greater than the first preset voiceprint matching degree threshold and less than the second preset voiceprint matching degree threshold, then the probability that the voice-producing object corresponding to the first voiceprint feature is the same as the voice-producing object of the second target audio segment is relatively high, but lower than the above-mentioned case where the matching degree is greater than the second preset voiceprint matching degree threshold. Therefore, the voiceprint feature corresponding to the target voice-producing object is not updated using the target voiceprint feature.

[0156] If the matching degree between each voiceprint feature in the preset voice database and the target voiceprint feature is less than the fifth preset voiceprint matching degree threshold, then the probability that the voice-producing object corresponding to each voiceprint feature in the preset voice database is not the voice-producing object corresponding to the target voiceprint feature is very high. Therefore, the correspondence between the target voiceprint feature and the voice-producing object corresponding to the target voiceprint feature is stored in the preset voice database, that is, the target voiceprint feature and the voice-producing object corresponding to the target voiceprint feature are registered.

[0157] If the matching degree between each voiceprint feature in the preset sound database and the target voiceprint feature is less than the seventh preset voiceprint matching degree threshold and greater than the fifth preset voiceprint matching degree threshold, then the probability that the voice-producing object corresponding to each voiceprint feature in the preset sound database is not the voice-producing object corresponding to the target voiceprint feature is relatively high, but lower than the case where the matching degree is less than the fifth preset voiceprint matching degree threshold. Therefore, the second target audio segment is discarded, and the target voiceprint feature and the voice-producing object corresponding to the target voiceprint feature are not registered.

[0158] See also Figure 3 When the matching degree between the target angle and the first angle is less than the second preset position matching degree threshold, that is, the target angle is different from the first angle, a higher preset voiceprint matching degree threshold is used for voiceprint comparison.

[0159] In other words, if there is a first voiceprint feature in the preset sound database that matches the target voiceprint feature with a degree greater than the fourth preset voiceprint matching threshold, then the probability that the voice-producing object corresponding to the first voiceprint feature is the same as the voice-producing object of the second target audio segment is very high. Therefore, the voiceprint feature corresponding to the target voice-producing object is updated using the target voiceprint feature.

[0160] If a first voiceprint feature exists in the preset voice database with a matching degree greater than the third preset voiceprint matching degree threshold and less than the fourth preset voiceprint matching degree threshold, then the probability that the voice-producing object corresponding to the first voiceprint feature is the same as the voice-producing object of the second target audio segment is relatively high, but lower than the above-mentioned case where the matching degree is greater than the fourth preset voiceprint matching degree threshold. Therefore, the voiceprint feature corresponding to the target voice-producing object is not updated using the target voiceprint feature.

[0161] If the matching degree between each voiceprint feature in the preset voice database and the target voiceprint feature is less than the eighth preset matching degree threshold, then the probability that the voice-producing object corresponding to each voiceprint feature in the preset voice database is not the voice-producing object corresponding to the target voiceprint feature is very high. Therefore, the correspondence between the target voiceprint feature and the voice-producing object corresponding to the target voiceprint feature is stored in the preset voice database, that is, the target voiceprint feature and the voice-producing object corresponding to the target voiceprint feature are registered.

[0162] If the matching degree between each voiceprint feature in the preset sound database and the target voiceprint feature is less than the tenth preset matching degree threshold and greater than the eighth preset matching degree threshold, then the probability that the voice-producing object corresponding to each voiceprint feature in the preset sound database is not the voice-producing object corresponding to the target voiceprint feature is relatively high, but lower than the case where the matching degree is less than the sixth preset voiceprint matching degree threshold. Therefore, the second target audio segment is discarded, and the target voiceprint feature and the voice-producing object corresponding to the target voiceprint feature are not registered.

[0163] In the embodiments of this application, by combining the matching degree between the target location information and the first location information, and using two sets of preset voiceprint matching degree thresholds for voiceprint matching, the accuracy of determining the voice source can be achieved.

[0164] In some embodiments of this application, in order to improve the accuracy of determining the sound source, before step 220, the sound source determination method provided in this application further includes: filtering the position information of the sound source corresponding to the audio frame in the first target audio segment to obtain filtered position information; and determining the first position information based on the filtered position information.

[0165] In some embodiments, the position information of the sound-producing object corresponding to each audio frame in the first target audio segment can be filtered by a median filter to obtain the filtered position information of the sound-producing object corresponding to each audio frame.

[0166] The idea behind median filtering is that the location information of the sound-producing object corresponding to each audio frame can be replaced by the statistical median of the location information of the sound-producing objects corresponding to all audio frames within a neighborhood of a preset size.

[0167] As an example, if the position information of the sound-emitting object is the angle between the sound source and the microphone array, then the first position information is the average value of the angle between the sound-emitting object and the microphone array for each audio frame after filtering.

[0168] In the embodiments of this application, by filtering the position information of the sound-producing object corresponding to each audio frame in the first target audio segment, some noise and glitches can be filtered out, making the position information more stable and obtaining smooth position information, thereby improving the accuracy of determining the sound-producing object.

[0169] In some other embodiments of this application, other methods can be used to filter the position information of the sound-producing object corresponding to each audio frame in the first target audio segment, such as using a mean filter.

[0170] Figure 4 This diagram illustrates a flowchart of the method for determining the sound-emitting object provided in the second aspect of this application. Figure 4 As shown, the second aspect of this application provides a method 400 for determining the sound source, which includes:

[0171] Step 410: Obtain the target audio frame emitted by the second sound source and the target location information of the second sound source;

[0172] Step 420: If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then extract the target voiceprint features of the first target audio segment. The first target audio segment includes all continuous audio frames emitted by the first voice object, and the first voice object is the voice object that emitted the audio frame preceding the target audio frame. The end point of the first target audio segment is the audio frame preceding the target audio frame.

[0173] Step 430: Determine the target speaker of the first target audio segment based on the target voiceprint characteristics.

[0174] In the embodiments of this application, the specific implementation of step 410 is similar to that of step 210, and will not be described again here.

[0175] In the embodiments of this application, the specific implementation of step 420 is similar to that of step 220. However, the difference between step 420 and step 220 is that the first target audio segment includes all consecutive audio frames emitted by the first sound-emitting object, the first sound-emitting object being the sound-emitting object of the audio frame preceding the target audio frame; and the endpoint of the first target audio segment is the audio frame preceding the target audio frame.

[0176] In step 220, the first target audio segment includes the first N audio frames emitted by the first sound-emitting object before the target audio frame, which is not necessarily all the consecutive audio frames emitted by the first sound-emitting object.

[0177] In the embodiments of this application, since the first position information corresponding to all consecutive audio frames emitted by the first voice-emitting object includes richer position information, the first position information can more accurately reflect the position information of the first voice-emitting object. Therefore, matching the first position information with the target position information can more accurately determine whether the voice-emitting object has switched, thereby improving the accuracy of character separation.

[0178] In the embodiments of this application, the specific implementation of step 430 is similar to that of step 230. The identity of the first voice-emitting object can be determined based on the target voiceprint features of the first target audio segment, which will not be described in detail here.

[0179] In an embodiment of the present invention, if the target location information of the second voice-emitting object emitting the target audio frame does not match the first location information corresponding to the first target audio segment, it can be determined that the first voice-emitting object emitting the first target audio segment and the second voice-emitting object of the target audio frame are different. Next, the target voiceprint features of the first target audio segment are extracted, and the target voice-emitting object of the second target audio segment is determined based on the target voiceprint features, i.e., the identity of the first voice-emitting object is determined, thereby achieving role separation. By combining voiceprint recognition technology and sound source localization technology, the accuracy of voice-emitting object determination can be improved, thereby increasing the accuracy of voice-emitting object determination.

[0180] In the embodiments of this application, the specific implementation of the sound-emitting object determination method provided by the second aspect is similar to the specific implementation of the sound-emitting object method provided by the first aspect, and will not be described again here.

[0181] In the embodiments of this application, if role separation is to be achieved, it is necessary to determine the start and end points of the voice content of each voice-speaking object in order to separate the voice content of each voice-speaking object and thus achieve role separation. Therefore, this application provides a method for determining the start point of voice content. Figure 5 This diagram illustrates a flowchart of the method for determining the starting point of vocal content provided in the third aspect of this application. Figure 5 As shown, the method 500 for determining the starting point of vocal content provided in the third aspect of this application includes:

[0182] Step 510: Obtain the target audio frame emitted by the second sound source and the target location information of the second sound source;

[0183] Step 520: If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target audio frame is determined as the starting point of the sound content of the second sound object.

[0184] The first target audio segment includes all consecutive audio frames emitted by the first sound-emitting object, which is the sound-emitting object of the audio frame preceding the target audio frame; the end point of the first target audio segment is the audio frame preceding the target audio frame.

[0185] In the embodiments of this application, the specific implementation of step 510 is similar to that of step 210, and will not be described again here.

[0186] In step 520, if it is determined that the target location information does not match the first location information corresponding to the first target audio segment, it is determined that the sound-emitting object has switched, that is, the sound-emitting object has switched from the first sound-emitting object to the second sound-emitting object. Therefore, the audio frame preceding the target audio frame can be used as the end point of the sound content of the first sound-emitting object, and the target audio frame can be used as the starting point of the sound content of the second sound-emitting object, so as to extract all the sound content of the second sound-emitting object in the subsequent process.

[0187] In the embodiments of this application, by matching the target position information of the voice-emitting object of the target audio frame with the first position information corresponding to all consecutive audio frames of the first voice-emitting object, it can be determined whether the voice-emitting object has switched, thereby determining the start and end points of the voice content of each voice-emitting object, and thus realizing the determination of the voice content of each voice-emitting object and achieving role separation.

[0188] In some scenarios, different individuals may speak, such as in a meeting room where different participants may speak. To improve meeting efficiency, the current speaker's identity can be indicated to other users. Therefore, a method for changing the speaker's identifier is needed to indicate the current speaker's identity. Figure 6 This diagram illustrates a flowchart of the method for changing the identifier of a sound-emitting object provided in the fourth aspect of this application. Figure 6 As shown, the fourth aspect of this application provides a method 600 for changing the identifier of a sound-emitting object, which includes:

[0189] Step 610: Obtain the target audio frame emitted by the second sound source and the target location information of the second sound source;

[0190] Step 620: If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then extract the target voiceprint features of the second target audio segment. The first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frame in the first target audio segment is the same as the sound-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment.

[0191] Step 630: Determine the target speaker of the second target audio segment based on the target voiceprint characteristics;

[0192] Step 640: Change the identifier of the second sound-emitting object and the identifier of the target sound-emitting object. The identifier is used to characterize the sound-emitting state of the sound-emitting object.

[0193] In the embodiments of this application, the specific implementation of step 610 is similar to that of step 210, the specific implementation of step 620 is similar to that of step 220, and the specific implementation of step 630 is similar to that of step 230, and will not be described again here.

[0194] In the embodiments of this application, when it is determined that the target location information does not match the first location information corresponding to the first target audio segment, it can be determined that the sound-emitting object has been switched. That is, the sound-emitting object has switched from the target sound-emitting object of the second target audio segment to the second sound-emitting object of the target audio frame.

[0195] Therefore, the identifiers of the second and target voice objects can be changed to indicate that the voice object has switched from the target to the second voice object. The identifier of the voice object is used to characterize its voice output state.

[0196] As an example, the identifier of a sound-emitting object can be the brightness of its image. For instance, if the brightness of the sound-emitting object's image is a first preset brightness, it indicates that the sound-emitting object is currently in a sound-emitting state. If the brightness of the sound-emitting object's image is a second preset brightness, it indicates that the sound-emitting object is currently in a non-sound-emitting state.

[0197] In embodiments of this application, if the current sound-emitting object is switched from the target sound-emitting object to the second sound-emitting object, the brightness of the image of the target sound-emitting object is changed from a first preset brightness to a second preset brightness to indicate that the target sound-emitting object has stopped emitting sound. The brightness of the image of the second sound-emitting object is changed from the second preset brightness to the first preset brightness to indicate that the second sound-emitting object has started emitting sound.

[0198] In some other embodiments of this application, the identifier of the sound-emitting object can be a label for the sound-emitting object. For example, when the label of the sound-emitting object is a first preset label, it is used to indicate that the sound-emitting object is currently in a sound-emitting state. If the label of the sound-emitting object is a second preset label, it is used to indicate that the sound-emitting object is currently in a non-sound-emitting state.

[0199] In embodiments of this application, if the current sound-emitting object is switched from the target sound-emitting object to the second sound-emitting object, the label of the target sound-emitting object is changed from the first preset label to the second preset label to indicate that the target sound-emitting object has stopped emitting sound. The label of the second sound-emitting object is changed from the second preset label to the first preset label to indicate that the second sound-emitting object has started emitting sound.

[0200] In the embodiments of this application, by matching the target location information of the sound-emitting object in the target audio frame with the first location information corresponding to the first target audio segment, it can be determined whether the second sound-emitting object is the same as the sound-emitting object of the second target audio segment, that is, whether the current sound-emitting object has been switched. If it is determined that the sound-emitting object has been switched, changing the identifier of the second sound-emitting object and the identifier of the target sound-emitting object can indicate the identity of the current sound-emitting object.

[0201] In some conversational scenarios, after acquiring the audio conversation data for that scenario, it is necessary to process the audio conversation data to obtain a conversation record in order to record the content of that conversation. Therefore, this application provides a method for generating conversation records. Figure 7 This diagram illustrates a flowchart of the session record generation method provided in the fifth aspect of this application. Figure 7 As shown, the session record generation method 700 provided in the fifth aspect of this application includes:

[0202] Step 710: Obtain the target audio frame emitted by the second speaker and the target location information of the second speaker in the audio session data;

[0203] Step 720: If it is determined that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, then extract the target voiceprint features of the second target audio segment in the audio session data. The first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frame in the first target audio segment is the same as the voice-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment.

[0204] Step 730: Determine the target speaker of the second target audio segment based on the target voiceprint characteristics;

[0205] Step 740: Associate the target voice object with the text content corresponding to the second target audio segment to obtain the conversation record of the target voice object.

[0206] In the embodiments of this application, the specific implementation of step 710 is similar to that of step 210, the specific implementation of step 720 is similar to that of step 220, and the specific implementation of step 730 is similar to that of step 230, and will not be described again here.

[0207] It should be noted that an audio acquisition device can be used to collect audio session data in a conversational scenario. To facilitate the generation of session records, the location information of the sound-producing object in each audio frame of the audio session data is also obtained during the audio session data acquisition process.

[0208] In the embodiments of this application, when it is determined that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, it is determined that the sound-emitting object of the target audio frame is different from the sound-emitting object of the first target audio segment, and it is necessary to extract the sound content of the sound-emitting object of the first target audio segment. Since the sound-emitting object switches, the target audio frame can be used as the starting point of the sound content of the second sound-emitting object, and the frame preceding the target audio frame can be used as the ending point of the sound content of the sound-emitting object of the first target audio segment. Regarding the relationship between the first target audio segment and the second target audio segment, please refer to the description of the embodiments of the sound-emitting object determination method provided in the first aspect.

[0209] Once the target speaker of the second target audio segment is determined based on the target voiceprint features, the text content corresponding to the second target audio segment is associated with the target speaker to obtain the meeting minutes of the target speaker.

[0210] In the embodiments of this application, the second target audio segment corresponding to each voicer in the audio session data can be extracted by the above method, and thus the meeting record of the audio session data can be obtained.

[0211] To improve the integrity of the meeting transcript, the second target audio segment may include all consecutive audio frames emitted by the first speaker, which is the speaker of the audio frame preceding the target audio frame.

[0212] In the embodiments of this application, when it is determined that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, it is determined that the sound source of the target audio frame is different from the sound source of the first target audio segment. Therefore, the text content corresponding to the second target audio segment that is the same as the sound source of the first target audio segment can be associated with the target object to form a meeting record, so as to extract the audio session data record in the future, which improves convenience.

[0213] In the embodiments of this application, the executing entity of the sound-emitting object determination method provided in the embodiments of this application can be a sound-emitting object determination device. It should be noted that, in the embodiments of this application, the sound-emitting object determination device executing the sound-emitting object determination method is used as an example to illustrate the sound-emitting object determination device provided in the embodiments of this application.

[0214] Figure 8 A schematic diagram of the structure of the sound-generating object determination device provided for the sixth aspect. (See diagram below.) Figure 8 As shown, the sound-emitting object determining device 800 includes:

[0215] The acquisition module 810 is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object;

[0216] Extraction module 820 is used to determine that the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the second target audio segment. The first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frame in the first target audio segment is the same as the sound-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment; N is an integer greater than or equal to 1.

[0217] The first determining module 830 is used to determine the target sound source of the second target audio segment based on the target voiceprint characteristics.

[0218] According to an embodiment of the present invention, if the target location information of the second voice-emitting object emitting the target audio frame does not match the first location information corresponding to the first target audio segment, it can be determined that the first voice-emitting object emitting the first target audio segment and the second voice-emitting object of the target audio frame are different. Next, the target voiceprint features of the second target audio segment emitted by the first voice-emitting object are extracted, and the target voice-emitting object of the second target audio segment is determined based on the target voiceprint features, i.e., the identity of the first voice-emitting object is determined, thereby achieving role separation. By combining voiceprint recognition technology and sound source localization technology, the accuracy of voice-emitting object determination can be improved, thereby increasing the accuracy of voice-emitting object determination.

[0219] In some embodiments of this application, the target location information includes the relative location information between the second sound-emitting object and the preset audio collector.

[0220] In some embodiments of this application, the sound-emitting object determination device 800 further includes:

[0221] The filtering module is used to filter the position information of the sound-producing object corresponding to the audio frame in the first target audio segment to obtain the filtered position information.

[0222] The second determining module is used to determine the first position information based on the filtered position information.

[0223] In some embodiments of this application, the first determining module 830 is used for:

[0224] If a first voiceprint feature exists in the preset sound database and its matching degree with the target voiceprint feature meets the first preset matching condition, the voice-producing object corresponding to the first voiceprint feature is determined as the target voice-producing object of the second target audio segment.

[0225] If the matching degree between the first voiceprint feature and the target voiceprint feature meets the second preset matching condition, the voiceprint feature corresponding to the target voice-emitting object in the preset sound database is updated using the target voiceprint feature.

[0226] Among them, the matching degree required to be met by the second preset matching condition is greater than the matching degree required to be met by the first preset matching condition.

[0227] In some embodiments of this application, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold but greater than a second preset location matching degree threshold, the first preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the first preset voiceprint matching degree threshold; the second preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the second preset voiceprint matching degree threshold, wherein the second preset voiceprint matching degree threshold is greater than the first preset voiceprint matching degree threshold.

[0228] When the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the first preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the third preset voiceprint matching degree threshold; the second preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the fourth preset voiceprint matching degree threshold.

[0229] Among them, the fourth preset voiceprint matching threshold is greater than the third preset voiceprint matching threshold, the first preset voiceprint matching threshold is less than the second preset voiceprint matching threshold, and the second preset voiceprint matching threshold is less than the fourth preset voiceprint matching threshold.

[0230] In some embodiments of this application, the sound-emitting object determination device 400 further includes:

[0231] The processing module is used to store the correspondence between the target voiceprint feature and the corresponding voice-emitting object in the preset voice database when the matching degree between each voiceprint feature in the preset voice database and the target voiceprint feature meets the third preset matching condition, and to determine the voice-emitting object corresponding to the target voiceprint feature as the target voice-emitting object of the second target audio segment.

[0232] The third preset matching condition is used to characterize the mismatch between the voiceprint features in the preset voice database and the target voiceprint features.

[0233] In some embodiments of this application, when the matching degree between the target location information and the first location information is less than the first preset location matching degree threshold but greater than the second preset location matching degree threshold, the third preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is less than the fifth preset voiceprint matching degree threshold.

[0234] If the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the third preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is less than the sixth preset voiceprint matching degree threshold.

[0235] Among them, the fifth preset voiceprint matching threshold is less than the sixth preset voiceprint matching threshold.

[0236] Other details of the sound-emitting object determination device 800 according to embodiments of the present invention are similar to the sound-emitting object determination method provided in the first aspect above, and will not be repeated here.

[0237] Figure 9 A schematic diagram of the structure of the sound-emitting object determination device provided for the seventh aspect. (See diagram below.) Figure 9 As shown, the sound-emitting object determining device 900 includes:

[0238] The acquisition module 910 is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object;

[0239] Extraction module 920 is used to determine that the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the first target audio segment. The first target audio segment includes all continuous audio frames emitted by the first voice object, and the first voice object is the voice object that emitted the audio frame preceding the target audio frame; the end point of the first target audio segment is the audio frame preceding the target audio frame.

[0240] The first determining module 930 is used to determine the target sound source of the first target audio segment based on the target voiceprint characteristics.

[0241] In an embodiment of the present invention, if the target location information of the second voice-emitting object emitting the target audio frame does not match the first location information corresponding to the first target audio segment, it can be determined that the first voice-emitting object emitting the first target audio segment and the second voice-emitting object of the target audio frame are different. Next, the target voiceprint features of the first target audio segment are extracted, and the target voice-emitting object of the second target audio segment is determined based on the target voiceprint features, i.e., the identity of the first voice-emitting object is determined, thereby achieving role separation. By combining voiceprint recognition technology and sound source localization technology, the accuracy of voice-emitting object determination can be improved, thereby increasing the accuracy of voice-emitting object determination.

[0242] Other details of the sound-emitting object determination device 900 according to embodiments of the present invention are similar to the sound-emitting object determination method provided in the second aspect above, and will not be repeated here.

[0243] Figure 10 A schematic diagram of the device for determining the starting point of the sound content provided for the eighth aspect. (For example...) Figure 10 As shown, the sound content starting point determining device 1000 includes:

[0244] The acquisition module 1010 is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object;

[0245] The first determining module 1020 is used to determine that if the target location information does not match the first location information corresponding to the first target audio segment, then the target audio frame is determined as the starting point of the sound content of the second sound object.

[0246] The first target audio segment includes all consecutive audio frames emitted by the first sound-emitting object, which is the sound-emitting object of the audio frame preceding the target audio frame; the endpoint of the first target audio segment is the audio frame preceding the target audio frame.

[0247] In the embodiments of this application, by matching the target position information of the voice-emitting object of the target audio frame with the first position information corresponding to all consecutive audio frames of the first voice-emitting object, it can be determined whether the voice-emitting object has switched, thereby determining the start and end points of the voice content of each voice-emitting object, and thus realizing the determination of the voice content of each voice-emitting object and achieving role separation.

[0248] Other details of the sound-emitting object determination device 1000 according to the embodiments of the present invention are similar to the sound content starting point determination method provided in the third aspect above, and will not be repeated here.

[0249] Figure 11 A schematic diagram of the structure of the sound-emitting object identification change device provided for the ninth aspect. (See diagram below.) Figure 11 As shown, the sound-emitting object identification changing device 1100 includes:

[0250] The acquisition module 1110 is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object;

[0251] Extraction module 1120 is used to determine that the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the second target audio segment, wherein the first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frame in the first target audio segment is the same as the sound-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment;

[0252] The first determining module 1130 is used to determine the target sound source of the second target audio segment based on the target voiceprint characteristics.

[0253] The modification module 1140 is used to modify the identifier of the second sound-emitting object and the identifier of the target sound-emitting object. The identifier is used to characterize the sound-emitting state of the sound-emitting object.

[0254] In the embodiments of this application, by matching the target location information of the sound-emitting object in the target audio frame with the first location information corresponding to the first target audio segment, it can be determined whether the second sound-emitting object is the same as the sound-emitting object of the second target audio segment, that is, whether the current sound-emitting object has been switched. If it is determined that the sound-emitting object has been switched, changing the identifier of the second sound-emitting object and the identifier of the target sound-emitting object can indicate the identity of the current sound-emitting object.

[0255] Other details of the sound object identification changing device 1100 according to the embodiments of the present invention are similar to the sound object identification changing method provided in the fourth aspect above, and will not be repeated here.

[0256] Figure 12 A schematic diagram of the session record generation device provided for the tenth aspect. (See diagram below.) Figure 12 As shown, the session record generation device 1200 includes:

[0257] The acquisition module 1210 is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object in the audio session data;

[0258] Extraction module 1220 is used to determine that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, and then extract the target voiceprint features of the second target audio segment in the audio session data, wherein the first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frame in the first target audio segment is the same as the sound-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment;

[0259] The first determining module 1230 is used to determine the target sound source of the second target audio segment based on the target voiceprint characteristics.

[0260] The association module 1240 is used to associate the target voice object with the text content corresponding to the second target audio segment to obtain the conversation record of the target voice object.

[0261] In the embodiments of this application, when it is determined that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, it is determined that the sound source of the target audio frame is different from the sound source of the first target audio segment. Therefore, the text content corresponding to the second target audio segment that is the same as the sound source of the first target audio segment can be associated with the target object to form a meeting record, so as to extract the audio session data record in the future, which improves convenience.

[0262] Other details of the session record generation apparatus 1200 according to embodiments of the present invention are similar to the session record generation method provided in the fourth aspect above, and will not be repeated here.

[0263] Combination Figures 2 to 12 The methods provided in any one of the first, second, third, fourth, and fifth aspects, as well as the apparatus provided in any one of the sixth, seventh, eighth, ninth, and tenth aspects, can be implemented by a computing device. Figure 13 This is a schematic diagram of the hardware structure of a computing device 1300 according to an embodiment of the invention.

[0264] like Figure 13 As shown, the computing device 1300 includes an input device 1301, an input interface 1302, a processor 1303, a memory 1304, an output interface 1305, and an output device 1306. The input interface 1302, processor 1303, memory 1304, and output interface 1305 are interconnected via a bus 1310. The input device 1301 and output device 1306 are connected to the bus 1310 via the input interface 1302 and output interface 1305, respectively, and are thus connected to other components of the computing device 1300.

[0265] Specifically, input device 1301 receives input information from the outside and transmits the input information to processor 1303 through input interface 1302; processor 1303 processes the input information based on computer-executable instructions stored in memory 1304 to generate output information, temporarily or permanently stores the output information in memory 1304, and then transmits the output information to output device 1306 through output interface 1305; output device 1306 outputs the output information to the outside of computing device 1300 for user use.

[0266] The processor 1303 may include processors such as a central processing unit (CPU), a network processing unit (NPU), a tensor processing unit (TPU), a field programmable gate array (FPGA) chip, or an artificial intelligence (AI) chip. The accompanying figure is for illustrative purposes only and is not limited to the types of processors listed in the text.

[0267] In other words, Figure 13The computing device shown may also be implemented as including: a memory storing computer-executable instructions; and a processor that, when executing the computer-executable instructions, can implement any embodiment of any of the first to tenth aspects.

[0268] This invention also provides a computer storage medium storing computer program instructions; when the computer program instructions are executed by a processor, they implement the sound-generating object determination method provided in this invention.

[0269] The functional blocks shown in the above structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0270] It should also be noted that the exemplary embodiments mentioned in this invention describe methods or systems based on a series of steps or apparatus. However, this invention is not limited to the order of the steps described above; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0271] The above are merely specific embodiments of the present invention. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A method for determining the source of a sound, comprising: Obtain the target audio frame emitted by the second sound-emitting object and the target location information of the second sound-emitting object; If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target voiceprint features of the second target audio segment are extracted, wherein the first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frames in the first target audio segment is the same as the voice-producing object of the audio frames in the second target audio segment; the first target audio segment is at least a part of the second target audio segment; N is an integer greater than or equal to 1; The voice object corresponding to the first voiceprint feature in the preset sound database is determined as the target voice object of the second target audio segment. Specifically, if the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold but greater than a second preset location matching degree threshold, then the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold. Alternatively, if the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, then the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, where the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold.

2. The method according to claim 1, wherein, The target location information includes the relative position information between the second sound-emitting object and the preset audio collector.

3. The method according to claim 1, wherein, Before determining that the target location information does not match the first location information corresponding to the first target audio segment, and before extracting the target voiceprint features of the second target audio segment, the method further includes: The location information of the sound-producing object corresponding to the audio frame in the first target audio segment is filtered to obtain the filtered location information. Based on the filtered location information, the first location information is determined.

4. The method according to claim 1, wherein, The step of determining the voice source corresponding to the first voiceprint feature in the preset sound database as the target voice source of the second target audio segment includes: If a first voiceprint feature exists in the preset sound database and its matching degree with the target voiceprint feature satisfies the first preset matching condition, the voice-producing object corresponding to the first voiceprint feature is determined as the target voice-producing object of the second target audio segment. The method further includes: If the matching degree between the first voiceprint feature and the target voiceprint feature meets the second preset matching condition, the voiceprint feature corresponding to the target voice-emitting object in the preset sound database is updated using the target voiceprint feature. Wherein, the matching degree that the second preset matching condition needs to satisfy is greater than the matching degree that the first preset matching condition needs to satisfy.

5. The method according to claim 4, wherein, When the matching degree between the target location information and the first location information is less than the first preset location matching degree threshold and greater than the second preset location matching degree threshold, the first preset matching condition includes that the matching degree between the voiceprint features in the preset voice database and the target voiceprint features is greater than the first preset voiceprint matching degree threshold; the second preset matching condition includes that the matching degree between the voiceprint features in the preset voice database and the target voiceprint features is greater than the second preset voiceprint matching degree threshold, wherein the second preset voiceprint matching degree threshold is greater than the first preset voiceprint matching degree threshold; When the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the first preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the third preset voiceprint matching degree threshold; the second preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is greater than the fourth preset voiceprint matching degree threshold. Wherein, the fourth preset voiceprint matching threshold is greater than the third preset voiceprint matching threshold, the first preset voiceprint matching threshold is less than the second preset voiceprint matching threshold, and the second preset voiceprint matching threshold is less than the fourth preset voiceprint matching threshold.

6. The method according to claim 4, wherein, The method further includes: If the matching degree between each voiceprint feature in the preset sound database and the target voiceprint feature satisfies the third preset matching condition, then the correspondence between the target voiceprint feature and the voice-producing object corresponding to the target voiceprint feature is stored in the preset sound database, and the voice-producing object corresponding to the target voiceprint feature is determined as the target voice-producing object of the second target audio segment. The third preset matching condition is used to characterize the mismatch between the voiceprint features in the preset voice database and the target voiceprint features.

7. The method according to claim 6, wherein, When the matching degree between the target location information and the first location information is less than the first preset location matching degree threshold and greater than the second preset location matching degree threshold, the third preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is less than the fifth preset voiceprint matching degree threshold. When the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the third preset matching condition includes that the matching degree between the voiceprint features in the preset sound database and the target voiceprint features is less than the sixth preset voiceprint matching degree threshold. The fifth preset voiceprint matching threshold is less than the sixth preset voiceprint matching threshold.

8. A method for determining the source of a sound, comprising: Obtain the target audio frame emitted by the second sound-emitting object and the target location information of the second sound-emitting object; If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target voiceprint features of the first target audio segment are extracted. The first target audio segment includes all continuous audio frames emitted by the first voice object, and the first voice object is the voice object that emitted the audio frame preceding the target audio frame. The endpoint of the first target audio segment is the audio frame preceding the target audio frame. The voice object corresponding to the first voiceprint feature in the preset sound database is determined as the target voice object of the second target audio segment. Specifically, if the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold but greater than a second preset location matching degree threshold, then the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold. Alternatively, if the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, then the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, where the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold.

9. A method for determining the starting point of vocal content, comprising: Obtain the target audio frame emitted by the second sound-emitting object and the target location information of the second sound-emitting object; If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target audio frame is determined as the starting point of the sound content of the second sound-emitting object; Wherein, the first target audio segment includes all continuous audio frames emitted by the first voice-emitting object, the first voice-emitting object being the voice-emitting object of the preceding audio frame of the target audio frame; the end point of the first target audio segment is the preceding audio frame of the target audio frame, the starting point of the voice content emitted by the second voice-emitting object is used to determine the target voice-emitting object of the second target audio segment, the target voice-emitting object of the second target audio segment is determined based on the voice-emitting object corresponding to the first voiceprint feature in the preset sound database, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold and greater than a second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold; or, when the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold, the voice-emitting object of the audio frames in the first target audio segment is the same as the voice-emitting object of the audio frames in the second target audio segment, and the first target audio segment is at least a part of the second target audio segment.

10. A method for changing the identifier of a sound-emitting object, comprising: Obtain the target audio frame emitted by the second sound-emitting object and the target location information of the second sound-emitting object; If it is determined that the target location information does not match the first location information corresponding to the first target audio segment, then the target voiceprint features of the second target audio segment are extracted, wherein the first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frame in the first target audio segment is the same as the sound-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment; The voice object corresponding to the first voiceprint feature in the preset sound database is determined as the target voice object of the second target audio segment. Specifically, if the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold but greater than a second preset location matching degree threshold, then the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold. Alternatively, if the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, then the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, where the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold. Change the identifier of the second vocal object and the identifier of the target vocal object, wherein the identifier is used to characterize the vocal state of the vocal object.

11. A method for generating session records, comprising: Obtain the target audio frame emitted by the second speaker in the audio session data and the target location information of the second speaker; If it is determined that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, then the target voiceprint features of the second target audio segment in the audio session data are extracted, wherein the first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frame in the first target audio segment is the same as the voice-producing object of the audio frame in the second target audio segment; and the first target audio segment is at least a part of the second target audio segment. The voice object corresponding to the first voiceprint feature in the preset sound database is determined as the target voice object of the second target audio segment. Specifically, if the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold but greater than a second preset location matching degree threshold, then the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold. Alternatively, if the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, then the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, where the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold. The target voice object is associated with the text content corresponding to the second target audio segment to obtain the conversation record of the target voice object.

12. A device for determining a sound-emitting object, wherein, The device includes: The acquisition module is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object; An extraction module is used to determine if the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the second target audio segment. The first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frames in the first target audio segment is the same as the voice-producing object of the audio frames in the second target audio segment; the first target audio segment is at least a part of the second target audio segment; and N is an integer greater than or equal to 1. The first determining module is used to determine the voice object corresponding to the first voiceprint feature in the preset sound database as the target voice object of the second target audio segment, wherein, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold and greater than a second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold; or, when the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, and the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold.

13. A device for determining a sound-emitting object, comprising: The acquisition module is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object; An extraction module is used to determine that the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the first target audio segment. The first target audio segment includes all continuous audio frames emitted by the first voice object, and the first voice object is the voice object that emitted the audio frame preceding the target audio frame. The end point of the first target audio segment is the audio frame preceding the target audio frame. The first determining module is used to determine the voice object corresponding to the first voiceprint feature in the preset sound database as the target voice object of the second target audio segment, wherein, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold and greater than a second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold; or, when the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, and the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold.

14. A device for determining the starting point of a sound content, comprising: The acquisition module is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object; The first determining module is used to determine that if the target location information does not match the first location information corresponding to the first target audio segment, then the target audio frame is determined as the starting point of the sound content of the second sound-emitting object. Wherein, the first target audio segment includes all continuous audio frames emitted by the first voice-emitting object, the first voice-emitting object being the voice-emitting object of the preceding audio frame of the target audio frame; the end point of the first target audio segment is the preceding audio frame of the target audio frame, the starting point of the voice content emitted by the second voice-emitting object is used to determine the target voice-emitting object of the second target audio segment, the target voice-emitting object of the second target audio segment is determined based on the voice-emitting object corresponding to the first voiceprint feature in the preset sound database, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold and greater than a second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold; or, when the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold, the voice-emitting object of the audio frames in the first target audio segment is the same as the voice-emitting object of the audio frames in the second target audio segment, and the first target audio segment is at least a part of the second target audio segment.

15. A device for changing the identifier of a sound-emitting object, comprising: The acquisition module is used to acquire the target audio frame emitted by the second sound-emitting object and the target position information of the second sound-emitting object; An extraction module is used to determine that the target location information does not match the first location information corresponding to the first target audio segment, and then extract the target voiceprint features of the second target audio segment, wherein the first target audio segment includes the first N audio frames of the target audio frame; the sound-producing object of the audio frame in the first target audio segment is the same as the sound-producing object of the audio frame in the second target audio segment; the first target audio segment is at least a part of the second target audio segment; The first determining module is used to determine the voice object corresponding to the first voiceprint feature in the preset sound database as the target voice object of the second target audio segment, wherein, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold and greater than a second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold; or, when the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, and the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold; The modification module is used to modify the identifier of the second sound-emitting object and the identifier of the target sound-emitting object, wherein the identifier is used to characterize the sound-emitting state of the sound-emitting object.

16. A session record generation apparatus, comprising: The acquisition module is used to acquire the target audio frame emitted by the second speaker in the audio session data and the target location information of the second speaker. An extraction module is used to determine that the target location information does not match the first location information corresponding to the first target audio segment in the audio session data, and then extract the target voiceprint features of the second target audio segment in the audio session data, wherein the first target audio segment includes the first N audio frames of the target audio frame; the voice-producing object of the audio frames in the first target audio segment is the same as the voice-producing object of the audio frames in the second target audio segment; and the first target audio segment is at least a part of the second target audio segment. The first determining module is used to determine the voice object corresponding to the first voiceprint feature in the preset sound database as the target voice object of the second target audio segment, wherein, when the matching degree between the target location information and the first location information is less than a first preset location matching degree threshold and greater than a second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than the first preset voiceprint matching degree threshold; or, when the matching degree between the target location information and the first location information is less than the second preset location matching degree threshold, the matching degree between the first voiceprint feature and the target voiceprint feature is greater than a third preset voiceprint matching degree threshold, and the first preset voiceprint matching degree threshold is less than the third preset voiceprint matching degree threshold; The association module is used to associate the target voice object with the text content corresponding to the second target audio segment to obtain the conversation record of the target voice object.

17. A computing device, wherein, The computing device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the method as described in any one of claims 1-11.

18. A computer storage medium, wherein, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the method as described in any one of claims 1-11.