Speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering
The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering solves the problem of indiscriminate segmentation in existing technologies, and achieves more efficient speaker segmentation and resource utilization.
Patent Information
- Application Number
- CN202511255871.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2026-01-09
AI Technical Summary
Existing speaker segmentation technologies fail to dynamically focus on the target speaker based on user needs, resulting in indiscriminate segmentation that increases computational burden and error rate, and may also obscure key speech information.
A multi-scale feature fusion and voiceprint perception density clustering method is adopted. Voiceprints are extracted by setting extraction points, voiceprint consistency and environmental information are judged, speech segments are segmented and speakers are classified, and transition points are confirmed by health information and background noise probability.
It reduces segmentation errors caused by environmental and physiological factors, improves the intelligence and accuracy of speaker segmentation, and reduces resource waste and ineffective segmentation.
Smart Images

Figure CN121306169A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speaker segmentation, and in particular to a speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering. Background Technology
[0002] Speaker segmentation technology, by accurately separating and labeling different speakers in an audio stream, provides core technical support for various scenarios such as meeting recording and medical consultation, and is of great value in improving the efficiency of structured speech information processing and privacy security management. However, existing speaker segmentation technologies typically segment all detected speakers indiscriminately, failing to dynamically focus on the target speaker based on user needs. This "one-size-fits-all" processing mode not only increases the burden of unnecessary computation but may also cause critical speech information to be buried by non-target speaker segments, exacerbating the error rate of subsequent speech recognition and analysis systems. These problems result in bottlenecks such as low resource utilization and poor adaptability to real-world applications of existing technologies. Summary of the Invention
[0003] The purpose of this invention is to provide a speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering to solve the problems mentioned in the background art.
[0004] This application provides a speaker segmentation method based on multi-scale feature fusion and voiceprint-aware density clustering, which adopts the following technical solution: Multiple extraction points are set for the target speech, and the corresponding voiceprints are extracted from the extraction points and recorded as the extraction point voiceprints. The extraction point that has been analyzed and is closest to the extraction point in real-time analysis is used as the reference point. The reference speaker corresponding to the reference point is collected, and the voiceprint of the reference speaker is recorded as the reference voiceprint. Compare whether the extracted voiceprint is consistent with the reference voiceprint. If they are inconsistent, collect the voiceprint perception density of the extracted point. The system determines whether there is overlap in the speech of multiple people based on the voiceprint perception density. If there is no overlap in the speech of multiple people, the system collects the health information and environmental information of the reference speaker. Based on health and environmental information, determine whether the extracted voiceprint is a reference voiceprint; if not, use the extracted voiceprint as a speech conversion point. If multiple people are speaking and there is overlap, the speech content and background noise probability of the extraction point are collected to determine whether the extraction point should be used as a speech conversion point. Speech segments are obtained by dividing speech based on speech conversion points. Speaker segmentation is achieved by classifying the speech segments by speaker.
[0005] Preferably, the step of determining whether the extracted voiceprint is a reference voiceprint based on health and environmental information, and using the extracted voiceprint as a speech conversion point if it is not, is as follows: Collect the speech content of the extraction point, determine whether the speech content is physiological speech, and if the speech content is physiological speech, then determine the voiceprint of the extraction point as the reference voiceprint. If the speech content is not physiological speech, the initial background sound of the reference speaker is extracted, and the extracted background sound of the extracted points is collected; the background similarity between the initial background sound and the extracted background sound is compared. Determine whether the background similarity meets the preset background similarity standard. If it does, extract the environment-related speech content from the target speech. Determine whether the extracted voiceprint is a reference voiceprint based on the environmental context and health information. If the preset background similarity standard is not met, the voiceprint feature change information of the entire speech segment corresponding to the extraction point is collected, and the voiceprint feature change information is used to determine whether the voiceprint of the extraction point is a reference voiceprint.
[0006] Preferably, the step of determining whether the extracted voiceprint is a reference voiceprint based on the environmentally relevant speech content and health information is as follows: Based on the environmentally relevant speech content, extract all speech content from which the speaker is uncomfortable with the environment as environmentally uncomfortable content; The frequency of the reference speaker's expression of environmental discomfort was counted, the proportion of the reference speaker's physiological speech characteristics was collected, and the probability of the reference speaker's physical discomfort was obtained by combining the data. The health score of the reference speaker is extracted based on the health information of the reference speaker, and the voiceprint change is obtained by combining the probability of physical discomfort. Compare the voiceprint difference between the extracted point and the reference voiceprint to determine if the voiceprint difference is greater than the voiceprint change. If it is greater than the voiceprint change, then the extracted point voiceprint is not the reference voiceprint.
[0007] Preferably, the step of collecting the voiceprint feature change information of the entire speech segment corresponding to the extraction point, and determining whether the voiceprint of the extraction point is a reference voiceprint based on the voiceprint feature change information, specifically includes: The extracted voiceprint is compared with all known voiceprints to obtain multiple similarities. The similarity between the extracted voiceprint and the reference voiceprint is recorded as the reference similarity. It is then determined whether the reference similarity is the maximum. If the reference similarity is the maximum, then the background sound features shared by the extraction point and the reference point are collected, and the similarity of the shared background sound features is obtained and recorded as the background feature similarity. Calculate the difference between the background feature similarity and the reference similarity, and determine whether the difference is within the preset standard. If it is within the preset standard, then determine that the extracted point voiceprint is the reference voiceprint. If the extraction point and the reference point do not share common background sound features, then collect the voiceprint changes of other extraction points to determine whether the voiceprint of the extraction point is the reference voiceprint. If the reference similarity is not the maximum, then the extracted voiceprint is determined to be not the reference voiceprint.
[0008] Preferably, the step of collecting voiceprint changes from other extraction points and determining whether the voiceprint at an extraction point is a reference voiceprint is as follows: Collect unanalyzed extraction points and record them as test points. Extract the test voiceprints of the test points and compare the test voiceprints with all known voiceprints to obtain the test similarity. Extract the maximum similarity scores corresponding to all voiceprints to be tested, set similarity intervals, and classify all maximum similarity scores according to the similarity intervals; Select the most numerous similarity intervals and calculate the variance of the maximum test similarity within each similarity interval; Determine whether the variance is within the preset variance standard. If it is within the preset variance standard, then determine that the extracted voiceprint is the reference voiceprint.
[0009] Preferably, the step of collecting the speech content and background noise probability of the extraction point and determining whether the extraction point should be used as a speech conversion point is as follows: The speech content of the extraction points is collected and segmented according to the voiceprint features to obtain multiple user voices; Collect user voiceprints, compare all user voiceprints with reference voiceprints to obtain similarity scores and record them as user similarity scores, and determine whether there is a reference speaker among the users based on user similarity scores; If a reference speaker exists, then users who are not reference speakers are recorded as test users, and the overlapping speech content of the test users is extracted. Based on the overlapping speech content, it is determined whether the extraction point can be used as a speech conversion point. If no reference speaker exists, then the similarity score is used to determine whether a known speaker exists among the users. If a known speaker exists, the extracted point is used as a speech conversion point; if no known speaker exists, the probability of background noise determines whether the extracted point should be used as a speech conversion point.
[0010] Preferably, the step of extracting overlapping speech content from the user under test and determining whether the extracted point should be used as a speech conversion point based on the overlapping speech content is as follows: Obtain the ratio of speech overlap duration to total speech duration and record it as the overlap ratio; determine whether the overlap ratio meets the preset overlap standard. If the preset overlap standard is met, the voice content of the reference speaker is extracted and recorded as the reference content, and the voice content of the user to be tested is collected and recorded as the test content. Assess the degree of correlation between the reference content and the content to be tested, and determine whether the extraction point should be used as a speech conversion point based on the degree of correlation; If the preset overlap standard is not met, the emotional values of the reference speaker and the user being tested are extracted based on the voice content. Determine whether the emotion value meets the preset emotion standard. If the preset emotion standard is met, then the extraction point is determined as the speech conversion point. If the preset emotional standard is not met, the probability of background noise in the user's voice content is estimated, and the extraction point is used as a voice conversion point based on the background noise probability.
[0011] Preferably, the step of estimating the background noise probability of the user's speech content is as follows: Collect the voice content of the users to be tested, check whether the voice content is an existing work, and filter the users whose voice content is not an existing work as the basic users; Collect the average voice volume of basic users, collect the average voice volume of known speakers, and calculate the difference between the average voice volume of basic users and known speakers as the volume difference; The average number of basic users is extracted from the overlapping speech, and the background noise probability is obtained by combining the volume difference and the overlap ratio.
[0012] Preferably, if no known speaker exists, the step of determining whether the extracted point should be used as a speech conversion point based on the background noise probability is as follows: The user with the longest voice duration in the voice content extracted from the voiceprint collection point is recorded as the first user; Determine whether the duration of the first user's voice covers the duration of the voice at the extraction point. If it does, extract the first user's voice content and evaluate the correlation between the first user's voice content and the reference content. If not covered, collect the distribution of user voice duration in the voice content of the extraction point, and find the user combination whose voice duration distribution covers the voice duration of the extraction point; Determine whether the content of the user group is coherent. If it is coherent, extract the combined content of the user group and evaluate the relevance between the combined content and the reference content. Determine whether an extraction point should be used as a speech conversion point based on its relevance. If there is no correlation, the extraction point is determined as a speech conversion point based on the probability of background noise.
[0013] In summary, this application includes at least one of the following beneficial technical effects: 1. Multiple extraction points are set for the target speech. The corresponding voiceprints are extracted from each extraction point and recorded as the extraction point voiceprint. The nearest extraction point that has been confirmed as a conversion point is used as a reference point. The speaker corresponding to the reference point is designated as the reference speaker, and the voiceprint corresponding to the reference speaker is designated as the reference voiceprint. The extraction point voiceprint is compared with the reference voiceprint. If they do not match, the voiceprint perception density of the extraction point is collected. Based on the voiceprint perception density, it is determined whether there is overlapping speech by multiple speakers. If there is no overlapping speech by multiple speakers, health information and environmental information are used to determine whether the extraction point voiceprint is a reference voiceprint. If not, the extraction point is used as a speech conversion point. If there is overlapping speech by multiple speakers, the speech content and background noise probability of the extraction point are used to determine whether the extraction point is a speech conversion point. The speech is segmented into multiple speech segments based on the speech conversion points. Speaker classification is performed on the speech segments to achieve speaker segmentation. Confirming conversion points through health information and environmental information can reduce segmentation errors caused by voiceprint changes due to environmental and physiological factors. By identifying transition points based on speech content and background noise probability, unnecessary noise segmentation can be reduced, thus improving the intelligence of speaker segmentation based on multi-scale feature fusion and voiceprint perception density clustering.
[0014] 2. Collect the speech content of the extraction point and determine whether the speech content is physiological speech. If the speech content is physiological speech, the extracted point voiceprint is determined to be the reference voiceprint. If the speech content is not physiological speech, the initial background sound of the reference speaker is extracted, and the extracted background sound of the speech is collected. The background similarity between the initial background sound and the extracted background sound is compared. It is determined whether the background similarity meets the preset background similarity standard. If the preset background similarity standard is met, the environmentally related speech content in the target speech is extracted. Based on the number of times the reference speaker expresses discomfort, the proportion of physiological speech features, and the speaker's health status in the environmentally related speech content, the voiceprint variability is estimated. Combined with the voiceprint difference between the extracted point voiceprint and the reference voiceprint, it is determined that the extracted point voiceprint is not the reference voiceprint. If the preset background similarity standard is not met, it is determined whether the extracted point voiceprint is the reference voiceprint based on the multiple similarities between the extracted point voiceprint and all known voiceprints, and the similarity of shared background sound features. If the extracted point and the reference point do not have common background sound features, the similarity variance between voiceprints is used for judgment. Determining whether the changed voiceprint belongs to the aforementioned speaker, and thus identifying the defect transition point, helps reduce segmentation errors caused by environmental and physiological reasons, and improves the accuracy of speaker segmentation based on multi-scale feature fusion and voiceprint perception density clustering.
[0015] 3. Segment the speech content of the extracted points based on voiceprint features to obtain multiple user voices. Determine if a reference speaker exists among the users. If a reference speaker exists, the user who is not a reference speaker is recorded as the test user. Determine if the speech overlap duration meets a preset standard. If it does, determine whether the extracted point should be used as a speech conversion point based on the correlation between the reference speaker's speech content and the test user's speech content. If the preset overlap standard is not met, determine whether the extracted point should be used as a speech conversion point based on the emotional values extracted from the reference speaker and the test user's speech content. If the preset emotional standard is not met, estimate the background noise probability of the test user's speech content and determine whether the extracted point should be used as a speech conversion point based on the background noise probability. Obtain the background noise probability based on the voice volume of different speakers, the number of overlapping speakers, and the proportion of overlapping content. If no reference speaker exists, determine if a known speaker exists among the users based on similarity. If a known speaker exists, the extracted point is used as a speech conversion point. If no known speaker exists, find users or user groups whose speech duration can cover the complete speech segment based on the voiceprint extraction point, evaluate the correlation of their content, and determine whether the extracted point should be used as a speech conversion point. By confirming whether there is a reference speaker or a known speaker, and by using different methods to determine whether to segment newly emerging speaker speech segments, unnecessary noise extraction and segmentation can be effectively reduced, and the resource utilization rate of speaker segmentation based on multi-scale feature fusion and voiceprint perception density clustering can be improved. Attached Figure Description
[0016] Figure 1 This is a schematic diagram illustrating the specific steps of an embodiment of the speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to the present invention. Detailed Implementation
[0017] The following examples and... Figure 1 The present invention will be described in further detail, but the embodiments of the present invention are not limited thereto.
[0018] This invention discloses a speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering, specifically including the following steps: Step S1: Set multiple extraction points for the target speech, extract the corresponding voiceprints at the extraction points and record them as extraction point voiceprints.
[0019] Users can set the time length and set extraction points based on the time length. For example, if the user sets the time length to 1 second, there will be an extraction point every 1 second to extract the voiceprint at the corresponding time.
[0020] Step S2: Take the extraction point that has been analyzed and is closest to the extraction point in real-time analysis as the reference point, collect the reference speaker corresponding to the reference point, and record the voiceprint of the reference speaker as the reference voiceprint.
[0021] The extraction points are determined sequentially to determine whether each point is a speech conversion point. Therefore, the extraction point preceding the current extraction point is used as a reference point. For example, if the user sets the extraction point to be taken every 1 second, then the reference point for 12 seconds is the extraction point at 11 seconds, and the reference point for 20 seconds is the extraction point at 19 seconds.
[0022] Step S3: Compare whether the extracted voiceprint is consistent with the reference voiceprint. If they are inconsistent, collect the voiceprint perception density of the extracted point.
[0023] Voiceprint comparison can be achieved by extracting Mel-frequency cepstral coefficients (MFCC) or linear-frequency cepstral coefficients (LFCC) from speech and calculating the distance between feature vectors (such as cosine similarity or Euclidean distance).
[0024] Step S4: Determine whether there is overlap in multiple people's speech based on the voiceprint perception density. If there is no overlap in multiple people's speech, collect the health information and environmental information of the reference speaker.
[0025] Voiceprint perceptual density (VPD) in speaker segmentation is a concept describing the density of voiceprint features in a speech signal. It is mainly used to quantify the complexity and overlap of multi-person speech interactions. VPD refers to the degree of aggregation of voiceprint features per unit time or frequency spectrum, reflecting the mixed state of multiple speakers' voices in the time or frequency domain. By setting a VPD threshold, it is possible to determine whether there is overlap in multiple speakers' speech.
[0026] Step S5: Determine whether the extracted voiceprint is a reference voiceprint based on health information and environmental information. If not, use the extracted voiceprint as a speech conversion point.
[0027] Step S6: If multiple people are speaking and there is overlap, collect the speech content and background noise probability of the extraction point, and determine whether the extraction point should be used as a speech conversion point.
[0028] Step S7: Segment the speech based on the speech conversion point to obtain multiple speech segments, and classify the speakers of the speech segments to achieve speaker segmentation.
[0029] In practical applications, speaker speech segmentation is often achieved through voiceprint segmentation, as voiceprints do not change easily, resulting in relatively high accuracy. However, special circumstances can cause changes in a speaker's voiceprint, such as when the speaker has a cold, is sick, or is clearing phlegm. Comparing this to the previous voiceprint in such cases would lead to segmentation errors. Environmental factors also affect voiceprints; combining speaker health information with environmental data allows for more accurate speaker segmentation. Furthermore, in some scenarios with significant background noise, segmenting the background noise not only increases the workload and wastes resources but also introduces invalid data, causing inconvenience to users. Therefore, determining the segmentation content based on the extracted speech content and background noise probability minimizes resource waste.
[0030] The process of determining whether the extracted voiceprint is a reference voiceprint based on health and environmental information, and if not, using the extracted voiceprint as a speech conversion point, is as follows: Step S51: Collect the speech content of the extraction point and determine whether the speech content is physiological speech. If the speech content is physiological speech, then determine the voiceprint of the extraction point as the reference voiceprint.
[0031] Physiological speech includes sounds such as coughing, yawning, and sneezing; these are the sounds made by the speaker, but do not contain specific speech content. By setting features for physiological speech and comparing whether the speech contains physiological characteristics, it can be confirmed whether the speech content is physiological speech.
[0032] Step S52: If the speech content is not physiological speech, extract the initial background sound of the reference speaker and collect the extracted background sound of the extracted points. Compare the background similarity between the initial background sound and the extracted background sound.
[0033] Step S53: Determine whether the background similarity meets the preset background similarity standard. If it meets the preset background similarity standard, extract the environment-related speech content from the target speech.
[0034] Step S54: Determine whether the extracted voiceprint is a reference voiceprint based on the environmental associated speech content and health information.
[0035] Step S55: If the preset background similarity standard is not met, the voiceprint feature change information of the entire speech segment corresponding to the extraction point is collected, and the voiceprint feature change information is used to determine whether the voiceprint of the extraction point is a reference voiceprint.
[0036] In practical applications, physiological speech (coughing, sneezing, yawning, etc.) can cause significant changes in voiceprint characteristics. The core principle is that abrupt changes in the physical state of the vocal organs disrupt the steady-state process of vocal cord vibration and vocal tract resonance. In existing speaker segmentation technologies, segmentation based on voiceprint easily overlooks the interference from physiological speech. Voiceprint changes caused by physiological speech do not contain valid speech content, nor are they due to speaker transitions, so segmentation is unnecessary. Even if it's not the voice of the reference speaker, it's merely a physiological reaction beyond the speaker's control, without any actual speech transition, and therefore not considered a transition point. If it's not physiological speech, but the extracted voiceprint changes, it could indicate a change in the speaker or environmental factors, requiring further analysis and confirmation.
[0037] The steps for determining whether an extracted voiceprint is a reference voiceprint based on environmentally relevant speech content and health information are as follows: Step S541: Extract all speaker speech content that is uncomfortable with the environment based on the environment-related speech content as environment-uncomfortable content.
[0038] Speech-to-text technology can be used to obtain environment-related speech content. Keywords indicating environmental discomfort can be set, such as "hot," "cold," and "dry." By searching for keywords, sentences containing those keywords can be set as environmentally inappropriate content.
[0039] Step S542: Count the number of times the reference speaker expresses content related to environmental discomfort, collect the proportion of the reference speaker's physiological speech features, and obtain the probability of the reference speaker's physical discomfort.
[0040] The process involves collecting data from the previous reference speech segment, focusing on the number of times the speaker mentioned environmental discomfort before the extraction point, and the proportion of physiological sounds associated with that environment. For example, coughing and sneezing are physiological sounds corresponding to a cold environment, while yawning is not. Therefore, the proportion of physiological sound features counted refers to the ratio of sentences containing physiological sounds associated with environmental discomfort to the total speech. A weighted summation method is used to obtain the probability of the speaker's physical discomfort, as the environment may cause physical discomfort, thus affecting the voiceprint.
[0041] Step S543: Extract the health score of the reference speaker based on the health information of the reference speaker, and obtain the voiceprint change score by combining the probability of physical discomfort.
[0042] If the speaker's identity is clear, their recent health information, such as their current immunity level, can be obtained. The probability of illness is then added to this immunity level to calculate the voiceprint variation. This is because a lower immunity level increases the speaker's likelihood of getting sick due to environmental factors, and the more severe the illness, the greater the voiceprint variation.
[0043] Step S544: Compare the voiceprint difference between the extracted point voiceprint and the reference voiceprint, and determine whether the voiceprint difference is greater than the voiceprint change. If it is greater than the voiceprint change, then determine that the extracted point voiceprint is not the reference voiceprint.
[0044] In practical applications, the changes in voiceprint caused by the environment are limited, and changes in voiceprint caused by the speaker's own physical condition are also related to the severity of the disease. Environmental interference and physiological lesions can indeed distort voiceprint characteristics, but due to the stability of the physical structure of the human vocal organs, such changes always have insurmountable physiological boundaries—even in extreme cases, the voiceprint still maintains a traceable similarity to its origin. For example, a 40-year-old bronchitis patient originally had a clear voice (fundamental frequency 120Hz, formant F1 at 550Hz), but after the disease, vocal cord edema caused a hoarse and low-pitched voice (fundamental frequency dropped to 95Hz, F1 dropped to 450Hz). In this case, the cosine similarity of the voiceprint verification may have declined, and the system will incorrectly reject it. However, illness or throat clearing will not cause the similarity to drop to 20%. Therefore, the degree of voiceprint change is estimated based on the speaker's health condition and the health impact of environmental changes, and then a judgment is made. If the speaker does not feel any discomfort and is in very good health, but the voiceprint changes significantly, then it can be determined that the voiceprint change is not caused by the speaker's illness, but rather by a change in the speaker. Therefore, the extraction point is used as the speech conversion point.
[0045] The steps for collecting and extracting the voiceprint feature change information of the entire speech segment corresponding to the extraction point, and determining whether the voiceprint of the extraction point is a reference voiceprint based on the voiceprint feature change information, are as follows: Step S551: Compare the extracted voiceprint with all known voiceprints to obtain multiple similarities. Record the similarity between the extracted voiceprint and the reference voiceprint as the reference similarity. Determine whether the reference similarity is the maximum.
[0046] Because voiceprint changes are limited, even if environmental interference or physical illness causes significant changes in voiceprint characteristics, the physiological structure of the human vocal organs is inherently stable. Therefore, the similarity between the changed voiceprint of the target speaker and its original voiceprint will still be higher than the similarity between the voiceprint of the target speaker and any other speaker.
[0047] Step S552: If the reference similarity is the maximum, collect the background sound features shared by the extraction point and the reference point, compare them to obtain the similarity of the shared background sound features and record it as the background feature similarity.
[0048] Step S553: Calculate the difference between the background feature similarity and the reference similarity, and determine whether the difference is within the preset standard. If it is within the preset standard, then determine that the extracted point voiceprint is the reference voiceprint.
[0049] In typical conversational scenarios, there are usually other background sounds besides the speaker's voice. For example, if the speaker is playing music, there will be music playing in addition to the speaker's voice. Shared background sound characteristics can reflect whether the environment has changed. For instance, when a speaker moves from an open outdoor area to a bathroom, the echo will cause a change in their voiceprint. The same background sound will also change during the environmental transition. Therefore, when the background sound characteristics and the speaker's voice undergo almost identical changes, and the voiceprint is most similar to that of the reference speaker, it is considered that the voiceprint change is caused by the environment, and therefore it is still the reference speaker speaking.
[0050] Step S554: If the extraction point and the reference point do not have common background sound features, then collect the voiceprint changes of other extraction points and determine whether the voiceprint of the extraction point is the reference voiceprint.
[0051] In step S555, if the reference similarity is not the maximum, then it is determined that the extracted voiceprint is not the reference voiceprint.
[0052] In practical applications, when the extracted point and the reference point do not share common background sound features, it is difficult to determine whether the environment has changed based on the background alone. Therefore, the judgment must be made based on the voiceprint characteristics of other extracted points. If the reference similarity is not the highest, as mentioned earlier, even if the voiceprint changes due to the environment, there are limits; the similarity between the voiceprint and the speaker's own voiceprint should be the highest. Therefore, if the similarity with the reference voiceprint is not the highest, regardless of whether the environment changes, the speaker has changed, and this point can be identified as a speech conversion point. Because there is a shared background sound, it means that the background sound changes with the speaker's environment. For example, if a user plays music through speakers and moves from the living room to the bathroom, the music similarity becomes 90%, while the reference similarity is 60%, a large difference. Clearly, the environment has not caused a significant change in the voiceprint, but the reference similarity has changed significantly, indicating that the extracted point's voiceprint is not the reference voiceprint, and the speaker has changed. If the reference similarity is 89%, the difference is within the preset range, confirming that the change in the speaker's voiceprint is due to environmental factors.
[0053] The steps for collecting voiceprint changes from other extraction points and determining whether the voiceprint at an extraction point is a reference voiceprint are as follows: Step S5541: Collect the unanalyzed extraction points and record them as test points, extract the test voiceprints of the test points, and compare the test voiceprints with all known voiceprints to obtain the test similarity.
[0054] The points to be tested include extraction points that are being analyzed in real time, i.e. extraction points that have not yet been determined as speech conversion points.
[0055] Step S5542: Extract the maximum similarity of all voiceprints to be tested, set the similarity interval, and classify all the maximum similarities according to the similarity interval.
[0056] Step S5543: Select the similarity interval with the largest number of similarities and calculate the variance of the maximum test similarity within the similarity interval.
[0057] Step S5544: Determine whether the variance is within the preset variance standard. If it is within the preset variance standard, then determine that the extracted voiceprint is the reference voiceprint.
[0058] In practical applications, if there are no shared background sound features, it's impossible to judge based on background sound alone. In such cases, judgment can be made based on the voiceprints of other speakers. If the voiceprint at the extraction point changes, and we suspect an environmental change, we assume the environment has changed. Therefore, the voiceprints of subsequent speakers will also change due to this environmental change at the extraction point. Since the environment has changed uniformly, the similarity changes in voiceprints should also be similar. The maximum similarity across all voiceprints is the original voiceprint that hasn't changed, representing the degree of voiceprint change. Excluding voiceprint changes caused by illness or other reasons, we select the most frequent similarity intervals and calculate the variance of these similarities. If the variance is too large, it indicates the change isn't due to the environment; otherwise, the similarity changes are similar. If the variance is small, it indicates an environmental change, and we assume the speaker is still the primary source of similarity.
[0059] The steps for collecting the speech content and background noise probability of the extraction point, and determining whether the extraction point should be used as a speech conversion point, are as follows: Step S61: Collect the speech content of the extraction point, and segment the speech content of the extraction point according to the voiceprint features to obtain multiple user voices.
[0060] The audio content extracted at each point refers to the audio content between the current extraction point and the next extraction point. This audio content is obtained through a speech conversion problem. Because it involves overlapping speech from multiple speakers, and each speaker has a different voiceprint, multiple user voices can be segmented based on their voiceprints.
[0061] Step S62: Collect the voiceprint of the user's voice, compare all user voiceprints with the reference voiceprint to obtain the similarity and record it as the user similarity, and determine whether there is a reference speaker among the users based on the user similarity.
[0062] Set a similarity threshold. If a user's similarity reaches the threshold, it is determined that there is a reference speaker.
[0063] Step S63: If a reference speaker exists, then the user who is not the reference speaker is recorded as the test user, the overlapping speech content of the test user is extracted, and the extraction point is determined as a speech conversion point based on the overlapping speech content.
[0064] Step S64: If there is no reference speaker, determine whether there is a known speaker among the users based on similarity.
[0065] Step S65: If a known speaker exists, the extracted point is used as a speech conversion point; if no known speaker exists, the extracted point is used as a speech conversion point based on the background noise probability.
[0066] The background noise probability is analyzed in steps S6361-S6363. When the background noise probability exceeds the probability threshold, the extraction point is determined to be background noise and is not used as a speech conversion point; otherwise, it is used as a speech conversion point.
[0067] In practical applications, if a speech segment contains no speakers, it's easily identified as a silent segment and doesn't require speaker segmentation. However, if multiple speakers are overlapping and the speech is mostly background noise, it could be considered a silent segment. For example, if a speaker is speaking in a classroom with many students talking, but the target speech is only for speakers A and B, then the other students' speech is background noise. Therefore, it's necessary to differentiate between background noise and other speech segments to reduce unnecessary segmentation and resource waste. If a reference speaker exists, but due to overlapping speech, other speakers may be included, further analysis is needed. If a known speaker is included, since they are not a reference speaker and are the speaker to be segmented, they are directly used as a speech transition point for segmentation. If no known speaker is present, it's necessary to determine if the speech segment is background noise. If it is background noise, it is not used as a speech transition point.
[0068] The steps for extracting overlapping speech content from the user under test and determining whether the extracted points should be used as speech conversion points based on the overlapping speech content are as follows: Step S631: Obtain the ratio of speech overlap duration to total speech duration and record it as the overlap ratio, and determine whether the overlap ratio meets the preset overlap standard.
[0069] In a normal conversation, one person usually speaks first, and then the conversation switches. However, other speakers may interrupt the current speaker, causing speech overlap. In a normal interruption, one party will stop speaking, so the overlap time will not be too long. When the overlap percentage is within the preset overlap standard, it means that the overlap has not lasted long and is considered normal in conversation. The overlap percentage refers to the ratio of speech overlap duration to the total speech duration of the sentence.
[0070] Step S632: If the preset overlap standard is met, the speech content of the reference speaker is extracted and recorded as the reference content, and the speech content of the user to be tested is collected and recorded as the test content.
[0071] Step S633: Evaluate the degree of correlation between the reference content and the content to be tested, and determine whether the extraction point should be used as a speech conversion point based on the degree of correlation.
[0072] By extracting keywords from the reference content and the content to be tested, and finding the co-occurrence probability of these keywords, the co-occurrence probability can be used as the degree of association. When the degree of association reaches a preset threshold, it is determined that another speaker has interrupted the reference speaker, thus introducing a new speaker, which needs to be considered a speech conversion point. If the degree of association does not reach the preset threshold, it is determined not to be a speech conversion point. This is because other users may interrupt, but they are not related to the current content, so it is unnecessary to segment that speech. For example, A and B are chatting. When A is speaking, a waiter interrupts A by serving food. At this moment, there is indeed a new speaker, but this new speaker is not someone who needs to be segmented; it is just an interlude, so it is unnecessary to convert or segment it.
[0073] Step S634: If the preset overlap standard is not met, extract the emotion values of the reference speaker and the user to be tested based on the voice content.
[0074] Emotion values can be used to perform speech emotion analysis by fusing acoustic features such as volume and speech rate with deep learning models. Emotion values of the reference speaker and the test user can be evaluated by relying on the temporal spectrum features of the audio signal and the emotional semantic context modeling.
[0075] Step S635: Determine whether the emotion value reaches the preset emotion standard. If the preset emotion standard is obtained, then the extraction point is determined as the speech conversion point.
[0076] Normally, when one person interrupts the other, the overlapping speech won't last too long. However, there are exceptions. When both speakers are emotionally charged, they may stop listening to each other, resulting in prolonged overlapping speech.
[0077] Step S636: If the preset emotion standard is not met, the background noise probability of the user's voice content is estimated, and the extraction point is determined as a voice conversion point based on the background noise probability.
[0078] In practical applications, the possibility of background noise cannot be ruled out in multi-person overlapping speech. If the overlap is small, it indicates that the speakers are communicating, and therefore the interruption is unlikely to be background noise. If the overlap is large, it may be due to prolonged overlap caused by the speaker's emotions, thus requiring analysis of the speaker's emotional state. If the speaker's emotions are not particularly agitated, then it is likely background noise; that is, there are indeed people speaking, but the speaker is not interacting with them, so the speech within the background noise is ignored. For example, a user is eating in a restaurant, and there is continuous human voice in the audio, but the speaker is not interacting with them, so the speaker's speech continues to be output, overlapping with the background noise. However, in practice, it is unnecessary to consider this background noise, reducing unnecessary speaker segmentation work and improving the speed of speaker segmentation.
[0079] The steps for estimating the probability of background noise in the voice content of a user being tested are as follows: Step S6361: Collect the voice content of the user to be tested, check whether the voice content is an existing work, and filter the users whose voice content is not an existing work as the basic users.
[0080] Some audio content, such as songs and stand-up comedy performances, contains human voices, but these are not actual speakers, so there's no need to segment this content. Therefore, audio clips found on the internet can be directly treated as background noise and filtered out.
[0081] Step S6362: Collect the average voice volume of the basic user, collect the average voice volume of the known speaker, and calculate the difference between the average voice volume of the basic user and the known speaker as the volume difference.
[0082] Step S6363: Extract the average number of basic users in the overlapping speech, and combine the volume difference and the overlap ratio to obtain the background noise probability.
[0083] In practical applications, the entropy weighting method is used to automatically calculate the objective weights of the basic number of users, volume difference, and overlap ratio. Based on the principle of information entropy, the weighted summation is used to output the background noise probability. The background noise probability reduces unnecessary speaker analysis, decreases the workload of subsequent similarity comparisons of extracted points, and minimizes resource waste. Furthermore, identifying background noise reduces the workload for users caused by invalid speaker segmentation, thereby improving segmentation speed and efficiency.
[0084] If no known speaker exists, the step of determining whether the extracted point should be used as a speech conversion point based on the background noise probability is as follows: Step S651: The user with the longest voice duration in the voice content extracted from the voiceprint collection point is recorded as the first user.
[0085] Step S652: Determine whether the voice duration of the first user covers the voice duration of the extraction point. If it covers the voice duration of the extraction point, extract the voice content of the first user and evaluate the correlation between the voice content of the first user and the reference content.
[0086] The relevance can be assessed by referring to the relevance evaluation of the content mentioned above, using the co-occurrence rate of keywords.
[0087] Step S653: If not covered, collect the distribution of user voice duration in the voice content of the extraction point, and find the user combination whose voice duration distribution covers the voice duration of the extraction point.
[0088] Step S654: Determine whether the content of the user combination is connected. If it is connected, extract the combined content of the user combination and evaluate the relevance between the combined content and the reference content.
[0089] The connection can be determined by the content relevance of user combinations, specifically whether the relevance of the voice content of two adjacent users meets a preset standard. For example, whether the content relevance between user A and user B meets the standard; if it does, a connection is determined. If not, the extraction point is directly not used as a voice conversion point.
[0090] Step S655: Determine whether the extracted point should be used as a speech conversion point based on the correlation.
[0091] You can set a correlation standard; if the correlation standard is met, it is used as a speech conversion point; otherwise, it is not used as a speech conversion point.
[0092] Step S656: If not associated, determine whether the extracted point should be used as a speech conversion point based on the background noise probability.
[0093] In practical applications, if there is no correlation, the method described above for determining whether an extraction point should be used as a speech conversion point is applied based on the probability of background noise. If the first user's speech duration covers the extraction point's speech duration, a new speaker may exist, and the speech segment was generated by this new speaker. However, if the first user cannot cover the entire extraction point's speech duration, it is examined whether several users can combine to cover it. For example, if the extraction point's speech duration is from 1:00 to 1:10, user A's speech duration is from 1:00 to 1:02, user B's speech duration is from 1:03 to 1:07, and user C's speech duration is from 1:08 to 1:10, then the combination of users A, B, and C covers the complete extraction point's speech duration. If a new speaker is involved, it is first determined whether the three users' content is connected. If not, it means the three people are not communicating, and therefore the extraction point's speech duration cannot be covered. If the extraction point's speech duration cannot be covered, the segment is considered background noise and not used as a conversion point, because human communication is continuous. If there is a connection, the judgment is based on the actual relevance of the content. If the content discussed by the three people is completely unrelated to the reference content, it is also judged as background noise and not used as a speech conversion point.
[0094] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering, characterized in that, Includes the following steps: Multiple extraction points are set for the target speech, and the corresponding voiceprints are extracted from the extraction points and recorded as the extraction point voiceprints. The extraction point that has been analyzed and is closest to the extraction point in real-time analysis is used as the reference point. The reference speaker corresponding to the reference point is collected, and the voiceprint of the reference speaker is recorded as the reference voiceprint. Compare whether the extracted voiceprint is consistent with the reference voiceprint. If they are inconsistent, collect the voiceprint perception density of the extracted point. The system determines whether there is overlap in the speech of multiple people based on the voiceprint perception density. If there is no overlap in the speech of multiple people, the system collects the health information and environmental information of the reference speaker. Based on health and environmental information, determine whether the extracted voiceprint is a reference voiceprint; if not, use the extracted voiceprint as a speech conversion point. If multiple people are speaking and there is overlap, the speech content and background noise probability of the extraction point are collected to determine whether the extraction point should be used as a speech conversion point. Speech segments are obtained by dividing speech based on speech conversion points. Speaker segmentation is achieved by classifying the speech segments by speaker.
2. The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to claim 1, characterized in that, The process of determining whether the extracted voiceprint is a reference voiceprint based on health and environmental information, and if not, using the extracted voiceprint as a speech conversion point, is as follows: Collect the speech content of the extraction point, determine whether the speech content is physiological speech, and if the speech content is physiological speech, then determine the voiceprint of the extraction point as the reference voiceprint. If the speech content is not physiological speech, the initial background sound of the reference speaker is extracted, and the extracted background sound of the extracted points is collected; the background similarity between the initial background sound and the extracted background sound is compared. Determine whether the background similarity meets the preset background similarity standard. If it does, extract the environment-related speech content from the target speech. Determine whether the extracted voiceprint is a reference voiceprint based on the environmental context and health information. If the preset background similarity standard is not met, the voiceprint feature change information of the entire speech segment corresponding to the extraction point is collected, and the voiceprint feature change information is used to determine whether the voiceprint of the extraction point is a reference voiceprint.
3. The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to claim 2, characterized in that, The steps for determining whether an extracted voiceprint is a reference voiceprint based on environmentally relevant speech content and health information are as follows: Based on the environmentally relevant speech content, extract all speech content from which the speaker is uncomfortable with the environment as environmentally uncomfortable content; The frequency of the reference speaker's expression of environmental discomfort was counted, the proportion of the reference speaker's physiological speech characteristics was collected, and the probability of the reference speaker's physical discomfort was obtained by combining the data. The health score of the reference speaker is extracted based on the health information of the reference speaker, and the voiceprint change is obtained by combining the probability of physical discomfort. Compare the voiceprint difference between the extracted point and the reference voiceprint to determine if the voiceprint difference is greater than the voiceprint change. If it is greater than the voiceprint change, then the extracted point voiceprint is not the reference voiceprint.
4. The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to claim 2, characterized in that, The steps for collecting and extracting the voiceprint feature change information of the entire speech segment corresponding to the extraction point, and determining whether the voiceprint of the extraction point is a reference voiceprint based on the voiceprint feature change information, are as follows: The extracted voiceprint is compared with all known voiceprints to obtain multiple similarities. The similarity between the extracted voiceprint and the reference voiceprint is recorded as the reference similarity. It is then determined whether the reference similarity is the maximum. If the reference similarity is the maximum, then the background sound features shared by the extraction point and the reference point are collected, and the similarity of the shared background sound features is obtained and recorded as the background feature similarity. Calculate the difference between the background feature similarity and the reference similarity, and determine whether the difference is within the preset standard. If it is within the preset standard, then determine that the extracted point voiceprint is the reference voiceprint. If the extraction point and the reference point do not share common background sound features, then collect the voiceprint changes of other extraction points to determine whether the voiceprint of the extraction point is the reference voiceprint. If the reference similarity is not the maximum, then the extracted voiceprint is determined to be not the reference voiceprint.
5. The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to claim 4, characterized in that, The steps for collecting voiceprint changes from other extraction points and determining whether the voiceprint at an extraction point is a reference voiceprint are as follows: Collect unanalyzed extraction points and record them as test points. Extract the test voiceprints of the test points and compare the test voiceprints with all known voiceprints to obtain the test similarity. Extract the maximum similarity scores corresponding to all voiceprints to be tested, set similarity intervals, and classify all maximum similarity scores according to the similarity intervals; Select the most numerous similarity intervals and calculate the variance of the maximum test similarity within each similarity interval; Determine whether the variance is within the preset variance standard. If it is within the preset variance standard, then determine that the extracted voiceprint is the reference voiceprint.
6. The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to claim 1, characterized in that, The steps for collecting the speech content and background noise probability of the extraction point, and determining whether the extraction point should be used as a speech conversion point, are as follows: The speech content of the extraction points is collected and segmented according to the voiceprint features to obtain multiple user voices; Collect user voiceprints, compare all user voiceprints with reference voiceprints to obtain similarity scores and record them as user similarity scores, and determine whether there is a reference speaker among the users based on user similarity scores; If a reference speaker exists, then users who are not reference speakers are recorded as test users, and the overlapping speech content of the test users is extracted. Based on the overlapping speech content, it is determined whether the extraction point can be used as a speech conversion point. If no reference speaker exists, then the similarity score is used to determine whether a known speaker exists among the users. If a known speaker exists, the extracted point is used as a speech conversion point; if no known speaker exists, the probability of background noise determines whether the extracted point should be used as a speech conversion point.
7. The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to claim 6, characterized in that, The steps for extracting overlapping speech content from the user under test and determining whether the extracted points should be used as speech conversion points based on the overlapping speech content are as follows: Obtain the ratio of speech overlap duration to total speech duration and record it as the overlap ratio; determine whether the overlap ratio meets the preset overlap standard. If the preset overlap standard is met, the voice content of the reference speaker is extracted and recorded as the reference content, and the voice content of the user to be tested is collected and recorded as the test content. Assess the degree of correlation between the reference content and the content to be tested, and determine whether the extraction point should be used as a speech conversion point based on the degree of correlation; If the preset overlap standard is not met, the emotional values of the reference speaker and the user being tested are extracted based on the voice content. Determine whether the emotion value meets the preset emotion standard. If the preset emotion standard is met, then the extraction point is determined as the speech conversion point. If the preset emotional standard is not met, the probability of background noise in the user's voice content is estimated, and the extraction point is used as a voice conversion point based on the background noise probability.
8. The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to claim 7, characterized in that, The steps for estimating the probability of background noise in the voice content of a user being tested are as follows: Collect the voice content of the users to be tested, check whether the voice content is an existing work, and filter the users whose voice content is not an existing work as the basic users; Collect the average voice volume of basic users, collect the average voice volume of known speakers, and calculate the difference between the average voice volume of basic users and known speakers as the volume difference; The average number of basic users is extracted from the overlapping speech, and the background noise probability is obtained by combining the volume difference and the overlap ratio.
9. The speaker segmentation method based on multi-scale feature fusion and voiceprint perception density clustering according to claim 7, characterized in that, If no known speaker exists, the step of determining whether the extracted point should be used as a speech conversion point based on the background noise probability is as follows: The user with the longest voice duration in the voice content extracted from the voiceprint collection point is recorded as the first user; Determine whether the duration of the first user's voice covers the duration of the voice at the extraction point. If it does, extract the first user's voice content and evaluate the correlation between the first user's voice content and the reference content. If not covered, collect the distribution of user voice duration in the voice content of the extraction point, and find the user combination whose voice duration distribution covers the voice duration of the extraction point; Determine whether the content of the user group is coherent. If it is coherent, extract the combined content of the user group and evaluate the relevance between the combined content and the reference content. Determine whether an extraction point should be used as a speech conversion point based on its relevance. If there is no correlation, the extraction point is determined as a speech conversion point based on the probability of background noise.
Citation Information
Cited By
Electronic mediation protocol tamper-proofing verification method based on dual hash and block chain
CN122333548A
Electronic mediation protocol tamper-proofing verification method based on double hashing and blockchain
CN122333548B