Marketing scene speaker distinguishing method and device based on multi-modal fusion algorithm
By using a multimodal fusion algorithm to segment, denoise, perform initial classification, and semantic analysis on audio, the problem of traditional speech recognition technology being unable to distinguish between unregistered new consultants and customers in real estate marketing scenarios is solved, achieving more efficient voice role differentiation and more accurate customer profiling.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional voice recognition technology cannot effectively distinguish between unregistered new consultants and customers in real estate marketing scenarios, and lacks the ability to adapt to unregistered scenarios.
A speaker differentiation method based on a multimodal fusion algorithm is adopted. This method involves segmenting the audio to be recognized, noise reduction and enhancement, initial classification, voiceprint feature extraction and semantic analysis, combined with adaptive processing of registered and unregistered modes, to achieve complex voice role differentiation.
It improves the efficiency of distinguishing complex voice roles in marketing scenarios, provides adaptive processing capabilities for non-registered scenarios, and enhances the accuracy of customer profiling and sales service efficiency.
Smart Images

Figure CN121768403A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition, and in particular to a method and apparatus for speaker differentiation in marketing scenarios based on a multimodal fusion algorithm. Background Technology
[0002] In real estate sales scenarios, it's crucial to accurately distinguish between sales staff and customers to effectively identify and evaluate customer profiles and sales services. However, in real estate sales offices and delivery sites, traditional voice recognition technology can only match existing voiceprint data to confirm the speaker's role, lacking adaptive processing capabilities for unregistered scenarios. For example, 30% of sales offices may have unregistered new consultants or temporary support consultants; in such cases, traditional recognition technology cannot differentiate between new consultants and customers. Summary of the Invention
[0003] To overcome the shortcomings of the prior art, this invention provides a method and apparatus for speaker differentiation in marketing scenarios based on a multimodal fusion algorithm, which can improve the efficiency of distinguishing complex voice roles in marketing scenarios.
[0004] An embodiment of the present invention provides a speaker differentiation method for marketing scenarios based on a multimodal fusion algorithm, comprising the following steps: The audio to be identified is segmented to obtain several sub-audio segments and the initial speaker classification results corresponding to each sub-audio segment; The system can detect the existence of a pre-registered voiceprint database in real time and determine the speaker's role based on the detection results; wherein, the pre-registered voiceprint database includes pre-registered voiceprint features corresponding to several pre-registered speakers; When a preset registered voiceprint library exists, the voiceprint features of the sub-audio are extracted, and the extraction results are matched with the registered voiceprint library to obtain a matching result; When the preset registered voiceprint library does not exist or the matching result is a match failure, the speaker corresponding to the sub-audio is determined according to the initial classification result, and semantic analysis is performed on all sub-audio corresponding to the speaker to determine the role determination result of the speaker.
[0005] Furthermore, the segmentation of the audio to be identified, resulting in several sub-audio segments and a preliminary speaker classification result corresponding to each sub-audio segment, specifically includes: The audio to be identified is subjected to noise reduction and enhancement processing to obtain preprocessed audio; The preprocessed audio is segmented using a preset ASR technique to obtain several sub-audio files; The pre-defined ASR technology is used to perform initial classification on the several sub-audio files, and the initial speaker classification result corresponding to each sub-audio file is obtained.
[0006] Furthermore, the extraction of the voiceprint features of the sub-audio specifically includes: The spectral features of the sub-audio are extracted by a preset residual network, and the temporal features of the sub-audio are extracted by a preset temporal convolutional network. The spectral features and the temporal features are fused to obtain the voiceprint features.
[0007] Furthermore, the step of matching the extracted results with the registered voiceprint database to obtain matching results specifically includes: The extracted results are matched with all the pre-registered voiceprint features to obtain the matching results; When the matching result is successful, the speaker corresponding to the sub-audio is determined to be the speaker corresponding to the pre-registered voiceprint feature that matches the extraction result, and the role of the speaker corresponding to the sub-audio is the role of the speaker corresponding to the pre-registered voiceprint feature that matches the extraction result.
[0008] Furthermore, determining the speaker corresponding to the sub-audio based on the initial classification result specifically includes: Based on the initial classification result, a sub-audio is selected as a reference audio from the initial group corresponding to the sub-audio; wherein, one initial group corresponds to one speaker, and the initial group includes several sub-audio that were classified as the same speaker during the initial classification; Based on the reference audio, the sub-audio is subjected to same-group voiceprint verification to obtain the same-group similarity between the sub-audio and the reference audio. When the voiceprint similarity is greater than a preset same-group similarity threshold, the speaker corresponding to the sub-audio is determined to be the speaker corresponding to the reference audio. When the voiceprint similarity is less than the preset same-group similarity threshold, cross-group voiceprint verification is performed on the sub-audio to obtain the cross-group similarity between the sub-audio and the reference audios of the other initial groups. Based on the comparison result between the cross-group similarity and the preset cross-group similarity threshold, the speaker corresponding to the sub-audio is determined. When the similarity between the sub-audio and all the reference audios in the initial group does not exceed a preset cross-group similarity threshold, a new group is created, and the sub-audio is included in the new group.
[0009] Furthermore, the step of performing semantic analysis on all sub-audio segments corresponding to the speaker to determine the role determination result of the speaker corresponding to the sub-audio segment specifically includes: Extract the text content of all sub-audio segments corresponding to the speaker, and preprocess the text content to obtain valid sentences; wherein, the preprocessing includes, but is not limited to, word segmentation, noise reduction and standardization. The valid statements are then filtered to obtain the filtering results. A confidence analysis is performed on the filtering results to determine the role determination result.
[0010] Furthermore, the term filtering of the valid statements to obtain the filtering results specifically includes: Construct a domain terminology library for the scene corresponding to the audio to be identified, the domain terminology library including several professional terms; The text content is encoded using a pre-defined domain BERT model to obtain the encoding result; The encoding result is standardized and filtered according to the domain terminology database to obtain the filtered result.
[0011] Furthermore, the step of performing confidence analysis on the filtering results to determine the role determination result specifically includes: Calculate the similarity between the filtering results and the preset sales corpus and the preset customer corpus, respectively; When the similarity between the filtering result and the preset sales corpus is greater than the preset sales threshold, the role determination result is determined to be sales. When the similarity between the filtering result and the preset customer corpus is greater than the preset customer threshold, the role determination result is determined to be a customer.
[0012] Another embodiment of the present invention provides a speaker differentiation device for marketing scenarios based on a multimodal fusion algorithm, comprising: a segmentation module, a detection module, a matching module, and an analysis module; The segmentation module is used to segment the audio to be identified, and obtain several sub-audio files and the initial speaker classification results corresponding to each sub-audio file; The detection module is used to detect in real time whether a pre-registered voiceprint database exists and to determine the speaker role based on the detection results; wherein, the registered voiceprint database includes pre-registered voiceprint features corresponding to several pre-registered speakers; The matching module is used to extract the voiceprint features of the sub-audio when a preset registered voiceprint library exists, and to match the extraction result with the registered voiceprint library to obtain a matching result; The analysis module is used to determine the speaker corresponding to the sub-audio based on the initial classification result when the preset registered voiceprint library does not exist or the matching result is a match failure, and to perform semantic analysis on all sub-audio corresponding to the speaker to determine the speaker's role determination result.
[0013] Furthermore, the segmentation module is used to segment the audio to be recognized, obtaining several sub-audio segments and the initial speaker classification result corresponding to each sub-audio segment, specifically including: The audio to be identified is subjected to noise reduction and enhancement processing to obtain preprocessed audio; The preprocessed audio is segmented using a preset ASR technique to obtain several sub-audio files; The pre-defined ASR technology is used to perform initial classification on the several sub-audio files, and the initial speaker classification result corresponding to each sub-audio file is obtained.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: By establishing a dual-mode speaker classification method with both registered and unregistered modes, and a seamless degradation processing method between the two modes, the recognition system can be provided with adaptive processing capabilities for unregistered scenarios, thereby improving the efficiency of distinguishing complex voice roles in marketing scenarios. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating a speaker differentiation method for marketing scenarios based on a multimodal fusion algorithm, as provided in an embodiment of the present invention.
[0016] Figure 2 This is a flowchart illustrating a dual-mode audio processing method according to an embodiment of the present invention.
[0017] Figure 3 This is a schematic diagram of a multi-model voiceprint fusion process provided in an embodiment of the present invention.
[0018] Figure 4 This is a schematic diagram of a semantic analysis process provided in an embodiment of the present invention.
[0019] Figure 5 This is a schematic diagram of a speaker differentiation device for a marketing scenario based on a multimodal fusion algorithm, provided as another embodiment of the present invention. Detailed Implementation
[0020] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent. It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.
[0021] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0022] Reference Figure 1 The following is a flowchart illustrating a speaker differentiation method for marketing scenarios based on a multimodal fusion algorithm, provided by an embodiment of the present invention, including the following steps: S1: Segment the audio to be identified to obtain several sub-audio segments and the initial speaker classification results corresponding to each sub-audio segment; S2: Real-time detection of the existence of a pre-registered voiceprint database, and determination of the speaker role based on the detection results; wherein, the registered voiceprint database includes pre-registered voiceprint features corresponding to several pre-registered speakers; S3: When a preset registered voiceprint library exists, extract the voiceprint features of the sub-audio and match the extraction result with the registered voiceprint library to obtain a matching result; S4: When the preset registered voiceprint library does not exist or the matching result is a match failure, determine the speaker corresponding to the sub-audio based on the initial classification result, and perform semantic analysis on all sub-audio corresponding to the speaker to determine the speaker's role determination result.
[0023] For step S1, specifically, the segmentation of the audio to be identified to obtain several sub-audio segments and the initial speaker classification result corresponding to each sub-audio segment includes: The audio to be identified is subjected to noise reduction and enhancement processing to obtain preprocessed audio; The preprocessed audio is segmented using a preset ASR technique to obtain several sub-audio files; The pre-defined ASR technology is used to perform initial classification on the several sub-audio files, and the initial speaker classification result corresponding to each sub-audio file is obtained.
[0024] In a preferred embodiment, the audio to be identified originates from a complex noisy environment and contains speech from multiple people conversing. Therefore, after obtaining the audio to be identified, it needs to undergo noise reduction and enhancement processing first. Subsequently, the preprocessed audio is segmented using ASR technology. The processing results include the start time, end time, speaker, and speech content of each sub-audio (here, the speaker classification is the initial classification result, which has insufficient accuracy).
[0025] For step S2, specifically, before determining the speaker's role, it is necessary to first check whether the preset registered voiceprint database exists in the system. The preset registered voiceprint database consists of the voiceprint features of several pre-recorded sales personnel, used to improve the efficiency of speaker and role identification based on voiceprint features. (Refer to...) Figure 2 This is a flowchart illustrating a dual-mode audio processing method according to an embodiment of the present invention. Figure 2 It can be seen that by detecting whether a pre-registered voiceprint database exists in the system, it is possible to determine whether to enter the registration mode or the non-registration mode.
[0026] For step S3, specifically, extracting the voiceprint features of the sub-audio audio includes: The spectral features of the sub-audio are extracted by a preset residual network, and the temporal features of the sub-audio are extracted by a preset temporal convolutional network. The spectral features and the temporal features are fused to obtain the voiceprint features.
[0027] In a preferred embodiment, multi-model voiceprint fusion processing during the voiceprint feature extraction stage is mainly used to improve the robustness and discriminativeness of voiceprint features.
[0028] Reference Figure 3 This is a schematic diagram of a multi-model voiceprint fusion process according to an embodiment of the present invention. First, the Resemblyzer model is used to extract the spectral features of the sub-audio, obtaining a 512-dimensional noise-resistant voiceprint feature. Simultaneously, the PyAnnote model is used to capture the long-term dependencies in the sub-audio, obtaining a 256-dimensional temporal feature, which has good temporal modeling capabilities. Finally, the 512-dimensional noise-resistant voiceprint feature and the 256-dimensional temporal feature are fused using a weighted fusion method, and the fused feature is normalized to obtain a 768-dimensional fused feature, which is the voiceprint feature. The feature fusion formula is: E_fused = 0.6 × E_resemblyzer + 0.4 × E_pyannote, where E_fused is the voiceprint feature, E_resemblyzer is the 512-dimensional noise-resistant voiceprint feature, and E_pyannote is the 256-dimensional temporal feature.
[0029] This step combines noise resistance and temporal modeling capabilities through heterogeneous feature complementarity, balancing accuracy and efficiency, thereby improving the robustness and discriminativeness of voiceprint features.
[0030] Furthermore, the step of matching the extracted results with the registered voiceprint database to obtain matching results specifically includes: The extracted results are matched with all the pre-registered voiceprint features to obtain the matching results; When the matching result is successful, the speaker corresponding to the sub-audio is determined to be the speaker corresponding to the pre-registered voiceprint feature that matches the extraction result, and the role of the speaker corresponding to the sub-audio is the role of the speaker corresponding to the pre-registered voiceprint feature that matches the extraction result.
[0031] In a preferred embodiment, by Figure 2 It is understood that when the system has a pre-set registered voiceprint database containing voiceprint data of several service advisors, the registration mode is activated, and then the extracted results are matched according to the registered voiceprint database. Specifically, matching is performed by calculating the weighted distance between the extracted results and each voiceprint data in the voiceprint database. If the match is successful, the speaker corresponding to the extracted result is directly labeled as the advisor corresponding to the voiceprint data that matches the extracted result; if the match fails, the system switches to the no-registration mode for processing.
[0032] Preferably, the specific process of the registration mode can be summarized as follows: # Voiceprint Feature Extraction audio_segment = extract_audio(audio_data, start_time, end_time) voiceprint = multi_model_extract(audio_segment) # Registry matching best_match = None min_distance = float('inf') for consultant_id, registered_voiceprint in voiceprint_db.items(): distance = weighted_distance(voiceprint, registered_voiceprint) if distance <min_distance and distance<0.25: min_distance = distance best_match = consultant_id # Output Results if best_match: return f"Sales Consultant({best_match})" else: # Downgrade to no registration mode return unregistered_mode_processing(audio_segment) For step S3, specifically, the step of grouping the sub-audio segments by speaker based on the initial classification result and determining the speaker corresponding to each sub-audio segment includes: Based on the initial classification result, a sub-audio is selected as a reference audio from the initial group corresponding to the sub-audio; wherein, one initial group corresponds to one speaker, and the initial group includes several sub-audio that were classified as the same speaker during the initial classification; Based on the reference audio, the sub-audio is subjected to same-group voiceprint verification to obtain the same-group similarity between the sub-audio and the reference audio. When the voiceprint similarity is greater than a preset same-group similarity threshold, the speaker corresponding to the sub-audio is determined to be the speaker corresponding to the reference audio. When the voiceprint similarity is less than the preset same-group similarity threshold, cross-group voiceprint verification is performed on the sub-audio to obtain the cross-group similarity between the sub-audio and the reference audios of the other initial groups. Based on the comparison result between the cross-group similarity and the preset cross-group similarity threshold, the speaker corresponding to the sub-audio is determined. When the similarity between the sub-audio and all the reference audios in the initial group does not exceed a preset cross-group similarity threshold, a new group is created, and the sub-audio is included in the new group.
[0033] In a preferred embodiment, refer to Figure 2 When the registered voiceprint library does not exist in the system, or matching fails in registration mode, the no-registration mode is activated. First, a dynamic reference voiceprint selection strategy is used to select samples from the speaker grouping results of the preliminary ASR results corresponding to the sub-audio as the reference audio. In this preferred embodiment, the reference audio is the longest speech segment in each group.
[0034] Once the reference audio is obtained, the sub-audio and the reference audio are compared and verified, and the similarity between them within the same group is calculated. When the similarity within the same group is greater than a preset similarity threshold, the initial classification result can be determined to be correct. In this preferred embodiment, the preset similarity threshold is set to 0.6.
[0035] When the same-group similarity is less than the preset same-group similarity threshold, cross-group comparison verification is performed, that is, the sub-audio is sequentially compared with the reference audio of each other group for cross-group similarity calculation. When the cross-group similarity between the sub-audio and a certain group of reference audio is greater than the preset cross-group similarity threshold, it can be determined that the initial classification result is incorrect, and the speaker corresponding to the sub-audio should be the speaker corresponding to that group.
[0036] When the similarity between the sub-audio and all the reference audios of the initial group does not meet the condition, a new group is created, corresponding to a new speaker, and the sub-audio becomes the reference audio of the new group.
[0037] Furthermore, the step of performing semantic analysis on all sub-audio segments corresponding to the speaker to determine the role determination result of the speaker corresponding to the sub-audio segment specifically includes: Extract the text content of all sub-audio segments corresponding to the speaker, and preprocess the text content to obtain valid sentences; wherein, the preprocessing includes, but is not limited to, word segmentation, noise reduction and standardization. The valid statements are then filtered to obtain the filtering results. A confidence analysis is performed on the filtering results to determine the role determination result.
[0038] Furthermore, the term filtering of the valid statements to obtain the filtering results specifically includes: Construct a domain terminology library for the scene corresponding to the audio to be identified, the domain terminology library including several professional terms; The text content is encoded using a pre-defined domain BERT model to obtain the encoding result; The encoding result is standardized and filtered according to the domain terminology database to obtain the filtered result.
[0039] In a preferred embodiment, in the no-registration mode, once the speaker corresponding to the sub-audio is determined, the role of the speaker can be determined through semantic analysis. Before the determination, it is necessary to construct a domain terminology library corresponding to the marketing scenario to which the audio to be identified belongs. The domain terminology library contains several professional terms in the corresponding domain of the marketing scenario. In this preferred embodiment, taking the real estate field as an example, the domain terminology library may contain 328 professional terms, such as: plot ratio, usable floor area ratio, LPR, etc.
[0040] Reference Figure 4 This is a schematic diagram of a semantic analysis process provided in an embodiment of the present invention. Figure 4 It is understood that after obtaining the content text, terminology filtering needs to be performed based on the domain terminology database. The terminology filtering process is as follows: 1. Construct a terminology database for the real estate sector (containing 328 professional terms). 2. Encode the input text using a domain-specific BERT model; 3. Identify and standardize technical terms in the text: For example, map "plot ratio" to "planning index", "usable floor area ratio" to "usable area ratio", and "LPR" to "loan interest rate". 4. Output the standardized text and terminology recognition results.
[0041] Furthermore, the step of performing confidence analysis on the filtering results to determine the role determination result specifically includes: Calculate the similarity between the filtering results and the preset sales corpus and the preset customer corpus, respectively; When the similarity between the filtering result and the preset sales corpus is greater than the preset sales threshold, the role determination result is determined to be sales. When the similarity between the filtering result and the preset customer corpus is greater than the preset customer threshold, the role determination result is determined to be a customer.
[0042] In a preferred embodiment, refer to Figure 4 After filtering is complete, confidence analysis is performed based on the filtered text, specifically: 1. Calculate two confidence scores separately: Sales Consultant Confidence Score: similarity between the text and the sales corpus; Customer Confidence Score: similarity between the text and the customer corpus; 2. Compare the two confidence scores: When the sales consultant's confidence score > the customer's confidence score, label the speaker as a sales consultant; when the customer's confidence score > the sales consultant's confidence score, label the speaker as a customer.
[0043] 3. Output the final role labeling results and confidence scores.
[0044] Reference Figure 5 The following is a schematic diagram of the structure of a speaker differentiation device for a marketing scenario based on a multimodal fusion algorithm, provided in another embodiment of the present invention, including: a segmentation module 101, a detection module 102, a matching module 103, and an analysis module 104; The segmentation module 101 is used to segment the audio to be identified, and obtain several sub-audio and the initial speaker classification result corresponding to each sub-audio; The detection module 102 is used to detect in real time whether a preset registered voiceprint library exists, and to determine the speaker role based on the detection results; wherein, the registered voiceprint library includes several pre-registered voiceprint features corresponding to pre-registered speakers; The matching module 103 is used to extract the voiceprint features of the sub-audio when a preset registered voiceprint library exists, and to match the extraction result with the registered voiceprint library to obtain a matching result; The analysis module 104 is used to group the sub-audio into speakers based on the initial classification results when the preset registered voiceprint library does not exist or the extraction result fails to match, determine the speaker corresponding to the sub-audio, and finally perform semantic analysis on the sub-audio to determine the role determination result of the speaker corresponding to the sub-audio.
[0045] Furthermore, the segmentation module 101 is used to segment the audio to be recognized, obtaining several sub-audio segments and the initial speaker classification result corresponding to each sub-audio segment, specifically including: The audio to be identified is subjected to noise reduction and enhancement processing to obtain preprocessed audio; The preprocessed audio is segmented using a preset ASR technique to obtain several sub-audio files; The pre-defined ASR technology is used to perform initial classification on the several sub-audio files, and the initial speaker classification result corresponding to each sub-audio file is obtained.
[0046] Finally, this embodiment of the invention provides a practical application scenario as an example to illustrate the superiority of the speaker differentiation method for marketing scenarios based on a multimodal fusion algorithm provided by this invention: The voice employee badge project for on-site marketing of a real estate project incorporates the method described in the embodiments of this invention. It is used to capture the dialogue content between customers and sales staff in real time in scenarios such as customer reception, property viewing, and consultation, and to perform semantic analysis and role recognition.
[0047] The system utilizes purchased voice ID badge hardware for recording, processes the recording files using this technology's algorithm to accurately distinguish between sales staff and customers, outputs translated text content, and then combines it with other technical models to analyze customer profiles and sales services.
[0048] The specific application process is as follows: 1. Scene Setup: Project size: 200 units for sale Daily customer reception capacity: 50-80 groups Sales staff: 15 (including 3 newly hired consultants) Hardware equipment: 30 sets of smart voice employee badges 2. Technical Implementation Details: (1) System initialization phase Register voiceprint information for 12 senior consultants Establish a terminology database for the real estate sector (328 terms). Loading pre-trained domain BERT models (2) Real-time processing stage Voice ID badges collect real-time conversation audio. The system automatically detects the status of the voiceprint database. Select processing mode based on consultant registration status. (3) Registration mode processing example # Voiceprint Feature Extraction audio_segment = extract_audio(audio_data, start_time, end_time) voiceprint = multi_model_extract(audio_segment) # Registry matching best_match = None min_distance = float('inf') for consultant_id, registered_voiceprint in voiceprint_db.items(): distance = weighted_distance(voiceprint, registered_voiceprint) if distance <min_distance and distance<0.25: min_distance = distance best_match = consultant_id # Output Results if best_match: return f"Sales Consultant({best_match})" else: # Downgrade to no registration mode return unregistered_mode_processing(audio_segment) (4) Example of handling the no-registration mode Dynamically select reference voiceprints (the longest speech segment in each group). Intra-group voiceprint verification (similarity threshold 0.6) Cross-group voiceprint matching (similarity threshold 0.7) New speaker group creation (5) Semantic Enhancement Decision Process ASR text transcription: "What is the usable floor area ratio of this apartment type?" Terminology filtering: "Usable floor area ratio" → "Practical area ratio" Semantic analysis: Calculating similarity with the client's corpus Confidence level assessment: Customer confidence level 0.72 > threshold 0.60 Decision output: labeled "Customer" 5.2 Implementation Effectiveness Verification A three-month field test was conducted at Yuexiu Property's Guangzhou Guanyue project: Test data: Total processing time: 1,200 hours Dialogue segments processed: 85,000 segments Speakers involved: 2,800 Performance results:
[0049] Commercial value: Improved customer profiling accuracy facilitates precision marketing. Increased efficiency and reduced labor costs; Increased conversion rates directly boost sales revenue; Increased customer satisfaction enhances brand influence.
[0050] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A marketing scene speaker diarization method based on a multi-modal fusion algorithm, characterized in that, The method comprises the following steps: segmenting the audio to be identified to obtain a plurality of sub-audios and speaker initial classification results corresponding to each of the sub-audios; wherein one speaker corresponds to a plurality of sub-audios; detecting in real time whether a preset registered voiceprint library exists, and determining a speaker role according to a detection result; wherein the registered voiceprint library comprises pre-registered voiceprint features corresponding to a plurality of pre-registered speakers; when the preset registered voiceprint library exists, extracting voiceprint features of the sub-audios, and matching the extraction result with the registered voiceprint library to obtain a matching result; when the preset registered voiceprint library does not exist or the matching result is a matching failure, determining a speaker corresponding to the sub-audios according to the initial classification result, and performing semantic analysis on all sub-audios corresponding to the speaker to determine a role determination result of the speaker.
2. The marketing scenario speaker diarization method based on multi-modal fusion algorithm of claim 1, wherein, The method of segmenting the audio to be identified to obtain a plurality of sub-audios and speaker initial classification results corresponding to each of the sub-audios specifically comprises: performing noise reduction and enhancement processing on the audio to be identified to obtain a preprocessed audio; segmenting the preprocessed audio by using a preset ASR technology to obtain the plurality of sub-audios; performing initial classification on the plurality of sub-audios by using a preset ASR technology to obtain speaker initial classification results corresponding to each of the sub-audios.
3. The marketing scenario speaker diarization method based on multi-modal fusion algorithm of claim 1, wherein, The method of extracting voiceprint features of the sub-audios specifically comprises: extracting spectral features of the sub-audios by using a preset residual network, and simultaneously extracting time sequence features of the sub-audios by using a preset time sequence convolution network; performing feature fusion on the spectral features and the time sequence features to obtain the voiceprint features.
4. The marketing scenario speaker diarization method based on multi-modal fusion algorithm of claim 1, wherein, The method of matching the extraction result with the registered voiceprint library to obtain a matching result specifically comprises: matching the extraction result with all the pre-registered voiceprint features to obtain a matching result; when the matching result is a matching success, determining that a speaker corresponding to the sub-audios is a speaker corresponding to a pre-registered voiceprint feature matched with the extraction result, and a role of the speaker corresponding to the sub-audios is a role of the speaker corresponding to the pre-registered voiceprint feature matched with the extraction result.
5. The marketing scenario speaker diarization method based on multi-modal fusion algorithm as claimed in claim 1, wherein, The method of determining a speaker corresponding to the sub-audios according to the initial classification result specifically comprises: selecting a sub-audio from an initial group corresponding to the sub-audios as a reference audio according to the initial classification result; wherein one initial group corresponds to one speaker, and the initial group comprises a plurality of sub-audios classified as the same speaker during initial classification; performing same-group voiceprint verification on the sub-audios according to the reference audio to obtain a same-group similarity between the sub-audios and the reference audio, and when the voiceprint similarity is greater than a preset same-group similarity threshold, determining that the speaker corresponding to the sub-audios is a speaker corresponding to the reference audio; when the voiceprint similarity is less than the preset same-group similarity threshold, performing cross-group voiceprint verification on the sub-audios to obtain cross-group similarities between the sub-audios and reference audios of the remaining initial groups, and determining the speaker corresponding to the sub-audios according to a comparison result of the cross-group similarities and a preset cross-group similarity threshold. When the similarity between the sub-audio and all the reference audios of the initial group does not exceed a preset cross-group similarity threshold, a new group is created, and the sub-audio is included in the new group.
6. The marketing scenario speaker diarization method based on multi-modal fusion algorithm as claimed in claim 1, wherein, The semantic analysis on all the sub-audios corresponding to the speaker determines a role determination result of the speaker corresponding to the sub-audios, and specifically includes: The text content of all the sub-audios corresponding to the speaker is extracted, and the text content is preprocessed to obtain effective sentences; wherein the preprocessing includes but is not limited to word segmentation processing, noise removal processing and standardization processing; The effective sentences are subjected to term filtering to obtain a filtering result; The filtering result is subjected to confidence analysis to determine the role determination result.
7. The marketing scenario speaker diarization method based on multi-modal fusion algorithm of claim 6, wherein, The term filtering on the effective sentences to obtain a filtering result specifically includes: A domain term library of a scene corresponding to the to-be-recognized audio is constructed, and the domain term library includes a plurality of professional terms; The text content is encoded by a preset domain BERT model to obtain an encoding result; The encoding result is subjected to standardization filtering processing according to the domain term library to obtain the filtering result.
8. The marketing scenario speaker diarization method based on multi-modal fusion algorithm of claim 6, wherein, The confidence analysis on the filtering result to determine the role determination result specifically includes: The similarity between the filtering result and a preset sales corpus and a preset customer corpus is calculated respectively; When the similarity between the filtering result and the preset sales corpus is greater than a preset sales threshold, it is determined that the role determination result is sales; When the similarity between the filtering result and the preset customer corpus is greater than a preset customer threshold, it is determined that the role determination result is customer.
9. A marketing scene speaker diarization device based on a multi-modal fusion algorithm, characterized in that, It includes: A segmentation module, a detection module, a matching module and an analysis module; The segmentation module is used to segment the to-be-recognized audio to obtain a plurality of sub-audios and speaker initial classification results corresponding to each of the sub-audios; The detection module is used to detect whether a preset registered voiceprint library exists in real time, and determine the role of the speaker according to the detection result; wherein the registered voiceprint library includes a plurality of pre-registered voiceprint features corresponding to pre-registered speakers; The matching module is used to extract the voiceprint features of the sub-audios when the preset registered voiceprint library exists, and match the extraction result with the registered voiceprint library to obtain a matching result; The analysis module is used to determine the speaker corresponding to the sub-audios according to the initial classification result when the preset registered voiceprint library does not exist or the matching result is a matching failure, and perform semantic analysis on all the sub-audios corresponding to the speaker to determine the role determination result of the speaker.
10. The marketing scenario speaker diarization apparatus based on multi-modal fusion algorithm as claimed in claim 9, wherein, The segmentation module is used to segment the to-be-recognized audio to obtain a plurality of sub-audios and speaker initial classification results corresponding to each of the sub-audios, specifically including: The to-be-recognized audio is subjected to noise reduction and enhancement processing to obtain a preprocessed audio; The preprocessed audio is segmented by a preset ASR technology to obtain the plurality of sub-audios; The plurality of sub-audios are subjected to initial classification by a preset ASR technology respectively to obtain speaker initial classification results corresponding to each of the sub-audios.
Citation Information
Patent Citations
Speaker role recognition method and device thereof, electronic equipment and storage medium
CN112233680A
Land-air communication speaker identity recognition method and device
CN113066499A
Role identification method, device and system in dialogue scene
CN113744742A
Voice extraction method and device, neural network model training method and device and storage medium
CN115116448A
Role classification method and device, electronic equipment and storage medium
CN117351964A