Conference voice data processing method and device, electronic equipment and storage medium

CN121662055BActive Publication Date: 2026-07-24IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2026-02-05
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional conference speech recognition technology relies on pre-registered voiceprint databases, which leads to a lack of identity recognition and poor adaptability in scenarios with short speech, noise interference, and accent differences, making it unable to automatically infer identity.

Method used

By acoustically clustering conference audio to generate audio segment sequences with speaker labels, large language models are used to analyze dialogue context and correct errors. Based on text sequence inference, a closed-loop iterative system of perception-cognition-feedback is constructed to dynamically optimize clustering processing.

Benefits of technology

It enables the automatic and accurate association of anonymous voices with real identities without the need for pre-recorded voiceprint databases, significantly improving the system's accuracy and usability in complex scenarios and outputting high-quality, real-name meeting records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121662055B_ABST
    Figure CN121662055B_ABST
Patent Text Reader

Abstract

The application discloses a conference voice data processing method and device, electronic equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: detecting conference voice data, extracting acoustic features and performing clustering processing, and generating a voice sequence with a speaker label; after voice recognition, the speaker label is corrected and the real identity is inferred based on context analysis and identity indication information of the text sequence; the correction and inference results are fed back to the clustering processing as constraints, and iterative optimization is performed until the termination condition is met, and finally, conference transcription text with real identity information is output. The application does not require a pre-recorded voiceprint library, effectively solves the problems of identity recognition missing and errors in short voice, noise and new speaker scenarios, and realizes the output of high-quality real-name conference records.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a method, apparatus, electronic device, and storage medium for processing conference voice data. Background Technology

[0002] Current technologies primarily use acoustic models to separate speakers and extract features from conference audio, generating anonymous tags. For identity verification, the extracted acoustic features are matched against a pre-registered voiceprint database to determine identity. This approach has several drawbacks: First, it heavily relies on pre-recorded voiceprint databases, leading to high maintenance costs, privacy concerns, and cold-start issues with new speakers. Second, the acoustic model's feature extraction is unstable under short speech, environmental noise, or accent interference, resulting in separation failures or misidentification. Third, the output is limited to anonymous tags, unable to automatically infer identity, requiring manual secondary matching, thus limiting its practicality. Summary of the Invention

[0003] This invention provides a method, apparatus, electronic device, and storage medium for processing conference voice data, aiming to solve the problems of identity recognition failure caused by the reliance on pre-registered voiceprint databases in traditional conference voice recognition technology, as well as poor adaptability to scenarios such as short speech, noise interference, and accent differences of speakers.

[0004] Firstly, a method for processing conference voice data is provided, including:

[0005] The input conference audio data is subjected to speech detection, audio segments are extracted, and acoustic features are extracted and clustered on the audio segments. Speaker labels are assigned to each audio segment to generate a sequence of audio segments with speaker labels.

[0006] The speech segment sequence is subjected to speech recognition to obtain a text sequence with speaker tags; based on the text sequence, dialogue context analysis is performed to correct errors in the speaker tags to obtain corrected speaker tags; and based on the identity indication information in the text sequence, the real identity corresponding to each speaker tag is inferred to obtain the identity inference result.

[0007] The corrected speaker labels and the identity reasoning results are fed back to the clustering process as optimization constraints to perform iterative processing until a preset iteration termination condition is reached, and the meeting transcript with real identity information is output.

[0008] Secondly, a conference voice data processing device is also provided, comprising:

[0009] The speech processing module is used to perform speech detection on the input conference speech data, extract speech segments, perform acoustic feature extraction and clustering processing on the speech segments, assign speaker labels to each speech segment, and generate a speech segment sequence with speaker labels.

[0010] The identity reasoning module is used to perform speech recognition on the speech segment sequence to obtain a text sequence with speaker tags; perform dialogue context analysis based on the text sequence to correct errors in the speaker tags, obtain the corrected speaker tags, and infer the real identity corresponding to each speaker tag based on the identity indication information in the text sequence to obtain the identity reasoning result.

[0011] The iterative optimization module is used to feed back the corrected speaker labels and the identity inference results as optimization constraints to the clustering process to perform iterative processing until a preset iteration termination condition is reached, and output the meeting transcript with identity information.

[0012] Thirdly, this application also provides an electronic device, including at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the conference voice data processing method described in any one of the first aspects.

[0013] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the conference voice data processing method described in any one of the first aspects.

[0014] Beneficial effects:

[0015] This application proposes a conference speech data processing method that does not rely on pre-registered speaker databases and achieves iterative optimization. This addresses the problems of identity recognition failure and poor adaptability to scenarios involving short speech, noise interference, and accent differences inherent in traditional conference speech recognition technologies. Specifically, the method first performs acoustic clustering on the conference speech to generate a sequence of speech segments labeled with the speaker, which is then converted into a text sequence. Subsequently, the speaker labels are corrected using dialogue context analysis, and the speaker's identity is inferred based on identity indicators in the text (such as self-identification, other-identification, and role keywords). Finally, the corrected and inferred results are fed back to the front-end acoustic clustering process as optimization constraints, dynamically adjusting the clustering process. This iterative process gradually overcomes the instability of acoustic features caused by short speech, noise, or accents, ultimately achieving accurate identity inference and real-name transcription for any new speaker, thus eliminating the reliance on pre-registered speaker databases. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic flowchart of the conference voice data processing method provided in the embodiments of this application;

[0018] Figure 2 This is a schematic diagram of the architecture of the conference voice data processing device provided in the embodiments of this application;

[0019] Figure 3 This is a structural block diagram of the conference voice data processing device provided in the embodiments of this application;

[0020] Figure 4 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0023] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.

[0024] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated.

[0025] In this application, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0026] To address the shortcomings of traditional conference speech recognition technology, such as the lack of identity recognition due to reliance on pre-registered voiceprint databases and poor adaptability to scenarios involving short speech, noise interference, and accent differences, this application provides a conference speech data processing method based on semantic understanding and closed-loop optimization. By converting acoustic clustering results into text sequences and leveraging the deep semantic understanding capabilities of large language models, the method directly infers the speaker's identity from the dialogue context. Simultaneously, the semantic inference results are fed back to acoustic processing to dynamically optimize the clustering effect. This enables the automatic and accurate association of anonymous speech with real identities without the need for pre-recorded voiceprint databases, significantly improving the system's accuracy and practicality in complex scenarios.

[0027] On the one hand, such as Figure 1 As shown, this embodiment provides a method for processing conference voice data, including:

[0028] S101, perform speech detection on the input conference speech data, extract speech segments, perform acoustic feature extraction and clustering processing on the speech segments, assign speaker labels to each speech segment, and generate a speech segment sequence with speaker labels.

[0029] Understandably, this step is the initial perceptual processing stage. Specifically, it processes the input raw conference speech through Voice Activity Detection (VAD), removing silent and background noise segments to extract valid speech fragments. Next, a lightweight acoustic model (such as x-vector) extracts acoustic features from these speech fragments and performs preliminary speaker clustering based on feature similarity. This assigns an initial, anonymous speaker label (e.g., SPK_0, SPK_1) to each speech fragment, ultimately generating a structured sequence of speech fragments with speaker labels. Its purpose is to provide foundational material for subsequent deep semantic analysis. This stage allows its output to be "noisy." The goal is to transform the continuous speech stream into a structured sequence that can be processed by the cognitive layer, that is, to segment the speech stream and attach preliminary speaker labels, laying the necessary input foundation for subsequent semantic-based correction and identity reasoning.

[0030] S102, perform speech recognition on the speech segment sequence to obtain a text sequence with speaker tags; perform dialogue context analysis based on the text sequence to correct errors in the speaker tags, obtain the corrected speaker tags, and infer the real identity corresponding to each speaker tag based on the identity indication information in the text sequence to obtain the identity inference result.

[0031] Understandably, this step is the cognitive reasoning stage, specifically: First, speech recognition (ASR) is performed on the sequence of speech segments with initial speaker labels, converting them into a text sequence with the same speaker labels. Then, utilizing the semantic understanding capabilities of the large-scale cognitive model, the text sequence undergoes dialogue context analysis. By understanding the logical flow, topic evolution, and question-and-answer relationships of the entire dialogue, incorrect speaker labels in the initial acoustic clustering are identified and corrected (e.g., merging speech segments acoustically misclassified as different speakers but semantically belonging to the same person). Simultaneously, based on identity indicators in the text sequence (such as self-reference, other-reference, and role-related expressions like "I am Zhang San" or "Mr. Li, what do you think?"), the true identity (such as name and position) corresponding to each anonymous speaker label (e.g., SPK_0) is inferred. Its purpose is to utilize high-level semantic information to compensate for and correct the deficiencies of low-level acoustic perception. This stage realizes the transition from "anonymous speech streams" to "named meeting minutes," which not only solves the speaker discrimination errors caused by acoustic models in scenarios such as noise and short speech, but also eliminates the dependence on pre-recorded voiceprint databases, enabling direct identity recognition through content reasoning.

[0032] S103, the corrected speaker label and the identity reasoning result are fed back to the clustering process as optimization constraints to perform iterative processing until the preset iteration termination condition is reached, and the meeting transcript with real identity information is output.

[0033] Understandably, this step is the feedback optimization and closed-loop control stage. Specifically, the high-confidence correction results generated by the cognitive layer (e.g., "SPK_1 and SPK_3 should be the same person") and the identity reasoning results are used as strong constraint signals (i.e., "must link" or "must separate" constraint pairs) and transmitted back to the front-end acoustic clustering process. This dynamically adjusts the parameters or decision criteria of the clustering model, thereby re-merging incorrectly segmented speech segments in subsequent iterations within the acoustic feature space. This process is repeated until the speaker label sequence stabilizes (e.g., changes are below a threshold) or the maximum number of iterations is reached, ultimately outputting an optimized meeting transcript with real identity information. Its purpose is to construct a closed-loop optimization mechanism of "perception-cognition-feedback." This mechanism enables the system to self-correct and iteratively improve, using reliable results from high-level semantic reasoning to guide and optimize the unstable acoustic processing at the lower level. This significantly enhances the system's robustness to complex scenarios such as short speech, noise, and accents, ultimately achieving high-accuracy real-name transcription output.

[0034] Therefore, this application significantly improves the accuracy and robustness of meeting transcription by constructing a closed-loop iterative system that includes perception, cognition, and feedback. Specifically, it first extracts and clusters acoustic features from the original speech to initially distinguish speakers and generate labels; then, after converting the speech into text, it uses dialogue context analysis to correct potential labeling errors from acoustic clustering and directly infers the speaker's true identity from the text content; and it uses the results of semantic correction and identity inference as optimization constraints, feeding them back to the front-end acoustic clustering process in real time. Through multiple iterations, it dynamically optimizes the clustering effect, thereby overcoming the problem of missing or incorrect identity recognition in short speech, noisy, and new speaker scenarios without the need for a pre-recorded voiceprint database. Finally, it outputs high-quality, practical, and real-name meeting transcripts.

[0035] In some embodiments, the preset iteration termination condition is: the rate of change of the speaker label sequence generated in two adjacent iterations is lower than a set threshold, or the number of iterations reaches a preset maximum value.

[0036] Understandably, the preset iteration termination conditions can include two parallel criteria: first, a quality convergence criterion, where the rate of change of the speaker label sequences generated between two adjacent iterations in a single iteration is lower than a set threshold, indicating that the system's output has stabilized and further optimization has no significant effect; second, a resource protection criterion, where the system's iteration count reaches a preset maximum value. This is a mandatory safety measure to prevent the system from falling into an infinite loop in extreme cases, ensuring the effective use of computing resources. Its function is to provide an automatic and objective termination rule for the dynamic process of "perception-cognition-feedback." It ensures that the system can stop in time when "sufficient optimization results have been obtained" or "the maximum allowed computing resources have been consumed," thereby outputting the current optimal meeting transcript, ensuring both result quality and system efficiency and reliability.

[0037] In some embodiments, each speech in the transcribed meeting text is associated with the speaker's identity information, which includes at least one of name, position, or department.

[0038] Understandably, the final result processed by the aforementioned "perception-cognition-feedback" closed-loop optimization mechanism is not a traditional transcribed text merely labeled with anonymous tags (such as "speaker A"). Instead, each speech segment in the text is bound to a specific, real identity information. This identity information is obtained through cognitive layer reasoning and includes at least one or more of the speaker's real name, position, or department. Its significance lies in directly demonstrating the fundamental difference between this application and existing technologies: transforming anonymous speech streams into named, structured texts that can be directly used for meeting minutes, content review, and accountability tracing. This solves the problem of "missing identity information" in traditional technologies, achieving a qualitative leap from "who is speaking" to "who is the speaker," greatly improving the usability and automation level of meeting transcription results.

[0039] In some embodiments, step S101, which involves extracting acoustic features and clustering the speech segments and assigning speaker labels to each speech segment, includes: extracting acoustic features of the speech segments using a preset acoustic model; clustering the speech segments based on the similarity of the acoustic features; and assigning speaker labels to the clustering results.

[0040] Understandably, a pre-defined lightweight acoustic model (such as x-vector) can be used to extract acoustic feature vectors representing the speaker's identity from speech segments. Subsequently, cluster analysis (such as hierarchical clustering or spectral clustering) is performed based on the similarity (such as cosine similarity) between these acoustic feature vectors, grouping speech segments with similar features into the same category. Each clustering result represents a different speaker, and each category is assigned an initial, anonymous speaker label (such as SPK_0, SPK_1). Its purpose is to complete a preliminary speaker separation at a purely acoustic level.

[0041] In some embodiments, step S102, which involves performing dialogue context analysis based on the text sequence to correct errors in the speaker tags, obtaining corrected speaker tag results, and inferring the real identity corresponding to each speaker tag based on identity indication information in the text sequence, to obtain identity inference results, includes: analyzing the logical flow, topic evolution, and question-answer relationships of the text sequence based on a preset large language model, correcting the speaker tags according to semantic consistency, and obtaining corrected speaker tag results; identifying self-identification, other-identification, and role-related expressions based on the text sequence, and inferring the identity corresponding to each speaker tag by combining domain knowledge, to obtain identity inference results corresponding to each speaker tag.

[0042] Understandably, firstly, speaker label correction involves analyzing the dialogue logic flow, topic evolution, and question-and-answer relationships of the text sequence using a pre-defined large language model (e.g., determining whether a statement directly refutes or supports the previous statement). Based on the principle of semantic consistency, it verifies and corrects speaker label errors generated by the initial acoustic clustering (e.g., this application re-merges consecutive statements from the same person that were incorrectly segmented by the acoustic model in traditional solutions due to short pauses or noise). Secondly, identity reasoning, based on the same text sequence, identifies self-identification (e.g., "I am Zhang San"), other-identification (e.g., "Mr. Li, what do you think?"), and expressions reflecting role responsibilities (e.g., discussing technical architecture, budget approval). Combined with a domain knowledge graph, it infers the real identity information (e.g., name, position, department) corresponding to each anonymous speaker label. Its purpose is to achieve a fundamental leap from "acoustic differentiation" to "semantic recognition." This step utilizes the deep understanding capabilities of the large language model to solve the problems of semantic coherence judgment and identity referencing reasoning that pure acoustic models cannot handle.

[0043] In some embodiments, step S103, feeding back the corrected speaker label result and the identity reasoning result as optimization constraints to the clustering process to perform iterative processing, includes: converting the corrected speaker label result into constraint signals in the acoustic feature space; and dynamically adjusting the parameters or decision criteria of the clustering process using the constraint signals to optimize the iterative clustering effect.

[0044] Understandably, the high-confidence speaker label correction results generated by the cognitive layer (e.g., determining that the speech segments "SPK_1" and "SPK_2" belong to the same speaker) are transformed into specific constraint signals in the acoustic feature space. In subsequent iterative processing, these constraint signals are used to dynamically adjust the parameters or decision criteria of the clustering algorithm (e.g., forcibly reducing the distance between constrained segments when recalculating acoustic feature distances), thereby optimizing the clustering process and ensuring that its output is consistent with the conclusions of semantic reasoning. Its role is to establish a reverse optimization path from "cognition" to "perception," forming a closed loop. This step transforms the reliable conclusions of high-level semantic reasoning into operable low-level acoustic processing instructions, enabling the system to self-correct and iteratively improve. Furthermore, it overcomes the limitations of acoustic models in short speech and noisy scenarios, utilizing semantic stability to guide and improve the accuracy of acoustic clustering, thereby ensuring that the final output of the entire system possesses both acoustic discriminative power and semantic logical consistency.

[0045] In some other embodiments, the method further includes: when the confidence level of the inference for a speaker label in the identity inference result exceeds a preset threshold, establishing a mapping relationship between the speaker label and its corresponding real identity, and updating it to a dynamic identity mapping table; in subsequent processing, querying the dynamic identity mapping table to assist or simplify the clustering process or the identity inference.

[0046] Understandably, when the confidence level of the identity reasoning module infers the identity information (e.g., "Zhang San") of a speaker tag (e.g., SPK_0) exceeds a preset threshold, the system confirms the validity of the reasoning result and automatically establishes and maintains a dynamic identity-voiceprint mapping table internally, persistently storing the correspondence between the speaker tag and its real identity (name, position, etc.). In subsequent processing (whether it's a subsequent iteration of the current meeting or a new meeting), when performing clustering or identity reasoning, the system will first query this mapping table. If it finds that the acoustic features of the current speech segment match the acoustic features of a registered identity in the table, it can directly apply the existing identity information without having to perform complex semantic reasoning again. Its function is to enable the system's "self-learning" and capability evolution. This mechanism transforms a high-confidence cognitive reasoning result into reusable perceptual knowledge, enabling the system to transition from a cold-start mode where "reasoning must be repeated for each meeting" to an optimized mode where "one-time recognition leads to recognition everywhere." This not only improves the efficiency and accuracy of subsequent processing but also lowers the threshold for speaker recognition, further enhancing the practicality and robustness of the method.

[0047] In other specific examples, such as Figure 2 As shown, Figure 2 This diagram illustrates the architecture of a conference voice data processing device according to an embodiment of this application. The device constructs a closed-loop optimized architecture comprising a "perception layer," a "cognitive layer," and an "application layer," as well as multiple feedback loops. It achieves the conversion from raw mixed speech to real-name transcribed conference text without relying on a pre-registered speaker database. See below:

[0048] (1) Primary processing of input and perception layer.

[0049] The input is a raw, mixed speech stream. First, the perception layer performs basic speech signal processing, with the explicit goal of "providing preliminary, coarse material for the upper layers," allowing the output to be "noisy." The perception layer comprises two modules: a speech activity detection (VAD) module and an acoustic feature extraction and primary clustering module. The VAD module removes silence segments, extracts valid speech fragments, and performs preliminary segmentation of the continuous speech stream. The acoustic feature extraction and primary clustering module uses a pre-defined lightweight acoustic model (such as x-vector) to extract acoustic features from valid speech fragments and performs preliminary speaker clustering based on feature similarity. This module does not pursue high precision; it assigns anonymous initial speaker labels (such as SPK_0, SPK_1) to each clustering result, generating a "speech fragment sequence with preliminary acoustic labels." Its purpose is to structure the continuous speech stream, providing processable input for the cognitive layer.

[0050] (2) Speech recognition and text conversion.

[0051] The speech segment sequence generated by the perceptual layer, with preliminary speaker labels, is sent to the speech recognition module (ASR module) for transcription, resulting in a text sequence with preliminary acoustic labels. This text sequence retains the speaker labels output by the perceptual layer, but these speaker labels may contain errors due to limitations in acoustic processing.

[0052] (3) Deep semantic understanding and reasoning at the cognitive level.

[0053] The cognitive layer receives the above text sequence and performs high-level semantic understanding, which includes two sub-modules: a semantic understanding and speaker differentiation module and an identity information reasoning module.

[0054] The semantic understanding and speaker differentiation module takes as input a sequence of text segments with preliminary acoustic labels from the speech recognition module. Its processing begins with context reconstruction, where the larger model understands the logical flow, topic evolution, and question-and-answer relationships of the entire dialogue. Based on semantic consistency (e.g., the topic should be coherent when the same person speaks), dialogue behaviors (e.g., asking questions, answering, agreeing), and referential relationships, it identifies and corrects incorrect speaker labels in the preliminary acoustic clustering. For example, by analyzing referential relationships and speaking logic (supporting and opposing stances), the semantic understanding and speaker differentiation module confidently corrects the segment with the acoustic label SPK_2 to the same person as SPK_1, thus outputting an intermediate text with more accurate speaker differentiation and clearer dialogue logic. For example:

[0055] Input example: [

[0057] {

[0058] "timestamp": "00:01:05",

[0059] "acoustic_speaker_tag": "SPK_0",

[0060] "text": "Regarding the third quarter budget, I believe we need to increase investment."

[0061] "asr_confidence": 0.95

[0062] },

[0063] {

[0064] "timestamp": "00:01:12",

[0065] "acoustic_speaker_tag": "SPK_1",

[0066] I disagree because we are currently facing a tight cash flow situation.

[0067] "asr_confidence": 0.87

[0068] },

[0069] {

[0070] "timestamp": "00:01:20",

[0071] "acoustic_speaker_tag": "SPK_0", / / The acoustic model may be mislabeled due to noise.

[0072] "text": "But market opportunities are fleeting, and we must seize them."

[0073] "asr_confidence": 0.92

[0074] },

[0075] {

[0076] "timestamp": "00:01:25",

[0077] "acoustic_speaker_tag": "SPK_2",

[0078] "text": "Mr. Wang, I support your viewpoint. The risks are manageable."

[0079] "asr_confidence": 0.98

[0080] } ]

[0082] Output:

[0083] {

[0084] "global_topic": "Discussion on whether to increase spending in the third quarter budget",

[0085] "turn_analysis": [

[0086] {

[0087] "timestamp": "00:01:05",

[0088] "original_speaker": "SPK_0",

[0089] "proposed_speaker": "SPK_0",

[0090] "confidence": "high",

[0091] "reasoning": The statement starts a new topic without contradiction.

[0092] },

[0093] {

[0094] "timestamp": "00:01:12",

[0095] "original_speaker": "SPK_1",

[0096] "proposed_speaker": "SPK_1",

[0097] "confidence": "high",

[0098] "reasoning": "Clearly opposes the previous speaker (SPK_0), taking a clear stance, and should be considered a different speaker."

[0099] },

[0100] {

[0101] "timestamp": "00:01:20",

[0102] "original_speaker": "SPK_0",

[0103] "proposed_speaker": "SPK_0", / / Note: The model is correctly identified here, and the acoustic label is correct.

[0104] "confidence": "high",

[0105] "reasoning": "This statement is a direct rebuttal to SPK_1's objections, echoing SPK_0's viewpoint and maintaining a consistent style. The acoustic label is correct."

[0106] },

[0107] {

[0108] "timestamp": "00:01:25",

[0109] "original_speaker": "SPK_2",

[0110] "proposed_speaker": "SPK_1", / / The model infers that SPK_2 is SPK_1

[0111] "confidence": "high",

[0112] "reasoning": "The statement 'Mr. Wang, I support your viewpoint' clearly indicates that 'Mr. Wang' is SPK_0, while the speaker refers to himself as 'I'. Furthermore, this statement supports SPK_0 and opposes SPK_1, but logically, SPK_1 just expressed opposition and cannot immediately switch to support. Therefore, the most reasonable explanation is that the acoustic model incorrectly segmented and labeled SPK_1's two statements as SPK_1 and SPK_2. In reality, SPK_2 is SPK_1, who continued speaking in support after a short pause."

[0113] }

[0114] ],

[0115] "summary_of_changes": "Change the speaker of the last segment from SPK_2 to SPK_1 because referential relations and logic indicate that they are the same person."

[0116] }

[0117] Based on the output of this semantic understanding and speaker discrimination module, the speaker discrimination in this sequence is already much more semantically reasonable than the original acoustic clustering results.

[0118] The identity information reasoning module, based on text content, identifies self-identification, other-identification, and role / responsibilities. Combining domain knowledge, it infers the true identity information corresponding to each corrected speaker tag, such as name, position, or department, achieving intelligent conversion from "anonymous acoustic tags" to "named identity information." Specifically, it analyzes text sequences, using a "self-identification and other-identification analyzer" to identify identity indicators such as "I am Zhang San" or "Mr. Li, please look"; a "role / responsibility inferencer" to infer the speaker's role based on the content (e.g., someone discussing budgets might be from finance); and a "contextual consistency check" to ensure logical coherence of statements under the same tag. Combining contextual clues such as the meeting topic and a known list of participants, this module can infer the true identity (e.g., name: "Zhang San," position: "Product Director") corresponding to each optimized speaker tag (e.g., SPK_0_opt) and assess the confidence level.

[0119] For example:

[0120] Input data format:

[0121] {

[0122] "session_id": "meeting_20240615_1400",

[0123] "optimized_dialog": [

[0124] {

[0125] "timestamp": "00:05:23",

[0126] "optimized_speaker_tag": "SPK_0_opt",

[0127] "text": "Regarding the third quarter budget, I believe we need to increase marketing investment."

[0128] "semantic_confidence": 0.92

[0129] },

[0130] {

[0131] "timestamp": "00:05:31",

[0132] "optimized_speaker_tag": "SPK_1_opt",

[0133] "I disagree with Mr. Zhang's view; we are currently facing a tight cash flow."

[0134] "semantic_confidence": 0.88

[0135] }

[0136] ],

[0137] "contextual_clues": {

[0138] "known_attendees": ["Zhang San", "Li Si", "Wang Wu"], / / Optional, retrieved from meeting invitation

[0139] "meeting_topic": "Q3 Product Planning Review Meeting",

[0140] "detected_roles": ["Product", "Technology", "Finance"] / / Roles initially identified from the dialogue

[0141] }

[0142] }

[0143] This identity information reasoning module adopts a multi-clue fusion reasoning architecture, comprehensively analyzing various identity cues. These include a self-identification and other-identification analyzer (identifying direct and indirect identity indicators in the dialogue), a role and responsibility inferencer (inferring the speaker's organizational role based on the content of their speech), a domain knowledge correlator (associating the dialogue content with a domain knowledge graph), and a contextual consistency analyzer (analyzing the consistency of a speaker's multiple statements in terms of topic, viewpoint, and language to ensure that statements under the same tag are logically coherent).

[0144] For example, the cognitive big model cue word is designed as follows: Assume you are a professional identity reasoning expert. Based on the following meeting dialogue, infer the true identity of each speaker.

[0145] Available information includes: meeting topic: {meeting_topic}, known attendees (optional): {known_attendees}, and dialogue history: {dialog_history}.

[0146] Reasoning tasks include:

[0147] Acoustic features for identity cue extraction: identifying self-reference, other-reference, and role indicator words in dialogue;

[0148] Role responsibility analysis: Infer the scope of each person's responsibilities based on their speech content;

[0149] Identity hypothesis generation: Propose the most likely identity hypothesis;

[0150] Confidence assessment: Evaluate the credibility of each hypothesis;

[0151] The rules of reasoning include:

[0152] If someone says "I am Zhang San", then the high confidence level is marked as Zhang San;

[0153] If someone is repeatedly referred to as "General Manager Li," they are likely a member of the management team with the surname Li.

[0154] If someone leads the technical discussions, they are likely the technical lead;

[0155] Consider the subject matter and vocabulary of the speech;

[0156] Consistency in topic, viewpoint, and wording across multiple statements by the speaker.

[0157] For example, the output format is:

[0158] {

[0159] "identity_mappings": {

[0160] "SPK_0_opt": {

[0161] "primary_identity": {

[0162] "name": "Zhang San",

[0163] Title: Product Director

[0164] "department": "Product Department"

[0165] "confidence": 0.94

[0166] },

[0167] "alternative_identities": [

[0168] {"name": "Zhang Ming", "title": "Product Manager", "confidence": 0.3}

[0169] ],

[0170] "reasoning_chain": [

[0171] {

[0172] "step": "Self-identification",

[0173] "evidence": "The speaker said, 'I am Zhang San from the product department.'"

[0174] "Impact": "Decisive evidence"

[0175] },

[0176] {

[0177] "step": "Responsibility Reasoning",

[0178] "evidence": "Discussion on key product requirements",

[0179] "impact": Supporting evidence

[0180] }

[0181] ],

[0182] "temporal_evolution": {

[0183] "first_mentioned": "00:05:10",

[0184] "confidence_trend": "Steady rise"

[0185] }

[0186] }

[0187] },

[0188] "cross_speaker_relations": {

[0189] "reporting_lines": [

[0190] {"from": "SPK_1_opt", "to": "SPK_0_opt", "relation": "reporting relationship", "confidence": 0.8}

[0191] ],

[0192] "department_affiliations": {

[0193] "Technical Department": ["SPK_1_opt", "SPK_3_opt"],

[0194] Product Department: ["SPK_0_opt"]

[0195] }

[0196] }

[0197] }

[0198] (4) The feedback loop mechanism realizes dynamic optimization.

[0199] This application establishes multiple feedback loops, which constitute a dynamic, bidirectional optimization process. This process transforms the high-level semantic reasoning results from the cognitive layer into specific constraints, guiding the front-end processing and achieving iterative optimization of the overall system. The multiple feedback loops include:

[0200] Feedback Loop 1: Semantic Correction Suggestion. This is the "semantic-guided clustering optimization" path. When the semantic understanding and speaker differentiation module makes a high-confidence correction judgment (e.g., determining that SPK_A and SPK_B should be the same person), it generates a "must-link" constraint pair, which is passed back to the acoustic feature extraction and primary clustering module of the perception layer through the "semantic correction suggestion of feedback loop 1". In subsequent processing, this acoustic feature extraction and primary clustering module will force the constrained speech segments to be closer in the acoustic feature space, thereby optimizing the clustering effect.

[0201] Feedback Loop 2: Semantic Role Constraint. This path feeds back the semantic analysis results from the cognitive layer (such as speaker roles and dialogue structure) to the speech recognition module, providing it with richer contextual information to help improve transcription accuracy in complex dialogue scenarios.

[0202] Feedback Loop 3: Identity Information Constraints. This is the "identity information-guided global mapping" path. After the identity information inference module infers the true identity of a speaker tag with high confidence, it generates an "identity-voiceprint tag mapping". This feedback loop updates a dynamic identity mapping table maintained internally by the system, enabling the system to directly call this mapping to achieve rapid real-name recognition in subsequent processing or new meetings, eliminating the dependence on pre-registered voiceprint databases.

[0203] (5) Iterative optimization process and final output.

[0204] The initial rounds involve preliminary processing and reasoning by the perception and cognition layers. The generated feedback signals are applied to the perception layer, resulting in higher-quality intermediate results. The optimized results are then fed back into the cognition layer, triggering more accurate reasoning. This "processing-reasoning-feedback-optimization" cycle continues until a preset iteration termination condition is met (e.g., the change in speaker labels falls below a threshold, or the maximum number of iterations is reached). Finally, the process enters the application layer. The result generation and output module integrates the final speaker label correction results with the identity reasoning results, generating and outputting transcribed text with real identity labels. Each speech segment in this transcribed text is associated with the speaker's specific identity information (e.g., name, position, department), achieving the goal of this application: converting anonymous speech into real-name knowledge text.

[0205] visible, Figure 2 This application demonstrates the technical path of obtaining initial acoustic hypotheses through the perception layer, performing semantic correction and identity reasoning through the cognitive layer, and then continuously optimizing the perception results through a feedback loop, thereby achieving end-to-end automated processing from anonymous speech to real-name text in scenarios without a registered voiceprint database.

[0206] In summary, this application aims to address the shortcomings of traditional conference speech processing technologies, such as the lack of identity recognition due to reliance on pre-registered speaker databases, and poor adaptability to scenarios involving short speech, noise interference, and accent differences. To solve this problem, this application proposes a conference speech data processing method based on a closed-loop optimization of "perception-cognition-feedback": First, the perception layer performs preliminary acoustic clustering on the original speech, outputting preliminary speaker labels with noise assumptions; then, the cognition layer uses a large language model to perform deep semantic analysis on the transcribed text, correcting erroneous speaker labels and directly inferring the speaker's true identity from the dialogue content; furthermore, the feedback mechanism transforms the high-confidence corrections and identity inference results obtained by the cognition layer into constraint signals in the acoustic feature space, inversely optimizing the clustering process of the perception layer, and continuously improving the system through iteration. Ultimately, this achieves the effect of automatically and accurately producing real-name transcribed conference text without the need for pre-recorded speaker databases, significantly improving the system's robustness, practicality, and automation level in complex real-world scenarios.

[0207] On the other hand, such as Figure 3 As shown, this embodiment provides a conference voice data processing device 300, including a voice processing module 301, an identity reasoning module 302, and an iterative optimization module 303.

[0208] For example, the speech processing module 301 is used to perform speech detection on the input conference speech data, extract speech segments, perform acoustic feature extraction and clustering processing on the speech segments, and assign speaker labels to each speech segment to generate a speech segment sequence with speaker labels.

[0209] For example, the identity reasoning module 302 is used to perform speech recognition on the speech segment sequence to obtain a text sequence with speaker tags; perform dialogue context analysis based on the text sequence to correct errors in the speaker tags, obtain the corrected speaker tags, and infer the real identity corresponding to each speaker tag based on the identity indication information in the text sequence to obtain the identity reasoning result.

[0210] For example, the iterative optimization module 303 is used to feed back the corrected speaker label and the identity reasoning result as optimization constraints to the clustering process to perform iterative processing until a preset iteration termination condition is reached, and output the meeting transcript with identity information.

[0211] For example, the speech processing module 301 is further configured to extract the acoustic features of the speech segment using a preset acoustic model; perform speaker clustering on the speech segment based on the similarity of the acoustic features; and assign speaker labels to the clustering results.

[0212] For example, the identity reasoning module 302 is also used to analyze the logical flow, topic evolution and question-answer relationship of the text sequence based on a preset large language model, correct the speaker tags according to semantic consistency, and obtain the corrected results of the speaker tags; based on the text sequence, identify self-identification, other-identification and role-related expressions, and combine domain knowledge to infer the identity corresponding to each speaker tag, and obtain the identity reasoning results corresponding to each speaker tag.

[0213] For example, the iterative optimization module 303 is further configured to convert the corrected speaker label into a constraint signal in the acoustic feature space; and to dynamically adjust the parameters or decision criteria of the clustering process using the constraint signal to optimize the clustering effect of the iteration.

[0214] For example, the conference voice data processing device 300 is further configured to establish a mapping relationship between the speaker label and its corresponding real identity when the reasoning confidence of the speaker label in the identity reasoning result exceeds a preset threshold, and update it to a dynamic identity mapping table; in subsequent processing, the dynamic identity mapping table is queried to assist or simplify the clustering process or the identity reasoning.

[0215] For example, the preset iteration termination condition is: the rate of change of the speaker label sequence generated in two adjacent iterations is lower than a set threshold, or the number of iterations reaches a preset maximum value.

[0216] For example, each speech in the transcribed meeting text is associated with the speaker's identity information, which includes at least one of name, position, or department.

[0217] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute the conference voice data processing method, the method including:

[0218] The input conference audio data is subjected to speech detection, audio segments are extracted, and acoustic features are extracted and clustered on the audio segments. Speaker labels are assigned to each audio segment to generate a sequence of audio segments with speaker labels.

[0219] The speech segment sequence is subjected to speech recognition to obtain a text sequence with speaker tags; based on the text sequence, dialogue context analysis is performed to correct errors in the speaker tags to obtain corrected speaker tags; and based on the identity indication information in the text sequence, the real identity corresponding to each speaker tag is inferred to obtain the identity inference result.

[0220] The corrected speaker labels and the identity reasoning results are fed back to the clustering process as optimization constraints to perform iterative processing until a preset iteration termination condition is reached, and the meeting transcript with real identity information is output.

[0221] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0222] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to execute the conference voice data processing methods provided by the above methods.

[0223] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the aforementioned conference voice data processing methods.

[0224] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0225] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0226] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing conference voice data, characterized in that, include: The input conference audio data is subjected to speech detection, audio segments are extracted, and acoustic features are extracted and clustered on the audio segments. Speaker labels are assigned to each audio segment to generate a sequence of audio segments with speaker labels. The speech segment sequence is subjected to speech recognition to obtain a text sequence with speaker tags; based on the text sequence, dialogue context analysis is performed to correct errors in the speaker tags to obtain corrected speaker tags; and based on the identity indication information in the text sequence, the real identity corresponding to each speaker tag is inferred to obtain the identity inference result. The corrected speaker labels and the identity reasoning results are fed back to the clustering process as optimization constraints to perform iterative processing until the preset iteration termination condition is reached, and the meeting transcript with real identity information is output. The step of feeding back the corrected speaker label results and the identity reasoning results as optimization constraints to the clustering process to perform iterative processing includes: converting the corrected speaker label results into constraint signals in the acoustic feature space; and dynamically adjusting the parameters or decision criteria of the clustering process using the constraint signals to optimize the iterative clustering effect.

2. The conference voice data processing method according to claim 1, characterized in that, The process of extracting acoustic features and clustering the speech segments, and assigning speaker labels to each speech segment, includes: The acoustic features of the speech segment are extracted using a preset acoustic model; The speech segments are clustered by speaker based on the similarity of acoustic features, and the clustering results are assigned speaker labels.

3. The conference voice data processing method according to claim 1, characterized in that, The process of performing dialogue context analysis based on the text sequence to correct errors in the speaker labels, obtaining corrected speaker label results, and inferring the true identity corresponding to each speaker label based on identity indication information in the text sequence, yields identity inference results including: Based on a pre-defined large language model, the logical flow, topic evolution, and question-and-answer relationships of the text sequence are analyzed. The speaker labels are then corrected based on semantic consistency to obtain the corrected speaker labels. Based on the text sequence, self-identification, other-identification, and role-related statements are identified, and combined with domain knowledge, the identities corresponding to each speaker label are inferred, resulting in identity inference results for each speaker label.

4. The conference voice data processing method according to claim 1, characterized in that, The method further includes: When the confidence level of the inference for a speaker label in the identity inference result exceeds a preset threshold, a mapping relationship between the speaker label and its corresponding real identity is established and updated to a dynamic identity mapping table. In subsequent processing, the dynamic identity mapping table is queried to assist or simplify the clustering process or the identity reasoning.

5. The conference voice data processing method according to claim 1, characterized in that, The preset iteration termination condition is: The rate of change of the speaker label sequence generated in two adjacent iterations is lower than the set threshold, or the number of iterations reaches the preset maximum value.

6. The conference voice data processing method according to claim 1, characterized in that, Each speech in the transcribed meeting text is associated with the speaker's identity information, which includes at least one of name, position, or department.

7. A conference voice data processing device, characterized in that, include: The speech processing module is used to perform speech detection on the input conference speech data, extract speech segments, perform acoustic feature extraction and clustering processing on the speech segments, assign speaker labels to each speech segment, and generate a speech segment sequence with speaker labels. The identity reasoning module is used to perform speech recognition on the speech segment sequence to obtain a text sequence with speaker tags; perform dialogue context analysis based on the text sequence to correct errors in the speaker tags, obtain the corrected speaker tags, and infer the real identity corresponding to each speaker tag based on the identity indication information in the text sequence to obtain the identity reasoning result. The iterative optimization module is used to feed back the corrected speaker labels and the identity reasoning results as optimization constraints to the clustering process to perform iterative processing until a preset iteration termination condition is reached, and output the meeting transcript with identity information. The iterative optimization module is further configured to: convert the corrected speaker labels into constraint signals in the acoustic feature space; and dynamically adjust the parameters or decision criteria of the clustering process using the constraint signals to optimize the iterative clustering effect.

8. An electronic device, characterized in that, The method includes at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the conference voice data processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the conference voice data processing method according to any one of claims 1-6.