Real-time voice data collection method and system for multilingual mixed scene

By combining voiceprint recognition and speaker segmentation technologies with real-time statistics of language usage rate and accent deviation, the problems of speaker identity matching and accent interference in multi-person speaking scenarios are solved, achieving adaptive optimization of multilingual speech acquisition and recognition, and improving the accuracy of speech transcription and recognition performance.

CN122116876APending Publication Date: 2026-05-29CHENGDU ZHONGQI YILIAN TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU ZHONGQI YILIAN TECH CO LTD
Filing Date
2026-03-04
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In multilingual scenarios, existing technologies struggle to accurately match speaker identities and effectively address accent interference when multiple people are speaking, leading to a decline in the accuracy and practicality of language recognition.

Method used

By generating uniquely identified speech segments through voiceprint recognition and speaker segmentation, and combining real-time statistics of language usage rate and accent deviation value, the language processing model is dynamically adjusted to adapt to the speaker's language switching and accent changes.

Benefits of technology

It improves the accuracy and reliability of multilingual speech acquisition, ensures the precision of conference transcription and the reliability of speaker matching, and enhances the adaptive capability of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116876A_ABST
    Figure CN122116876A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech data processing, in particular to a real-time speech data acquisition method and system for a multi-language mixed scene, comprising continuously acquiring original speech data under a current conference environment, performing voiceprint recognition and speaker segmentation on the original speech data to generate a plurality of speech segments, the speech segments carrying identity numbers; reading the speech segments and the identity numbers carried thereby, identifying the identity numbers or the speech segments to obtain final language tags or accent features of the current speech segments; selecting a corresponding language processing model according to the final language tags, processing the speech segments using the language processing model or the accent features to obtain final speech data. The present application solves the technical problem of language misjudgment caused by speaker identity ambiguity and accent interference when multiple people speak in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice data processing technology, and more specifically, to a real-time voice data acquisition method and system for multilingual mixed scenarios. Background Technology

[0002] In current applications of speech technology, real-time speech data acquisition in multilingual scenarios has become a research hotspot. With the increasing prevalence of long-duration conferences such as multinational conferences and international seminars, the demand for speech acquisition has expanded from simple recording to intelligent services such as automatic transcription, real-time translation, speaker identification, and language habit analysis. These applications rely on high-quality, multi-dimensionally annotated speech data to support core tasks such as language recognition, speakerprint segmentation, and automatic speech recognition. However, speech data acquisition in long-duration conference scenarios still faces challenges: the same speaker may switch languages ​​at different times or even mix multiple languages ​​within a single sentence, making it difficult for traditional systems to dynamically adapt.

[0003] A Chinese invention patent, titled "A Multilingual Full-Speech Processing Method, Device, and Medium Based on Speech Recognition" (publication number CN121506100A), improves the continuity of language recognition to some extent by constructing explicit and implicit language trajectories to generate language weight trajectories and dividing nodes for recognition. This solves the problem of language recognition when the same speaker uses multiple languages ​​in a single sentence and enhances recognition continuity in language switching scenarios. However, this patent still has the following shortcomings: it cannot accurately identify the corresponding speaker when multiple people are speaking, lacking the ability to distinguish speaker identities; it does not consider the impact of accent interference on language recognition, leading to a sharp drop in the recognition rate of the general speech model in accent scenarios; furthermore, a speaker's voiceprint and accent can drift due to fatigue, positional changes, etc., and this patent only uses historical language preference information in a one-way manner, lacking a closed-loop adaptive mechanism and unable to update and optimize the model online to adapt to these changes. These deficiencies severely restrict the accuracy and practicality of applications such as conference transcription and cross-language communication.

[0004] Therefore, in long-duration meeting scenarios, how to accurately match the identities of speakers in multi-person speaking scenarios, effectively deal with accent interference, and achieve adaptive optimization of multilingual speech acquisition and recognition remains a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0005] The purpose of this application is to provide a real-time speech data acquisition method and system for multilingual mixed scenarios. It solves the technical problems of ambiguous speaker identification and language misjudgment caused by accent interference when multiple people are speaking in the prior art. It realizes adaptive and optimized multilingual speech acquisition and recognition, and effectively ensures the accuracy of speech transcription and the reliability of speaker matching in conference scenarios.

[0006] To solve the above-mentioned technical problems, the solution adopted in this application is as follows:

[0007] This invention provides a real-time voice data acquisition method for multilingual mixed scenarios, characterized by the following steps:

[0008] Step S1: Continuously collect raw voice data in the current conference environment, perform voiceprint recognition and speaker segmentation on the raw voice data, and generate multiple voice segments, each voice segment carrying an identity number; the identity number represents the speaker's unique identifier.

[0009] Step S2: Read the speech segment and its carried identification number, identify the identification number or speech segment, and obtain the final language label or accent feature of the current speech segment; the final language label represents the language type of the speech segment; the final language label represents at least one language type;

[0010] Step S3: Select the corresponding language processing model according to the final language label, and process the speech segment using the language processing model or the accent features to obtain the final speech data;

[0011] The implementation of step S2 above includes the following steps:

[0012] Step S21: Read the identity number, and read the usage rate of each language type based on the identity number. The usage rate refers to the frequency or number of times each language type is used.

[0013] Step S22: Compare the usage rate and usage threshold of each language type, and generate intermediate language labels based on the comparison results; the intermediate language labels represent the language type of the speech segment;

[0014] Step S23: Determine whether the intermediate language tag contains a language type; if the intermediate language tag contains at least one language type, proceed to step S24; if the intermediate language tag does not contain any language type, proceed to step S26.

[0015] Step S24: Read the speech segment, perform accent recognition on the speech segment according to the intermediate language tag, and generate accent features based on the recognition results; proceed to step S25;

[0016] Step S25: The intermediate language tags are used to generate the final language tags;

[0017] Step S26: Add all language types to the intermediate language tags to generate the final language tags.

[0018] In some embodiments, the comparison of usage rates and usage thresholds for each language type in step S22 is achieved through the following steps:

[0019] Step S221: Sort each language type in descending order of usage rate to generate a language priority sequence; label each language type according to the language priority sequence;

[0020] Step S222: Read the usage rate of the language type of the first label, and compare the usage rate with a preset usage threshold;

[0021] Step S223: If the usage rate is greater than or equal to the usage threshold, proceed to step S224; if the usage rate is less than the usage threshold, proceed to step S225.

[0022] Step S224: Add the language type under the current label to the intermediate language label; read the usage rate of the language type of the next label, compare the usage rate with the preset usage threshold, and execute step S223;

[0023] Step S225: Add the language type under the current label to the intermediate language label and output the intermediate language label.

[0024] In some embodiments, the process of performing voiceprint recognition and speaker segmentation on the original speech data to generate multiple speech segments in step S1 is achieved through the following steps:

[0025] Step S11: Continuously collect raw audio data in the current conference environment, perform voice activity detection on the raw audio data, and extract audio segments containing valid audio.

[0026] Step S12: Extract voiceprint features for each speech segment and compare the voiceprint features with the voiceprint feature database; if the match is successful, assign the corresponding identity number from the voiceprint feature database to the speech segment; if the match fails, create a new identity number, assign the new identity number to the speech segment, and add the identity number to the voiceprint feature database.

[0027] Step S13: Concatenate the voice segments of the same identity number in chronological order to generate a voice segment carrying the identity number.

[0028] In some embodiments, step S24, which involves performing accent recognition on the speech segment based on the intermediate language tag and generating accent features based on the recognition result, is achieved through the following steps:

[0029] Step S241: Read the language type in the intermediate language tag and perform phoneme recognition to obtain the standard pronunciation phonemes; perform phoneme recognition on the speech segment to obtain the acoustic features of the current speech segment;

[0030] Step S242: Process the speech segment according to the standard phonemes and the acoustic features to obtain the accent deviation value of the speech segment;

[0031] Step S243: Determine whether an accent exists based on the accent deviation value of the speech segment; if the accent deviation value is lower than the accent correction threshold, then generate an accent feature; if the accent deviation value is higher than the accent correction threshold, then do not generate an accent feature.

[0032] In some embodiments, the speech segment is processed according to the standard phonemes and the acoustic features, using the following formula:

[0033]

[0034] Where p represents the standard phoneme; X is the acoustic feature of the current speech segment; Q is the set of all possible phonemes; P(p|X) is the probability that the pronunciation is the standard phoneme p given the acoustic feature X; max q∈Q P(q|X) is the probability of the most likely phoneme when the model freely recognizes it; the phoneme is the smallest unit of speech in the language; GOP(p) is the accent deviation value of the current speech segment;

[0035] This yields the accent deviation value for each speech segment.

[0036] In some embodiments, step S3 involves processing the speech segment using the language processing model or the accent features, including any of the following methods:

[0037] The accent features are used to correct the accent of the speech segment, and the language processing model processes the corrected speech segment.

[0038] The speech segment is processed using the language processing model.

[0039] In some embodiments, the usage rate includes the global usage rate of each language type in the current conference environment or the usage rate of the speaker's historical language type corresponding to the identity number.

[0040] The present invention also provides a real-time voice data acquisition system for multilingual mixed scenarios, characterized in that, in order to implement the above-mentioned real-time voice data acquisition method for multilingual mixed scenarios, it includes an identity recognition module, a language accent recognition module, and a language processing module; the identity recognition module, the language accent recognition module, and the language processing module are connected in sequence.

[0041] The identity recognition module is used to perform voiceprint recognition and speaker segmentation on the original voice data, and transmit the generated multiple voice segments carrying identity numbers to the language accent recognition module.

[0042] The language accent recognition module receives the speech segment and its carried identification number, identifies the identification number or speech segment, and obtains the final language label or accent feature of the current speech segment; the language accent recognition module transmits the final language label or accent feature to the language processing module;

[0043] The language processing module selects the corresponding language processing model based on the final language label, and processes the speech segment according to the language processing model or the accent features to obtain the final speech data.

[0044] In some embodiments, the language accent recognition module includes a usage rate comparison module, an accent recognition module, and a language tag integration module. The usage rate comparison module is connected to the accent recognition module, the language tag integration module, and the identity recognition module, respectively. The accent recognition module and the language tag integration module are connected to the language processing module, respectively.

[0045] The usage rate comparison module receives multiple voice segments carrying identity numbers transmitted by the identity recognition module. The usage rate comparison module reads the usage rate of each language type based on the identity number, compares the usage rate of each language type with a usage threshold, and generates an intermediate language label based on the comparison result. The usage rate comparison module transmits the intermediate language label to the language label integration module and the accent recognition module respectively.

[0046] The accent recognition module performs accent recognition on the speech segment based on the intermediate language tag, and generates accent features based on the recognition results; the accent recognition module transmits the accent features to the language processing module;

[0047] The language tag integration module determines whether there is a language type in the intermediate language tag, generates the final language tag based on the determination result, and transmits the final language tag to the language processing module.

[0048] In some embodiments, the language accent recognition module further includes an output module, which is connected to the accent recognition module and the language tag integration module respectively; the accent recognition module and the language tag integration module output the accent features and the final language tags to the language processing module through the output module respectively.

[0049] The technical solution of this application has at least the following advantages and beneficial effects:

[0050] 1. This invention utilizes voiceprint recognition and speaker segmentation technology to decompose aliased conference audio into multiple individual voice segments carrying unique identification numbers, thus distinguishing the speaker's identity and solving the problem in existing technologies where multiple speakers cannot be accurately identified to their corresponding speakers. Simultaneously, it establishes independent language usage statistics for each speaker, enabling subsequent language recognition and speech processing to be optimized based on the speaker's personalized historical data, significantly improving the accuracy of speaker matching in conference transcription.

[0051] 2. This invention introduces an accent deviation value calculation, which quantifies the degree of accent deviation by comparing the actual pronunciation with standard phonemes. When the accent deviation value is lower than a preset threshold, the system automatically generates accent features and triggers an accent correction mechanism, enabling the language processing model to adaptively adjust its recognition of accented speech. When the accent deviation value is higher than the threshold, the standard model is used directly. This mechanism avoids invalid calculations in the absence of accents and effectively solves the problems of language misjudgment and decreased recognition rate caused by accent interference.

[0052] 3. This invention utilizes real-time statistics on language usage rates, including the global usage rate of each language type in the current meeting environment or the historical language usage rate of the speaker corresponding to the ID number. This feedback mechanism based on historical usage habits enables the system to dynamically adapt to users' language switching behavior, especially for changes in speaker language preferences during long meetings. When speaker historical data is insufficient, the system automatically uses the global usage rate or adds all language types as alternatives to ensure the reliability of speech recognition throughout the meeting. Simultaneously, as the meeting progresses, the system continuously accumulates speech data for each speaker, updating language usage rate statistics in real time to make the speaker's language profile increasingly accurate and the understanding of each person's language preferences more precise, thereby improving the overall performance of speech acquisition and recognition in multilingual scenarios. Attached Figure Description

[0053] Figure 1 This is an overall flowchart of the present invention;

[0054] Figure 2 This is a flowchart of step S2 in this embodiment;

[0055] Figure 3 This is a schematic diagram of the overall structure of the present invention;

[0056] Figure 4 This is a schematic diagram of the language accent recognition module in this embodiment. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. The terms "center," "upper," "lower," "inner," and "outer," indicating orientation or positional relationships based on the orientation or positional relationships shown in the figures, or the orientation or positional relationships commonly used when the product is in use, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed or operated in a specific orientation, and therefore should not be construed as a limitation on this application. It should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," and "connect" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two elements. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0059] Example

[0060] To solve the above-mentioned technical problems, the solution adopted in this application is as follows:

[0061] This invention provides a real-time speech data acquisition method for multilingual mixed scenarios, such as... Figure 1 As shown, its implementation includes the following steps:

[0062] Step S1: Continuously collect raw voice data in the current conference environment, perform voiceprint recognition and speaker segmentation on the raw voice data, generate multiple voice segments, and each voice segment carries an identity number; the identity number represents the speaker's unique identifier.

[0063] Step S2: Read the speech segment and its associated identification number, identify the identification number or speech segment to obtain the final language label or accent feature of the current speech segment; the final language label represents the language type of the speech segment; the final language label represents at least one language type.

[0064] Step S3: Select the corresponding language processing model based on the final language label, and process the speech segment using the language processing model or accent features to obtain the final speech data.

[0065] It should be noted that the language processing models mentioned include, but are not limited to, Chinese language processing models, French language processing models, English language processing models, and minority language processing models.

[0066] It should be explained that the language processing model described is existing technology and will not be repeated in this invention. The main improvements of this invention are as follows: To address the problem of inaccurate speaker identification in multi-person speaking scenarios, a unique identifier is generated for each speaker through voiceprint recognition and speaker segmentation; to address the problem of language misjudgment and decreased recognition rate caused by accent interference, an accent deviation detection and feature generation mechanism is introduced; and to address the problem that speaker voiceprints and accents drift with fatigue and positional changes, while existing technologies only use historical information unidirectionally and lack adaptive capabilities, a closed-loop feedback mechanism based on dynamic statistics of language usage is constructed. Through these improvements, this invention effectively solves the core defects of existing technologies in multilingual mixed scenarios.

[0067] Furthermore, such as Figure 2 As shown, the implementation of step S2 above includes the following steps:

[0068] Step S21: Read the identity number, and read the usage rate of each language type based on the identity number. The usage rate refers to the frequency or number of times each language type is used.

[0069] Step S22: Compare the usage rate and usage threshold of each language type, and generate intermediate language labels based on the comparison results; the intermediate language labels represent the language type of the speech segment;

[0070] Step S23: Determine whether there is a language type in the intermediate language tag; if the intermediate language tag contains at least one language type, proceed to step S24; if the intermediate language tag does not contain any language type, proceed to step S26.

[0071] Step S24: Read the speech segment, perform accent recognition on the speech segment based on the intermediate language label, and generate accent features based on the recognition results; proceed to step S25;

[0072] Step S25: Generate final language tags from intermediate language tags;

[0073] Step S26: Add all language types to the intermediate language tags to generate the final language tags.

[0074] It should be noted that the usage rate includes the global usage rate of each language type in the current meeting environment or the usage rate of the speaker's historical language type corresponding to the ID number.

[0075] In this embodiment, during the initial stages of a meeting, when new speakers join, when personal historical data is sparse, or when a speaker's accent is heavy, resulting in insufficient confidence in language recognition, the global usage rate in the current meeting environment is used as the basis for language selection. Once sufficient personal data has been accumulated and language habits have stabilized, the system switches to the usage rate of the speaker's historical language type corresponding to that ID number to achieve accurate personalized matching. As the meeting progresses, the system continuously accumulates the voice data of each speaker, updating language usage statistics in real time to make the language profile of each person increasingly accurate and the grasp of each person's language preferences more precise, thereby improving the overall performance of voice acquisition and recognition in multilingual mixed scenarios.

[0076] Furthermore, in step S22, the usage rate and usage threshold of each language type are compared, which is achieved through the following steps:

[0077] Step S221: Sort each language type in descending order of usage rate to generate a language priority sequence; label each language type according to the language priority sequence;

[0078] Step S222: Read the usage rate of the language type of the first label and compare the usage rate with the preset usage threshold;

[0079] Step S223: If the usage rate is greater than or equal to the usage threshold, proceed to step S224; if the usage rate is less than the usage threshold, proceed to step S225.

[0080] Step S224: Add the language type under the current label to the intermediate language label; read the usage rate of the language type of the next label, compare the usage rate with the preset usage threshold, and execute step S223;

[0081] Step S225: Add the language type under the current label to the intermediate language label and output the intermediate language label.

[0082] It should be explained that the purpose of setting the usage threshold is to select high-frequency languages ​​as candidates from the language usage rate, avoid making indiscriminate identification attempts for all languages, and thus improve the system processing efficiency.

[0083] Further, in step S24, accent recognition is performed on the speech segment based on the intermediate language label, and accent features are generated based on the recognition results through the following steps:

[0084] Step S241: Read the language type in the intermediate language tag and perform phoneme recognition to obtain the standard pronunciation phonemes; perform phoneme recognition on the speech segment to obtain the acoustic features of the current speech segment;

[0085] Step S242: Process the speech segment according to the standard pronunciation phonemes and acoustic features to obtain the accent deviation value of the speech segment;

[0086] Step S243: Determine whether an accent exists based on the accent deviation value of the speech segment; if the accent deviation value is lower than the accent correction threshold, then generate accent features; if the accent deviation value is higher than the accent correction threshold, then do not generate accent features.

[0087] It should be explained that if the accent deviation value is lower than the preset accent correction threshold, it means that the accent deviation of the current speech segment is large and the accent is serious. In this case, accent features need to be generated for subsequent correction. If the accent deviation value is higher than the preset accent correction threshold, it means that the current speech segment is pronounced clearly and the accent is light or there is no accent. In this case, accent features are not generated and the language processing model is used directly for recognition.

[0088] Specifically, speech segments are processed based on standard phonemes and acoustic features using the following formula:

[0089]

[0090] Where p represents the standard phoneme; X is the acoustic feature of the current speech segment; Q is the set of all possible phonemes; P(p|X) is the probability that the pronunciation is the standard phoneme p given the acoustic feature X; max q∈Q P(q|X) is the probability of the most likely phoneme when the model freely recognizes it; a phoneme is the smallest unit of speech in a language; GOP(p) is the accent deviation value of the current speech segment;

[0091] This yields the accent deviation value for each speech segment.

[0092] Furthermore, in step S1, voiceprint recognition and speaker segmentation are performed on the original speech data to generate multiple speech segments, which is achieved through the following steps:

[0093] Step S11: Continuously collect raw audio data in the current meeting environment, perform voice activity detection on the raw audio data, and extract audio segments containing valid speech.

[0094] Step S12: Extract voiceprint features for each speech segment and compare the voiceprint features with the voiceprint feature database; if the match is successful, assign the corresponding identity number from the voiceprint feature database to the speech segment; if the match fails, create a new identity number, assign the new identity number to the speech segment, and add the identity number to the voiceprint feature database.

[0095] Step S13: Concatenate the audio segments of the same identity number in chronological order to generate an audio segment carrying the identity number.

[0096] It should be noted that in step S3, the speech segment is processed using a language processing model or accent features, including any of the following methods:

[0097] Accent features are used to correct the accent of speech segments, and the language processing model processes the corrected speech segments.

[0098] Speech segments are processed using a language processing model.

[0099] It should be noted that when the accent is determined to be severe, the speech segment is corrected by accent features and then processed by the language processing model, ensuring the recognition accuracy in accent scenarios; when the accent is determined to be mild or non-accented, the speech segment is processed directly by the language processing model, avoiding unnecessary computational overhead and ensuring the system's processing efficiency.

[0100] Specifically, this method uses accent features to correct the accent of a speech segment, and a language processing model processes the corrected speech segment. The steps include: generating an accent adapter based on the accent features; inserting the accent adapter into a selected language processing model; and using the language processing model with the accent adapter inserted to recognize the speech segment, thus obtaining the corrected final speech data. This method effectively solves the problem of decreased recognition rate caused by accent interference through personalized accent adaptation.

[0101] This invention also provides a real-time voice data acquisition system for multilingual mixed scenarios, to implement the aforementioned real-time voice data acquisition method for multilingual mixed scenarios, such as... Figure 3 As shown, it includes an identity recognition module, a language accent recognition module, and a language processing module; the identity recognition module, language accent recognition module, and language processing module are connected in sequence.

[0102] The identity recognition module is used to perform voiceprint recognition and speaker segmentation on the raw voice data, and transmits the generated multiple voice segments carrying identity numbers to the language accent recognition module.

[0103] The language accent recognition module receives a speech segment and its associated identification number, identifies the identification number or speech segment, and obtains the final language label or accent feature of the current speech segment; the language accent recognition module transmits the final language label or accent feature to the language processing module;

[0104] The language processing module selects the corresponding language processing model based on the final language label, and processes the speech segment according to the language processing model or accent features to obtain the final speech data.

[0105] Furthermore, such as Figure 4As shown, the language accent recognition module includes a usage rate comparison module, an accent recognition module, and a language tag integration module. The usage rate comparison module is connected to the accent recognition module, the language tag integration module, and the identity recognition module, respectively. The accent recognition module and the language tag integration module are connected to the language processing module, respectively.

[0106] The usage rate comparison module receives multiple voice segments carrying identity numbers from the identity recognition module. The usage rate comparison module reads the usage rate of each language type based on the identity number, compares the usage rate of each language type with the usage threshold, and generates intermediate language tags based on the comparison results. The usage rate comparison module transmits the intermediate language tags to the language tag integration module and the accent recognition module respectively.

[0107] The accent recognition module identifies the accent of the speech segment based on the intermediate language tag and generates accent features based on the recognition results; the accent recognition module transmits the accent features to the language processing module;

[0108] The language tag integration module determines whether there is a language type in the intermediate language tag, generates the final language tag based on the determination result, and transmits the final language tag to the language processing module.

[0109] In this embodiment, the language accent recognition module further includes an output module, which is connected to the accent recognition module and the language tag integration module respectively. The accent recognition module and the language tag integration module output accent features and final language tags to the language processing module through the output module respectively.

[0110] The various embodiments of the present invention have now been described in detail. To avoid obscuring the concept of the invention, some details known in the art have not been described. Those skilled in the art will fully understand how to implement the technical solutions of this invention based on the above description, and the scope of the invention is defined by the appended claims.

Claims

1. A real-time speech data acquisition method for multilingual mixed scenarios, characterized in that, Its implementation includes the following steps: Step S1: Continuously collect raw voice data in the current conference environment, perform voiceprint recognition and speaker segmentation on the raw voice data, and generate multiple voice segments, each voice segment carrying an identity number; the identity number represents the speaker's unique identifier. Step S2: Read the speech segment and its carried identification number, identify the identification number or speech segment, and obtain the final language label or accent feature of the current speech segment; the final language label represents the language type of the speech segment; the final language label represents at least one language type; Step S3: Select the corresponding language processing model according to the final language label, and process the speech segment using the language processing model or the accent features to obtain the final speech data; The implementation of step S2 above includes the following steps: Step S21: Read the identity number, and read the usage rate of each language type based on the identity number. The usage rate refers to the frequency or number of times each language type is used. Step S22: Compare the usage rate and usage threshold of each language type, and generate intermediate language labels based on the comparison results; the intermediate language labels represent the language type of the speech segment; Step S23: Determine whether the intermediate language tag contains a language type; if the intermediate language tag contains at least one language type, proceed to step S24; if the intermediate language tag does not contain any language type, proceed to step S26. Step S24: Read the speech segment, perform accent recognition on the speech segment according to the intermediate language tag, and generate accent features based on the recognition results; proceed to step S25; Step S25: The intermediate language tags are used to generate the final language tags; Step S26: Add all language types to the intermediate language tags to generate the final language tags.

2. The real-time voice data acquisition method for multilingual mixed scenarios according to claim 1, characterized in that, In step S22, the usage rate and usage threshold of each language type are compared, which is achieved through the following steps: Step S221: Sort each language type in descending order of usage rate to generate a language priority sequence; label each language type according to the language priority sequence; Step S222: Read the usage rate of the language type of the first label, and compare the usage rate with a preset usage threshold; Step S223: If the usage rate is greater than or equal to the usage threshold, proceed to step S224; if the usage rate is less than the usage threshold, proceed to step S225. Step S224: Add the language type under the current label to the intermediate language label; read the usage rate of the language type of the next label, compare the usage rate with the preset usage threshold, and execute step S223; Step S225: Add the language type under the current label to the intermediate language label and output the intermediate language label.

3. The real-time voice data acquisition method for multilingual mixed scenarios according to claim 1, characterized in that, In step S1, voiceprint recognition and speaker segmentation are performed on the original speech data to generate multiple speech segments, which is achieved through the following steps: Step S11: Continuously collect raw audio data in the current conference environment, perform voice activity detection on the raw audio data, and extract audio segments containing valid audio. Step S12: Extract voiceprint features for each speech segment and compare the voiceprint features with the voiceprint feature database; if a match is successful, assign the corresponding identity number from the voiceprint feature database to the speech segment; If a match fails, a new identity number is created, the new identity number is assigned to the voice segment, and the identity number is added to the voiceprint feature database. Step S13: Concatenate the voice segments of the same identity number in chronological order to generate a voice segment carrying the identity number.

4. The real-time voice data acquisition method for multilingual mixed scenarios according to claim 1, characterized in that, In step S24, accent recognition is performed on the speech segment based on the intermediate language label, and accent features are generated based on the recognition results through the following steps: Step S241: Read the language type in the intermediate language tag and perform phoneme recognition to obtain the standard pronunciation phonemes; perform phoneme recognition on the speech segment to obtain the acoustic features of the current speech segment; Step S242: Process the speech segment according to the standard phonemes and the acoustic features to obtain the accent deviation value of the speech segment; Step S243: Determine whether an accent exists based on the accent deviation value of the speech segment; If the accent deviation value is lower than the accent correction threshold, then an accent feature is generated; If the accent deviation value is higher than the accent correction threshold, then no accent feature is generated.

5. A real-time voice data acquisition method for multilingual mixed scenarios according to claim 4, characterized in that, The speech segment is processed according to the standard phonemes and acoustic features, using the following formula: Where p represents the standard phoneme; X is the acoustic feature of the current speech segment; Q is the set of all possible phonemes; P(p|X) is the probability that the pronunciation is the standard phoneme p given the acoustic feature X; max q∈Q P(q|X) is the probability of the most likely phoneme when the model freely recognizes it; the phoneme is the smallest unit of speech in the language; GOP(p) is the accent deviation value of the current speech segment; This yields the accent deviation value for each speech segment.

6. The real-time voice data acquisition method for multilingual mixed scenarios according to claim 1, characterized in that, In step S3, the speech segment is processed using the language processing model or the accent features, including any of the following methods: The accent features are used to correct the accent of the speech segment, and the language processing model processes the corrected speech segment. The speech segment is processed using the language processing model.

7. The real-time voice data acquisition method for multilingual mixed scenarios according to claim 1, characterized in that, The usage rate includes the global usage rate of each language type in the current meeting environment or the usage rate of the speaker's historical language type corresponding to the identity number.

8. A real-time voice data acquisition system for multilingual mixed scenarios, characterized in that, To realize the real-time voice data acquisition method for multilingual mixed scenarios as described in claims 1 to 7, the method includes an identity recognition module, a language accent recognition module, and a language processing module; the identity recognition module, the language accent recognition module, and the language processing module are connected in sequence. The identity recognition module is used to perform voiceprint recognition and speaker segmentation on the original voice data, and transmit the generated multiple voice segments carrying identity numbers to the language accent recognition module. The language accent recognition module receives the speech segment and its carried identification number, identifies the identification number or speech segment, and obtains the final language label or accent feature of the current speech segment; the language accent recognition module transmits the final language label or accent feature to the language processing module; The language processing module selects the corresponding language processing model based on the final language label, and processes the speech segment according to the language processing model or the accent features to obtain the final speech data.

9. A real-time voice data acquisition system for multilingual mixed scenarios according to claim 8, characterized in that, The language accent recognition module includes a usage rate comparison module, an accent recognition module, and a language tag integration module. The usage rate comparison module is connected to the accent recognition module, the language tag integration module, and the identity recognition module, respectively. The accent recognition module and the language tag integration module are connected to the language processing module, respectively. The usage rate comparison module receives multiple voice segments carrying identity numbers transmitted by the identity recognition module. The usage rate comparison module reads the usage rate of each language type based on the identity number, compares the usage rate of each language type with a usage threshold, and generates an intermediate language label based on the comparison result. The usage rate comparison module transmits the intermediate language label to the language label integration module and the accent recognition module respectively. The accent recognition module performs accent recognition on the speech segment based on the intermediate language tag, and generates accent features based on the recognition results; the accent recognition module transmits the accent features to the language processing module; The language tag integration module determines whether there is a language type in the intermediate language tag, generates the final language tag based on the determination result, and transmits the final language tag to the language processing module.

10. A real-time voice data acquisition system for multilingual mixed scenarios according to claim 9, characterized in that, The language accent recognition module also includes an output module, which is connected to the accent recognition module and the language tag integration module respectively. The accent recognition module and the language tag integration module output the accent features and the final language tags to the language processing module through the output module.