Multilingual voice input recognition and switching method

By using a multilingual adaptive recognition algorithm based on timbre features and semantic difference quantification, the language switching problem of speech recognition systems in multilingual scenarios is solved, enabling real-time and accurate recognition and adaptive switching of multilingual speech input, thus improving the efficiency of voice interaction.

CN121075313APending Publication Date: 2025-12-05SHANGHAI MAIJUN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511579607.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing speech recognition systems struggle to achieve real-time and accurate recognition and adaptive switching of multilingual speech input in multilingual scenarios, leading to increased language recognition error rates and reduced voice interaction efficiency.

Method used

A multilingual adaptive recognition algorithm based on timbre feature statistics, semantic difference quantification, and language standard score calculation is adopted. The default input language is set by the user's location, the semantic difference score and voice intensity of the user group are analyzed, and the voice signal is collected in real time to make language switching decisions.

Benefits of technology

It achieves real-time accurate recognition and adaptive switching of multilingual voice input, improving the efficiency of voice interaction and reducing language recognition errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075313A_ABST
    Figure CN121075313A_ABST
Patent Text Reader

Abstract

The invention discloses a multilingual voice input recognition and switching method, relates to the technical field of voice recognition and switching, is used for solving the problem that the language recognition error rate is increased, and comprises the following steps: after receiving incoming call information through a multilingual smart phone customer service, setting a default language according to the position of a user and sending instruction information; receiving voice fed back by the users, counting the timbre number, determining the number of the users, forming a user group, generating a semantic difference score through signal processing, judging whether to switch languages according to the semantic difference score, if switching is needed, setting detection time, collecting voice intensity, and comparing voice signals in a detection period to obtain the same semantic repetition times, and calculating a language standard score by synthesizing the voice intensity and the number of repetitions, selecting a language corresponding to the maximum value of the language standard score as a language switching standard and executing switching, determining a default language through a user position, and distinguishing a multi-user group in combination with timbre analysis to realize semantic difference quantification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition switching, more particularly, to a multilingual speech input recognition and switching method. BACKGROUND

[0002] With the rapid development of global communication, multilingual communication has become a key requirement in intelligent speech service systems (such as telephone customer service, online voice assistants, and cross-border business call centers). Existing speech recognition technology generally builds a speech recognition engine based on a single language model and achieves recognition of input speech through acoustic feature extraction and semantic matching algorithms.

[0003] The prior art has the following disadvantages: Currently, in a multilingual scenario, different users may use different languages, dialects, or accents for communication. Existing speech recognition systems have a significant contradiction between recognition accuracy and response speed, making it difficult to achieve real-time accurate recognition and adaptive switching of multilingual speech input, resulting in increased language recognition error rate and reduced speech interaction efficiency. Therefore, a multilingual speech input recognition and switching method is proposed.

[0004] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present disclosure, and therefore it can include information that does not constitute the prior art known to those of ordinary skill in the art. SUMMARY

[0005] To overcome the above-mentioned defects of the prior art, embodiments of the present application provide a multilingual speech input recognition and switching method that uses a multilingual adaptive recognition algorithm based on timbre feature statistics, semantic difference quantization, and language standard score calculation to solve the problems raised in the background art.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solution, a multilingual speech input recognition and switching method, comprising the following steps: Step S1: When the multilingual intelligent telephone customer service receives incoming call information, sets a default input language according to the user location in the incoming call information, and sends instruction information using the default input language to wait for the user's feedback speech; Step S2: Receive the user's feedback speech and count the number of timbres of the feedback speech to determine the number of users consulting, and count the number of users to obtain a user group and generate a semantic difference score by signal processing and analyzing the voice signal of the user group; Step S3: Analyze whether to perform language switching according to the semantic difference score, set a detection time when language switching is performed, and real-time collect the voice intensity of the user group, call the voice signal of the user group and the voice signal within the detection time to calculate the number of times of the same semantics; Step S4: Synthesize the voice intensity of the user groups and the number of repetitions of the same semantics to obtain a language standard score, select the language corresponding to the maximum language standard score as the language switching standard and perform switching.

[0007] In a preferred embodiment, in step S1, the telephone number of the user is formatted according to international standards to obtain the area code of the user number; Match the area code of the user number with the area code database to obtain the area information corresponding to the area code of the user number; Determine the area where the user is located according to the area information corresponding to the area code, and set the default input language according to the area where the user is located.

[0008] In a preferred embodiment, in step S1, select the instruction information corresponding to the default input language, and convert the instruction information into a voice signal through a speech synthesis unit; Send the voice signal to the incoming user through the voice channel in time sequence, and receive the feedback voice of the user.

[0009] In a preferred embodiment, in step S2, the feedback voice of each user is obtained through the voice receiving unit of the telephone call-in interface; Preprocess the feedback voice of the user through endpoint detection and basic noise suppression to obtain the sound signal of each user; Obtain the formant frequency of the sound signal of each user through the hardware DSP unit built-in real-time voice analyzer; Subtract the formant frequencies of the sound signals of different users to obtain the timbre feature value; If the timbre feature value is less than the preset timbre feature threshold, it is determined to be the same timbre cluster; Otherwise, it is determined to be different timbre clusters.

[0010] In a preferred embodiment, in step S2, the number of different timbre clusters is counted to obtain the number of users consulting; The users in the same timbre cluster are integrated into a user group; Integrate the sound signals of each user in each user group into the sound signal of each user group; Cut the sound signal of each user group into voice frames of each user group according to a fixed time length.

[0011] In a preferred embodiment, in step S2, integrate the voice frames of each user group into a set of user group voice data; Generate the semantic vector of each user group by identifying the set of user group voice data through a semantic embedding model; The similarity of the semantic vectors of different users in each user group is calculated by average value, and the absolute value of 1 minus the semantic difference score is obtained.

[0012] In a preferred embodiment, in step S3, if the semantic difference score is greater than or equal to the preset difference score threshold, it is determined to perform language switching; On the contrary, it is determined not to perform language switching; When language switching is performed, set the detection time, and obtain the voice intensity of each user group through the signal amplitude detection unit of the telephone call-in interface; The voice signals of each user group are continuously collected by the voice receiving unit. The voice signals of each user group in the detection time are stored in the cache area and integrated into the detection voice data segment.

[0013] In a preferred embodiment, in step S3, the voice signals corresponding to each user group are called to match the semantic content with the detection voice data segment; If the voice signals of each user group and the detection voice data segment in the detection time are detected to have the same short sentence or key word in the word sequence, it is determined to be a semantic repetition; The number of times of the same semantic appearance is counted, and the number of times is taken as the same semantic repetition number.

[0014] In a preferred embodiment, in step S4, the voice intensity and the same semantic repetition number of each user group are standardized to obtain the voice intensity factor and the repetition number factor; The language standard score is calculated by integrating the voice intensity factor and the repetition number factor; The language corresponding to the maximum value of the language standard score is taken as the language switching standard and the switching is performed.

[0015] The technical effects and advantages of the present application are: The application receives the incoming call information through the multi-language intelligent telephone customer service, sets the default input language according to the user location in the incoming call information, and sends instruction information using the language, waits for user feedback voice, receives the user feedback voice and counts the number of voice tones to determine the number of consulting users, then obtains the user group according to the number statistics, generates a semantic difference score by signal processing and analyzing the voice signal of the user group, judges whether to perform language switching according to the semantic difference score, sets the detection time and collects the voice intensity of the user group in real time if switching is required, simultaneously calls the voice signal of the user group and the voice signal within the detection time for comparison to obtain the same semantic repetition number, calculates the language standard score by integrating the voice intensity of the user group and the same semantic repetition number, selects the language corresponding to the maximum value of the language standard score as the language switching standard and performs switching; the default language is determined by the user location and the multi-user group is distinguished by combining the voice tone analysis, the semantic difference is quantified, and the language switching is judged based on objective data. BRIEF DESCRIPTION OF DRAWINGS

[0016] Fig. 1 The implementation flowchart of the multi-language voice input recognition and switching method of the application.

[0017] Fig. 2 The step schematic diagram of the multi-language voice input recognition and switching method of the application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.

[0019] The application receives the incoming call information through the multi-language intelligent telephone customer service, sets the default input language according to the user location in the incoming call information, and sends instruction information using the language, waits for user feedback voice, receives the user feedback voice and counts the number of voice tones to determine the number of consulting users, then obtains the user group according to the number statistics, generates a semantic difference score by signal processing and analyzing the voice signal of the user group, judges whether to perform language switching according to the semantic difference score, sets the detection time and collects the voice intensity of the user group in real time if switching is required, simultaneously calls the voice signal of the user group and the voice signal within the detection time for comparison to obtain the same semantic repetition number, calculates the language standard score by integrating the voice intensity of the user group and the same semantic repetition number, selects the language corresponding to the maximum value of the language standard score as the language switching standard and performs switching; the default language is determined by the user location and the multi-user group is distinguished by combining the voice tone analysis, the semantic difference is quantified.

[0020] Embodiment 1, a multilingual voice input recognition and switching method, as shown, comprising the following steps: Figs. 1-2 Step S1: When the multi-language intelligent telephone customer service receives the incoming call information, set the default input language according to the user location in the incoming call information, and send the instruction information using the default input language to wait for the user's feedback voice; Step S2: Receive the user's feedback voice and count the number of voice tones to determine the number of users consulting, and count the user groups according to the number of users and generate semantic difference scores by analyzing the voice signals of the user groups through signal processing; Step S3: According to the semantic difference score, analyze whether to perform language switching, set the detection time when language switching is performed, and real-time collect the voice intensity of the user group, call the voice signal of the user group and the voice signal in the detection time to calculate the same semantic repetition times; Step S4: Synthesize the voice intensity of the user group and the same semantic repetition times to get the language standard score, select the language corresponding to the maximum language standard score as the language switching standard and perform switching.

[0021] The specific implementation is as follows: In step S1, the user's phone number is formatted according to the international standard to get the area code of the user's number; Match the area code of the user's number with the area code database to get the area information corresponding to the area code of the user's number; Determine the user's area according to the area information corresponding to the area code, and set the default input language according to the user's area; Match the default input language with the instruction information database to get the instruction information, and convert the instruction information into a voice signal through a voice synthesis unit; Send the voice signal to the incoming user through the voice channel in time sequence, and receive the user's feedback voice.

[0022] ​It needs to be explained that the international standard refers to the E.164 standard formulated by the International Telecommunication Union, which is a global telephone number numbering standard used to ensure the uniformity and identifiability of telephone number formats across countries and operators; the area code database refers to a data set used to store and manage global or specific regional telephone number area code information, which can quickly identify the geographical area or country to which the user belongs according to the area code part of the telephone number; the area code refers to a digital number used to distinguish telephone lines in different geographical areas, which is part of the telephone number and is used to identify the area or country where the user is located; the instruction information database refers to a data set storing multi-lingual voice instruction texts and their related parameters, which is used to match the default input language to obtain instruction information; the instruction information refers to the voice instruction text matched with the default input language; the voice synthesis unit refers to a unit that converts text instructions into playable voice signals; the voice channel refers to a logical or physical transmission path used to transmit user voice signals during telephone call-in or communication; The above process realizes the rapid identification of the user's geographical location by formatting the user's telephone number according to the international standard and extracting the area code, and automatically sets the default input language based on the user's location to ensure that multi-lingual telephone customer service can interact with the user in the most appropriate language.

[0023] In step S2, the voice receiving unit of the telephone call-in interface obtains the feedback voice of each user; The feedback voice of the user is preprocessed by endpoint detection and basic noise suppression to obtain the sound signal of each user; The formant frequency of the sound signal of each user is obtained by the built-in hardware DSP unit of the real-time voice analyzer; The formant frequencies of the sound signals of different users are subtracted to obtain the timbre feature value; According to the comparison between the timbre feature value and the preset timbre feature threshold value, it is determined that: If the timbre feature value is less than the preset timbre feature threshold value, it is determined to be the same timbre cluster; If the timbre feature value is greater than or equal to the preset timbre feature threshold value, it is determined to be different timbre clusters; The number of different timbre clusters is counted to obtain the number of consultation users corresponding to the feedback voice; Users in the same timbre cluster are integrated into a user group; Repeat the above steps to generate a corresponding number of user groups according to the number of consultation users; The sound signals of each user in each user group are integrated into the sound signal of each user group; It should be noted that the voice receiving unit of the telephone inbound interface refers to the unit used to receive the voice signal sent by the user via telephone and convert the telephone signal into a digital audio sequence; endpoint detection refers to the technology of identifying the start and end positions of valid voice segments in a continuous voice signal; basic noise suppression refers to the processing method of removing environmental noise or background interference from the audio signal and retaining the main voice components; a real-time voice analyzer is an electronic device used to collect, process, and analyze voice signals, capable of real-time detection of voice characteristics such as pitch, timbre, loudness, and spectrum; the built-in hardware DSP unit is a microprocessor specifically used for high-speed digital signal processing to obtain the formant frequencies of the feedback voice; the preset timbre feature threshold is an important parameter for determining whether different feedback voices belong to the same timbre cluster. By collecting voice samples from different users in historical inbound calls, the mean and standard deviation of the timbre feature values ​​of the same user are calculated to obtain the range of the same user; the mean and standard deviation of the timbre feature values ​​of another user are calculated to obtain the range of another user; the average value of the midpoint of the distribution interval of the two is selected as the preset timbre feature threshold to balance false positives and false negatives.

[0024] The audio signals of each user group are divided into audio frames of each user group according to a fixed time length; The voice frames of each user group are integrated into a voice data set for each user group; Semantic embedding models are used to identify speech data sets from various user groups and generate semantic vectors for each user group. The specific process is as follows: The semantic embedding model is built on a deep learning architecture and uses a convolutional neural network. The input is a set of user speech data after endpoint detection and noise suppression processing. Each speech segment is first subjected to short-time Fourier transform to extract acoustic features. Through pre-training, the input acoustic features are mapped into semantic vectors of fixed dimensions; The pairwise similarity of semantic vectors within each user group is calculated using the following formula: ,in For the first Semantic vectors of individual users For the first Semantic vectors of individual users The similarity of semantic vectors among different users within each user group; The semantic similarity of the semantic vectors of different users within each user group is averaged, and the absolute value is subtracted from 1 to obtain the semantic difference score.

[0025] It needs to be explained that the purpose of the semantic difference score calculation is to quantify the degree of semantic consistency among members of the same user group. If all users are highly consistent in semantics, the similarity is close to 1, and the difference score tends to 0, indicating that the user group may use the same language or express consensus, and there is no need to switch languages; on the contrary, if the semantic difference score is high, it indicates that there may be language understanding differences or different languages used within the user group, which is a key signal to trigger language switching analysis; the fixed time length refers to dividing the continuous user group voice signal into small segments, according to the fact that the fundamental frequency of human voice is generally between 80 and 400 Hz, and the voice information does not change much within 20 to 50 milliseconds, therefore, 20 to 40 milliseconds are usually selected as the voice frame length, which can ensure the stability of the voice characteristics within the frame and will not lose important information; the voice frame is a small segment of continuous data with a fixed time length in the voice signal; the convolutional neural network is a kind of deep learning model, which is used to process data with grid structure, such as images, voice or time series signals, and can automatically extract features from raw data without manual design of features by simulating the processing mechanism of the biological visual system; the short-time Fourier transform is a tool for analyzing the time-varying spectrum of non-stationary signals; the method realizes fast analysis of semantic consistency and difference in a multi-user and multi-language environment, which can automatically complete user number recognition, timbre distinction and semantic analysis without human intervention, improve the recognition accuracy and response efficiency of multi-language telephone customer service in complex call scenarios, and ensure the real-time and stability of voice processing, which is helpful for subsequent language switching and service optimization.

[0026] In step S3, the semantic difference score is compared with the preset difference score threshold to determine whether to perform language switching: If the semantic difference score is greater than or equal to the preset difference score threshold, it is determined to perform language switching; If the semantic difference score is less than the preset difference score threshold, it is determined not to perform language switching; It needs to be explained that the preset difference score threshold is an important parameter for determining whether to perform language switching, by collecting user voice data of multiple actual incoming calls and calculating the semantic difference score of each call to obtain a set of historical difference score data, calculating the average and standard deviation of the historical difference score data, and the sum of the average and standard deviation of the historical difference score data as the preset difference score threshold.

[0027] When language switching is performed, set a detection time, and obtain the voice intensity of each user group through the signal amplitude detection unit of the telephone incoming interface; During the detection time, continuously collect the voice signals of each user group through the voice receiving unit; Store the voice signals of each user group in the detection time to the cache area and integrate them into the detection voice data segment; The voice signal of each user group is matched with the detected voice data segment in semantic content; If the voice signal of each user group and the detected voice data segment in the detection time have the same short sentence or key word in the word sequence, it is determined that there is semantic repetition; The number of times of the same semantic content is counted, and the number of times is taken as the number of times of the same semantic repetition; The purpose of counting the number of times of the same semantic repetition is to capture the behavior pattern of the user group when the communication is blocked. When the user finds that the system cannot understand his intention, he will often repeat or change the way to express the same meaning. Therefore, the number of times is an important behavior indicator, which reflects the degree to which the current language cannot meet the user's needs. The higher the number of times, the greater the communication obstacle, and the stronger the urgency of switching languages.

[0028] It needs to be explained that the detection time is a time window for collecting and analyzing the voice intensity of the user group. The voice is usually processed in short time frames, each frame being 20 to 30 milliseconds. The detection time should cover enough frames, such as 30 to 50 frames, to ensure the stability of the voice intensity statistics. The signal amplitude detection unit of the telephone incoming interface is to calculate the instantaneous amplitude of the digital audio sequence to obtain the voice intensity of the user group after receiving the audio through the analog front-end circuit built-in the telephone interface or communication board card. The voice receiving unit is an audio acquisition unit built-in the call center telephone interface or communication card, which acquires digital audio frames through audio API for continuous acquisition of the voice signal of the user group. The semantic content matching is to compare the existing voice signal of the user group with the voice data collected in the detection time period to determine whether they express the same meaning or semantic content. The specific process is as follows: the voice data is converted into text by the embedded speech recognition chip and output in the form of word sequence. The voice signal of the user group and the voice signal in the detection time are compared to determine whether they have the same short sentence or key word in the word sequence.

[0029] In step S4, the voice intensity of each user group and the number of times of the same semantic repetition are standardized to obtain the voice intensity factor and the number of times factor; The voice intensity factor and the number of times factor are combined to calculate the language standard score, and the calculation formula is: wherein, is the voice intensity factor, is the number of times factor, and is a preset weighting coefficient, is the language standard score; It should be noted that the preset weighting coefficient is a parameter for balancing the influence degree of the speech intensity factor and the repetition factor on the language standard score, and the initial value is set according to historical data and actual experience, and is continuously optimized through a dynamic adjustment method to balance the contribution of speech intensity and repetition times to the language switching decision; the greater the speech intensity factor and the greater the repetition factor, the louder the sound of the user group and the more times the same semantics are repeatedly expressed, and the greater the language standard score; the smaller the speech intensity factor and the smaller the repetition factor, the less prominent the sound of the user group and the fewer times the semantics are repeated, and the smaller the language standard score.

[0030] The language standard scores of each user group are compared to obtain a maximum language standard score; The language corresponding to the maximum language standard score is taken as the language switching standard and switching is performed.

[0031] It should be explained that the standardization processing mode includes but is not limited to a standard linear transformation based on interval scaling, a Z-Score standardization method based on statistics, or a normalization method based on a nonlinear mapping function, and the application method of the standardization processing is not described here; the most suitable language is selected according to the speech characteristics of the user group in the actual call to reduce the communication obstacles caused by language misjudgment, and based on objective quantitative indicators, the language switching is more scientific, controllable, and unnecessary switching operations are reduced.

[0032] Finally, it should be noted that in this document, relational terms such as first and second and the like can only be used to distinguish one entity or action from another entity or action, without necessarily requiring or implying that these entities or actions exist in any such actual relationship or order.

[0033] Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the sentence "includes a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0034] In this document, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. It will be further understood that the terms "comprises," "comprising," "includes," and "including," or the like, when used in this specification, specify the presence of stated features, integers, steps, operations, components, parts, or the like, but do not preclude the presence or addition of one or more other features, integers, steps, operations, components, parts, or the like.

[0035] Various embodiments described in this specification are described with reference to particular implementations. Embodiments can be practiced with other systems, components, materials, acts, operations, and steps than those described and / or shown in this specification, and not solely with the particular implementations described in this specification. The terms "comprise," "comprising," "include," "including," and "includes" or the like are used herein to indicate the presence of stated features, integers, steps, operations, components, parts, or the like, but not to the exclusion of the presence or addition of one or more other features, integers, steps, operations, components, parts, or the like.

[0036] modifications of the embodiments disclosed in this application that can be made by those skilled in the art, without departing from the spirit or scope of the application. Accordingly, although specific embodiments have been illustrated and described herein, it should be appreciated that the application is not limited to the exact construction described above. Rather, it is the following claims, including all equivalents that should be afforded the broadest scope of protection.

Claims

1. A method for multilingual voice input recognition and switching, characterized in that: The method comprises the following steps: Step S1: After the multi-language intelligent telephone customer service receives the call information, the default input language is set according to the user location in the call information, and instruction information is sent out using the default input language to wait for the feedback voice of the user; Step S2: The feedback voice of the user is received and the number of timbres of the feedback voice is counted to determine the number of users for consultation, the user group is obtained by counting according to the number of users, and the semantic difference score is generated by analyzing the sound signals of the user group through signal processing; Step S3: Whether to perform language switching is analyzed according to the semantic difference score, the voice intensity of the user group is collected in real time when the language switching is performed, and the same semantic repetition number is obtained by calculating the sound signals of the user group and the sound signals within the detection time; Step S4: The language standard score is obtained by comprehensively considering the voice intensity of the user group and the same semantic repetition number, the language switching standard is selected as the language corresponding to the maximum value of the language standard score, and the switching is performed.

2. The multi-language voice input recognition and switching method according to claim 1, wherein: in step S1, the telephone number of the user is formatted according to the international standard to obtain the area code of the user number; The area code of the user number is matched with the area code database to obtain the area information corresponding to the area code of the user number; The area where the user is located is determined according to the area information corresponding to the area code, and the default input language is set according to the area where the user is located.

3. The multi-language voice input recognition and switching method according to claim 2, wherein: In step S1, the instruction information corresponding to the default input language is selected, and the instruction information is converted into a voice signal through a voice synthesis unit; The voice signal is sent to the incoming user in time sequence through the voice channel, and the feedback voice of the user is received.

4. The multi-language voice input recognition and switching method according to claim 1, wherein: In step S2, the feedback voice of each user is obtained through the voice receiving unit of the telephone call-in interface; The feedback voice of the user is preprocessed to obtain the sound signal of each user through endpoint detection and basic noise suppression; The formant frequency of the sound signal of each user is obtained through the built-in hardware DSP unit of the real-time voice analyzer; The formant characteristic values of the sound signals of different users are obtained by subtracting the formant frequencies of the sound signals of different users; If the formant characteristic value is less than the preset formant characteristic threshold value, it is determined that it is the same timbre cluster; Otherwise, it is determined that it is a different timbre cluster.

5. The multi-language voice input recognition and switching method according to claim 4, wherein: In step S2, the number of different timbre clusters is counted to obtain the number of users for consultation; The users in the same timbre cluster are integrated into a user group; The sound signals of each user in each user group are integrated into the sound signals of each user group; The sound signals of each user group are cut into voice frames of each user group according to a fixed time length.

6. The multi-language voice input recognition and switching method according to claim 5, wherein: In step S2, the voice frames of each user group are integrated into a set of user group voice data. ​ The semantic vector of each user group is generated by recognizing the voice data set of each user group through a semantic embedding model; The similarity of the semantic vectors of different users in each user group is calculated by averaging, and the semantic difference score is obtained by taking the absolute value of 1 minus the similarity.

7. The multilingual voice input recognition and switching method of claim 1, wherein: In step S3, if the semantic difference score is greater than or equal to the preset difference score threshold, it is determined to perform language switching; Otherwise, it is determined not to perform language switching; When language switching is performed, set a detection time, and obtain the voice intensity of each user group through the signal amplitude detection unit of the telephone call-in interface; The voice signals of each user group are continuously collected through the voice receiving unit. The voice signals of each user group within the detection time are stored in the cache area and integrated into the detection voice data segment.

8. The multilingual voice input recognition and switching method of claim 7, wherein: In step S3, the voice signals of each user group are matched with the detection voice data segment in terms of semantic content; If the voice signals of each user group are detected to have the same short sentence or key word in the word sequence as the detection voice data segment within the detection time, it is determined to be a semantic repetition; The number of times the same semantic appears is counted, and the number of times is taken as the number of times of the same semantic repetition.

9. The multilingual voice input recognition and switching method of claim 1, wherein: In step S4, the voice intensity and the number of times of the same semantic repetition of each user group are standardized to obtain the voice intensity factor and the repetition number factor; The language standard score is calculated by integrating the voice intensity factor and the repetition number factor; The language corresponding to the maximum language standard score is taken as the language switching standard and the switching is performed.