Multi-language support method and device for language communication auxiliary training instrument

By constructing a three-dimensional evaluation system and a categorized scenario library, the problems of single input and lack of hierarchical evaluation in existing equipment have been solved, enabling multilingual support and efficient training, and improving the pertinence and effectiveness of language communication-assisted training.

CN121366569APending Publication Date: 2026-01-20SHENZHEN HANNIKON TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511569761.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing language communication assistance training devices have limited input methods, insufficient data preprocessing, static and unstratified assessment systems, and a lack of prioritization in correction schemes, resulting in poor training targeting and difficulty in improving cross-language communication skills.

Method used

We construct a three-dimensional evaluation system covering pronunciation, grammar, and semantics, establish a dedicated standard expression benchmark library for each language, receive speech and text data through dual input ports, perform noise reduction, filtering, and format unification processing, prioritize according to the degree of impact of the problem, conduct reinforcement training in combination with historical records, construct a scenario library classified by type and difficulty, and create an audiovisual interactive environment.

Benefits of technology

It significantly reduces recognition errors, improves training efficiency, simulates scenarios that closely resemble real communication, enhances user immersion, helps users overcome language difficulties in a targeted manner, and steadily improves cross-language communication skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366569A_ABST
    Figure CN121366569A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-language support method and device for a language communication auxiliary training instrument, relates to the technical field of language auxiliary training, and aims to solve the problem that the auxiliary training effect of different language types is poor. According to the method, a three-dimensional evaluation system covering pronunciation, grammar and semantics is constructed, an exclusive standard expression benchmark library is established for each language, an evaluation threshold value is set, three-level priority is divided according to the degree of influence of questions on communication, historical records can be called to strengthen repeated question training, the scheme is made to be more suitable for user shortages, targeted attacking of language problems is assisted, and the user experience is improved. Voice and character data are received through double input ports, input habits of different users are adapted, meanwhile, preprocessing is carried out in a targeted mode, the voice data are subjected to noise reduction, segmentation and normalization to eliminate environmental interference and data differences, and the character data are subjected to invalid character filtering and format unification to eliminate redundant information; and features can be extracted according to data type differentiation, so that subsequent identification errors are greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of language auxiliary training, in particular to a multi-language support method and device for a language communication auxiliary training instrument. BACKGROUND

[0002] In the current language communication auxiliary training field, the existing technology has obvious shortcomings. On the one hand, most training devices have single input mode, only supporting one of voice or text, which is difficult to adapt to different user habits, and the data preprocessing link is simplified, the voice input is not fully denoised, and the text input is not effectively filtered invalid characters, resulting in large recognition error of subsequent content, easy misjudgment of language type judgment, and inaccurate semantic analysis. On the other hand, the evaluation system is mostly static fixed standard, without establishing a dedicated benchmark library for different language characteristics, the evaluation dimension is extensive, and it is difficult to fully capture the details of pronunciation, grammar, and semantics; the correction scheme lacks priority division, without formulating intensive training combined with user historical problems, and the scene simulation and correction demand are disconnected, the analysis only stays in single dimension, without historical data comparison and common problem association, the training is not targeted, and it is difficult to help users efficiently improve cross-language communication ability. SUMMARY

[0003] The purpose of the present application is to provide a multi-language support method and device for a language communication auxiliary training instrument, to build a three-dimensional evaluation system covering pronunciation, grammar, and semantics, to establish a dedicated standard expression benchmark library for each language and set an evaluation threshold, to divide three levels of priority according to the influence degree of problems on communication, and to retrieve historical records to intensify repeated problem training, so that the scheme is more suitable for user shortcomings, helps to target language difficulties, and steadily improves training efficiency. Through the double input ports, voice and text data are received to adapt to different user input habits, and preprocessing is carried out in a targeted manner. The voice data is denoised, segmented, and normalized to eliminate environmental interference and data differences. The text data is filtered for invalid characters, unified in format, and redundant information is removed. It can also extract features according to data type differences, greatly reduce subsequent recognition error, and solve the problems in the prior art.

[0004] To achieve the above purpose, the present application provides the following technical scheme: A multi-language support method for a language communication auxiliary training instrument, comprising: First, the original language data of the user is received and processed; the processed original language data is subjected to content recognition, and the content is analyzed according to the recognized content; after the original language data is analyzed, a dynamic evaluation system is constructed, and a correction scheme is generated according to the constructed dynamic evaluation system; the user simulates a practical language communication scene according to the generated correction scheme; and the user's pronunciation or expression is analyzed in multiple dimensions according to the simulation result.

[0005] Preferably, the original language data of the user is received and processed, including: The original language data of the user is received according to a data receiving port, and the original language data includes voice input data and text input data; The received original language data is pre-processed, wherein the voice input data is subjected to noise reduction processing, voice segmentation and normalization processing; and the text input data is subjected to invalid character filtering and format unification.

[0006] Preferably, the processed original language data is subjected to content recognition, and the content is parsed according to the recognized content, including: The original language data after data pre-processing is subjected to feature extraction, wherein the voice input data is subjected to feature extraction according to phoneme distribution, syllable structure and prosody characteristics; and the text input data is subjected to feature extraction according to character set characteristics, vocabulary characteristics and grammar characteristics; The extracted feature data is subjected to language matching, which is similarity comparison of the extracted feature data with language templates in a feature library, and the language type of the extracted feature data is determined according to the comparison result; After the language type is confirmed, the content is recognized, which is voice-to-text conversion and semantic analysis of the voice input data; and voice analysis of the text input data; The specific content of the original language data of the user is obtained after the parsing is completed.

[0007] Preferably, after the original language data is parsed, a dynamic evaluation system is constructed, and a correction scheme is generated according to the constructed dynamic evaluation system, including: First, the evaluation dimensions of the dynamic evaluation system are confirmed, including pronunciation dimension, grammar dimension and semantic dimension, wherein the pronunciation dimension includes phoneme accuracy, tone and intonation matching degree, stress position and speech rhythm; the grammar dimension includes word combination rationality, sentence structure standardization, time and aspect accuracy and punctuation symbol use correctness; and the semantic dimension includes semantic integrity, semantic accuracy and scene adaptability; According to the evaluation dimensions, a standard expression benchmark library is established for each language type, and the evaluation threshold range is set; According to the evaluation dimensions and the standard expression benchmark library, the parsed original language data is evaluated; Among them, according to different evaluation dimensions, the parsed original language data is compared with sample materials in the standard expression benchmark library, and the comparison results are comprehensively scored according to the evaluation threshold range.

[0008] Preferably, after the original language data is parsed, a dynamic evaluation system is constructed, and a correction scheme is generated according to the constructed dynamic evaluation system, further including: According to the comprehensive evaluation result, priority ranking is performed, wherein the priority ranking includes first priority, second priority and third priority, the first priority is a problem that seriously affects understanding and core grammar error, the second priority is a problem that does not affect understanding but does not conform to the standard, and the third priority is a detail problem; According to different problems, correction measures are made, including pronunciation problem correction, grammar problem correction and semantic problem correction; At the same time, the historical evaluation record of the user is called, and if there are repeated problems, targeted reinforcement training is performed; Finally, the correction scheme is visualized to generate a report, and the user's correction scheme is obtained after generation.

[0009] Preferably, the user simulates a practical language communication scene according to the generated correction scheme, including: The communication scene of the practical scene library in the database is called; the scene corresponding to the problem in the correction scheme is adapted; The communication scene of the practical scene library includes daily communication, business in the workplace, academic education and emergency handling, each scene includes a primary, intermediate and advanced difficulty level, the primary scene contains 5-8 basic expressions; the intermediate scene contains 10-15 expressions with logical association; the advanced scene contains complex interaction; The problem in the correction scheme is adapted according to the communication scene, and the practical language communication scene in the correction scheme is obtained after the adaptation is completed; The practical language communication scene in the correction scheme is created in an interactive environment, which includes a visual environment, an auditory environment and an interactive prompt; After the interactive environment is created, the interactive logic and the dialogue role are designed, wherein 1-3 interactive roles are set according to the practical language communication scene in the correction scheme, the interactive roles include a preset language style and a response logic, and the dialogue process is set, the dialogue process includes an initial dialogue, a core interaction and an advanced interaction; After the interactive logic and the dialogue role are designed, a complete voice communication scene is obtained; The user simulates the scene according to the complete voice communication scene, and the simulation process is recorded in real time during the simulation process.

[0010] Preferably, the pronunciation or expression of the user is analyzed in multiple dimensions according to the simulation result, including: The simulation process data recorded in real time during the simulation process is summarized and standardized; The simulated process data after summarization and standardization is subjected to deep analysis, wherein the deep analysis includes pronunciation dimension analysis, grammar dimension analysis and semantic dimension analysis, the pronunciation dimension analysis is basic pronunciation accuracy and scene adaptive pronunciation analysis, the scene adaptive pronunciation is basic grammar rule implementation and sentence structure and complexity analysis, and the semantic dimension analysis is semantic transmission quality and scene adaptation and style unification analysis; The deep analysis result is combined with the historical simulation data of the user, and after the combination, overall judgment is performed, wherein the overall judgment includes longitudinal comparison and transverse correlation, the longitudinal comparison is to compare the deep analysis result with the historical simulation data, and after the comparison, the improvement track of the key problems in the correction scheme is tracked; the transverse correlation is to analyze the common problems repeatedly occurring in different scenes, and the relationship between the common problems and the scene difficulty is correlated; After the combination, the final multi-dimensional analysis result is obtained, the multi-dimensional analysis result is subjected to comprehensive conclusion generation, and the generated comprehensive conclusion is subjected to visual report presentation.

[0011] Preferably, the construction method of the practical scene library comprises: Obtain four types of original communication scene data of daily communication, business in the workplace, academic education and emergency handling marked with difficulty levels; the difficulty levels are defined based on preset difficulty standards, and the preset difficulty standards include dialogue turns, vocabulary complexity, grammar structure complexity and cultural background depth; respectively extract features from the four types of original communication scene data to obtain daily features, workplace features, academic features and emergency features including difficulty features; Based on the preset feature type classification system, perform full connection cross-class matching on all four types of features: for any two types of features, extract feature pairs with the same feature type, calculate a matching degree value, and generate a common target feature type; wherein the target feature type is a standardized scene feature type in the preset feature reference library; Based on the preset feature reference library, determine the contribution coefficient corresponding to the target feature type and the matching degree value, and associate the contribution coefficient with the feature pairs involved in the matching result; Based on the preset weight, calculate the comprehensive contribution coefficient of the original scene data; when the comprehensive contribution coefficient is greater than or equal to a preset contribution coefficient threshold, extract the original scene data as target scene data; Establish a scene timeline, arrange the target scene data in ascending order of difficulty level, and cluster them according to scene feature similarity in the same difficulty level; record multiple scene sub-records corresponding to the difficulty level of the daily communication category on the scene timeline, and obtain multiple scene record items; similarly, perform the expansion operation on the scene data corresponding to the difficulty level of the business in the workplace category, the academic education category and the emergency handling category to obtain their respective scene record items; Extracting scene record items satisfying a preset first screening condition, scene record items satisfying a preset second screening condition and scene record items satisfying a preset third screening condition in the scene record items respectively, as primary scene candidates, intermediate scene candidates and advanced scene candidates; For each scene type and each difficulty level, the following operations are performed: Extracting a primary scene scheme from the primary scene candidates corresponding to the scene type-difficulty level pair, and setting a first weight of the primary scene scheme; Extracting an intermediate scene scheme from the intermediate scene candidates corresponding to the scene type-difficulty level pair, and setting a second weight of the intermediate scene scheme; Extracting an advanced scene scheme from the advanced scene candidates corresponding to the scene type-difficulty level pair, and setting a third weight of the advanced scene scheme; wherein the first weight < the second weight < the third weight; Obtaining a preset scene quality evaluation model, inputting the primary scene scheme with the first weight, the intermediate scene scheme with the second weight and the advanced scene scheme with the third weight into the scene quality evaluation model, and obtaining an optimal scene scheme under the corresponding scene type-difficulty level pair; Obtaining a preset blank scene database, the blank scene database having a preset hierarchical storage structure of scene type-difficulty level pairs, forming independent storage slots; Inputting the optimal scene schemes corresponding to each scene type-difficulty level pair into the corresponding storage slots of the blank scene database according to the corresponding relationship of the scene type-difficulty level pairs; When the optimal scene schemes of all scene type-difficulty level pairs are inputted, the blank scene database is used as a practical scene library, and the construction is completed.

[0012] Preferably, the parsed original language data is compared with the sample materials in the standard expression benchmark library according to different evaluation dimensions, and the comparison results are comprehensively scored according to the evaluation threshold range, including: Comparing the parsed original language data with the sample materials in the standard expression benchmark library to determine the evaluation value of each parsed sample data in the original language data in each sub-dimension after being truncated by the evaluation threshold range; Obtaining the standard expression benchmark value of each sub-dimension based on the standard expression benchmark library; Taking the ratio of the evaluation value and the standard expression benchmark value of the corresponding dimension as the sample dimension relative ratio; , wherein represents the sample dimension relative ratio; represents the evaluation value of the tth parsed sample data in the original language data in the dth sub-dimension after being truncated by the evaluation threshold range; a standard expression benchmark value of the dth sub-dimension; a dimension evaluation value of each dimension in the original language data is calculated by weighting the cumulative value of the sample dimension relative ratio of all samples under each dimension in the original language data after weighting; The sum of the dimension evaluation values of all dimensions in the original language data is globally averaged as the normalized consistency score of the original language data; ; wherein The normalized consistency score of the original language data is represented by; The total number of parsed samples of the original language data is represented by; The total number of sub-dimensions of the evaluation system is represented by; The weight value of the dth sub-dimension is represented by; The absolute difference between the evaluation value and the standard expression benchmark value of the corresponding dimension is calculated to obtain the sample dimension deviation value; , wherein, The sample dimension deviation value is represented by; The sum of the deviation values of all sample data in the original language data is calculated as the deviation adjustment factor of the original language data; , wherein The deviation adjustment factor of the original language data is represented by; The ratio of the normalized consistency score of the original language data to the deviation adjustment factor is taken as the comprehensive score of the original language data.

[0013] A multi-language support device for a language communication auxiliary training instrument, comprising: A workbench is installed with a microphone, a speaker, a display screen and a printer, wherein the microphone receives voice input data of the user, the display screen is double-operated by touch screen and keyboard, the display screen receives text input data of the user, and the bottom end of the display screen is connected with the upper end of a rotating rod, the bottom end of the rotating rod is rotatably connected with the upper end of the workbench, the speaker broadcasts the voice simulated in the language communication scene, the printer prints the final visual report, and a bottom cabinet is installed under the workbench, and a host is installed in the bottom cabinet.

[0014] Compared with the prior art, the beneficial effects of the present application are as follows: 1. The multi-language support method and device for the language communication auxiliary training instrument provided by the present application receives voice and text data through double input ports, adapts to different user input habits, and carries out targeted preprocessing, voice data is subjected to noise reduction, segmentation and normalization to eliminate environmental interference and data differences, text data is subjected to invalid character filtering, format unification to remove redundant information, and can also extract features according to data type differences, greatly reducing subsequent recognition errors and improving data usability.

[0015] 2.The application provides a multi-language support method and device for a language communication auxiliary training instrument, a three-dimensional evaluation system covering pronunciation, grammar and semantics is constructed, a special standard expression benchmark library is established for each language and an evaluation threshold is set, three priority levels are divided according to the influence degree of problems on communication, and historical records can be called to strengthen repeated problem training. This design breaks the limitations of "multi-language sharing standards" and "problems without priorities", the evaluation is more objective and comprehensive, the correction scheme focuses on core problems, and at the same time, a training closed loop is formed through historical data linkage, avoiding the low-efficiency training of "one-size-fits-all", making the scheme more suitable for user weaknesses, helping to overcome language difficulties in a targeted manner, and steadily improving training efficiency.

[0016] 3.The application provides a multi-language support method and device for a language communication auxiliary training instrument, a scene library is constructed according to types and difficulties, an audio-visual interactive environment and multi-role dialogue logic are created, a microphone, a double-control display screen, a loudspeaker and a printer are integrated on the device side, and a main machine is integrated in a bottom cabinet. This not only makes the scene simulation close to real communication and improves the user's immersion, but also solves the problems of single input, inconvenient operation and difficult report storage of traditional devices. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 A multi-language support method for a language communication auxiliary training instrument of the application is shown in the figure; Figure 2 A structure diagram of a language communication auxiliary training instrument of the application is shown in the figure; In the figure: 1, workbench; 2, microphone; 3, loudspeaker; 4, display screen; 5, printer; 6, bottom cabinet; 7, main machine; 8, rotating rod. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the application will be described in detail below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.

[0019] In order to solve the problems in the prior art that language data processing is often single input, pre-processing is insufficient, language type judgment and analysis are not accurate in content recognition, the evaluation system is static and has no hierarchical benchmark, correction has no priority and is not combined with historical problems, resulting in poor training targeting and difficulty in accurately improving language ability, please refer to Figure 1 The embodiment provides the following technical solutions: A multi-language support method for a language communication auxiliary training instrument, comprising: The original language data of a user is received and processed first; the processed original language data is subjected to content recognition, and content analysis is performed according to the recognized content; a dynamic evaluation system is constructed after the original language data is analyzed, and a correction scheme is generated according to the constructed dynamic evaluation system; the user simulates a practical language communication scene according to the generated correction scheme; and the pronunciation or expression of the user is analyzed in multiple dimensions according to the simulation result.

[0020] Specifically, by automatically recognizing the content of the original language data, the operation threshold is reduced, and users of different ages and technical bases are adapted, especially for language learning beginners or the elderly, avoiding interrupting the training rhythm due to complex operations. The dynamic evaluation system constructed after language conversion can break free from the limitations of fixed evaluation standards, adjust the evaluation dimensions in real time according to the actual language level of the user, and generate correction schemes that are more personalized, such as focusing on phonetic symbol correction for users with pronunciation deviations and strengthening sentence structure guidance for users with rigid expressions, avoiding inefficient training of "one size fits all". Based on the correction scheme, practical language communication scene simulation is carried out, breaking the training dilemma of "paper tiger", allowing users to use language in a realistic scenario, strengthening language application ability, shortening the gap between "learning" and "using", and analyzing the simulation results in multiple dimensions, not only covering pronunciation accuracy and expression fluency, but also considering grammar correctness, vocabulary adaptability and other details, providing comprehensive feedback on user weaknesses, helping users accurately identify improvement directions, and providing data support for subsequent training optimization, continuously improving training effectiveness and efficiently helping users improve cross-language communication ability.

[0021] The original language data of a user is received and processed, including: The original language data of a user is received according to a data receiving port, and the original language data includes voice input data and text input data; The received original language data is subjected to data preprocessing, wherein the voice input data is subjected to noise reduction processing, voice segmentation and normalization processing; and the text input data is subjected to invalid character filtering and format unification.

[0022] The processed original language data is subjected to content recognition, and content analysis is performed according to the recognized content, including: The original language data after data preprocessing is subjected to feature extraction, wherein the voice input data is subjected to feature extraction according to phoneme distribution, syllable structure and prosody features; and the text input data is subjected to feature extraction according to character set features, vocabulary features and grammar features; The extracted feature data is subjected to language matching, which is similarity comparison of the extracted feature data with language templates in a feature library, and the language type of the extracted feature data is determined according to the comparison result; After the language type is confirmed, content recognition is performed, which is to convert the voice input data into text and perform semantic analysis; and to convert the text input data into voice and perform semantic analysis; After the analysis, the specific content of the user's original language data is obtained.

[0023] Specifically, the two types of input data, voice and text, are covered through multiple ports to adapt to different input habits of users, such as voice input for those who express fluently in spoken language and text input for those who need accurate information transmission, avoiding the limitations of a single input method, expanding the scope of user application, and pre-processing links processing two types of data: voice noise reduction, segmentation and normalization to eliminate environmental interference and data differences, text invalid character filtering, format unification to remove redundant information, effectively reducing subsequent recognition errors, improving data usability, and extracting features according to the characteristics of voice and text data: voice focusing on phonemes, syllables and other phonetic core dimensions, and text focusing on character sets and grammar and other linguistic key features, accurately capturing the essential properties of data to provide reliable basis for language matching. Instead of relying on a single identifier, it can handle different language feature differences, reduce misjudgment rate, and ensure the accuracy of language type judgment. Voice-to-text combined with semantic analysis and text supplemented with voice analysis not only realizes data format conversion but also excavates content meaning, providing accurate content support for subsequent language conversion and evaluation system construction, avoiding the impact of content recognition bias on training effectiveness.

[0024] After the original language data is analyzed, a dynamic evaluation system is constructed, and a correction scheme is generated according to the constructed dynamic evaluation system, including: First, confirm the evaluation dimensions of the dynamic evaluation system, including pronunciation dimension, grammar dimension and semantic dimension, wherein the pronunciation dimension includes phoneme accuracy, tone and intonation matching degree, stress position and rhythm; the grammar dimension includes word combination rationality, sentence structure specification, time and aspect accuracy and punctuation symbol usage correctness; the semantic dimension includes semantic integrity, semantic accuracy and scene adaptability; According to the evaluation dimensions, standard expression benchmark library is established for each language type, and at the same time, the evaluation threshold range is set; According to the evaluation dimensions and the standard expression benchmark library, the original language data after analysis is evaluated; Among them, according to different evaluation dimensions, the original language data after analysis is compared with the sample materials in the standard expression benchmark library, and the comparison results are comprehensively scored according to the evaluation threshold range.

[0025] According to the comprehensive evaluation results, the priority is sorted, wherein the priority sorting includes first priority, second priority and third priority, the first priority is the problem that seriously affects understanding and core grammar error; the second priority is the problem that does not affect understanding but does not conform to the standard; the third priority is the detail problem; According to different problems, correction measures are formulated, including pronunciation problem correction, grammar problem correction and semantic problem correction; At the same time, the user's historical evaluation record is called, and if there are repeated problems, targeted reinforcement training is carried out; Finally, the correction scheme is visualized to generate a report.

[0026] Specifically, not only the three core dimensions of pronunciation, grammar and semantics are covered, but also each dimension is further subdivided, such as the pronunciation dimension including key indicators such as tone and intonation matching degree, stress position, the grammar dimension being refined to details such as tense and aspect accuracy, punctuation usage, etc. It can capture user language problems in all directions and without dead angles, neither missing the core defects affecting communication nor overlooking the details restricting expression quality, avoiding the problem of misjudgment caused by single or extensive evaluation. A standard expression benchmark library is established for each language type, and an evaluation threshold range is set, breaking the limitation of "one set of standards for multiple languages". For example, the pronunciation rules and grammar structures of English and Japanese are significantly different, and the exclusive benchmark library can ensure that the evaluation standard fits the characteristics of different languages. Combined with the threshold quantization comparison result, it replaces subjective experience judgment, making the evaluation result more objective and accurate, providing a reliable basis for subsequent correction. According to the impact of problems on communication, it is divided into three levels of priority, guiding users to solve the first-level problems that seriously affect understanding first, then handle the second-level problems that do not affect understanding but do not meet the standard, and finally optimize the third-level details, avoiding users from "grabbing everything in one go" in improvement, focusing on key pain points, and greatly improving training efficiency. The user's historical evaluation record is called, and a reinforcement training scheme is developed for problems that repeatedly occur, forming a closed loop of "evaluation, correction, re-evaluation, and reinforcement". It can not only target the user's long-standing language weaknesses, but also avoid the recurrence of similar problems, helping users steadily improve their language skills. The final visual report clearly presents the problem type and corresponding correction measures, and users can quickly and clearly identify their weaknesses and improvement direction without professional interpretation. Whether it is the specific practice method for pronunciation correction or the rule explanation for grammar correction, it can be directly implemented, reducing the use threshold of the scheme and ensuring efficient correction training. Among them, the table of dynamic evaluation system is as follows: Core parameters Threshold range / parameter value Parameter description Pronunciation dimension - phoneme accuracy ≥ 80% is qualified, 60-79% is to be improved, < 60% is to be focused on correction Used to judge the degree of matching between a single phoneme in the user's pronunciation and a standard phoneme sample, such as the pronunciation similarity of English phonemes " / 0 / " and " / s / " Pronunciation dimension - tone and intonation matching degree ≥ 85% is qualified, 70-84% is to be improved, < 70% is to be focused on correction For tone languages such as Chinese, the degree of coincidence between the user's tone curve and the standard tone template is evaluated; for intonation languages such as English, the degree of coincidence of the rising intonation of a question sentence and the smooth intonation of a statement sentence is evaluated Grammar dimension - tense and voice accuracy ≥ 90% is qualified, 80-89% is to be improved, < 80% is to be focused on correction Statistically evaluate the correctness of the user's expression in terms of tense (such as English past tense and present tense) and voice (active voice and passive voice), such as "Yesterday I went" is correct, and "Yesterday I go" is incorrect Semantic dimension - scene adaptability ≥ 85% is qualified, 75-84% is to be improved, < 75% is to be focused on correction Judge whether the user's expression conforms to the current scene specification, such as using "Please reply as soon as possible" in a business scene is qualified, and using "Hurry up and reply to me" is unqualified In order to solve the problem in the prior art that language training scene simulation is often disconnected from correction schemes, the analysis dimension is single, there is a lack of historical data longitudinal comparison and common problem horizontal correlation, it is difficult to accurately track the improvement trajectory, and the training is not targeted and effective, please refer to Figure 1 The embodiment provides the following technical solutions: The user simulates practical language communication scenes according to the generated correction scheme, including: The communication scene of the practical scene library in the database is called; the scene corresponding to the problem in the correction scheme is adapted; The communication scene of the practical scene library includes daily communication, business in the workplace, academic education, and emergency handling. Each scene includes a primary, intermediate, and advanced difficulty level. The primary scene contains 5-8 basic expressions. The intermediate scene contains 10-15 expressions with logical connections. The advanced scene contains complex interactions. The problem in the correction scheme is adapted according to the communication scene. After the adaptation is completed, the practical language communication scene in the correction scheme is obtained. The practical language communication scene in the correction scheme is created in an interactive environment. The interactive environment includes a visual environment, an auditory environment, and interactive prompts. After the interactive environment is created, the interactive logic and dialogue roles are designed. According to the practical language communication scene in the correction scheme, 1-3 interactive roles are set, including preset language styles and response logic. The dialogue process is also set, including initial dialogue, core interaction, and advanced interaction. After the interactive logic and dialogue roles are designed, a complete voice communication scene is obtained. The user simulates the scene according to the complete voice communication scene. The simulation process is recorded in real time.

[0027] Specifically, the called practical scene library contains four high-frequency scenes: daily communication, business in the workplace, academic education, and emergency handling. It fully adapts to the real communication needs of users in their daily life, work, and study, avoiding the "training useless" problem caused by scenes that are not practical. It allows users to directly transfer the skills they learn in simulation to real-life communication. Each scene is divided into three difficulty levels: primary, intermediate, and advanced. The primary level focuses on basic expressions, the intermediate level emphasizes logical connections, and the advanced level emphasizes complex interactions. It can accurately match the user's language ability, avoiding frustration caused by too high a difficulty level and low efficiency caused by too low a difficulty level. Instead of randomly assigning scenes, it customizes scenes for specific problems in the correction scheme, allowing users to focus on improving their skills in simulation, enhancing the correction effect, and improving the training specificity. By creating a multi-dimensional interactive environment with visual, auditory, and interactive prompts, it breaks the monotony of traditional "pure text and voice practice" and allows users to feel immersed in the scene, making it easier for them to enter a real communication state. It sets 1-3 interactive roles with exclusive language styles and response logic, and pairs them with a complete process of "initial dialogue, core interaction, and advanced interaction" to restore the multi-role and progressive communication scenes in reality. It helps users improve their dialogue cohesion and expression skills, rather than mechanically memorizing them. During the simulation process, user performance is recorded in real time, providing data support for subsequent dynamic evaluation system updates and correction scheme optimization, forming a training closed loop of "simulation, feedback, and optimization."

[0028] According to the simulation results, the pronunciation or expression of the user is analyzed in multiple dimensions, including: The simulation process data recorded in real time during the scene simulation process is summarized and standardized; The simulation process data after summarization and standardization is analyzed in depth, wherein the depth analysis includes pronunciation dimension analysis, grammar dimension analysis and semantic dimension analysis, the pronunciation dimension analysis is basic pronunciation accuracy and scene adaptive pronunciation analysis; the scene adaptive pronunciation is basic grammar rule implementation and sentence structure and complexity analysis; the semantic dimension analysis is semantic transmission quality and scene adaptation and style uniformity analysis; The depth analysis result is combined with the historical simulation data of the user, and after combination, overall judgment is performed, wherein the overall judgment includes longitudinal comparison and horizontal correlation, the longitudinal comparison is to compare the depth analysis result with the historical simulation data, and after comparison, the improvement track of the key problems in the correction scheme is tracked; the horizontal correlation is to analyze the common problems repeatedly appearing in different scenes, and the relationship between the common problems and the scene difficulty is correlated; After combination, the final multi-dimensional analysis result is obtained, the multi-dimensional analysis result is used to generate a comprehensive conclusion, and the generated comprehensive conclusion is presented in a visual report.

[0029] Specifically, the simulation process data is summarized and standardized, which can eliminate data format differences and interference information in different scenes, ensure that the analysis data benchmark is unified and the quality meets the standard, avoid analysis deviation caused by data confusion, lay a precise data foundation for subsequent depth analysis, depth analysis covers three core dimensions of pronunciation, grammar and semantics, and each dimension is further refined: pronunciation not only focuses on basic accuracy, but also focuses on scene adaptability; grammar considers rule implementation and sentence structure complexity; semantics pays attention to transmission quality and scene style uniformity, which can capture the language shortcomings of the user in actual communication in all directions, avoid the one-sidedness of single dimension or surface analysis, make overall judgment combined with historical simulation data, longitudinal comparison can clearly track the improvement track of the key problems in the correction scheme, so that the user can directly see the progress, and it is also convenient to identify stubborn problems that have not been improved for a long time; horizontal correlation can excavate common problems in different scenes and correlate scene difficulty, accurately locate the problem source, convert multi-dimensional analysis results into comprehensive conclusions, and present them through visual reports, which can not only clearly summarize the advantages and points to be improved of the user, but also display analysis data in an intuitive form, without professional knowledge, the user can quickly understand, reduce the threshold of obtaining analysis information, and the analysis result can provide data basis for subsequent dynamic evaluation system update, correction scheme adjustment and scene simulation optimization, promote the continuous improvement of the training closed loop of "evaluation, correction, simulation, analysis and re-optimization", ensure that the language training always meets the user's needs, and improve the overall training effect.

[0030] The construction method of the practical scene library comprises: Obtain four types of original communication scene data of daily communication, business, academic education and emergency handling marked with difficulty levels; the difficulty levels are defined based on preset difficulty standards, and the preset difficulty standards include dialogue turns, vocabulary complexity, syntax structure complexity and cultural background depth; respectively extract features of the four types of original communication scene data to obtain daily features, business features, academic features and emergency features including difficulty features; Based on the preset feature type classification system, perform full connection cross-category matching on all four types of features: for any two types of features, extract feature pairs with the same feature type, calculate the matching degree value, and generate a common target feature type; wherein the target feature type is a standardized scene feature type in the preset feature reference library; Based on the preset feature reference library, determine the contribution coefficient corresponding to the target feature type and the matching degree value, and associate the contribution coefficient with the feature pairs involved in the matching result; Based on the preset weight, calculate the comprehensive contribution coefficient of the original scene data; when the comprehensive contribution coefficient is greater than or equal to the preset contribution coefficient threshold, extract the original scene data as target scene data; Establish a scene timeline, arrange the target scene data in ascending order of difficulty level, and cluster them according to scene feature similarity within the same difficulty level; spread the multiple scene sub-records corresponding to the difficulty level of the daily communication category on the scene timeline to obtain multiple scene record items; similarly, perform the spread operation on the scene data corresponding to the difficulty level of the business category, the academic education category and the emergency handling category to obtain their respective scene record items; In the scene record items, extract scene record items that meet the preset first screening condition, scene record items that meet the preset second screening condition, and scene record items that meet the preset third screening condition, which correspond to the primary scene candidate, the intermediate scene candidate and the advanced scene candidate; For each scene type and each difficulty level, perform the following operations: Extract the primary scene scheme from the primary scene candidate corresponding to the scene type-difficulty level, and set the first weight of the primary scene scheme; Extract the intermediate scene scheme from the intermediate scene candidate corresponding to the scene type-difficulty level, and set the second weight of the intermediate scene scheme; Extract the advanced scene scheme from the advanced scene candidate corresponding to the scene type-difficulty level, and set the third weight of the advanced scene scheme; wherein the first weight < the second weight < the third weight; Obtain a preset scene quality evaluation model, input the primary scene scheme with the first weight, the intermediate scene scheme with the second weight, and the advanced scene scheme with the third weight into the scene quality evaluation model, and obtain the optimal scene scheme corresponding to the scene type-difficulty level; Obtain a preset blank scene database, which has a hierarchical storage structure of preset scene type-difficulty levels, and forms independent storage slots; Input the optimal scene scheme corresponding to each scene type-difficulty level into the corresponding storage slot of the blank scene database according to the corresponding relationship between the scene type-difficulty level; After the optimal scene schemes of all scene type-difficulty level combinations are input, the blank scene database is used as a practical scene library, and the construction is completed.

[0031] In this embodiment, the daily features include expression sentence number features, spoken language naturalness features, and interactive simplicity features; the job features include business expression professionalism features, etiquette standardization features, and logic rigor features; the academic features include subject expression accuracy features, thinking reasoning features, and knowledge level features; and the emergency features include information transmission clarity features, disposal orderliness features, and interactive strategy features.

[0032] In this embodiment, the comprehensive contribution coefficient = Σ (each feature contribution coefficient × feature weight), and the feature weight is preset according to the importance of the scene type.

[0033] In this embodiment, the practical scene library supports subsequent user language level correction schemes, and realizes accurate retrieval and adaptation of communication scenes.

[0034] In this embodiment, the scene record items meeting the preset first screening condition, the preset second screening condition, and the preset third screening condition are extracted from the scene record items, and correspond to the primary scene candidate, the intermediate scene candidate, and the advanced scene candidate, including: The scene record items meeting the preset first screening condition are extracted from the scene record items as the primary scene candidate; the first screening condition includes that the scene record item is an independent scene, and the corresponding difficulty feature meets the primary difficulty requirement; The scene record items meeting the preset second screening condition are extracted from the scene record items as the intermediate scene candidate; the second screening condition includes that the scene record item has moderate dependency, and the corresponding difficulty feature meets the intermediate difficulty requirement; The scene record items meeting the preset third screening condition are extracted from the scene record items as the advanced scene candidate; the third screening condition includes that the scene record item has multi-round complex interaction features, and the corresponding difficulty feature meets the advanced difficulty requirement.

[0035] The working principle and beneficial effects of the above technical solution are: by obtaining four types of original communication scene data labeled with difficulty levels in daily communication, business communication, academic education, and emergency handling, different fields of communication scenes can be fully covered to ensure the comprehensiveness and practicality of the scene library and meet diversified actual needs; feature extraction is performed on the four types of original communication scene data, not only considering the difficulty characteristics, but also comprehensively considering multi-dimensional factors such as dialogue turns, vocabulary complexity, syntax structure complexity, and cultural background depth, which can accurately depict the characteristics of each scene and provide rich and accurate information for subsequent analysis; based on the preset feature type classification system, full connection cross-category matching is performed, which can discover potential links and common features between different scene types, help to break down the barriers between scene categories, realize knowledge migration and fusion, and improve the universality of the scene library; by calculating the comprehensive contribution coefficient of the original scene data and comparing it with the preset threshold, target scene data is extracted, and this screening method can ensure that the scene data entering the scene library has high quality and representativeness, avoiding the interference of useless or low-value data; a scene timeline is established, arranged in ascending order of difficulty level and clustered, and different filtering conditions of primary, intermediate, and advanced are set, which can finely divide scenes of different difficulty levels to meet the needs of different users in different learning or application stages; different weights are set for primary, intermediate, and advanced scene schemes, and the first weight < the second weight < the third weight, which can highlight the importance of advanced scene schemes in scene quality evaluation, guide the scene library to optimize to more complex and advanced scenes, and improve the overall quality of the scene library; the preset scene quality evaluation model is used to evaluate scene schemes with different weights, which can objectively and scientifically select the optimal scene scheme under the corresponding scene type - difficulty level, and ensure that each scene type and difficulty level in the scene library has a high-quality representative scene.

[0036] The parsed original language data is compared with the sample materials in the standard expression benchmark library according to different evaluation dimensions, and the comparison results are comprehensively scored according to the evaluation threshold range, including: The parsed original language data is compared with the sample materials in the standard expression benchmark library to determine the evaluation value of each parsed sample data in the original language data in each sub-dimension after being truncated by the evaluation threshold range; The standard expression benchmark value of each sub-dimension is obtained based on the standard expression benchmark library; The ratio of the evaluation value to the standard expression benchmark value of the corresponding dimension is taken as the sample dimension relative ratio; , wherein represents the sample dimension relative ratio; represents the evaluation value of the tth parsed sample data in the original language data in the dth sub-dimension after being truncated by the evaluation threshold range; a standard expression reference value of the dth sub-dimension; calculating the cumulative value of the sample dimension relative ratio of all samples in each dimension of the original language data after being weighted by the weight, as the dimension evaluation value of each dimension of the original language data; globally averaging the sum of the dimension evaluation values of all dimensions in the original language data as the normalized consistency score of the original language data; ; wherein the normalized consistency score of the original language data; the total number of parsed sample data of the original language data; the total number of sub-dimensions of the evaluation system; the weight value of the dth sub-dimension; calculating the absolute difference between the evaluation value and the standard expression reference value of the corresponding dimension to obtain the sample dimension deviation value; , wherein, the sample dimension deviation value; calculating the sum of the deviation values of all sample data in the original language data as the deviation adjustment factor of the original language data; , wherein the deviation adjustment factor of the original language data; the ratio of the normalized consistency score of the original language data to the deviation adjustment factor as the comprehensive score of the original language data.

[0037] In this embodiment, each parsed sample data can be a parsed sentence or a parsed fragment.

[0038] In this embodiment, the evaluation value of each parsed sample data in the original language data in each sub-dimension can be generated by an NLP model or expert scoring.

[0039] In this embodiment, the comprehensive score is: , wherein, the comprehensive score of the original language data; the total number of parsed sample data of the original language data; the total number of sub-dimensions of the evaluation system; the weight value of the dth sub-dimension; the evaluation value of the tth parsed sample data in the original language data in the dth sub-dimension after being truncated by the threshold range; the standard expression reference value of the dth sub-dimension.

[0040] The working principle and beneficial effects of the technical solution are as follows: by comparing the parsed original language data with the standard expression benchmark library, each parsed sample data is evaluated in each sub-dimension, which can comprehensively analyze the original language data from multiple angles; different sub-dimensions can represent different language characteristics or evaluation indexes, such as grammatical accuracy, semantic coherence, and vocabulary richness, thereby avoiding the limitations of single-dimension evaluation and more accurately reflecting the overall quality of the original language data; the deviation adjustment factor reflects the deviation of the original language data from the standard expression benchmark value in each sub-dimension. By taking the ratio of the normalized consistency score and the deviation adjustment factor as the comprehensive score, the influence of consistency and deviation can be balanced, and the accuracy of the original language data evaluation can be improved.

[0041] In order to solve the problems of single input mode, fixed display screen angle, inconvenient operation, lack of scene simulation voice broadcast, training report cannot be printed immediately, and low device integration of the language communication auxiliary training device in the prior art, please refer to Figure 1 and 2 The embodiment provides the following technical solutions: A multi-language support device for a language communication auxiliary training instrument, comprising: A workbench 1, a microphone 2, a loudspeaker 3, a display screen 4, and a printer 5 are installed on the workbench 1, wherein the microphone 2 receives voice input data of a user, the display screen 4 is dual-operated by a touch screen and a keyboard, the display screen 4 receives text input data of the user, and the bottom end of the display screen 4 is connected with the upper end of a rotating rod 8, the bottom end of the rotating rod 8 is rotatably connected with the upper end of the workbench 1, the loudspeaker 3 broadcasts voice of a language communication scene simulation, the printer 5 prints a final visual report, a bottom cabinet 6 is installed below the workbench 1, and a host 7 is installed in the bottom cabinet 6.

[0042] Specifically, the microphone 2 accurately receives user voice input data, meeting the voice type training scene; the display screen 4 adopts dual control of touch screen and keyboard, which not only adapts to fast text input, but also supports touch screen convenient operation, covering different user input habits; the loudspeaker 3 plays the scene simulation voice, restores the real communication auditory environment, and helps the user to immerse in the training; the printer 5 can print the final visual report, which is convenient for the user to retain the paper version, check the correction scheme and analysis result at any time, avoids the problem of easy loss of electronic report, and the core functional components are installed in the workbench 1, so that the user does not need to switch between multiple devices, from voice, text input, scene simulation listening to report printing, which can be completed in the same operation area, greatly simplifying the use process, reducing the operation complexity, improving the training efficiency, the bottom cabinet 6 under the workbench 1 not only provides installation space for the host 7, avoids the host 7 from being exposed to the outside and being affected by collision and dust, prolongs the service life of the device, but also allows the workbench 1 surface to only retain the core interactive components, keeps the table top clean, reduces the interference of clutter layout on user operation, and deeply matches the device components and the whole training process, from the original data collection of voice reception and text input, to the voice broadcast of scene simulation, to the result retention of report printing, forming a complete device support chain of "collection, training, feedback and retention", ensuring efficient landing of each link of multi-language training, and providing the user with continuous and stable training hardware support.

[0043] It should be noted that, in this document, the terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that these entities or operations have any such actual relationship or order. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device.

[0044] Although the embodiments of the present application have been shown and described, it can be understood by those of ordinary skill in the art that various changes, modifications, replacements and variations can be made to the embodiments without departing from the principles and spirits of the present application.

Claims

1. A multi-language support method for a language communication aid training device, characterized by, The application comprises the following steps: First, the user's original language data is received and processed; The processed original language data is content-identified, and content analysis is performed according to the identified content; after the original language data is analyzed, a dynamic evaluation system is constructed, and a correction scheme is generated according to the constructed dynamic evaluation system; the user simulates a practical language communication scene according to the generated correction scheme; the user's pronunciation or expression is analyzed in multiple dimensions according to the simulation results.

2. The multi-language support method for the language communication aid training apparatus according to claim 1, wherein, The user's original language data is received and processed, including: The user's original language data is received according to the data receiving port, and the original language data includes voice input data and text input data; The received original language data is pre-processed, wherein the voice input data is pre-processed by noise reduction, voice segmentation and normalization; the text input data is pre-processed by invalid character filtering and format unification.

3. The multi-language support method for the language communication aid training apparatus according to claim 2, wherein, The processed original language data is content-identified, and content analysis is performed according to the identified content, including: The original language data after data preprocessing is feature-extracted, wherein the voice input data is feature-extracted according to phoneme distribution, syllable structure and prosodic features; the text input data is feature-extracted according to character set features, lexical features and grammatical features; The extracted feature data is language-matched, which is a similarity comparison between the extracted feature data and the language templates in the feature library, and the language type of the extracted feature data is determined according to the comparison result; After the language type is confirmed, content identification is performed, which includes voice-to-text conversion and semantic analysis of voice input data; and voice analysis of text input data; The specific content of the user's original language data is obtained after analysis.

4. The multi-language support method for the language communication aid training apparatus according to claim 3, wherein After the original language data is analyzed, a dynamic evaluation system is constructed, and a correction scheme is generated according to the constructed dynamic evaluation system, including: First, the evaluation dimensions of the dynamic evaluation system are confirmed, including pronunciation dimension, grammar dimension and semantic dimension, wherein the pronunciation dimension includes phoneme accuracy, tone and intonation matching degree, stress position and rhythm; the grammar dimension includes word combination rationality, sentence structure specification, time and aspect accuracy and punctuation symbol use correctness; the semantic dimension includes semantic integrity, semantic accuracy and scene adaptability; According to the evaluation dimensions, standard expression benchmark library is established for each language type, and the evaluation threshold range is set; According to the evaluation dimensions and the standard expression benchmark library, the analyzed original language data is evaluated; Among them, according to different evaluation dimensions, the analyzed original language data is compared with the sample materials in the standard expression benchmark library, and the comparison results are scored according to the evaluation threshold range.

5. The multi-language support method for the language communication aid training apparatus according to claim 4, wherein, After the original language data is analyzed, a dynamic evaluation system is constructed, and a correction scheme is generated according to the constructed dynamic evaluation system, which also includes: According to the comprehensive evaluation results, the priority is sorted, wherein the priority sorting includes first priority, second priority and third priority, the first priority is the problem of seriously affecting understanding and core grammar error; the second priority is the problem of not affecting understanding but not conforming to the standard; the third priority is the detail problem; According to different problems, correction measures are formulated, including pronunciation correction, grammar correction and semantic correction; At the same time, the user's historical evaluation record is called, and if there are repeated problems, targeted training is carried out; Finally, the correction scheme is visualized to generate a report, and the user's correction scheme is obtained.

6. The multi-language support method for the language communication aid training apparatus according to claim 5, wherein, The user simulates the practical language communication scene according to the generated correction scheme, including: The communication scene of the practical scene library in the database is called; the scene corresponding to the problem in the correction scheme is adapted; The communication scene of the practical scene library includes daily communication, business, academic education and emergency handling, each scene includes primary, intermediate and advanced difficulty levels, and the primary scene contains 5-8 basic expressions; the intermediate scene contains 10-15 expressions with logical association; the advanced scene contains complex interaction; The problem in the correction scheme is adapted according to the communication scene, and the practical language communication scene in the correction scheme is obtained after the adaptation is completed; The practical language communication scene in the correction scheme is created in an interactive environment, which includes visual environment, auditory environment and interactive prompt; After the interactive environment is created, the interactive logic and dialogue role are designed, wherein 1-3 interactive roles are set according to the practical language communication scene in the correction scheme, the interactive roles include preset language style and response logic, and the dialogue process is set, including initial dialogue, core interaction and advanced interaction; The complete voice communication scene is obtained after the interactive logic and dialogue role are designed; The user simulates the scene according to the complete voice communication scene, and records the simulation process in real time during the simulation process.

7. The multi-language support method for the language communication aid training apparatus according to claim 6, wherein, According to the simulation results, the pronunciation or expression of the user is analyzed in multiple dimensions, including: The simulation process data recorded in real time during the simulation process is summarized and standardized; The simulation process data after summarizing and standardizing is analyzed in depth, wherein the depth analysis includes pronunciation dimension analysis, grammar dimension analysis and semantic dimension analysis, the pronunciation dimension analysis is basic pronunciation accuracy and scene adaptability pronunciation analysis; the scene adaptability pronunciation is the implementation of basic grammar rules and sentence structure and complexity analysis; the semantic dimension analysis is semantic transmission quality and scene adaptation and style uniformity analysis; The depth analysis result is combined with the historical simulation data of the user, and after the combination, the overall judgment is carried out, including longitudinal comparison and horizontal correlation, the longitudinal comparison is to compare the depth analysis result with the historical simulation data, and after the comparison, the improvement track of the key problem in the correction scheme is tracked; the horizontal correlation is to analyze the common problems repeatedly appearing in different scenes, and the relationship between the common problems and the scene difficulty is associated; The final multi-dimensional analysis result is obtained after combination, the multi-dimensional analysis result is synthesized to generate a conclusion, and the generated comprehensive conclusion is visualized to present a report.

8. The multi-language support method for the language communication aid training apparatus according to claim 6, wherein, The construction method of the practical scene library, comprising: Obtain four types of original communication scene data of daily communication, business, academic education and emergency treatment marked with difficulty levels; the difficulty levels are defined based on preset difficulty standards, and the preset difficulty standards include dialogue turns, vocabulary complexity, syntax structure complexity and cultural background depth; extract features of the four types of original communication scene data respectively to obtain daily features, business features, academic features and emergency features including difficulty features; Based on a preset feature type classification system, perform full connection cross-category matching on all four types of features: for any two types of features, extract feature pairs with the same feature type, calculate a matching degree value, and generate a common target feature type; wherein the target feature type is a standardized scene feature type in a preset feature reference library; Based on the preset feature reference library, determine a contribution coefficient corresponding to the target feature type and the matching degree value, and associate the contribution coefficient with the feature pairs involved in the matching result; Based on a preset weight, calculate a comprehensive contribution coefficient of the original scene data; when the comprehensive contribution coefficient is greater than or equal to a preset contribution coefficient threshold, extract the original scene data as target scene data; Establish a scene time axis, arrange the target scene data in ascending order of difficulty levels, and cluster them according to scene feature similarity within the same difficulty level; expand multiple scene sub-records corresponding to the difficulty levels of the daily communication type on the scene time axis to obtain multiple scene record items; similarly, perform expansion operations on the scene data corresponding to the difficulty levels of the business, academic education and emergency treatment types to obtain respective scene record items; In the scene record items, extract scene record items that meet a preset first screening condition, scene record items that meet a preset second screening condition and scene record items that meet a preset third screening condition, which correspond to primary scene candidates, intermediate scene candidates and advanced scene candidates, respectively; For each scene type and each difficulty level, perform the following operations: Extract a primary scene scheme from the primary scene candidates corresponding to the scene type and the difficulty level, and set a first weight for the primary scene scheme; Extract an intermediate scene scheme from the intermediate scene candidates corresponding to the scene type and the difficulty level, and set a second weight for the intermediate scene scheme; Extract an advanced scene scheme from the advanced scene candidates corresponding to the scene type and the difficulty level, and set a third weight for the advanced scene scheme; wherein the first weight < the second weight < the third weight; Obtain a preset scene quality evaluation model, input the primary scene scheme with the first weight, the intermediate scene scheme with the second weight and the advanced scene scheme with the third weight into the scene quality evaluation model, and obtain an optimal scene scheme under the corresponding scene type and difficulty level; Obtain a preset blank scene database, which has a hierarchical storage structure of preset scene types and difficulty levels, forming independent storage slots; Input the optimal scene schemes corresponding to each scene type and difficulty level into the corresponding storage slots of the blank scene database according to the corresponding relationship between the scene type and the difficulty level. When the optimal scene scheme of all scene type-difficulty level combinations is completed, the blank scene database is taken as a practical scene library, and the construction is completed.

9. The multi-language support method for the language communication aid training apparatus according to claim 4, wherein, The parsed original language data is compared with the sample materials in the standard expression benchmark library according to different evaluation dimensions, and the comparison results are comprehensively scored according to the evaluation threshold range, including: The parsed original language data is compared with the sample materials in the standard expression benchmark library, and the evaluation value of each parsed sample data in the original language data is determined in each sub-dimension after being truncated by the evaluation threshold range; The standard expression benchmark value of each sub-dimension is obtained based on the standard expression benchmark library; a ratio of the evaluation value and a standard expression benchmark value of the corresponding dimension as a sample dimension relative ratio; wherein represents a sample dimension relative ratio; represents an evaluation value of the tth parsed sample data in the original language data in the dth sub-dimension after being truncated by the evaluation threshold range; represents a standard expression benchmark value of the dth sub-dimension; The cumulative value of the sample dimension relative comparison value of all samples in each dimension in the original language data after being weighted by the weight is calculated as the dimension evaluation value of each dimension in the original language data. The sum of the dimension evaluation values of all dimensions in the original language data is globally averaged as a normalized consistency score of the original language data. ; wherein represents the normalized consistency score of the original language data. represents the total number of parsed samples of the original language data. represents the total number of sub-dimensions of the evaluation system. represents the weight value of the dth sub-dimension. calculating an absolute difference between the evaluation value and a standard expression benchmark value of a corresponding dimension, to obtain a sample dimension deviation value; , wherein represents a sample dimension deviation value; a sum of bias values of all sample data in the original language data is calculated as a bias adjustment factor of the original language data; wherein represents the bias adjustment factor of the original language data; The ratio of the normalized consistency score of the original language data and the deviation adjustment factor is taken as the comprehensive score of the original language data.

10. A multi-language support device for a language communication aid training apparatus, which is applied in the multi-language support method for a language communication aid training apparatus according to claim 7, characterized by, It comprises a workbench (1), a microphone (2), a loudspeaker (3), a display screen (4) and a printer (5) installed on the workbench (1), wherein the microphone (2) receives voice input data of a user, the display screen (4) is double-operated by a touch screen and a keyboard, the display screen (4) receives text input data of the user, and the bottom end of the display screen (4) is connected with the upper end of a rotating rod (8), the bottom end of the rotating rod (8) is rotatably connected with the upper end of the workbench (1), the loudspeaker (3) broadcasts voice of a language communication scene simulation, the printer (5) prints a final visual report, and a bottom cabinet (6) is installed below the workbench (1), and a host (7) is installed in the bottom cabinet (6). ​