Intelligent translation processing method and system based on multi-language learning platform

By breaking down and analyzing the audio features of mixed-language audio data, identifying language and spoken language preferences, and combining this with intelligent translation devices to generate accurate translation audio data, the problem of translation errors in different languages ​​and dialects has been solved, improving the accuracy of multilingual translation and user experience.

CN121365671APending Publication Date: 2026-01-20SUZHOU UNIV OF SCI & TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511522074.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing translation equipment is prone to translation errors when dealing with different languages ​​and dialects, resulting in low accuracy in multilingual translation.

Method used

The audio data of mixed languages ​​is split into sub-audio data corresponding to each audio source through an audio segmentation strategy. The audio features of each sub-audio data are extracted, the language type and spoken language preference information are identified, and the translated audio data is generated using an intelligent audio translation device. The audio is then proofread and fused using spoken language preference information.

Benefits of technology

It improves the comprehensiveness of analysis of multilingual audio content, enhances translation efficiency and accuracy, and enables users to accurately identify details such as emotion, tone, and voice of the audio source, as well as locate the position of surrounding sound sources in real time, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365671A_ABST
    Figure CN121365671A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent translation processing method and system based on a multi-language learning platform, and the method comprises the steps: obtaining audio data of a mixed language, and splitting the audio data into sub-audio data corresponding to each sound source; in each piece of sub-audio data, extracting audio characteristics of a sampling audio segment in each piece of sub-audio data, and based on the audio characteristics of each sampling audio segment, identifying a language type corresponding to each piece of sub-audio data and spoken language preference information corresponding to each piece of sub-audio data through a language characteristic analysis strategy; and based on the language type corresponding to each piece of sub-audio data, generating translated audio data corresponding to each piece of sub-audio data through the intelligent audio translation equipment, and performing audio proofreading fusion processing on the translated audio data corresponding to each piece of sub-audio data through the spoken language preference information corresponding to each piece of sub-audio data, thereby obtaining the spoken language preference information corresponding to each piece of sub-audio data. And obtaining the translated audio data of the mixed language. By adopting the scheme, the accuracy of simultaneous translation of multiple languages can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and language translation technology, and in particular to a multi-language learning platform-based translation intelligent processing method and system. BACKGROUND

[0003] With the development of intelligent translation technology, translation technology gradually shifts from traditional single-language translation to multi-language translation technology. However, with the development of translation technology, the requirement for translation accuracy is even higher. Especially when different languages and different dialects need to be translated at the same time, the translation device often fails, makes translation errors, or even translation abnormalities. Therefore, how to accurately translate under multi-language conditions is the current research focus.

[0004] The traditional translation method is to translate a single language through a translation device or to switch languages. However, when different languages and different dialects need to be translated, the translation device often makes translation errors, affecting the translation effect and resulting in low accuracy of simultaneous translation of multiple languages. SUMMARY

[0006] The main purpose of the present application is to provide a multi-language learning platform-based translation intelligent processing method and system, which aims to solve the problem of low accuracy of simultaneous translation of multiple languages due to translation errors of translation devices when different languages and different dialects need to be translated.

[0007] To achieve the above purpose, the present application provides a multi-language learning platform-based translation intelligent processing method, which comprises: Obtaining audio data of a mixed language, and splitting the audio data into sub-audio data corresponding to each audio source through an audio splitting strategy; In each of the sub-audio data, extracting the audio features of the sampled audio segments in each sub-audio data, and based on the audio features of each of the sampled audio segments, identifying the language type corresponding to each sub-audio data and the spoken language preference information corresponding to each sub-audio data through a language characteristic analysis strategy; Based on the language type corresponding to each sub-audio data, generating the translation audio data corresponding to each sub-audio data through an intelligent audio translation device, and performing audio proofreading fusion processing on the translation audio data corresponding to each sub-audio data through the spoken language preference information corresponding to each sub-audio data to obtain the translation audio data of the mixed language.

[0008] Optionally, the splitting of the audio data into sub-audio data corresponding to each audio source through the audio splitting strategy comprises: splitting the audio data into initial sub-audio data by an audio splitting program; identifying an audio source direction of each initial sub-audio data, and pre-processing each initial sub-audio data to obtain sub-audio data; marking the audio source direction of each sub-audio data as a corresponding sound source of each sub-audio data to obtain sub-audio data corresponding to each sound source.

[0009] Optionally, the audio features of the sampled audio segments in each sub-audio data are extracted from the sub-audio data, including: For each sub-audio data, a sampling audio segment is collected in the sub-audio data according to a preset sampling strategy; identifying sound features of each sound feature type corresponding to the sampled audio segment, and extracting audio variation features corresponding to the sampled audio segment and audio distribution features corresponding to the sampled audio segment through an audio feature extraction network; The sound features of each sound feature type, the audio variation features corresponding to the sampled audio segment, and the audio distribution features corresponding to the sampled audio segment are used as the audio features of the sampled audio segment.

[0010] Optionally, the audio features of each sampled audio segment are used to identify the language type corresponding to each sub-audio data and the spoken language preference information corresponding to each sub-audio data through a language characteristic analysis strategy, including: For each sampled audio segment, a text translation program of each language type is used to identify translated texts of each language type corresponding to the sampled audio segment, and a semantic logic detection program is used to identify a logical reasonableness value corresponding to each translated text; The language type corresponding to the translated text with the maximum logical reasonableness value is selected as the language type of the sub-audio data corresponding to the sampled audio segment; Based on the sound features of each sound feature type, the audio variation features corresponding to the sampled audio segment, and the audio distribution features corresponding to the sampled audio segment, a user pronunciation preference of the sampled audio segment and a user spoken language preference of the sampled audio segment are identified through an audio preference analysis network; The user pronunciation preference of the sampled audio segment and the user spoken language preference of the sampled audio segment are used as the spoken language preference information corresponding to the sub-audio data.

[0011] Optionally, the translated audio data corresponding to each sub-audio data is generated through an intelligent audio translation device based on the language type corresponding to each sub-audio data, including: For each sub-audio data, a translation association mark between a sound source corresponding to the sub-audio data and a language type corresponding to the sub-audio data is constructed; Based on a user pronunciation preference of the sub-audio data, a translation program of the language type corresponding to the sub-audio data in the intelligent audio translation device is adjusted to obtain a target translation program corresponding to the sub-audio data; Based on each sub-audio data and the translation association mark, a translation text data corresponding to the sub-audio data is generated through the target translation program corresponding to the sub-audio data, and the translation text data is subjected to data conversion processing to obtain a translation audio data corresponding to the translation text data.

[0012] Optionally, the translation audio data corresponding to each sub-audio data is subjected to audio proofreading fusion processing through the spoken language preference information corresponding to each sub-audio data to obtain the translation audio data of the mixed language, including: For each sub-audio data, a sound preference feature of each sound type corresponding to the sub-audio data is identified based on the spoken language preference information corresponding to the sub-audio data; Based on the sound preference feature of each sound type, the target translation audio data corresponding to the sub-audio data is obtained through an audio adjustment network based on the sound preference feature of each sound type; Based on the audio source direction corresponding to each sub-audio data, the target translation audio data corresponding to each sub-audio data is subjected to data fusion processing through an audio mixing program to obtain the translation audio data of the mixed language.

[0013] In addition, to achieve the above-mentioned purpose, the present application also provides a multi-language learning platform-based intelligent translation processing system, which comprises: An acquisition module is configured to acquire audio data of a mixed language, and split the audio data into sub-audio data corresponding to each sound source through an audio splitting strategy; An identification module is configured to extract an audio feature of a sampled audio segment in each sub-audio data in each sub-audio data, and identify a language type corresponding to each sub-audio data and spoken language preference information corresponding to each sub-audio data through a language characteristic analysis strategy based on the audio feature of each sampled audio segment; A fusion module is configured to generate translation audio data corresponding to each sub-audio data through an intelligent audio translation device based on the language type corresponding to each sub-audio data, and perform audio proofreading fusion processing on the translation audio data corresponding to each sub-audio data through the spoken language preference information corresponding to each sub-audio data to obtain the translation audio data of the mixed language.

[0014] Optionally, the acquisition module is specifically configured to: split the audio data into initial sub-audio data through an audio splitting program; identify an audio source direction of each initial sub-audio data, and pre-process each initial sub-audio data to obtain sub-audio data; mark the audio source direction of each sub-audio data as a corresponding sound source of the sub-audio data, and obtain sub-audio data corresponding to the sound source.

[0015] Optionally, the identification module is specifically configured to: for each sub-audio data, collect a sampling audio segment in the sub-audio data according to a preset sampling strategy; identify a sound feature corresponding to each sound feature type of the sampling audio segment, and extract an audio variation feature corresponding to the sampling audio segment and an audio distribution feature corresponding to the sampling audio segment through an audio feature extraction network; use the sound features of each sound feature type, the audio variation feature corresponding to the sampling audio segment, and the audio distribution feature corresponding to the sampling audio segment as audio features of the sampling audio segment.

[0016] Optionally, the identification module is specifically configured to: for each sampling audio segment, identify translated texts of each language type corresponding to the sampling audio segment through a text translation program of each language type, and identify a logical reasonableness value corresponding to each translated text through a semantic logic detection program; select a language type corresponding to a translated text with a maximum logical reasonableness value as a language type of sub-audio data corresponding to the sampling audio segment; based on the sound features of each sound feature type, the audio variation feature corresponding to the sampling audio segment, and the audio distribution feature corresponding to the sampling audio segment, identify a user pronunciation preference of the sampling audio segment and a user spoken language preference of the sampling audio segment through an audio preference analysis network; use the user pronunciation preference of the sampling audio segment and the user spoken language preference of the sampling audio segment as spoken language preference information corresponding to the sub-audio data.

[0017] Optionally, the fusion module is specifically configured to: for each sub-audio data, construct a translation association mark between a corresponding sound source of the sub-audio data and a corresponding language type of the sub-audio data; based on a user pronunciation preference of the sub-audio data, adjust a translation program of the language type corresponding to the sub-audio data in an intelligent audio translation device to obtain a target translation program corresponding to the sub-audio data. Based on each of the sub-audio data and the translation association mark, the target translation program corresponding to the sub-audio data is used to generate the translation text data corresponding to the sub-audio data, and the translation text data is subjected to data conversion processing to obtain the translation audio data corresponding to the translation text data.

[0018] Optionally, the fusion module is specifically used for: For each sub-audio data, based on the spoken language preference information corresponding to the sub-audio data, the sound preference features of each sound type corresponding to the sub-audio data are identified; Based on the sound preference features of each of the sound types, the sub-audio data is adjusted through an audio adjustment network to obtain the target translation audio data corresponding to the sub-audio data; Based on the audio source direction corresponding to each of the sub-audio data, the target translation audio data corresponding to each of the sub-audio data is subjected to data fusion processing through an audio mixing program to obtain the translation audio data of the mixed language.

[0019] In a third aspect, the present application provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the intelligent translation processing method based on the multi-language learning platform when executing the computer program.

[0020] In a fourth aspect, the present application provides a computer readable storage medium. A computer program is stored thereon, and the computer program implements the steps of the intelligent translation processing method based on the multi-language learning platform when executed by a processor.

[0021] In a fifth aspect, the present application provides a computer program product. The computer program product comprises a computer program, and the computer program implements the steps of the intelligent translation processing method based on the multi-language learning platform when executed by a processor.

[0022] The application provides a multi-language learning platform-based intelligent translation processing method and system. The method comprises the following steps: acquiring mixed language audio data, and splitting the audio data into sub-audio data corresponding to each audio source through an audio splitting strategy; in each sub-audio data, extracting the audio features of a sample audio segment in each sub-audio data, and based on the audio features of each sample audio segment, identifying the language type corresponding to each sub-audio data and the spoken language preference information corresponding to each sub-audio data through a language characteristic analysis strategy; based on the language type corresponding to each sub-audio data, generating the translation audio data corresponding to each sub-audio data through an intelligent audio translation device, and performing audio proofreading fusion processing on the translation audio data corresponding to each sub-audio data through the spoken language preference information corresponding to each sub-audio data to obtain the translation audio data of the mixed language. Through the data splitting of the mixed language audio data, the audio feature analysis of different audio sources, the language type corresponding to each sub-audio data and the spoken language preference information are obtained, thereby effectively improving the analysis comprehensiveness of each audio source in multi-language audio content. Then, in order to improve the translation efficiency, the sub-audio data needs to be sampled and analyzed to obtain the language type and the spoken language preference information, thereby improving the analysis efficiency of each sub-audio data. Then, after the audio translation, the audio is corrected in combination with the spoken language preference information, so that the user can not only obtain the translated audio, but also accurately identify the audio source corresponding to each audio, thereby improving the recognition effect of the user on the emotional, tone, voice and other detailed information of the audio source. Finally, the audio data is fused, so that the user can locate the correct position of the surrounding audio source and the current multi-audio source environment according to the audio translation content in real time, thereby improving the user experience effect and the accuracy of multi-language simultaneous translation. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the schemes in the application, the drawings needed in the description of the embodiments of the application will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0025] Figure 1 is a flowchart of the multi-language learning platform-based intelligent translation processing method provided by the embodiments of the application; Figure 2 is a structural schematic diagram of the multi-language learning platform-based intelligent translation processing system provided by the embodiments of the application; Figure 3The internal structure diagram of the computer device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] The intelligent processing method for translation based on the multi-language learning platform provided by the embodiment of the present application is applied to an intelligent processing system for translation based on the multi-language learning platform. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terms used in the specification of the application are only for the purpose of describing specific embodiments and are not intended to limit the application. The terms "include" and "have" and any variations thereof in the specification and claims of the application and the above description of the drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of the application or the above description of the drawings are used to distinguish different objects and are not used to describe a specific order.

[0028] In this document, reference to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily mutually exclusive of other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0029] In order for those skilled in the art to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings.

[0030] The method provided by the embodiment of the application can be applied to an intelligent translation processing application environment based on a multi-language learning platform. The method can be applied to a terminal, a server, or a system including a terminal and a server, and is implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, notebook computers, and the like. The terminal splits the mixed language audio data, analyzes the audio features of different sound sources, obtains the language type corresponding to each sub-audio data and the spoken language preference information, thereby effectively improving the analysis comprehensiveness of each sound source in the multi-language audio content. Then, the method does not need to analyze the entire sub-audio data to improve the translation efficiency, but needs to sample and analyze the sub-audio data to obtain the language type and the spoken language preference information, thereby improving the analysis efficiency of each sub-audio data. Then, the method corrects the audio based on the spoken language preference information after audio translation, so that the user can not only obtain the translated audio, but also accurately identify the audio source corresponding to each audio, thereby improving the recognition effect of the user on the emotional, tone, and voice details of the audio source. Finally, the method fuses the audio data, so that the user can locate the correct position of the surrounding sound source and the current multi-sound source environment according to the audio translation content in real time, thereby improving the user experience effect and the accuracy of multi-language simultaneous translation.

[0031] In one embodiment, as shown in Figure 1 A method for intelligent translation processing based on a multi-language learning platform is provided. The method is applied to a terminal, and includes the following steps: In step S101, audio data of mixed languages is obtained, and the audio data is split into sub-audio data corresponding to each sound source through an audio splitting strategy.

[0032] In this embodiment, the terminal receives audio data collected by an audio device, and obtains audio data of mixed languages around the user in the current environment. The environment is an environment in which multiple persons communicate through more than one different language. The positions of different persons can be located at different directions of the user, such as left rear, left front, front, right front, right rear, and rear. Then, the terminal splits the audio data into sub-audio data corresponding to each sound source through an audio splitting program pre-installed in the terminal. The audio splitting program can be an audio splitting software based on AI intelligent technology. The specific splitting process will be described in detail later.

[0033] In step S102, the audio features of the sampled audio segments in each sub-audio data are extracted, and the language type corresponding to each sub-audio data and the spoken language preference information corresponding to each sub-audio data are identified based on the audio features of the sampled audio segments by a language characteristic analysis strategy.

[0034] In this embodiment, the terminal extracts the audio features of the sampled audio segments in each sub-audio data, and identifies the language type corresponding to each sub-audio data and the spoken language preference information corresponding to each sub-audio data based on the audio features of the sampled audio segments by a language characteristic analysis strategy. The sampled audio segment is an audio segment of a preset time length from the start time in each sub-audio data, and the preset time length can be 2S, 3S, 5S, etc. The language characteristic analysis strategy is an identification strategy for identifying the language type corresponding to the sub-audio data and the spoken language preference information of the user, and the specific identification process will be described in detail later. The language type includes but is not limited to English type, French type, Japanese type, Korean type, Russian type, etc. The spoken language preference information includes but is not limited to male spoken language type, female spoken language type based on gender division, child spoken language type, teenager spoken language type, youth spoken language type, middle-aged spoken language type, and middle-aged spoken language type based on age division, high-pitched spoken language type, medium-pitched spoken language type, low-pitched spoken language type, and undefined volume spoken language type based on pronunciation volume division, and front nasal sound pronunciation type, back nasal sound pronunciation type, ambiguous sound pronunciation type, and thick sound pronunciation type based on sound pronunciation preference classification. The specific identification process will be described in detail later.

[0035] In step S103, the translation audio data corresponding to each sub-audio data is generated based on the language type corresponding to each sub-audio data by an intelligent audio translation device, and the translation audio data corresponding to each sub-audio data is subjected to audio proofreading fusion processing based on the spoken language preference information corresponding to each sub-audio data to obtain mixed language translation audio data.

[0036] In this embodiment, the terminal generates the translation audio data corresponding to each sub-audio data based on the language type corresponding to each sub-audio data by an intelligent audio translation device, and subjects the translation audio data corresponding to each sub-audio data to audio proofreading fusion processing based on the spoken language preference information corresponding to each sub-audio data to obtain mixed language translation audio data. The audio angle fusion processing mode includes adjusting the sound of the translation audio data corresponding to each sub-audio data according to the spoken language preference information corresponding to each sub-audio data, so that the pronunciation of the obtained translation audio data after translation is closer to the sound source pronunciation / pronunciation user of the sub-audio data, and then the translation audio data is subjected to audio mixing processing to obtain mixed language translation audio data.

[0037] Based on the above scheme, by splitting the mixed language audio data, then analyzing the audio features of different sound sources, the language type and the spoken preference information corresponding to each sub-audio data are obtained, thereby effectively improving the analysis comprehensiveness of each sound source in the multi-language audio content. Then, in order to improve the translation efficiency, the sub-audio data does not need to be analyzed, and the language type and the spoken preference information are obtained by sampling and analyzing the sub-audio data, thereby improving the analysis efficiency of each sub-audio data. Then, after the audio translation, the audio is corrected combined with the spoken preference information, so that the user can not only obtain the translated audio, but also accurately identify the audio source corresponding to each audio, thereby improving the recognition effect of the user on the emotional, tone, voice and other detailed information of the audio source. Finally, the audio data is fused, so that the user can locate the correct position of the surrounding sound source and the current multi-sound source environment according to the audio translation content in real time, thereby improving the user experience effect and the accuracy of multi-language simultaneous translation.

[0038] Optionally, the audio data is split into sub-audio data corresponding to each sound source through an audio splitting strategy, including: splitting the audio data into initial sub-audio data through an audio splitting program; identifying the audio source direction of each initial sub-audio data and pre-processing each initial sub-audio data to obtain each sub-audio data; marking the audio source direction of each sub-audio data as the sound source corresponding to each sub-audio data to obtain sub-audio data corresponding to each sound source.

[0039] In this embodiment, the terminal splits the audio data into initial sub-audio data through an audio splitting program. Then, the terminal identifies the audio source direction of each initial sub-audio data and pre-processes each initial sub-audio data to obtain each sub-audio data. Finally, the terminal marks the audio source direction of each sub-audio data as the sound source corresponding to each sub-audio data to obtain sub-audio data corresponding to each sound source. The pre-processing method includes but is not limited to noise reduction processing, frame processing, normalization processing, and silence elimination processing. The identification method of the audio source direction of each sub-audio data is that the audio acquisition device for collecting audio data is a microphone array composed of multiple microphones, and the terminal locates the audio source direction of each sub-audio data through a multi-source positioning method based on the positions of the multiple microphones and the audio intensities collected by the multiple microphones.

[0040] Based on the above scheme, after audio splitting, the audio source direction is marked, so that not only the source of each sub-audio data can be accurately identified, but also when the translated audio data is played to the user after subsequent audio fusion, the audio source direction is played, which can make the user more immersive and improve the user experience effect.

[0041] Optionally, in each sub-audio data, the audio features of the sampled audio segment in each sub-audio data are extracted, including: for each sub-audio data, in the sub-audio data, according to a preset sampling strategy, a sampled audio segment in the sub-audio data is collected; identifying the sound features of each sound feature type corresponding to the sampled audio segment, and extracting the audio variation features corresponding to the sampled audio segment and the audio distribution features corresponding to the sampled audio segment through an audio feature extraction network; and taking the sound features of each sound feature type, the audio variation features corresponding to the sampled audio segment, and the audio distribution features corresponding to the sampled audio segment as the audio features of the sampled audio segment.

[0042] In this embodiment, the terminal collects a sampled audio segment in each sub-audio data in the sub-audio data according to a preset sampling strategy. The preset sampling strategy is to collect the sound features of each sound feature type corresponding to the sampled audio segment in each sub-audio data, and extract the audio variation features corresponding to the sampled audio segment and the audio distribution features corresponding to the sampled audio segment through an audio feature extraction network. The sound features of each sound feature type include but are not limited to pitch type, loudness type, and timbre type. The audio feature extraction network is a convolutional neural network based on Amphion technology.

[0043] Finally, the terminal takes the sound features of each sound feature type, the audio variation features corresponding to the sampled audio segment, and the audio distribution features corresponding to the sampled audio segment as the audio features of the sampled audio segment.

[0044] Based on the above scheme, by combining timbre, audio variation, and audio distribution, the audio features are comprehensively analyzed, and the comprehensiveness of audio extraction of the sampled audio segment is improved.

[0045] Optionally, based on the audio features of each sampled audio segment, the language type corresponding to each sub-audio data and the spoken language preference information corresponding to each sub-audio data are identified through a language characteristic analysis strategy, including: for each sampled audio segment, the translated texts of each language type corresponding to the sampled audio segment are identified through a text translation program of each language type, and the logical reasonableness value corresponding to each translated text is identified through a semantic logic detection program; the language type corresponding to the translated text with the maximum logical reasonableness value is selected as the language type of the sub-audio data corresponding to the sampled audio segment; based on the sound features of each sound feature type, the audio variation features corresponding to the sampled audio segment, and the audio distribution features corresponding to the sampled audio segment, the user pronunciation preference of the sampled audio segment and the user spoken language preference of the sampled audio segment are identified through an audio preference analysis network; and the user pronunciation preference of the sampled audio segment and the user spoken language preference of the sampled audio segment are taken as the spoken language preference information corresponding to the sub-audio data.

[0046] In this embodiment, the terminal identifies the translated texts of each language type corresponding to each sampled audio segment through a text translation program of each language type, and identifies the logical reasonableness value corresponding to each translated text through a semantic logic detection program. The semantic logic detection program is a large language model based on AI intelligent technology, which analyzes the semantic logic degree of each text from core elements such as entities, attributes, and relationships.

[0047] Then, the terminal selects the language type corresponding to the translated text with the maximum logical reasonableness value as the language type of the sub-audio data corresponding to the sampled audio segment.

[0048] The terminal identifies the user pronunciation preference of the sampled audio segment and the user spoken language preference of the sampled audio segment through an audio preference analysis network based on the sound features of each sound feature type, the audio variation features corresponding to the sampled audio segment, and the audio distribution features corresponding to the sampled audio segment. The audio preference analysis network is a convolutional neural network based on a self-attention mechanism. Specifically, the convolutional neural network includes a pronunciation preference analysis subnetwork and a spoken language preference analysis subnetwork. The neural architectures of the two subnetworks are the same, but the training methods are different. The terminal identifies the user pronunciation preference of the sampled audio segment through the pronunciation preference analysis subnetwork based on the sound features of each sound feature type, and identifies the user spoken language preference of the sampled audio segment through the spoken language preference analysis subnetwork based on the audio variation features corresponding to the sampled audio segment and the audio distribution features corresponding to the sampled audio segment.

[0049] Finally, the terminal takes the user pronunciation preference of the sampled audio segment and the user spoken language preference of the sampled audio segment as the spoken language preference information corresponding to the sub-audio data.

[0050] Based on the above scheme, the feature analysis is performed on different features, so as to give the user's pronunciation preference and the user's spoken language preference of the recognized sampling audio segment, and the pronunciation analysis comprehensiveness of the sub-audio data is improved.

[0051] Optionally, based on the language type corresponding to each sub-audio data, the intelligent audio translation device generates translation audio data corresponding to each sub-audio data, including: for each sub-audio data, constructing a translation association mark between the sound source corresponding to the sub-audio data and the language type corresponding to the sub-audio data; based on the user's pronunciation preference of the sub-audio data, adjusting the translation program of the language type corresponding to the sub-audio data in the intelligent audio translation device to obtain a target translation program corresponding to the sub-audio data; based on each sub-audio data and the translation association mark, generating translation text data corresponding to the sub-audio data through the target translation program corresponding to the sub-audio data, and performing data conversion processing on the translation text data to obtain translation audio data corresponding to the translation text data.

[0052] In this embodiment, the terminal constructs a translation association mark between the sound source corresponding to each sub-audio data and the language type corresponding to the sub-audio data. The translation association mark is used to directly translate the sub-audio data starting from this step without the above splitting recognition and preference recognition process after obtaining the sub-audio data of the sound source.

[0053] Then, the terminal adjusts the translation program of the language type corresponding to the sub-audio data in the intelligent audio translation device based on the user's pronunciation preference of the sub-audio data, to obtain a target translation program corresponding to the sub-audio data. Finally, the terminal generates translation text data corresponding to the sub-audio data through the target translation program corresponding to the sub-audio data based on each sub-audio data and the translation association mark, and performs data conversion processing on the translation text data to obtain translation audio data corresponding to the translation text data.

[0054] Based on the above scheme, by constructing the corresponding relationship between the sound source and the translation program, and optimizing the translation program through the user's pronunciation preference, the accuracy of translation and the subsequent audio translation efficiency are improved.

[0055] Optionally, for each sub-audio data corresponding to the translation audio data, the audio data is fused by the oral preference information corresponding to each sub-audio data to obtain mixed language translation audio data, including: for each sub-audio data, based on the oral preference information corresponding to the sub-audio data, identifying the sound preference features of each sound type corresponding to the sub-audio data; based on the sound preference features of each sound type, adjusting the sub-audio data through the audio adjustment network to obtain the target translation audio data corresponding to the sub-audio data; based on the audio source direction corresponding to each sub-audio data, through the audio mixing program, the target translation audio data corresponding to each sub-audio data is data fusion processing to obtain mixed language translation audio data.

[0056] In this embodiment, the terminal identifies the sound preference features of each sound type corresponding to each sub-audio data based on the oral preference information corresponding to the sub-audio data. Wherein, the sound preference features of each sound type are different from the sound preference features of each sound type identified above. The sound preference features of each sound type identified above are directly identified, while the sound preference features of each sound type identified here are the sound preference features of each sound type adapted in the preference feature identification database preset in the terminal through the gender division, age division, pronunciation volume division of each oral preference information identified. Wherein, the preference feature identification database includes the corresponding relationship between different oral preference information and the sound preference features of each sound type, which can be stored in the form of a tree diagram.

[0057] Then, the terminal adjusts the sub-audio data based on the sound preference features of each sound type through the audio adjustment network to obtain the target translation audio data corresponding to the sub-audio data. Wherein, the audio adjustment network is an audio data optimization program based on AI intelligent technology. Finally, the terminal adjusts the target translation audio data corresponding to each sub-audio data based on the audio source direction corresponding to each sub-audio data through the audio mixing program to obtain mixed language translation audio data. Wherein, the audio mixing program is the audio data obtained by sequentially sorting each sub-audio data according to the propagation time interval of each sub-audio data and then fusing through the audio fusion program.

[0058] Based on the above scheme, by optimizing and adjusting the audio adjustment network and then performing audio mixing, the user can not only obtain translated audio, but also accurately identify the audio source corresponding to each audio, thereby improving the user's recognition of the emotional, tone, voice and other detailed information of the audio source. Finally, the present scheme fuses the audio data, so that the user can locate the correct position of the surrounding sound source and the current multi-sound source environment according to the audio translation content in real time, thereby improving the user experience effect and the accuracy of multi-language simultaneous translation.

[0059] Based on the above scheme, by actively preventing and controlling the power transmission project, the noise of each noise source can be accurately prevented and controlled, thereby reducing the noise prevention and control cost and efficiently preventing and controlling the noise.

[0060] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.

[0061] Based on the same inventive concept, the embodiments of the present application also provide a multi-language learning platform translation intelligent processing system for implementing the above-mentioned multi-language learning platform translation intelligent processing method. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme described in the above method, so the specific limitations in one or more multi-language learning platform translation intelligent processing system embodiments provided below can refer to the limitations of the multi-language learning platform translation intelligent processing method in the above text, which will not be repeated here.

[0062] Further referring to Figure 2 , as an implementation of the method shown in the above Figure 1 , the present application provides an embodiment of a multi-language learning platform translation intelligent processing system 200, which includes an acquisition module 210, an identification module 220, and a fusion module 230, wherein: The acquisition module 210 is configured to acquire audio data of mixed languages, and split the audio data into sub-audio data corresponding to each sound source through an audio splitting strategy. The recognition module 220 is configured to extract audio features of a sampling audio segment in each of the sub-audio data, and identify a language type corresponding to each of the sub-audio data and spoken preference information corresponding to each of the sub-audio data based on the audio features of the sampling audio segment and through a language characteristic analysis strategy. The fusion module 230 is configured to generate translated audio data corresponding to each of the sub-audio data through an intelligent audio translation device based on the language type corresponding to each of the sub-audio data, and perform audio proofreading and fusion processing on the translated audio data corresponding to each of the sub-audio data through the spoken preference information corresponding to each of the sub-audio data, to obtain the translated audio data in the mixed language.

[0063] Optionally, the acquisition module 210 is specifically configured to: split the audio data into initial sub-audio data through an audio splitting program; identify an audio source direction of each of the initial sub-audio data, and pre-process each of the initial sub-audio data to obtain the sub-audio data; mark the audio source direction of each of the sub-audio data as a sound source corresponding to each of the sub-audio data, to obtain the sub-audio data corresponding to each of the sound sources.

[0064] Optionally, the recognition module 220 is specifically configured to: for each of the sub-audio data, collect a sampling audio segment in the sub-audio data according to a preset sampling strategy; identify a sound feature of each of the sound feature types corresponding to the sampling audio segment, and extract an audio variation feature corresponding to the sampling audio segment and an audio distribution feature corresponding to the sampling audio segment through an audio feature extraction network; use the sound feature of each of the sound feature types, the audio variation feature corresponding to the sampling audio segment, and the audio distribution feature corresponding to the sampling audio segment as the audio feature of the sampling audio segment.

[0065] Optionally, the recognition module 220 is specifically configured to: for each of the sampling audio segments, identify translated texts of each of the language types corresponding to the sampling audio segment through a text translation program of each of the language types, and identify a logical reason value corresponding to each of the translated texts through a semantic logic detection program; select a language type corresponding to a translated text with a maximum logical reason value as the language type of the sub-audio data corresponding to the sampling audio segment; The user pronunciation preference of the sample audio segment and the user spoken language preference of the sample audio segment are identified as the spoken language preference information corresponding to the sub-audio data. The user pronunciation preference of the sample audio segment and the user spoken language preference of the sample audio segment are identified as the spoken language preference information corresponding to the sub-audio data.

[0066] Optionally, the fusion module 230 is specifically configured to: For each sub-audio data, a translation association mark between the audio source corresponding to the sub-audio data and the language type corresponding to the sub-audio data is constructed; Based on the user pronunciation preference of the sub-audio data, the translation program of the language type corresponding to the sub-audio data in the intelligent audio translation device is adjusted to obtain a target translation program corresponding to the sub-audio data; Based on each sub-audio data and the translation association mark, the target translation program corresponding to the sub-audio data is used to generate translation text data corresponding to the sub-audio data, and the translation text data is subjected to data conversion processing to obtain translation audio data corresponding to the translation text data.

[0067] Optionally, the fusion module 230 is specifically configured to: For each sub-audio data, based on the spoken language preference information corresponding to the sub-audio data, a sound preference feature of each sound type corresponding to the sub-audio data is identified; Based on the sound preference feature of each sound type, an audio adjustment network is used to adjust the sub-audio data to obtain target translation audio data corresponding to the sub-audio data; Based on the audio source direction corresponding to each sub-audio data, an audio mixing program is used to perform data fusion processing on the target translation audio data corresponding to each sub-audio data to obtain translation audio data of the mixed language.

[0068] The above-mentioned various modules in the multi-language learning platform-based intelligent translation processing system can be all or part realized by software, hardware and combinations thereof. The above-mentioned various modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above-mentioned various modules.

[0069] In one embodiment, a computer device, which can be a terminal, is provided, and an internal structure diagram of the computer device can be as shown in Figure 3As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input system connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used for wired or wireless communication with external terminals. Wireless communication can be achieved through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement a multi-language learning platform translation intelligent processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input system of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad provided on the computer device shell, or an external keyboard, touchpad or mouse, etc.

[0070] Those skilled in the art can understand that, Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0071] In one embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method of any one of the first aspect.

[0072] In one embodiment, a computer readable storage medium is provided, having a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the method of any one of the first aspect.

[0073] In one embodiment, a computer program product is provided, comprising a computer program, and the computer program is executed by a processor to implement the steps of the method of any one of the first aspect.

[0074] It should be noted that the patient information (including but not limited to patient device information, patient personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the patient or fully authorized by all parties.

[0075] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0076] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0077] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for intelligent processing of translation based on a multi-language learning platform, characterized in that, The method comprises: acquiring mixed language audio data, and splitting the audio data into sub-audio data corresponding to each sound source through an audio splitting strategy; extracting the audio features of the sampled audio segments in each sub-audio data from the sub-audio data, and identifying the language type corresponding to each sub-audio data and the spoken language preference information corresponding to each sub-audio data based on the audio features of each sampled audio segment through a language characteristic analysis strategy; based on the language type corresponding to each sub-audio data, generating translated audio data corresponding to each sub-audio data through an intelligent audio translation device, and performing audio proofreading and fusion processing on the translated audio data corresponding to each sub-audio data through the spoken language preference information corresponding to each sub-audio data to obtain the translated audio data of the mixed language.

2. The method of claim 1, wherein, The audio data is split into sub-audio data corresponding to each sound source through an audio splitting strategy, which comprises: splitting the audio data into initial sub-audio data through an audio splitting program; identifying the audio source direction of each initial sub-audio data and preprocessing each initial sub-audio data to obtain sub-audio data; marking the audio source direction of each sub-audio data as the sound source corresponding to each sub-audio data to obtain sub-audio data corresponding to each sound source.

3. The method of claim 1, wherein, The audio features of the sampled audio segments in each sub-audio data are extracted from the sub-audio data, which comprises: for each sub-audio data, collecting the sampled audio segments in the sub-audio data according to a preset sampling strategy in the sub-audio data; identifying the sound features of each sound feature type corresponding to the sampled audio segments, and extracting the audio variation features and audio distribution features corresponding to the sampled audio segments through an audio feature extraction network; the sound features of each sound feature type, the audio variation features corresponding to the sampled audio segments, and the audio distribution features corresponding to the sampled audio segments are taken as the audio features of the sampled audio segments.

4. The method of claim 3, wherein, The language type corresponding to each sub-audio data and the spoken language preference information corresponding to each sub-audio data are identified based on the audio features of each sampled audio segment through a language characteristic analysis strategy, which comprises: for each sampled audio segment, identifying the translated text of each language type corresponding to the sampled audio segment through a text translation program of each language type, and identifying the logical reasonableness value corresponding to each translated text through a semantic logic detection program; selecting the language type corresponding to the translated text with the maximum logical reasonableness value as the language type of the sub-audio data corresponding to the sampled audio segment; based on the sound features of each sound feature type, the audio variation features corresponding to the sampled audio segments, and the audio distribution features corresponding to the sampled audio segments, identifying the user pronunciation preference of the sampled audio segment and the user spoken language preference of the sampled audio segment through an audio preference analysis network; the user pronunciation preference of the sampled audio segment and the user spoken language preference of the sampled audio segment are taken as the spoken language preference information corresponding to the sub-audio data.

5. The method of claim 4, wherein, The method comprises the following steps: For each sub-audio data, a translation association mark between the audio source corresponding to the sub-audio data and the language type corresponding to the sub-audio data is constructed; Based on the user pronunciation preference of the sub-audio data, a translation program of the language type corresponding to the sub-audio data in the intelligent audio translation device is adjusted to obtain a target translation program corresponding to the sub-audio data; Based on each sub-audio data and the translation association mark, the target translation program corresponding to the sub-audio data is used to generate translation text data corresponding to the sub-audio data, and the translation text data is subjected to data conversion processing to obtain translation audio data corresponding to the translation text data.

6. The method of claim 2, wherein, The method comprises the following steps: For each sub-audio data, based on the spoken language preference information corresponding to the sub-audio data, the sound preference features of each sound type corresponding to the sub-audio data are identified; Based on the sound preference features of each sound type, the sub-audio data is adjusted through an audio adjustment network to obtain target translation audio data corresponding to the sub-audio data; Based on the audio source direction corresponding to each sub-audio data, the target translation audio data corresponding to each sub-audio data is subjected to data fusion processing through an audio mixing program to obtain translation audio data of the mixed language.

7. A multi-lingual learning platform based translation intelligence processing system, characterized in that, The system comprises: An acquisition module is configured to acquire audio data of a mixed language, and split the audio data into sub-audio data corresponding to each audio source through an audio splitting strategy; An identification module is configured to extract audio features of a sampling audio segment in each sub-audio data from the sub-audio data, and identify a language type corresponding to each sub-audio data and spoken language preference information corresponding to each sub-audio data based on the audio features of each sampling audio segment through a language characteristic analysis strategy; A fusion module is configured to generate translation audio data corresponding to each sub-audio data based on the language type corresponding to each sub-audio data through an intelligent audio translation device, and perform audio collation fusion processing on the translation audio data corresponding to each sub-audio data through the spoken language preference information corresponding to each sub-audio data to obtain translation audio data of the mixed language.

8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 6.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-user multi-language recognition and translation method and device

    CN113299276A

  • Voice processing method and device, computer equipment and storage medium

    CN114220436A

  • Multilingual speech recognition method and device

    CN116453503A

  • Audio processing method and device, electronic equipment and readable storage medium

    CN119252243A

  • Translation processing method and device, electronic equipment and computer readable storage medium

    CN119886162A