Audio language conversion method, system and related devices
By obtaining audio features and semantic segmentation, the target feature fragment is generated, and the audio language conversion is used to use intelligent analysis models to solve the problem of high translation costs or poor results in the existing technology, and efficient and accurate cross-language conversion is achieved.
Patent Information
- Application Number
- CN202510278032.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-23
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Existing audio language conversion methods rely on manual translation or neural network models, resulting in high translation costs or poor conversion effects, making it difficult to improve the accuracy of cross-language conversion.
By obtaining the initial audio stream of the target object, extracting the initial audio features and the current language, generating the target feature fragment based on semantic segmentation, and using the intelligent analysis model to convert languages to generate converted audio matching the converted language.
It improves the real-time accuracy of audio language conversion, reduces dependence on context semantics, and improves the efficiency and quality of translation.
Smart Images

Figure CN119832897B_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 2024114869196 and the application title "Audio Language Conversion Method, System and Related Devices" submitted to the National Intellectual Property Administration on October 23, 2024, the entire content of which is incorporated herein by reference. Technical Field
[0002] This application relates to the technical field of audio processing, and particularly to an audio language conversion method, system and related devices. Background Art
[0003] With the continuous development of artificial intelligence, real-time audio language conversion has been applied in more and more scenarios. For example, in scenarios such as large conferences or lectures, there may be a situation where the language of the speaker and the participants is different, resulting in the participants being unable to accurately understand what the speaker is expressing. Currently, the language conversion methods mainly rely on human translators for manual translation or use neural network models to translate word by word according to the input audio stream. For the former, due to the short translation time, the quality requirements for translators are extremely high, and it requires a high labor cost. For the latter, due to the differences in different language expressions, neural network models are prone to produce unsmooth sentences during translation, resulting in poor conversion effects.
[0004] In view of this, how to improve the accuracy of audio cross-language conversion has become an urgent problem to be solved. Summary of the Invention
[0005] The main technical problem to be solved by this application is to provide an audio language conversion method, system and related devices, which can improve the accuracy of audio cross-language conversion.
[0006] To solve the above technical problem, one technical solution adopted by this application is: to provide an audio language conversion method, including: obtaining an initial audio stream of a target object, determining the initial audio features corresponding to the initial audio stream and the current language corresponding to the initial audio stream; based on the initial audio features and the current language, obtaining a target feature segment corresponding to the current conversion round; wherein, the target feature segments corresponding to different conversion rounds are segmented based on the semantics of the initial audio features; determining at least one conversion language, and generating a conversion audio matching the conversion language based on the current language and the target feature segment.
[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide an audio language conversion system, including: an acquisition module, configured to acquire an initial audio stream of a target object, determine an initial audio feature corresponding to the initial audio stream, and a current language corresponding to the initial audio stream; a processing module, configured to obtain a target feature segment corresponding to a current conversion round based on the initial audio feature and the current language; wherein, the target feature segments corresponding to different conversion rounds are segmented based on the semantics of the initial audio feature; a conversion module, configured to determine at least one conversion language, generate a conversion audio matching the conversion language based on the current language and the target feature segment, and broadcast the conversion audio.
[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, including: a memory and a processor coupled to each other, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the audio language conversion method mentioned in the above technical solution.
[0009] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the audio language conversion method mentioned in the above technical solution is implemented.
[0010] The beneficial effect of this application is: different from the prior art, the audio language conversion method proposed by this application obtains the initial audio stream of the target object in the target scenario in real time, and extracts features from the initial audio stream to obtain the corresponding initial audio feature and the current language. According to the semantics of the current language and the initial audio feature, the initial audio feature is semantically segmented to obtain the target feature segment. Moreover, when translating and converting the language of the target feature segment, there is no need to refer to the context semantics, thereby improving the accuracy of real-time conversion of the audio language. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings. Among them:
[0012] Figure 1 is a flowchart of an implementation manner of the audio language conversion method of this application;
[0013] Figure 2 is Figure 1 a flowchart of another implementation manner corresponding to step S101 in
[0014] Figure 3 is Figure 2 The flowchart corresponding to step S203 in [the other embodiment].
[0015] Figure 4 is Figure 1 The flowchart corresponding to step S102 in [the other embodiment].
[0016] Figure 5 is Figure 1 The flowchart corresponding to step S103 in [the other embodiment].
[0017] Figure 6 is Figure 5 The flowchart corresponding to step S502 in [the other embodiment].
[0018] Figure 7 is the structural schematic diagram of an embodiment of the audio language conversion system of the present application;
[0019] Figure 8 is the structural schematic diagram of an embodiment of the electronic device of the present application;
[0020] Figure 9 is the structural schematic diagram of an embodiment of the computer-readable storage medium of the present application. Specific Embodiments
[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments, and different embodiments can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0022] Please refer to Figure 1 , Figure 1 is the flowchart of an embodiment of the audio language conversion method of the present application. The method includes:
[0023] S101: Obtain the initial audio stream of the target object, and determine the initial audio features corresponding to the initial audio stream and the current language of the initial audio stream.
[0024] In one embodiment, the target object is determined from the target scene, and the initial audio stream generated when the target object makes a real-time statement is extracted. Feature extraction is performed on the initial audio stream to obtain the corresponding initial audio features. And, based on the obtained initial audio features, the current language expressed by the target object is determined. The current language is one of multiple language types, such as one of Chinese, English, Japanese, Korean, etc.
[0025] In another embodiment, after extracting the initial audio stream of the target object, real-time noise reduction processing is performed on the initial audio stream to remove environmental noise. Further, feature extraction is performed on the initial audio stream after removing the noise to obtain corresponding initial audio features, and the current language is determined based on the initial audio features.
[0026] S102: Obtain the target feature segment corresponding to the current conversion round based on the initial audio features and the current language; wherein, the target feature segments corresponding to different conversion rounds are segmented based on the semantics of the initial audio features.
[0027] In one embodiment, based on the obtained initial audio features and the current language, semantic analysis and segmentation are performed on the initial audio features to determine the target feature segment corresponding to the current conversion round.
[0028] Specifically, the above target feature segment includes an initial audio feature segment with a semantic correlation degree greater than the correlation degree threshold, that is, the semantic integrity degree of the target feature segment corresponding to the current conversion round is relatively high, so that no context semantics need to be relied on when translating according to the target feature segment, which helps to improve the accuracy of real-time audio language conversion.
[0029] S103: Determine at least one target language, and generate a converted audio that matches the target language based on the current language and the target feature segment.
[0030] In one embodiment, determine at least one target language specified by the user, and generate a converted audio that matches each target language according to the current language and the target feature segment.
[0031] In one implementation scenario, after determining the current language corresponding to the above initial audio stream, display multiple languages to be selected and prompt the user to make a selection. In response to obtaining the user's selection instruction, use the language to be selected corresponding to the selection instruction as the target language. Analyze the current language and the target feature segment using the fine-tuned intelligent analysis model, and prompt the intelligent analysis model to generate a converted audio that matches each target language. Among them, the above intelligent analysis model is a large language model with relatively good data analysis capabilities. By inputting the current language and the target feature segment into the intelligent analysis model, the intelligent analysis model is enabled to generate a translation text corresponding to the target feature segment, and generate a converted audio that matches the target language according to the translation text.
[0032] In a specific application scenario, the above-mentioned large language model may include, but is not limited to, deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTM), and generative pre-trained Transformer models, etc. There is no specific limitation on the specific structure and specific deployment of the large language model here. In addition, it should be noted that the specific structure and specific deployment of the intelligent analysis model mentioned in other embodiments of this application can all refer to this embodiment.
[0033] In another embodiment, the above-mentioned language conversion is determined according to the first languages of different objects in the target scenario.
[0034] Specifically, determine the first languages mastered by each object in the target scenario. For the target object in the target scenario, use the first languages of other objects as the conversion languages, so that after audio language conversion, it can help other objects in the target scenario understand the expression content of the target object.
[0035] The audio language conversion method proposed in this application obtains the initial audio stream of the target object in the target scenario in real time, extracts features from the initial audio stream to obtain the corresponding initial audio features and the current language. According to the semantics of the current language and the initial audio features, semantic segmentation is performed on the initial audio features to obtain target feature segments. And when translating and converting the language of the target feature segments, there is no need to refer to the context semantics, thereby improving the accuracy of real-time conversion of the audio language.
[0036] Please refer to Figure 2 , Figure 2 is Figure 1 The flowchart corresponding to step S101 in another embodiment. Specifically, the implementation process of step S101 includes:
[0037] S201: Obtain the initial audio stream of the target object and determine the initial audio features corresponding to the initial audio stream.
[0038] In one embodiment, in response to the audio language conversion method proposed in this application being applied to a target scenario including at least one target object, the scenario audio data in the target scenario is obtained in real time. Input the scenario audio data into the trained voice detection model, so that the voice detection model extracts the initial audio stream matching each target object from the scenario audio data according to the timbre difference. Input the initial audio into the trained voice encoding model to obtain the initial audio features.
[0039] In another embodiment, scene information matching the target scene is obtained. The scene information includes scene video data and scene audio data.
[0040] Specifically, in order to improve the accuracy of detecting the initial audio streams of different target objects, the scene video data and scene audio data corresponding to the target scene are obtained. The scene video data is used to determine the number of target objects in the target scene and the facial action information of the target objects, and the facial action information is used to determine whether the corresponding target object is making a statement. By combining the scene video data and the scene audio data, it helps to improve the accuracy of separating the audio of different target objects.
[0041] Further, based on the scene information, at least one target object in the target scene and the initial audio stream corresponding to the target object are determined. The initial audio is input into the trained speech encoding model to obtain initial audio features.
[0042] Specifically, according to the scene audio data, the different timbre information contained therein is determined. And according to the scene video data, the target objects in the target scene are determined. The above timbre information is matched with the target objects, and the scene audio data is separated based on the facial action information of different target objects in the scene video data, so as to obtain the initial audio stream corresponding to each target object. The obtained initial audio is audio-encoded using the trained speech encoding model to obtain initial audio features.
[0043] In yet another embodiment, the above scene information further includes microphone information, and the above scene audio data includes the audio collected by different microphones in the target scene. According to the audio collected by each microphone, the initial audio stream corresponding to the target object in the target scene and its corresponding initial audio features are determined.
[0044] Specifically, each target object in the target scene is matched with a corresponding microphone. For the audio collected by each microphone, the initial audio stream of the corresponding target object is extracted and audio-encoded to obtain initial audio features.
[0045] S202: Based on the current moment and the initial audio features, obtain the current window features corresponding to the preset duration.
[0046] In one embodiment, a preset duration is obtained, and based on the current moment and the preset duration, the current window features are determined.
[0047] Specifically, the initial audio features corresponding to the initial audio stream within the preset duration before the current moment are used as the current window features. The above preset duration can be set according to actual needs, such as 1 second, 2 seconds, or 3 seconds, etc.
[0048] S203: Determine the current language from multiple candidate languages based on the current window features.
[0049] In one embodiment, input the current window features into the trained language discriminant model, so that the language discriminant model outputs the current discriminant information corresponding to the current window features, and use the candidate language corresponding to the maximum reference score as the current language. Wherein, the current discriminant information includes reference scores matching each candidate language; and the higher the reference score, the higher the probability that the corresponding candidate language is the current language.
[0050] In another embodiment, after obtaining the reference scores of each candidate language corresponding to the current window features by using the language discriminant model, determine whether the maximum reference score is greater than or equal to a preset score threshold. If so, use the candidate language corresponding to the maximum reference score as the current language. If not, increase the preset duration, obtain the current window features matching the adjusted preset duration, and return to step S203 to re-obtain the reference scores of each candidate language corresponding to the current window features until the maximum reference score is greater than or equal to the score threshold, and use the candidate language corresponding to the maximum reference score as the current language.
[0051] In a specific application scenario, the preset duration before adjustment is 1 second, then use the initial audio features corresponding to the initial audio stream from the previous 1 second to the current moment as the current window features. In response to obtaining the reference scores of each candidate language corresponding to the current window features by using the language discriminant model, if the maximum reference score is less than the preset score threshold, increase the preset duration to 2 seconds, and use the initial audio features corresponding to the initial audio stream within the previous 2 seconds before the current moment as the current window features, and use the language discriminant model to determine the current language.
[0052] Further, after determining the current language according to the current window features in the above embodiment, update the current window features to historical window features, update the current discriminant information obtained according to the current window features to historical discriminant information, and continue to obtain the next window features.
[0053] The above solution determines the current window features in a sliding window manner, and identifies the current language expressed by the target object according to the current window features, so as to improve the effect of language conversion in the subsequent process of combining the identified current language for language conversion.
[0054] Please refer to Figure 3 , Figure 3 is Figure 2 the schematic flow chart of another embodiment corresponding to step S203 in
[0055] S301: Obtain historical window features adjacent to the current window features and the historical discrimination information corresponding to the historical window features; wherein, the historical discrimination information includes historical scores matching each candidate language.
[0056] In one embodiment, determine historical window features adjacent to the current window features and the historical discrimination information corresponding to the historical window features according to a preset duration.
[0057] S302: Obtain current discrimination information based on the current window features; wherein, the current discrimination information includes current scores matching each candidate language.
[0058] In one embodiment, input the obtained current window features into a trained language discrimination model, so that the language discrimination model outputs current scores matching each candidate language, and use each language and its matching current scores as the current discrimination information.
[0059] S303: Determine a first weight matching the current discrimination information and a second weight matching the historical discrimination information. Based on the first weight, the second weight, the historical discrimination information, and the current discrimination information, obtain target scores of the current window features matching each candidate language. Based on the target scores, determine the current language.
[0060] In one embodiment, obtain a pre-determined first weight and a second weight. For each candidate language, obtain a first product of the current score and the first weight, and a second product of the historical score and the second weight, and use the sum of the first product and the second product as the target score matching the candidate language.
[0061] Specifically, for any candidate language, the specific calculation formula for its corresponding target score is as follows:
[0062]
[0063] Wherein, represents the candidate language corresponding target score, represents the first weight, represents the candidate language corresponding current score, represents the second weight, represents the candidate language corresponding historical score.
[0064] In the above solution, by combining the current discrimination information with the historical discrimination information, the improvement is achieved, which helps to improve the accuracy of determining the current language.
[0065] Please refer to Figure 4 , Figure 4 isFigure 1 The flowchart diagram corresponding to step S102 in another embodiment. Specifically, the implementation process of step S102 includes:
[0066] S401: Input the initial audio feature and the current language into the trained semantic segmentation model to obtain the semantic information corresponding to the initial audio feature.
[0067] In one embodiment, obtain the trained semantic segmentation model, input the initial audio feature and the determined current language into the semantic segmentation model, so as to perform real-time semantic analysis on the initial audio feature by using the semantic segmentation model, obtain the corresponding semantic information in real time, and determine the correlation degree between the semantic information corresponding to different moments.
[0068] S402: In response to the semantic information satisfying the preset segmentation condition, assign the current segmentation identifier to the initial audio feature; wherein, the preset segmentation condition is related to the completeness of the semantic information.
[0069] In one embodiment, determine whether the correlation degree between the semantic information obtained at the current moment and the semantic information corresponding to the previous moment is greater than the correlation degree threshold. If it is less than the correlation degree threshold, it indicates that the correlation degree between the semantic information before the current moment and the speech information at and after the current moment is relatively small, then it is determined that the preset segmentation condition is satisfied, and the initial audio feature corresponding to the current moment is assigned the current segmentation identifier.
[0070] S403: Based on the historical segmentation identifier corresponding to the previous conversion round and the current segmentation identifier, determine the target feature segment corresponding to the current conversion round.
[0071] In one embodiment, according to the historical segmentation identifier corresponding to the previous conversion round and the current segmentation identifier, determine the target feature segment corresponding to the current conversion round; that is, take the initial audio feature between the historical segmentation identifier corresponding to the previous conversion round and the current segmentation identifier as the target feature segment.
[0072] In the above solution, the target feature segment with complete semantic information is obtained through the semantic segmentation model, so as to improve the accuracy of subsequent translation based on the target feature segment.
[0073] Please refer to Figure 5 , Figure 5 is Figure 1 The flowchart diagram corresponding to step S103 in another embodiment. Specifically, the implementation process of step S103 includes:
[0074] S501: Based on the target feature segment, use the intelligent analysis model to obtain the corresponding speech-level representation; wherein, the speech-level representation includes at least one of the language information corresponding to the current language, the timbre information matching the target object, the prosody information, and the speech rate information.
[0075] In one embodiment, the obtained target feature segment is input into an intelligent analysis model to obtain a corresponding speech-level representation by using the intelligent analysis model. The speech-level representation includes language information corresponding to the current language, timbre information matching the target object, prosody information, and speech rate information. Alternatively, the speech-level representation may also include at least one of the above-mentioned language information, timbre information, prosody information, and speech rate information.
[0076] In one implementation scenario, in response to the target feature segment being an audio feature, before inputting the target feature segment into the intelligent analysis model, a feature adapter is used to process the target feature segment so that the intelligent analysis model has the ability to process the target feature segment.
[0077] In another embodiment, the initial audio feature and the target feature segment are input into the intelligent analysis model to obtain a corresponding speech-level representation by using the intelligent analysis model to combine the initial audio feature and the target feature segment, avoiding the poor effect of speech-level representation extraction caused by the too short target feature segment.
[0078] In yet another embodiment, a reference feature is obtained based on the target feature segments corresponding to at least one historical conversion round and the current conversion round respectively. The reference feature is input into the intelligent analysis model to obtain a speech-level representation.
[0079] Specifically, the target feature segment corresponding to the current conversion round is fused with the target feature segments corresponding to at least one historical conversion round to obtain a reference feature. The reference feature is processed by a feature adapter and then input into the intelligent analysis model to obtain a speech-level representation.
[0080] In a specific application scenario, to improve the acquisition efficiency of the speech-level representation, the target feature segments corresponding to the current conversion round and the historical conversion round adjacent to the current conversion round are fused to obtain a reference feature, and the intelligent analysis model is used to obtain a speech-level representation according to the reference feature.
[0081] S502: Based on the speech-level representation and the target language, use the intelligent analysis model to generate a corresponding converted audio.
[0082] In one embodiment, a pre-determined task template is obtained, and a corresponding task text is generated according to the speech-level representation and the target language. The task text is input into the intelligent analysis model so that the intelligent analysis model generates a corresponding converted audio. Among them, the above-mentioned task text is used to prompt the intelligent analysis model to convert the target feature segment into text and translate it into a translation text corresponding to the target language, and combine the speech-level representation to generate a corresponding converted audio.
[0083] Specifically, the pre-determined task template is "You are an interpreter with the ability to perform real-time translation. When you input <phonetic-level representation> and <target feature segment>, please output the translation text and conversion audio corresponding to each target language, in the following format: <target language 1><translation text corresponding to target language 1><conversion audio corresponding to target language 1>, <target language 2><translation text corresponding to target language 2><conversion audio corresponding to target language 2>...". Fill in the determined phonetic-level representation, target feature segment, and specific types of target languages in the corresponding positions in the above task template to obtain the task text. Input this task text into the intelligent analysis model so that the intelligent analysis model outputs the conversion audio corresponding to each translation text, and the conversion audio has the same timbre as the initial audio stream.
[0084] S503: Broadcast the conversion audio.
[0085] In one embodiment, in response to generating the conversion audio that matches each target language, broadcast the conversion audio.
[0086] Please refer to Figure 6 , Figure 6 is Figure 5 a schematic flowchart of another embodiment corresponding to step S502 in
[0087] S601: Obtain the target language representation that matches the target language; wherein, the target language representation at least includes the pronunciation information corresponding to the target language.
[0088] In one embodiment, obtain the target language representation that matches the determined target language. The target language representation at least includes the pronunciation information corresponding to the target language, and the pronunciation information at least includes pronunciation habits.
[0089] In one implementation scenario, the above target language representation can be pre-determined using an intelligent analysis model.
[0090] S602: Generate a task text based on the phonetic-level representation, target feature segment, target language, and target language representation; wherein, the task text is used to prompt the intelligent analysis model to generate the conversion audio that matches the target language.
[0091] In one embodiment, obtain the pre-determined task template and generate the corresponding task text according to the phonetic-level representation, target feature segment, target language, and target language representation.
[0092] Specifically, the pre-determined task template is "You are an interpreter and you have the ability to perform real-time translation using the same intonation as the input speech. When you input <speech-level representation>, <target feature segment>, <conversion language>, and <conversion language representation>, please output the translation text and conversion audio corresponding to each conversion language in the following form: <conversion language 1><translation text corresponding to conversion language 1><conversion audio corresponding to conversion language 1>, <conversion language 2><translation text corresponding to conversion language 2><conversion audio corresponding to conversion language 2>...". Fill the determined speech-level representation, target feature segment, conversion language, and conversion language representation into the corresponding positions in the above task template to obtain the task text.
[0093] S603: Input the task text into the intelligent analysis model to obtain the conversion audio generated by the intelligent analysis model.
[0094] In one embodiment, input the task text into the intelligent analysis model so that after the intelligent analysis model analyzes the task text, it generates the corresponding conversion audio.
[0095] In the above solution, when using the intelligent analysis model to generate the conversion audio, the conversion language representation corresponding to the conversion language is combined so that the finally obtained conversion audio meets the pronunciation habits of the corresponding conversion language, thereby improving the usage experience of related objects.
[0096] The present application proposes an audio language conversion device. This audio language conversion device is applied to a target scenario to implement the audio language conversion method mentioned in any of the above embodiments. The audio language conversion device stores the above-mentioned speech detection model, speech coding model, language discrimination model, semantic segmentation model, and intelligent analysis model in the corresponding embodiments. Moreover, the specific structures of the above-mentioned speech detection model, speech coding model, language discrimination model, and semantic segmentation model can refer to the existing neural network model structures. Additionally, before using the audio language conversion device to perform audio language conversion, obtain a plurality of training samples and use these training samples to jointly train the above-mentioned models until a preset stop condition is met. Among them, the preset stop condition is also related to the number of training rounds or the degree of model convergence.
[0097] Please refer to Figure 7 , Figure 7 is a schematic structural diagram of a corresponding embodiment of the audio language conversion system of the present application. The audio language conversion system includes an acquisition module 10, a processing module 20, and a conversion module 30 that are mutually coupled.
[0098] The acquisition module 10 is used to acquire the initial audio stream of the target object, determine the initial audio features corresponding to the initial audio stream, and the current language corresponding to the initial audio stream.
[0099] The processing module 20 is configured to obtain a target feature segment corresponding to the current conversion round based on the initial audio feature and the current language; wherein, the target feature segments corresponding to different conversion rounds are segmented based on the semantics of the initial audio feature.
[0100] The conversion module 30 is configured to determine at least one conversion language, generate a conversion audio matching the conversion language based on the current language and the target feature segment, and broadcast the conversion audio.
[0101] In one embodiment, the obtaining module 10 obtains the initial audio stream of the target object, and determines the initial audio feature corresponding to the initial audio stream and the current language corresponding to the initial audio stream, including: obtaining the initial audio stream of the target object and determining the initial audio feature corresponding to the initial audio stream; based on the current moment and the initial audio feature, obtaining a current window feature corresponding to a preset duration; and determining the current language from multiple candidate languages based on the current window feature.
[0102] In one embodiment, the obtaining module 10 determines the current language from multiple candidate languages based on the current window feature, including: obtaining a historical window feature adjacent to the current window feature and historical discrimination information corresponding to the historical window feature; wherein, the historical discrimination information includes historical scores matching each of the candidate languages; obtaining current discrimination information based on the current window feature; wherein, the current discrimination information includes current scores matching each of the candidate languages; determining a first weight matching the current discrimination information and a second weight matching the historical discrimination information, obtaining a target score of the current window feature matching each of the candidate languages based on the first weight, the second weight, the historical discrimination information and the current discrimination information, and determining the current language based on the target score.
[0103] In one embodiment, the processing module 20 obtains a target feature segment corresponding to the current conversion round based on the initial audio feature and the current language, including: inputting the initial audio feature and the current language into a trained semantic segmentation model to obtain semantic information corresponding to the initial audio feature; in response to the semantic information satisfying a preset segmentation condition, assigning a current segmentation identifier to the initial audio feature; wherein, the preset segmentation condition is related to the completeness of the semantic information; and determining the target feature segment corresponding to the current conversion round based on the historical segmentation identifier corresponding to the previous conversion round and the current segmentation identifier.
[0104] In one embodiment, the conversion module 30 generates a converted audio that matches the converted language based on the current language and the target feature segment, including: obtaining a corresponding speech-level representation by using an intelligent analysis model based on the target feature segment; wherein, the speech-level representation includes at least one of language information corresponding to the current language, timbre information matching the target object, prosody information, and speech rate information; generating the corresponding converted audio by using the intelligent analysis model based on the speech-level representation and the converted language; and broadcasting the converted audio.
[0105] In one embodiment, the conversion module 30 obtains a corresponding speech-level representation by using an intelligent analysis model based on the target feature segment, including: obtaining a reference feature based on at least one historical conversion round and the target feature segments respectively corresponding to the current conversion round; and inputting the reference feature into the intelligent analysis model to obtain the speech-level representation.
[0106] In one embodiment, the conversion module 30 generates the corresponding converted audio by using the intelligent analysis model based on the speech-level representation and the converted language, including: obtaining a converted language representation that matches the converted language; wherein, the converted language representation includes at least pronunciation information corresponding to the converted language; generating a task text based on the speech-level representation, the target feature segment, the converted language, and the converted language representation; wherein, the task text is used to prompt the intelligent analysis model to generate the converted audio that matches the converted language; and inputting the task text into the intelligent analysis model to obtain the converted audio generated by the intelligent analysis model.
[0107] In one embodiment, the acquisition module 10 acquires the initial audio stream of the target object and determines the initial audio feature corresponding to the initial audio stream, including: acquiring scene information that matches the target scene; wherein, the scene information includes scene video data and scene audio data; determining at least one target object in the target scene and the initial audio stream corresponding to the target object based on the scene information; and inputting the initial audio stream into a trained speech coding model to obtain the initial audio feature.
[0108] Please refer to Figure 8 , Figure 8It is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device includes a memory 40 and a processor 50 that are coupled to each other. Program instructions are stored in the memory 40, and the processor 50 is configured to execute the program instructions to implement the methods mentioned in any of the above embodiments. Specifically, the electronic device includes, but is not limited to, a desktop computer, a laptop computer, a tablet computer, a server, etc., which are not limited herein. In addition, the processor 50 may also be referred to as a CPU (Central Processing Unit). The processor 50 may be an integrated circuit chip with signal processing capabilities. The processor 50 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 50 may be implemented jointly by integrated circuit chips.
[0109] Please refer to Figure 9 , Figure 9 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Program instructions 70 that can be run by a processor are stored on the computer-readable storage medium 60. When the program instructions 70 are executed by the processor, the methods mentioned in any of the above embodiments are implemented.
[0110] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0111] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0112] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, may exist physically separately for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0114] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. An audio language conversion method, characterized in that, Including: Obtain the initial audio stream of the target object, and determine the initial audio features corresponding to the initial audio stream and the current language corresponding to the initial audio stream; Based on the initial audio features and the current language, obtain the target feature segment corresponding to the current conversion round; wherein, the target feature segments corresponding to different conversion rounds are segmented based on the semantics of the initial audio features; Determine at least one target language, and based on the current language and the target feature segment, generate a converted audio that matches the target language; Among them, the generating a converted audio that matches the target language based on the current language and the target feature segment includes: based on the target feature segment, using an intelligent analysis model to obtain a corresponding speech-level representation; wherein, the intelligent analysis model is a large language model with data analysis capabilities, and the speech-level representation includes at least one of language information corresponding to the current language, timbre information matching the target object, prosody information, and speech rate information; based on the speech-level representation and the target language, using the intelligent analysis model to generate the corresponding converted audio; broadcast the converted audio; wherein, the speech-level representation is obtained based on the target feature segments corresponding to at least one historical conversion round and the current conversion round respectively; Among them, the generating the corresponding converted audio using the intelligent analysis model based on the speech-level representation and the target language includes: obtaining a target language representation that matches the target language; wherein, the target language representation at least includes pronunciation information corresponding to the target language; based on the speech-level representation, the target feature segment, the target language, and the target language representation, generate a task text; wherein, the task text is used to prompt the intelligent analysis model to generate the converted audio that matches the target language; input the task text into the intelligent analysis model to obtain the converted audio generated by the intelligent analysis model.
2. The audio language conversion method according to claim 1, wherein The obtaining the initial audio stream of the target object, and determining the initial audio features corresponding to the initial audio stream and the current language corresponding to the initial audio stream includes: Obtain the initial audio stream of the target object, and determine the initial audio features corresponding to the initial audio stream; Based on the current time and the initial audio features, obtain the current window features corresponding to a preset duration; Based on the current window features, determine the current language from multiple candidate languages.
3. The audio language conversion method according to claim 2, wherein The determining the current language from multiple candidate languages based on the current window features includes: Obtain historical window features adjacent to the current window features and historical discrimination information corresponding to the historical window features; wherein, the historical discrimination information includes historical scores matching each of the candidate languages; Based on the current window features, obtain current discrimination information; wherein, the current discrimination information includes current scores matching each of the candidate languages; Determine a first weight that matches the current discrimination information and a second weight that matches the historical discrimination information. Based on the first weight, the second weight, the historical discrimination information, and the current discrimination information, obtain the target scores of the current window feature matching each of the candidate languages. Based on the target scores, determine the current language.
4. The audio language conversion method according to claim 1, wherein The obtaining, based on the initial audio feature and the current language, of the target feature segment corresponding to the current conversion round includes: Input the initial audio feature and the current language into the trained semantic segmentation model to obtain the semantic information corresponding to the initial audio feature; In response to the semantic information satisfying a preset segmentation condition, assign a current segmentation identifier to the initial audio feature; wherein the preset segmentation condition is related to the completeness of the semantic information; Based on the historical segmentation identifier corresponding to the previous conversion round and the current segmentation identifier, determine the target feature segment corresponding to the current conversion round.
5. The audio language conversion method according to claim 1, characterized in that The obtaining, based on the target feature segment, of the corresponding speech-level representation using the intelligent analysis model includes: Obtain a reference feature based on the target feature segments corresponding to at least one historical conversion round and the current conversion round respectively; Input the reference feature into the intelligent analysis model to obtain the speech-level representation.
6. The audio language conversion method according to claim 2, characterized in that The obtaining of the initial audio stream of the target object and the determination of the initial audio feature corresponding to the initial audio stream include: Obtain scene information that matches the target scene; wherein the scene information includes scene video data and scene audio data; Based on the scene information, determine at least one target object in the target scene and the initial audio stream corresponding to the target object; Input the initial audio stream into the trained speech coding model to obtain the initial audio feature.
7. An audio language conversion system, characterized in that, Includes: An acquisition module, configured to acquire the initial audio stream of the target object, determine the initial audio feature corresponding to the initial audio stream, and the current language corresponding to the initial audio stream; A processing module, configured to obtain the target feature segment corresponding to the current conversion round based on the initial audio feature and the current language; wherein the target feature segments corresponding to different conversion rounds are segmented based on the semantics of the initial audio feature; A conversion module, configured to determine at least one conversion language, generate a conversion audio that matches the conversion language based on the current language and the target feature segment, and broadcast the conversion audio; Among them, generating a converted audio that matches the target language based on the current language and the target feature segment includes: obtaining a corresponding speech-level representation based on the target feature segment by using an intelligent analysis model; wherein, the intelligent analysis model is a large language model with data analysis capabilities, and the speech-level representation includes at least one of language information corresponding to the current language, timbre information matching the target object, prosody information, and speech rate information; generating the corresponding converted audio based on the speech-level representation and the target language by using the intelligent analysis model; and playing the converted audio; wherein, the speech-level representation is obtained based on the target feature segments corresponding to at least one historical conversion round and the current conversion round respectively. Among them, generating the corresponding converted audio based on the speech-level representation and the target language by using the intelligent analysis model includes: obtaining a converted language representation that matches the target language; wherein, the converted language representation includes at least pronunciation information corresponding to the target language; generating a task text based on the speech-level representation, the target feature segment, the target language, and the converted language representation; wherein, the task text is used to prompt the intelligent analysis model to generate the converted audio that matches the target language; and inputting the task text into the intelligent analysis model to obtain the converted audio generated by the intelligent analysis model.
8. An electronic device, characterized in that, Including: A memory and a processor that are mutually coupled, wherein program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the audio language conversion method according to any one of claims 1-6.
9. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the audio language conversion method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Speech translation method and device and storage medium
CN113591495A
Language recognition method and device
CN113744717A
Live broadcast voice simultaneous transmission method and device, medium and electronic equipment
CN114554238A
Audio processing method and device, storage medium and electronic equipment
CN115798459A
Speech recognition method and device, electronic equipment and storage medium
CN117854507A