Speech translation method and system based on artificial intelligence

Through the speech translation method based on artificial intelligence, the CNN-CTC model and third-party translation software are used to solve the problems of complex speech translation process and low recognition rate in the existing technology, and efficient and accurate cross-language communication is achieved.

CN120218088AInactive Publication Date: 2025-06-27GUANGDONG AISHI INTELLIGENT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510306455.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the speech translation process is complex and the recognition rate is low, making it difficult to achieve efficient and accurate cross-language communication.

Method used

Using a speech translation method based on artificial intelligence, voice signals are collected and preprocessed, voice features are extracted, and voice features are identified using CNN-CTC model, and translated in combination with third-party translation software, ultimately real-time voice broadcasting is realized.

Benefits of technology

It improves the accuracy and efficiency of pronunciation translation, reduces the requirements for pronunciation feature extraction, and enhances the effect of pronunciation translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218088A_ABST
    Figure CN120218088A_ABST
Patent Text Reader

Abstract

The invention discloses a speech translation method and system based on artificial intelligence, and belongs to the technical field of speech data processing. Speech features corresponding to a target speech signal are obtained by extracting data features in the preprocessed target speech signal; recognizing voice features corresponding to the target voice signal by adopting an artificial intelligence model, obtaining first target text information corresponding to the target voice signal, calling third-party text information translation software to translate the first target text information, obtaining second target text information, and sending the second target text information to a server; and finally, broadcasting the second target text information, thereby realizing rapid translation, improving the accuracy of speech translation by means of strong data processing capability of artificial intelligence, reducing the requirement for speech feature extraction, and further improving the effect of speech translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech data processing, and particularly relates to a speech translation method and system based on artificial intelligence. Background Art

[0002] Speech translation is a method that uses artificial intelligence technology to achieve real-time communication between different languages. Through deep learning algorithms, the system first recognizes the input source language speech and converts it into text, and then performs semantic analysis to understand the content. Using a pre-trained neural network translation model, the text is translated into the target language, and finally, through speech synthesis technology, the target language speech is output to achieve efficient, accurate, and real-time cross-language communication, providing strong support for language communication in the context of globalization. However, in the prior art, after speech recognition is often performed using a hidden Markov model and then translation is carried out, there are problems of complex processes and low recognition rates. Summary of the Invention

[0003] The present invention provides a speech translation method and system based on artificial intelligence to solve the problems of complex processes and low recognition rates existing in the prior art.

[0004] On the one hand, the present invention provides a speech translation method based on artificial intelligence, including:

[0005] Collecting a target speech signal to be translated, and preprocessing the target speech signal to obtain a preprocessed target speech signal;

[0006] Extracting data features from the preprocessed target speech signal to obtain a speech feature corresponding to the target speech signal;

[0007] Using an artificial intelligence model to recognize the speech feature corresponding to the target speech signal to obtain first target text information corresponding to the target speech signal;

[0008] Invoking a third-party text information translation software to translate the first target text information to obtain second target text information; wherein, the second target text information represents information of the text type required by the user;

[0009] Based on the second target text information, performing real-time speech broadcast to complete the speech translation process based on artificial intelligence.

[0010] Further, collecting a target speech signal to be translated, and preprocessing the target speech signal to obtain a preprocessed target speech signal, including:

[0011] Collect the target voice signal to be translated, and perform pre-emphasis, framing, and windowing processing on the target voice signal once to obtain the pre-processed target voice signal; wherein, the pre-processed target voice signal includes multiple frames of sub-voice signals.

[0012] Further, extract the data features in the pre-processed target voice signal to obtain the voice features corresponding to the target voice signal, including:

[0013] After performing Fourier transform, taking the power spectrum, squaring the amplitude, Mel filter bank, and taking the logarithm processing on each frame of sub-voice signal in the pre-processed target voice signal in sequence, obtain the FBank features corresponding to each frame of sub-voice signal;

[0014] For multiple frames of sub-voice signals within each processing period, first perform mean normalization processing on the FBank features corresponding to each frame of sub-voice signal, and then splice them in chronological order to obtain the voice features corresponding to the target voice signal.

[0015] Further, use an artificial intelligence model to identify the voice features corresponding to the target voice signal, and obtain the first target text information corresponding to the target voice signal, including:

[0016] Before voice translation, first construct a CNN-CTC model, and train the CNN-CTC model to obtain the trained CNN-CTC model, and use the trained CNN-CTC model as the artificial intelligence model;

[0017] Use the obtained voice features corresponding to the target voice signal as the output of the artificial intelligence model to obtain the first target text information corresponding to the target voice signal.

[0018] Further, train the CNN-CTC model to obtain the trained CNN-CTC model, including:

[0019] Initialize the model parameters of the CNN-CTC model and perform encoding to obtain model parameter encodings, and repeat to obtain multiple model parameter encodings;

[0020] Obtain the loss function value corresponding to each model parameter encoding, and determine the model parameter encoding with the smallest loss function value as the optimal encoding;

[0021] Use the initial search strategy that combines the information of learning the optimal encoding with Levy flight to perform a search on the model parameter encodings once to obtain the first model parameter encoding;

[0022] Perform a secondary search on the first model parameter encoding using a multi-information fusion search strategy that combines adaptive adjustment of the search range and multi-point information fusion to obtain the second model parameter encoding;

[0023] Perform a tertiary search on the second model parameter encoding using a joint contraction search strategy that combines a non-linear convergence factor and position wrapping search to obtain the third model parameter encoding;

[0024] Perform a quaternary search on the third model parameter encoding using an adaptive global search strategy that combines the optimal position as a reference and multi-search states to obtain the fourth model parameter encoding;

[0025] Repeat the execution of the initial search strategy, multi-information fusion search strategy, joint contraction search strategy, and adaptive global search strategy until the training end condition is met, and re-determine the optimal encoding according to the fourth model parameter encoding;

[0026] Obtain the CNN-CTC model after training according to the re-determined optimal encoding.

[0027] Furthermore, perform a primary search on the model parameter encoding using an initial search strategy that combines the information of learning the optimal encoding and Levy flight to obtain the first model parameter encoding, including:

[0028]

[0029] λ = λ min +(λ max -λ min )e (-t / T)

[0030] where, represents the k-th model parameter encoding in the t-th training process, represents the k-th first model parameter encoding, L represents the random step size generated by Levy flight, λ represents the non-linear step size adjustment factor, λ min represents the minimum value of the non-linear step size adjustment factor, λ max represents the maximum value of the non-linear step size adjustment factor, e represents the natural constant, T represents the preset maximum number of training times, represents the optimal encoding, k = 1, 2, …, K, and K represents the total number of model parameter encodings.

[0031] Furthermore, perform a secondary search on the first model parameter encoding using a multi-information fusion search strategy that combines adaptive adjustment of the search range and multi-point information fusion to obtain the second model parameter encoding, including:

[0032]

[0033]

[0034] ω = ω max -(ω max -ω min )(t / T) 2

[0035] wherein, represents the nth first model parameter encoding in the t-th training process, represents the nth second model parameter encoding, represents the search speed of the nth first model parameter encoding in the (t + 1)-th training process, represents the search speed of the nth second model parameter encoding in the (t + 1)-th training process, ω represents the search range adjustment factor, ω max represents the preset maximum value of the search range adjustment factor, ω min represents the preset minimum value of the search range adjustment factor, T represents the preset maximum number of training times, C1 represents the second information fusion factor, C2 represents the second information fusion factor, C3 represents the third information fusion factor, R1 represents the first random number between (0, 1), R2 represents the second random number between (0, 1), R3 represents the third random number between (0, 1), represents the optimal encoding, represents the first model parameter encoding with the second smallest loss function value, represents the first model parameter encoding with the third smallest loss function value.

[0036] Furthermore, a combined contraction search strategy combining a non-linear convergence factor and position-wrapping search is adopted to perform three searches on the second model parameter encoding to obtain the third model parameter encoding, including:

[0037]

[0038]

[0039] wherein, represents the i-th second model parameter encoding in the t-th training process, represents the i-th third model parameter encoding, ξ represents the non-linear convergence factor, R4 represents the fourth random number between (0, 1), represents the optimal encoding, π represents pi, T represents the preset maximum number of training times.

[0040] Furthermore, an adaptive global search strategy combining the optimal position as a reference and multiple search states is adopted to perform four searches on the third model parameter encoding to obtain the fourth model parameter encoding, including:

[0041] Randomly generate a fifth random number between (0, 1), and determine whether the fifth random number is greater than a preset decision factor. If so, perform a first global search on the third model parameter encoding to obtain the global search state corresponding to the third model parameter encoding; otherwise, perform a second global search on the third model parameter encoding to obtain the global search state corresponding to the third model parameter encoding;

[0042] Determine whether the loss function value of the global search state is less than the loss function value of the corresponding third model parameter encoding. If so, use the global search state as the fourth model parameter encoding; otherwise, directly use the original third model parameter encoding as the fourth model parameter encoding;

[0043] The first global search is:

[0044]

[0045] where, represents the d-th dimension parameter of the m-th third model parameter encoding in the t-th training process, d = 1, 2,..., D, and D represents the total dimension of the parameters, represents the d-th dimension parameter of the global search state corresponding to the m-th third model parameter encoding, represents the d-th dimension parameter of the optimal encoding, η represents the global search adjustment factor, ε represents a preset constant term, and ε < 0.0001; T represents the preset maximum number of training times;

[0046] The second global search is:

[0047] On the other hand, the present invention provides an artificial intelligence-based speech translation system, including: a speech signal preprocessing module, a speech feature extraction module, an artificial intelligence recognition module, a text information conversion module, and a text information broadcasting module;

[0048] The speech signal preprocessing module is used to collect a target speech signal to be translated and preprocess the target speech signal to obtain a preprocessed target speech signal;

[0049] The speech feature extraction module is used to extract data features from the preprocessed target speech signal to obtain speech features corresponding to the target speech signal;

[0050] The artificial intelligence recognition module is used to use an artificial intelligence model to recognize the speech features corresponding to the target speech signal to obtain first target text information corresponding to the target speech signal;

[0051] The text information conversion module is used to call a third-party text information translation software to translate the first target text information, and obtain second target text information; wherein, the second target text information represents the information of the text type required by the user.

[0052] The text information broadcast module is used to perform real-time voice broadcast based on the second target text information, and complete the voice translation process based on artificial intelligence.

[0053] A voice translation method and system based on artificial intelligence provided by the present invention extracts data features in the target voice signal after preprocessing to obtain the voice features corresponding to the target voice signal, uses an artificial intelligence model to identify the voice features corresponding to the target voice signal, obtains the first target text information corresponding to the target voice signal, calls a third-party text information translation software to translate the first target text information, obtains the second target text information, and finally broadcasts the second target text information, thereby realizing fast translation. By virtue of the powerful data processing ability of artificial intelligence, the accuracy of voice translation is improved, and at the same time, the requirement for voice feature extraction is reduced, further improving the effect of voice translation. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0055] Figure 1 It is a flowchart of a voice translation method based on artificial intelligence provided by an embodiment of the present invention.

[0056] Figure 2 It is a schematic structural diagram of a voice translation system based on artificial intelligence provided by an embodiment of the present invention.

[0057] Through the above drawings, specific embodiments of the present invention have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the inventive concept in any way, but to illustrate the concept of the present invention to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are only examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0059] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0060] As Figure 1 shown, an embodiment of the present invention provides a speech translation method based on artificial intelligence, including:

[0061] S11. Collect the target speech signal to be translated, and preprocess the target speech signal to obtain the target speech signal after preprocessing;

[0062] The preprocessing can be to denoise the target speech signal, so as to ensure that the extracted speech features are more easily recognized and improve the accuracy of speech recognition.

[0063] S12. Extract the data features in the target speech signal after preprocessing to obtain the speech features corresponding to the target speech signal;

[0064] The speech features corresponding to the target speech signal can be FBank features. However, it should be noted that the FBank features are only the preferred features of the embodiments of the present invention, and other speech features can also be used for recognition.

[0065] S13. Use an artificial intelligence model to recognize the speech features corresponding to the target speech signal to obtain the first target text information corresponding to the target speech signal;

[0066] The first target text information refers to the text information corresponding to the original language type in the target speech signal. By using an artificial intelligence model for recognition, the recognition accuracy can be effectively improved, thereby achieving accurate translation.

[0067] S14. Call a third-party text information translation software to translate the first target text information to obtain the second target text information; wherein, the second target text information represents the information of the text type required by the user;

[0068] S15. Based on the second target text information, perform real-time voice broadcast to complete the speech translation process based on artificial intelligence.

[0069] In the embodiments of the present invention, by using an artificial intelligence model for speech recognition and translation, the speech recognition accuracy can be effectively improved, and finally the translation accuracy can be improved.

[0070] In the embodiments of the present invention, collecting the target speech signal to be translated and preprocessing the target speech signal to obtain the target speech signal after preprocessing includes:

[0071] Collect the target voice signal to be translated, and perform pre-emphasis, framing, and windowing processing on the target voice signal once to obtain the target voice signal after preprocessing; wherein, the target voice signal after preprocessing includes multiple frames of sub-voice signals.

[0072] In the embodiment of the present invention, extracting the data features in the target voice signal after preprocessing to obtain the voice features corresponding to the target voice signal includes:

[0073] After performing Fourier transform, power spectrum extraction, amplitude squaring, Mel filter bank, and logarithm processing on each frame of sub-voice signal in the target voice signal after preprocessing in sequence, the FBank feature corresponding to each frame of sub-voice signal is obtained;

[0074] For multiple frames of sub-voice signals within each processing period (so as to ensure that the lengths of the voice features corresponding to the target voice signal are the same, which is convenient for artificial intelligence to process), the FBank features corresponding to each frame of sub-voice signal are first subjected to mean normalization processing, and then spliced in chronological order to obtain the voice features corresponding to the target voice signal.

[0075] For example: Extracting the FBank feature map from the voice signals in the training set may include: First, frame the voice signals with a frame length of 25 ms and a frame shift of 10 ms. During the framing process, in order to reduce spectral leakage, the Hamming window is used as the truncation function; then, a Mel filter bank containing 40 triangular band-pass filters is used to extract the 40-dimensional Fbank features of each frame of signal; finally, after performing mean normalization operations on the obtained FBank features of each frame and splicing them in chronological order, the conversion from the feature vector to the feature map is completed.

[0076] It should be noted that after extracting the first target text information, the repeated parts can also be removed in chronological order to ensure data coherence.

[0077] In the embodiment of the present invention, using an artificial intelligence model to identify the voice features corresponding to the target voice signal to obtain the first target text information corresponding to the target voice signal includes:

[0078] Before voice translation, first construct a CNN-CTC (Convolutional Neural Network-Connectionist Temporal Classification) model, and train the CNN-CTC model to obtain the trained CNN-CTC model, and use the trained CNN-CTC model as the artificial intelligence model;

[0079] Use the obtained voice features corresponding to the target voice signal as the output of the artificial intelligence model to obtain the first target text information corresponding to the target voice signal.

[0080] Under the CTC end-to-end framework, the convolutional neural network has greater application potential, mainly reflected in two aspects: First, in the selection of input features, compared with the HMM (Hidden Markov Model) framework that can only use a fixed number of adjacent frames for splicing, CTC can splice all the feature vectors of the entire speech signal into a large feature map as the input. Through the stacking of the CNN convolutional layer and pooling layer, the "receptive field" of the output layer neurons is improved as the number of layers deepens, so that the final output of the network can capture more context information between speech frames. Second, in the selection of modeling units, since CTC introduces the "blank" label and does not require strict classification of the categories of each speech frame, larger-grained modeling units such as syllables and Chinese characters can be selected. Moreover, the unique convolution and downsampling operations of CNN not only compress the input feature map but also integrate and filter the speech frame data, and can extract more essential features from the speech signals with larger-grained modeling units containing more complex information.

[0081] In an embodiment of the present invention, training the CNN-CTC model to obtain the trained CNN-CTC model includes:

[0082] Initializing the model parameters of the CNN-CTC model and performing encoding to obtain model parameter encodings, and repeating to obtain multiple model parameter encodings;

[0083] For example, it can be randomly initialized between the upper and lower limits of the model parameters, and then the model parameters are encoded as vectors to obtain model parameter encodings. After repeating multiple times, multiple different model parameter encodings are obtained.

[0084] Obtaining the loss function value (such as the CTC loss function value) corresponding to each model parameter encoding, and determining the model parameter encoding with the smallest loss function value as the optimal encoding; it should be noted that existing data sets can be used for training.

[0085] Performing a first search on the model parameter encodings using an initial search strategy that combines the information of learning the optimal encoding with Levy flight to obtain the first model parameter encoding;

[0086] Performing a second search on the first model parameter encoding using a multi-information fusion search strategy that combines adaptive adjustment of the search range and multi-point information fusion to obtain the second model parameter encoding;

[0087] Performing a third search on the second model parameter encoding using a joint contraction search strategy that combines a non-linear convergence factor and position wrapping search to obtain the third model parameter encoding;

[0088] The fourth model parameter encoding is obtained by performing four searches on the third model parameter encoding using an adaptive global search strategy that combines the optimal position as a reference with multiple search states;

[0089] The initial search strategy, multi-information fusion search strategy, joint contraction search strategy, and adaptive global search strategy are repeatedly executed until the training end condition is met, and the optimal encoding is re-determined according to the fourth model parameter encoding;

[0090] According to the re-determined optimal encoding, the CNN-CTC model after training is obtained.

[0091] The prior art generally uses a gradient descent optimization algorithm to train the model parameters of an artificial intelligence model, which has problems such as poor effects and often falling into local optimal values, ultimately resulting in poor speech recognition ability. Therefore, the embodiments of the present invention provide a training algorithm to improve the data learning ability and ultimately improve the speech recognition and speech translation abilities.

[0092] In the embodiments of the present invention, an initial search strategy that combines the information of learning the optimal encoding with Levy flight is used to perform a search on the model parameter encoding once to obtain the first model parameter encoding, including:

[0093]

[0094] λ = λ min +(λ max -λ min )e (-t / T)

[0095] where, represents the k-th model parameter encoding in the t-th training process, represents the k-th first model parameter encoding, L represents the random step size generated by Levy flight, λ represents the non-linear step size adjustment factor, λ min represents the minimum value of the non-linear step size adjustment factor, λ max represents the maximum value of the non-linear step size adjustment factor, e represents the natural constant, T represents the preset maximum number of training times, represents the optimal encoding, k = 1, 2,..., K, and K represents the total number of model parameter encodings.

[0096] The single search provided by the embodiments of the present invention can learn the optimal position information, and then introduce Levy flight and the non-linear step size adjustment factor, which can not only provide a certain degree of global search ability in the middle and early stages of the algorithm, but also turn into a fine search in the later stage of the algorithm.

[0097] In an embodiment of the present invention, a multi-information fusion search strategy combining adaptive adjustment of the search range and multi-point information fusion is adopted to perform a secondary search on the first model parameter encoding to obtain the second model parameter encoding, including:

[0098]

[0099]

[0100] ω = ω max -(ω max -ω min )(t / T) 2

[0101] Wherein, represents the nth first model parameter encoding in the tth training process, represents the nth second model parameter encoding, represents the search speed of the nth first model parameter encoding in the (t + 1)th training process, represents the search speed of the nth second model parameter encoding in the (t + 1)th training process, ω represents the search range adjustment factor, ω max represents the preset maximum value of the search range adjustment factor, ω min represents the preset minimum value of the search range adjustment factor, T represents the preset maximum number of training times, C1 represents the second information fusion factor (such as a constant between (0, 0.2), but it should be noted that other information fusion factors can also be set to numbers in the same interval or in a smaller interval), C2 represents the second information fusion factor, C3 represents the third information fusion factor, R1 represents the first random number between (0, 1), R2 represents the second random number between (0, 1), R3 represents the third random number between (0, 1), represents the optimal encoding, represents the first model parameter encoding with the second smallest loss function value, represents the first model parameter encoding with the third smallest loss function value.

[0102] In the secondary search of the embodiment of the present invention, the optimal several position information can be combined for the search, and the search range adjustment factor is introduced to adaptively adjust the search range, so that the encoding can not only search in a better area but also not fall into the local optimum, effectively improving the search effect.

[0103] In an embodiment of the present invention, a combined contraction search strategy combining a non-linear convergence factor and position wrapping search is adopted to perform a tertiary search on the second model parameter encoding to obtain the third model parameter encoding, including:

[0104]

[0105]

[0106] Among them, represents the encoding of the i-th second model parameter in the t-th training process, represents the encoding of the i-th third model parameter, ξ represents a non-linear convergence factor, and R4 represents a fourth random number between (0, 1), represents the optimal encoding, π represents pi, and T represents the preset maximum number of training times.

[0107] The three-time search provided by the embodiments of the present invention can perform adaptive position selection around the optimal position. All model parameter encodings perform an encircling search around the optimal position, improving the search speed while enhancing the local space search ability.

[0108] In the embodiments of the present invention, a fourth model parameter encoding is obtained by performing four-time search on the third model parameter encoding using an adaptive global search strategy that combines the optimal position as a reference with multiple search states, including:

[0109] Randomly generate a fifth random number between (0, 1), and determine whether the fifth random number is greater than a preset decision factor. If so, perform a first global search on the third model parameter encoding to obtain the global search state corresponding to the third model parameter encoding; otherwise, perform a second global search on the third model parameter encoding to obtain the global search state corresponding to the third model parameter encoding;

[0110] Determine whether the loss function value of the global search state is less than the loss function value of the corresponding third model parameter encoding. If so, use the global search state as the fourth model parameter encoding; otherwise, directly use the original third model parameter encoding as the fourth model parameter encoding;

[0111] The first global search is:

[0112]

[0113] Among them, represents the d-th dimension parameter of the m-th third model parameter encoding in the t-th training process, d = 1, 2,..., D, and D represents the total dimension of the parameters, represents the d-th dimension parameter of the global search state corresponding to the m-th third model parameter encoding, represents the d-th dimension parameter of the optimal encoding, η represents the global search adjustment factor, ε represents a preset constant term, and ε < 0.0001; T represents the preset maximum number of training times;

[0114] The second global search is:

[0115] The four - time search provided by the embodiments of the present invention can provide powerful global search capabilities. At the same time, a greedy strategy is introduced for search control, which can effectively ensure the search speed of the algorithm.

[0116] In summary, the training algorithm provided by the embodiments of the present invention, through the mutual cooperation of several search processes, improves the traversability of the solution space, enhances the global search ability, strengthens the training effect, and thus improves the accuracy of speech recognition and speech translation.

[0117] Optionally, after each search, out - of - bounds processing can also be performed on the model parameters to ensure the validity of the model parameters.

[0118] A speech translation method based on artificial intelligence provided by the present invention extracts data features from the pre - processed target speech signal to obtain the speech features corresponding to the target speech signal, uses an artificial intelligence model to recognize the speech features corresponding to the target speech signal to obtain the first target text information corresponding to the target speech signal, calls a third - party text information translation software to translate the first target text information to obtain the second target text information, and finally broadcasts the second target text information, thus achieving fast translation. By virtue of the powerful data - processing ability of artificial intelligence, the accuracy of speech translation is improved, and at the same time, the requirement for speech feature extraction is reduced, further improving the effect of speech translation.

[0119] As Figure 2 shown, an artificial - intelligence - based speech translation system provided by the embodiments of the present invention includes: a speech signal pre - processing module 21, a speech feature extraction module 22, an artificial intelligence recognition module 23, a text information conversion module 24, and a text information broadcasting module 25;

[0120] The speech signal pre - processing module 21 is used to collect the target speech signal to be translated and pre - process the target speech signal to obtain the pre - processed target speech signal;

[0121] The speech feature extraction module 22 is used to extract data features from the pre - processed target speech signal to obtain the speech features corresponding to the target speech signal;

[0122] The artificial intelligence recognition module 23 is used to use an artificial intelligence model to recognize the speech features corresponding to the target speech signal to obtain the first target text information corresponding to the target speech signal;

[0123] The text information conversion module 24 is used to call a third - party text information translation software to translate the first target text information to obtain the second target text information; wherein, the second target text information represents the information of the text type required by the user.

[0124] The text information broadcast module 25 is configured to perform real-time voice broadcast based on the second target text information, thereby completing the voice translation process based on artificial intelligence.

[0125] The voice translation system based on artificial intelligence provided by the embodiments of the present invention can execute the above voice translation method, and its principle and effect are similar, so details are not described herein again.

[0126] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. It should be understood that the present invention is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A speech translation method based on artificial intelligence, characterized in that: include: Collecting a target speech signal to be translated, and preprocessing the target speech signal to obtain the preprocessed target speech signal; Extracting data features from the target speech signal after the preprocessing to obtain speech features corresponding to the target speech signal; Using an artificial intelligence model to identify the speech feature corresponding to the target speech signal, and obtaining first target text information corresponding to the target speech signal; Invoking a third-party text information translation software to translate the first target text information to obtain second target text information; wherein the second target text information represents information of a text type required by the user; Based on the second target text information, real-time voice broadcast is performed to complete the artificial intelligence-based voice translation process.

2. The artificial intelligence-based speech translation method according to claim 1, characterized in that: Collecting a target speech signal to be translated, and preprocessing the target speech signal to obtain the preprocessed target speech signal, including: A target speech signal to be translated is collected, and the target speech signal is pre-emphasized, framed and windowed once to obtain a pre-processed target speech signal; wherein the pre-processed target speech signal includes a plurality of frame sub-speech signals.

3. The artificial intelligence-based speech translation method according to claim 2, characterized in that: Extracting data features from the target speech signal after the preprocessing to obtain speech features corresponding to the target speech signal includes: After performing Fourier transform, power spectrum, amplitude square, Mel filter bank and logarithm processing on each frame of sub-speech signal in the target speech signal after the preprocessing, the FBank feature corresponding to each frame of sub-speech signal is obtained; For the multiple frames of sub-speech signals in each processing cycle, the FBank features corresponding to each frame of sub-speech signals are first mean-normalized, and then spliced ​​in chronological order to obtain the speech features corresponding to the target speech signal.

4. The artificial intelligence-based speech translation method according to claim 3, characterized in that: Using an artificial intelligence model to identify the speech feature corresponding to the target speech signal to obtain first target text information corresponding to the target speech signal includes: Before speech translation, a CNN-CTC model is first constructed, and the CNN-CTC model is trained to obtain a trained CNN-CTC model, and the trained CNN-CTC model is used as an artificial intelligence model; The speech features corresponding to the obtained target speech signal are used as the output of the artificial intelligence model to obtain the first target text information corresponding to the target speech signal.

5. The artificial intelligence-based speech translation method according to claim 4, characterized in that: The CNN-CTC model is trained to obtain a trained CNN-CTC model, including: Initializing and encoding the model parameters of the CNN-CTC model, obtaining the model parameter encoding, and repeatedly obtaining multiple model parameter encodings; Obtain the loss function value corresponding to each model parameter encoding, and determine the model parameter encoding with the smallest loss function value as the optimal encoding; An initial search strategy combining the information of the optimal coding with Levy flight is used to search the model parameter coding once to obtain the first model parameter coding; A multi-information fusion search strategy combining adaptively adjusting the search range with multi-point information fusion is used to perform a secondary search on the first model parameter code to obtain a second model parameter code; The second model parameter encoding is searched three times by using a joint contraction search strategy combining a nonlinear convergence factor and a position surround search to obtain a third model parameter encoding; The third model parameter encoding is searched four times by using an adaptive global search strategy combining the optimal position as a reference and multiple search states to obtain a fourth model parameter encoding; Repeat the initial search strategy, the multi-information fusion search strategy, the joint contraction search strategy, and the adaptive global search strategy until the training end condition is met, and re-determine the optimal encoding according to the fourth model parameter encoding; According to the re-determined optimal encoding, the trained CNN-CTC model is obtained.

6. The artificial intelligence-based speech translation method according to claim 5, characterized in that: The model parameter encoding is searched once using the initial search strategy combining the information of the optimal encoding with the Levy flight to obtain the first model parameter encoding, including: λ=λ min +(λ max -l min )e (-t / T) in, represents the kth model parameter encoding during the tth training process, represents the kth first model parameter encoding, L represents the random step size generated by Levy flight, λ represents the nonlinear step size adjustment factor, λ min represents the minimum value of the nonlinear step size adjustment factor, λ max represents the maximum value of the nonlinear step size adjustment factor, e represents a natural constant, T represents the preset maximum number of training times, represents the optimal coding, k=1,2,…,K, K represents the total number of model parameter codings.

7. The artificial intelligence-based speech translation method according to claim 6, characterized in that: A multi-information fusion search strategy combining adaptively adjusting the search range with multi-point information fusion is used to perform a secondary search on the first model parameter code to obtain the second model parameter code, including: oh = oh max -(oh max -oh min (t / T) 2 in, represents the nth first model parameter encoding during the tth training process, represents the nth second model parameter encoding, represents the search speed of the nth first model parameter encoding during the t+1th training process, represents the search speed of the nth second model parameter encoding during the t+1th training process, ω represents the search range adjustment factor, ω max Represents the preset maximum value of the search range adjustment factor, ω min represents the preset minimum value of the search range adjustment factor, T represents the preset maximum number of training times, C1 represents the second information fusion factor, C2 represents the second information fusion factor, C3 represents the third information fusion factor, R1 represents the first random number between (0,1), R2 represents the second random number between (0,1), and R3 represents the third random number between (0,1). represents the optimal encoding, Indicates the first model parameter encoding with the second smallest loss function value, Indicates the first model parameter encoding with the third smallest loss function value.

8. The artificial intelligence-based speech translation method according to claim 7, characterized in that: The second model parameter encoding is searched three times using a joint contraction search strategy combining a nonlinear convergence factor and a position surround search to obtain a third model parameter encoding, including: in, represents the i-th second model parameter encoding during the t-th training process, represents the i-th third model parameter encoding, ξ represents the nonlinear convergence factor, R4 represents the fourth random number between (0,1), represents the optimal encoding, π represents the circumference of a circle, and T represents the preset maximum number of training times.

9. The artificial intelligence-based speech translation method according to claim 8, characterized in that: The third model parameter encoding is searched four times using an adaptive global search strategy combining the optimal position as a reference and multiple search states to obtain a fourth model parameter encoding, including: A fifth random number is randomly generated between (0, 1), and it is determined whether the fifth random number is greater than a preset decision factor; if so, a first global search is performed on the third model parameter code to obtain a global search state corresponding to the third model parameter code; otherwise, a second global search is performed on the third model parameter code to obtain a global search state corresponding to the third model parameter code; Determine whether the loss function value of the global search state is less than the loss function value of the corresponding third model parameter code. If so, use the global search state as the fourth model parameter code; otherwise, use the original third model parameter code directly as the fourth model parameter code. The first global search is: in, represents the d-th dimension parameter of the m-th third model parameter encoding in the t-th training process, d = 1, 2, ..., D, D represents the total dimension of the parameter, represents the d-th dimension parameter of the global search state corresponding to the m-th third model parameter encoding, represents the d-th dimension parameter of the optimal encoding, η represents the global search adjustment factor, ε represents the preset constant term, and ε<0.0001; T represents the preset maximum number of training times; The second global search is:

10. A speech translation system based on artificial intelligence, characterized in that: include: Speech signal preprocessing module, speech feature extraction module, artificial intelligence recognition module, text information conversion module and text information broadcasting module; The speech signal preprocessing module is used to collect the target speech signal to be translated, and preprocess the target speech signal to obtain the target speech signal after preprocessing; The speech feature extraction module is used to extract data features from the target speech signal after the preprocessing to obtain speech features corresponding to the target speech signal; The artificial intelligence recognition module is used to recognize the voice features corresponding to the target voice signal using an artificial intelligence model to obtain the first target text information corresponding to the target voice signal; The text information conversion module is used to call a third-party text information translation software to translate the first target text information to obtain second target text information; wherein the second target text information represents information of the text type required by the user; The text information broadcasting module is used to perform real-time voice broadcasting based on the second target text information to complete the voice translation process based on artificial intelligence.

Citation Information

Cited By

  • Signal processing method and system based on deep learning

    CN120856513A