Speech processing method and system, electronic device, and storage medium

WO2026200245A1PCT designated stage Publication Date: 2026-10-01ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/074031
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-01-21
Publication Date
2026-10-01

Smart Images

  • Figure CN2026074031_01102026_PF_FP_ABST
    Figure CN2026074031_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the fields of large model technology and data processing, and relates to a speech processing method and system, an electronic device, and a storage medium. The method comprises: acquiring speech to be translated; retrieving, from a preset storage area, original terminology information corresponding to the speech to be translated, the preset storage area being used for storing multimodal information of a plurality of candidate terms according to a preset data structure; and, on the basis of the original terminology information, performing associative translation of the speech to be translated, and obtaining a translation result. The present disclosure solves the technical problems in the related art of low translation efficiency and poor accuracy when performing speech translation.
Need to check novelty before this filing date? Find Prior Art

Description

Speech processing methods, systems, electronic devices and storage media Technical Field

[0001] This disclosure relates to large model technology and data processing, and more specifically, to a speech processing method, system, electronic device, and storage medium. Background Technology

[0002] In today's increasingly frequent cross-language communication landscape, end-to-end speech translation technology, with its convenience and efficiency, is becoming an important tool for breaking down language barriers. Accurate transmission of proper nouns and neologisms is crucial for communication in professional fields or specific scenarios. Traditional speech translation methods can achieve direct conversion from speech to text; however, when faced with proper nouns or neologisms, they often suffer from translation illusions due to a lack of contextual understanding, leading to erroneous understanding and translation of these words. This severely undermines the reliability and value of speech translation in practical applications. Related technologies can collect and integrate all relevant terminology translations through statistical and interventional methods, but this approach suffers from problems such as introducing a large amount of irrelevant information, limited terminology coverage, and modal inconsistencies. Other technologies can leverage the model's contextual learning capabilities by providing similar speech translation samples as examples through retrieval and example methods. However, inconsistent retrieval granularity, the inclusion of irrelevant parts, and the inability to effectively correlate audio features negatively impact translation quality and model performance. Therefore, these technologies suffer from low translation efficiency and poor accuracy, limiting the widespread application value of end-to-end speech translation technology in specific fields and scenarios.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This disclosure provides a speech processing method, system, electronic device, and storage medium to at least solve the technical problems of low translation efficiency and poor accuracy in speech translation in related technologies.

[0005] According to one aspect of the present disclosure, a speech processing method is provided, comprising: acquiring speech to be translated; retrieving original terminology information corresponding to the speech to be translated from a preset storage area, wherein the preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure; and performing associative translation on the speech to be translated based on the original terminology information to obtain a translation result.

[0006] According to another aspect of the embodiments of this disclosure, a speech processing method is also provided, including: acquiring a court hearing recording to be translated; retrieving original legal terminology information corresponding to the court hearing recording to be translated from a preset storage area, wherein the preset storage area is used to store multimodal information of multiple candidate legal terms according to a preset data structure; and performing associative translation on the court hearing recording to be translated based on the original legal terminology information to obtain a translation result.

[0007] According to another aspect of the embodiments of this disclosure, a voice processing method is also provided, comprising: obtaining a voice processing request through a first application programming interface, wherein the request data carried in the voice processing request includes: voice to be translated; and returning a voice processing response through a second application programming interface, wherein the response data carried in the voice processing response includes: a translation result, wherein the translation result is generated according to any one of the voice processing methods in the embodiments of this disclosure.

[0008] According to another aspect of the embodiments of this disclosure, a voice processing method is also provided, comprising: acquiring a currently input voice processing dialogue request, wherein the request data carried in the voice processing dialogue request includes: voice to be translated; responding to the voice processing dialogue request, returning a voice processing dialogue response, wherein the information carried in the voice processing dialogue response includes: a translation result, the translation result being generated according to any one of the voice processing methods in the embodiments of this disclosure; and displaying the translation result in a graphical user interface.

[0009] According to another aspect of the embodiments of this disclosure, a speech processing method is also provided, comprising: responding to an input command applied to an operation interface, selecting speech to be translated on the operation interface; responding to a processing command applied to the operation interface, displaying a translation result on the operation interface; wherein the translation result is generated according to any one of the speech processing methods in the embodiments of this disclosure.

[0010] According to another aspect of the embodiments of this disclosure, a speech processing system is also provided, including: a client for sending speech to be translated; a server connected to the client for retrieving original terminology information corresponding to the speech to be translated from a preset storage area, and performing associative translation on the speech to be translated based on the original terminology information to obtain a translation result, wherein the preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure; the client is also used to output the translation result.

[0011] According to another aspect of the present disclosure, a computing device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of the present disclosure when it runs.

[0012] According to another aspect of the embodiments of this disclosure, an electronic device is also provided, including: a memory storing an executable program; and a processor connected to the memory via a bus for running the program, wherein the program executes the methods in the various embodiments of this disclosure when it runs.

[0013] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is executed, it controls the device where the computer-readable storage medium is located to perform the methods of the various embodiments of the present disclosure.

[0014] According to another aspect of the embodiments of this disclosure, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this disclosure.

[0015] According to another aspect of the embodiments of this disclosure, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods of various embodiments of this disclosure.

[0016] According to another aspect of the embodiments of this disclosure, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this disclosure.

[0017] In this embodiment, by acquiring the speech to be translated, the original terminology information corresponding to the speech is retrieved from a preset storage area. The preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure. Finally, the speech to be translated is correlated based on the original terminology information to obtain the translation result. This allows for the retrieval of original terminology information related to the speech to be translated from the preset storage area, effectively avoiding the introduction of a large amount of irrelevant information, thereby reducing errors and misleading information in the speech translation process, and further improving translation efficiency and quality. The multimodal information of multiple candidate terms stored in the preset storage area further enriches the translation basis in the correlated translation process, helping to more accurately understand the context and meaning of professional terms in the speech to be translated, further improving the relevance and efficiency of the translation process, thus solving the technical problems of low translation efficiency and poor accuracy in related technologies when performing speech translation.

[0018] It is worth noting that the above general description and the following detailed description are merely for illustrative and explanatory purposes and do not constitute a limitation thereof. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:

[0020] Figure 1 is a schematic diagram of a speech translation method based on related technologies;

[0021] Figure 2 is a schematic diagram of another speech translation method based on related technologies;

[0022] Figure 3 is a schematic diagram of an application scenario of a speech processing method according to an embodiment of the present disclosure;

[0023] Figure 4 is a flowchart of a speech processing method according to an embodiment of the present disclosure;

[0024] Figure 5 is a schematic diagram of a speech processing method according to an embodiment of the present disclosure;

[0025] Figure 6 is a schematic diagram of another speech processing method according to an embodiment of the present disclosure;

[0026] Figure 7 is a flowchart of another speech processing method according to an embodiment of the present disclosure;

[0027] Figure 8 is a flowchart of another speech processing method according to an embodiment of the present disclosure;

[0028] Figure 9 is a flowchart of another speech processing method according to an embodiment of the present disclosure;

[0029] Figure 10 is a flowchart of another speech processing method according to an embodiment of the present disclosure;

[0030] Figure 11 is a structural block diagram of a speech processing device according to an embodiment of the present disclosure;

[0031] Figure 12 is a structural block diagram of another voice processing device according to an embodiment of the present disclosure;

[0032] Figure 13 is a structural block diagram of another voice processing device according to an embodiment of the present disclosure;

[0033] Figure 14 is a structural block diagram of another voice processing device according to an embodiment of the present disclosure;

[0034] Figure 15 is a structural block diagram of another voice processing device according to an embodiment of the present disclosure;

[0035] Figure 16 is a structural block diagram of a computing device according to an embodiment of the present disclosure;

[0036] Figure 17 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0037] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.

[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0039] The technical solution disclosed herein is primarily implemented using large-scale model technology. Here, "large-scale model" refers to a deep learning model with a massive number of parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of parameters. Large-scale models, also known as foundation models, are pre-trained using large-scale unlabeled corpora to produce pre-trained models with hundreds of millions of parameters. These models are adaptable to a wide range of downstream tasks and exhibit good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.

[0040] It should be noted that, in practical applications, large models can be fine-tuned using a small number of samples to adapt them to different tasks. For example, large models can be widely used in Natural Language Processing (NLP), computer vision, and speech processing. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. Therefore, the main application scenarios for large models include, but are not limited to, digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design. In this embodiment, the example of data processing using a target speech translation model in a speech processing scenario is used for explanation.

[0041] In the field of end-to-end speech translation, especially when dealing with proper nouns and neologisms, related technologies mainly employ two strategies: collect-and-integrate and retrieve-and-demonstrate. However, these methods have exposed some key technical problems in practical applications, affecting translation quality and efficiency.

[0042] Figure 1 illustrates a speech translation method based on related technologies. As shown in Figure 1, the statistical and interventional approach collects all relevant terminology dictionaries through data statistics and uses these dictionaries as external knowledge to warm up the model's context before translation, thereby improving the model's understanding and translation capabilities. However, this approach easily introduces a large number of irrelevant terms, causing the model's attention to be distracted and reducing translation accuracy. The size of the terminology dictionary is also difficult to balance; a dictionary that is too small has limited coverage, while a dictionary that is too large places excessive demands on the model's long text generation capabilities. Furthermore, this approach suffers from modal inconsistency; the dictionary is text-modal, while the input to be translated is audio-modal, making it difficult for the model to effectively utilize cross-modal information for associative translation.

[0043] Figure 2 illustrates another speech translation method based on related technologies. As shown in Figure 2, the retrieval and example method retrieves other speech translation samples similar to the current audio and containing common terms, providing them as examples to the model. This allows the model to leverage its context learning capabilities to assist in translation. However, since the retrieval granularity is at the sentence level while the retrieval target is at the term level, the inconsistent granularity leads to poor retrieval results. The examples ultimately provided to the model contain a large amount of sentence information unrelated to the terms, further increasing the redundancy and complexity of the translation. Furthermore, the retrieved audio differs from the current audio source, such as in timbre and intonation, making it difficult for the model to effectively align and utilize audio features, thus affecting the accuracy of term translation.

[0044] Therefore, related technologies suffer from low translation efficiency and poor accuracy when performing speech translation, which limits the widespread application value of end-to-end speech translation technology in specific fields and scenarios.

[0045] According to embodiments of this disclosure, a voice processing method is provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0046] Considering the large number of model parameters in the large model and the limited computing resources of mobile terminals, the method provided in this disclosure can be applied to the application scenario shown in Figure 3, but is not limited thereto. In the application scenario shown in Figure 3, the large model is deployed on server 10. Server 10 can connect to one or more client devices 20 via a local area network (LAN), wide area network (WAN), Internet, or other types of data networks. These client devices 20 may include, but are not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users through a graphical user interface to invoke the large model, thereby implementing the method provided in this disclosure.

[0047] In this embodiment, the system consisting of a client device and a server can perform the following steps: the client can send the speech to be translated to the server; after receiving the speech, the server retrieves the original terminology information corresponding to the speech from a preset storage area, and performs associative translation based on the original terminology information to obtain the translation result. The preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure. Finally, the translation result is output on the client's graphical user interface.

[0048] It should be noted that with the rapid development of high-performance computing units, the methods provided in this disclosure can also be applied to integrated model machines in other application scenarios. In one optional embodiment, the integrated model machine has multiple built-in models. Users can select one model to adjust as needed to obtain their own model. The high-performance computing unit built into the integrated model machine can then directly call the adjusted model to execute the methods provided in this disclosure. In another optional embodiment, the large integrated model machine has a pre-trained model built-in. Therefore, the high-performance computing unit built into the integrated model machine can directly call this model to execute the methods provided in this disclosure.

[0049] Furthermore, when users need to train their own models, they can upload their own datasets via the client. These datasets are then sent to the server, allowing the server to adjust the pre-trained model using the dataset to obtain the user's customized model, which can then be deployed to the production environment. To facilitate users' model adjustment needs, the server provides complete adjustment tools, development frameworks, and processes, supporting multiple adjustment strategies. This allows the adjusted model to better adapt to different application domains and achieve a high degree of customization.

[0050] Under the above operating environment, this disclosure provides a speech processing method as shown in Figure 4. Figure 4 is a flowchart of a speech processing method according to an embodiment of this disclosure. As shown in Figure 4, the method may include the following steps:

[0051] Step S41: Obtain the speech to be translated;

[0052] Step S42: Retrieve the original terminology information corresponding to the speech to be translated from the preset storage area, wherein the preset storage area is used to store the multimodal information of multiple candidate terms according to a preset data structure;

[0053] Step S43: Based on the original terminology information, perform associative translation on the speech to be translated to obtain the translation result.

[0054] The aforementioned speech to be translated can be user input or system-received audio signals that need to be converted into another language. The speech to be translated is the starting point for the speech translation task, containing the speaker's voice information, including but not limited to pronunciation, intonation, speech rate, and the audio representation of the content being spoken. In an end-to-end speech translation system, the speech to be translated is the raw input data for achieving speech-to-text translation.

[0055] There are various ways to acquire the audio to be translated, covering multiple scenarios from direct user interaction to system backend processing. Different acquisition methods correspond to different application scenarios and technical requirements, but for any acquisition method, ensuring audio quality and clarity is key to improving translation results.

[0056] For example, the methods for obtaining the speech to be translated include, but are not limited to: direct user input, pre-recorded audio files, real-time audio streams, and integration with mobile applications or devices. When obtaining the speech to be translated through direct user input, the speech can be recorded directly using a microphone or other recording devices, suitable for applications such as instant messaging, conference translation, and voice assistants. Selecting audio files from a pre-stored audio library or database as the speech to be translated is commonly used for offline translation tasks and post-processing translation of audio content, such as video subtitle translation and voice recording translation. Obtaining the speech to be translated from real-time audio streams such as web conferences, live broadcasts, and telephone calls is crucial for scenarios such as real-time speech transcription and remote simultaneous interpretation, requiring the system to have real-time processing and translation capabilities. By integrating the voice input function into mobile devices or applications such as smartphones, tablets, and smart speakers, capturing user speech as the speech to be translated makes the speech translation function more convenient, allowing users to use it anytime in different scenarios.

[0057] The aforementioned audio to be translated has a wide range of applications across various scenarios, each with its own unique characteristics and specific technical challenges, such as noisy environments, accent differences, and specialized terminology. When performing associative translation on the audio, appropriate technologies and strategies must be adopted based on the specific characteristics of the scenario to improve translation accuracy and user experience. For example, the audio to be translated can be, but is not limited to, courtroom recordings in courtroom settings, diagnostic audio in medical settings, lecture audio in educational and training settings, meeting audio in business communication settings, and conversational audio in e-commerce customer service settings.

[0058] After acquiring the speech to be translated, the original terminology information corresponding to the speech is retrieved from a preset storage area. The pre-review storage area can be a retrieval pool. In speech translation technology, a retrieval pool usually refers to a predefined database or storage area that contains a large amount of multimodal information that can be used to assist translation. Multimodal information can be used to retrieve specific terms or context-related information in the speech to be translated in order to improve the accuracy and professionalism of the translation.

[0059] The aforementioned preset storage area stores multimodal information of multiple candidate terms according to a preset data structure. Multimodal information represents candidate terms in different information modalities and may include, but is not limited to, audio samples, text forms, and translated versions of candidate data. In multimodal information, candidate terms exist in audio, text, and even visual images or video as auxiliary information. In speech translation, the combined use of multimodal information can improve the model's understanding and translation accuracy of terms, especially when dealing with proper nouns, new words, or domain-specific terms. For example, audio samples provide pronunciation information, while text provides spelling and grammatical structure; combining both can lead to more accurate term recognition and translation.

[0060] The aforementioned predefined data structures can be, but are not limited to, term triple structures, term graph structures, term embedding space structures, term knowledge base structures, term hash table structures, and term relationship network structures. The choice of which predefined data structure to use in a specific application depends primarily on the specific application scenario, the characteristics of the data, and the training and inference requirements of the model. For example, in scenarios requiring fast retrieval and high-precision matching, term embedding spaces and hash table structures are more efficient; while in complex tasks requiring the handling of term relationships and contextual information, term graph structures and term relationship networks offer better support.

[0061] It should be noted that the embodiments disclosed herein use the term triplet structure as an example for explanation and illustration, but this does not constitute a specific limitation on the preset data structure.

[0062] For example, taking a term triplet structure as the preset data structure, the term triplet structure typically contains the following three elements: source language audio of the term, source language text of the term, and target language text of the term. The source language audio is the phonetic form of the candidate term in the source language; it can be a recorded professional pronunciation or a speech fragment from actual use. The source language text is the written form of the candidate term in the source language, thus providing a precise written expression, which is crucial for understanding the correct meaning and spelling of the candidate term. The target language text is the translation of the candidate term in the target language, that is, the correct expression of the candidate term in the translated language. The above-mentioned original term information can be term triplets retrieved from a preset storage area based on the speech to be translated.

[0063] After the original term information corresponding to the speech to be translated is retrieved, associated translation is performed on the speech to be translated based on the original term information to obtain a translation result. Associated translation means that in the translation process, not only the meaning of the content to be translated itself, but also context information, domain knowledge, and relevant specific scenario information are fully combined to improve the accuracy and semantic coherence of the translation result. In term translation of end-to-end speech translation large models, associated translation can use the retrieved original term information, especially the original language audio of terms, original language text of terms and target language text of terms, to assist the model in more accurately identifying and translating professional vocabulary or new words in the speech to be translated. The key to associated translation lies in effectively combining the term knowledge in the retrieved original term information with the speech segment to be translated.

[0064] The above translation result may be a target language text processed by associated translation, and the translation result shall accurately reflect the content in the speech to be translated, including but not limited to professional terms, new words, specific personal names, place names, etc. Through associated translation, the translation result can have higher accuracy in professional fields or specific scenarios, term translation can more meet the requirements of professional fields, while maintaining the naturalness and coherence of the text.

[0065] For example, the speech to be translated includes the term "Blockchain". In direct translation, due to the lack of sufficient context information or professional domain knowledge, it cannot be accurately translated into "区块链" in Chinese. Without the assistance of external knowledge, the speech translation model in the related art will generate an inaccurate or general translation based on its own training data and understanding ability, such as "chain structure", however, this translation result is not equivalent to "Blockchain" in the professional field.

[0066] Through the speech processing method in the embodiments of the present disclosure for associated translation, when processing the speech to be translated containing "Blockchain", the original term information input with the following structure can be obtained through retrieval: <"Blockchain" audio clip, "Blockchain" text, Chinese translation "区块链">. The original term information can explicitly inform the speech translation model to translate the term "Blockchain" appearing in the current speech to be translated into Chinese "区块链".

[0067] In the process of related translation, the original term information can be preferentially used to locate terms and associate knowledge, so that "Blockchain" can be accurately translated as "blockchain" in the generated Chinese translation. By introducing external knowledge, especially multimodal information of terms including audio, text and translation, the related translation process enhances the translation ability of the speech translation model in professional fields or when facing new words, ensures the accuracy of professional terms and the coherence of context, thereby significantly improving the performance and reliability of end-to-end large speech translation models in practical applications.

[0068] Based on the foregoing step S41 to step S42, by acquiring the speech to be translated, then retrieving the original term information corresponding to the speech to be translated from a preset storage area, wherein the preset storage area is configured to store multimodal information of a plurality of alternative terms according to a preset data structure, and finally performing related translation on the speech to be translated based on the original term information to obtain a translation result, thus the original term information related to the speech to be translated can be retrieved from the preset storage area, which effectively avoids introducing a large amount of irrelevant information, thereby reducing errors and misleading in the speech translation process, and further improving translation efficiency and translation quality. The multimodal information of a plurality of alternative terms stored in the preset storage area further enriches the translation basis in the related translation process, helps to more accurately understand the context and meaning of professional terms in the speech to be translated, further improves the pertinence and translation efficiency in the translation process, thus solves the technical problem of low translation efficiency and poor accuracy existing in related technologies when performing speech translation.

[0069] The speech processing method in the embodiments of the present disclosure is further introduced below.

[0070] In an optional embodiment, retrieving the original term information corresponding to the speech to be translated from the preset storage area comprises:

[0071] performing speech encoding on the speech to be translated by using a target speech encoding model to obtain a first feature vector, and performing speech encoding on term audio corresponding to a plurality of alternative terms by using the target speech encoding model to obtain a second feature vector;

[0072] calculating the similarity between the first feature vector and the second feature vector to obtain the original term information.

[0073] The above target speech encoding model is a deep learning model for converting speech signals into feature vectors. The target speech encoding model is usually trained with a large amount of speech data, and can capture information such as acoustic features, pitch patterns and speech rate of speech, and convert them into digital representations that can be understood and processed by machines.

[0074] The target speech coding model encodes the speech to be translated, resulting in a first feature vector, which is a digital representation of the translated speech. In other words, after processing by the target speech coding model, the speech to be translated becomes a dense, continuous set of numerical vectors. The first feature vector captures the acoustic characteristics, pitch, rhythm, and other information of the speech, reflecting its semantic content and the speaker's expression. The dimension of the first feature vector can be M*H, where M is the length of the speech features, reflecting the duration of the speech to be translated; longer speech will result in a longer feature vector sequence, and thus a larger value for M. H is the dimension of the encoded feature vector, i.e., the number of elements in each feature vector. The size of H reflects the complexity and information content of the model's encoding, and is usually related to the model's architecture design and training objectives. For example, H can include a combination of frequency, time-domain, and spectral-domain features of the speech to be translated.

[0075] Each vector in the first feature vector represents a small segment of the speech to be translated or information at a specific point in time. Once the entire speech is encoded, the resulting sequence of M vectors comprehensively and thoroughly describes the features of the entire speech to be translated. The first feature vector not only carries the original information of the speech but, after processing by the target speech coding model, also incorporates a higher level of semantic understanding.

[0076] A target speech coding model is used to encode the audio of multiple candidate terms, resulting in a second feature vector. This second feature vector represents the features obtained after encoding the original language audio of each candidate term. The target speech coding model transforms each original language audio of a term into a dense sequence of feature vectors with dimension N*H. Here, N represents the length of the term's speech features, i.e., the sequence length of the encoded term audio feature vectors, reflecting the duration of the term in the speech; H is the dimension of the encoded feature vectors, representing the number of elements in each feature vector, reflecting the complexity of the model's encoding and the amount of information it covers.

[0077] Furthermore, similarity calculations are performed on the first and second feature vectors to obtain the original terminology information. Various methods exist for similarity calculation, each with its applicable scenarios and characteristics. For example, cosine similarity is one of the most commonly used similarity calculation methods, especially when dealing with vectors in high-dimensional space. It measures the similarity between the first and second feature vectors by calculating the cosine of the angle between the two vectors. The range of cosine similarity is [-1, 1], where a value of 1 indicates complete similarity, 0 indicates orthogonality, and -1 indicates complete opposites. Euclidean distance is another standard for measuring vector similarity, calculating the linear distance between two vectors in multidimensional space. The smaller the Euclidean distance, the more similar the two vectors are. However, for vectors in high-dimensional space, Euclidean distance is affected by the "curse of dimensionality," leading to inaccurate similarity calculation results. The sliding window similarity calculation method compares the first feature vector segment by segment with the second feature vector to locate the position of the term in the speech. This not only considers the direction of the vector but also overcomes the challenge of sequence length differences through local comparison.

[0078] Based on the above optional embodiments, the target speech coding model is used to encode the speech to be translated to obtain a first feature vector, and the target speech coding model is used to encode the audio of the terms corresponding to multiple candidate terms to obtain a second feature vector. Then, the similarity between the first feature vector and the second feature vector is calculated to quickly obtain the original term information. Thus, the occurrence position of the term in the speech to be translated can be precisely located through similarity calculation, reducing the interference of irrelevant information and further optimizing the efficiency and accuracy of term translation.

[0079] In one optional embodiment, similarity calculation is performed between the first feature vector and the second feature vector to obtain the original terminology information, including:

[0080] Obtain the length of the speech feature to be translated corresponding to the first feature vector and the length of the term feature corresponding to the second feature vector; based on the length of the speech feature to be translated and the length of the term feature, perform a sliding window similarity calculation on the first feature vector and the second feature vector to obtain the original term information.

[0081] Specifically, the length M of the speech feature corresponding to the first feature vector reflects the duration and complexity of the entire speech to be translated, and the length N of the term feature corresponding to the second feature vector reflects the duration of the original language audio of the candidate term.

[0082] Based on the feature length M of the speech to be translated and the feature length N of the terminology, a sliding window similarity calculation is performed on the first feature vector and the second feature vector. The sliding window similarity calculation can not only identify professional terms or new words in the speech to be translated, but also accurately locate the position of these words in the speech to be translated, thereby significantly improving the accuracy and reliability of terminology translation. Especially when dealing with speech translation tasks in specific fields or containing a large number of professional terms, it can effectively overcome the problems of mismatched retrieval granularity and the introduction of too much irrelevant information in related technologies.

[0083] Based on the above optional embodiments, by obtaining the length of the speech feature to be translated corresponding to the first feature vector and the length of the term feature corresponding to the second feature vector, and then performing a sliding window similarity calculation on the first feature vector and the second feature vector based on the length of the speech feature to be translated and the length of the term feature, the original term information can be obtained. This can significantly enhance the term recognition and translation performance through sliding window similarity calculation, and further improve the translation quality and translation efficiency.

[0084] In one optional embodiment, based on the length of the speech feature to be translated and the length of the term feature, a sliding window similarity calculation is performed on the first feature vector and the second feature vector to obtain the original term information, including:

[0085] The sliding window size is determined based on the term feature length, and the number of similarity calculations is determined based on the feature length of the speech to be translated and the term feature length.

[0086] The sliding window is moved along the speech features to be translated corresponding to the first feature vector according to the sliding window size and preset step size, and multiple similarity calculations are performed according to the number of similarity calculations to obtain multiple calculation results;

[0087] The original terminology information is obtained by comparing the results of multiple calculations.

[0088] Specifically, the term feature length N is used as the sliding window size, the preset step size is set to 1, and the number of similarity calculations is determined to be M-N+1 times based on the speech feature length M to be translated and the term feature length N. The sliding window is moved on the speech feature to be translated corresponding to the first feature vector according to the sliding window size N and the preset step size 1, and M-N+1 similarity calculations are performed to obtain multiple calculation results.

[0089] For example, multiple calculation results are compared, and the maximum similarity among the results is used as the similarity between the first and second feature vectors. The sliding position of the window corresponding to the maximum similarity among the multiple calculation results can indicate the location of the term in the speech to be translated. This allows for full consideration of the inclusion relationship between the term and the speech sentence, avoiding loss of feature representation accuracy and thus achieving better retrieval results. Based on the comparison results among multiple calculation results, specific segments related to the term in the speech to be translated can be quickly located.

[0090] For example, if the maximum similarity value determined by comparing multiple calculation results exceeds a preset threshold, it can be determined that an audio segment highly matching the second feature vector has been found in the speech to be translated. At this point, it can be determined that the speech features at a specific location in the first feature vector represent the original terminology information, i.e., the actual occurrence of the candidate terms in the speech to be translated. The preset threshold is set based on experiments and specific application scenarios, and is used to filter out possible false matches to ensure the accuracy of recognition.

[0091] After determining the location and matching degree of the identified candidate terms, the multimodal information corresponding to the candidate term can be output to obtain the original term information. For each identified candidate term, not only is there a matching location of its speech features, but also related text information and known translations. This not only enables the identification of technical terms or new words in the speech to be translated, but also provides its accurate representation in the source and target languages, thereby greatly improving the accuracy of term translation.

[0092] Based on the above optional embodiments, by determining the sliding window size based on the term feature length and determining the number of similarity calculations based on the length of the speech feature to be translated and the term feature length, the sliding window is then moved along the speech feature to be translated corresponding to the first feature vector according to the sliding window size and a preset step size, and multiple similarity calculations are performed according to the number of similarity calculations to obtain multiple calculation results. Finally, the original term information is obtained based on the comparison results between the multiple calculation results. Thus, the term recognition and translation performance can be significantly enhanced through sliding window similarity calculation, further improving translation quality and translation efficiency.

[0093] In an optional embodiment, the speech processing method in this disclosure further includes:

[0094] Obtain speech coding training samples, which include: speech coding training statements, positive example terms, and negative example terms. Positive example terms are terms actually contained in the speech coding training statements, and negative example terms are irrelevant terms not contained in the speech coding training statements.

[0095] The initial speech coding model is trained using speech coding training samples to obtain the similarity calculation loss;

[0096] The model parameters of the initial speech coding model are fine-tuned based on the similarity calculation loss to generate the target speech coding model.

[0097] The aforementioned speech coding training utterance can be an audio signal containing specific words or terms, used to train the model to recognize and encode the speech features of these technical terms or new words. The positive example terms are actual terms present in the speech coding training utterance, used to guide the model in learning how to correctly encode the speech features of these terms to achieve high scores in similarity calculations. The negative example terms are randomly sampled terms not present in the speech coding training utterance, used to help the model learn to distinguish and ignore irrelevant term features, avoiding false positives in similarity calculations.

[0098] For example, during the training of the initial speech coding model using speech coding training samples, the initial speech coding model receives speech coding training sentences and positive example terms as input, attempts to encode them, and calculates the similarity between the input data. Simultaneously, the initial speech coding model also receives negative example terms for comparison and learning how to assign low scores in similarity calculations, thereby distinguishing between positive and negative examples. By calculating the similarity between the model output and positive example terms, and the difference between the output and negative example terms, the similarity calculation loss can be obtained, which can then measure the model's performance in recognizing and distinguishing different terms. The calculation process for the similarity calculation loss is as follows:

[0099] As shown in Equation 1:

[0100] Where sim() is the similarity calculation function, L AE The similarity loss value is represented by u, which represents the feature vector of the speech encoding training utterance, and c represents the similarity loss value. + c represents the eigenvector of the positive example term. - The feature vector represents the negative example term. The similarity calculation function is calculated as shown in Formula 2:

[0101] Here, MP() represents the pooling operation, used to handle the problem of inconsistent feature vector lengths. Pooling transforms feature vectors of different lengths into vectors of the same length, allowing similarity calculations to be performed on the same dimension. During the calculation, the sliding window moves along the feature vector z of the speech to be translated. u Move from left to right, and at each window position i, compare with the term feature vector z. iPooling and cosine similarity calculations are performed. By taking the maximum cosine similarity value calculated from all sliding windows, the location of the audio segment in the speech to be translated that matches the term features can be determined, thus achieving the localization of the term in the audio.

[0102] After obtaining the similarity calculation loss, the model parameters of the initial speech coding model are fine-tuned based on the similarity calculation loss to generate the target speech coding model. The similarity calculation loss is used to guide the update of the model parameters. Optimization algorithms such as backpropagation are used to adjust the weights of the initial speech coding model, so that the initial speech coding model outputs higher similarity when processing positive terms and lower similarity when processing negative terms. This fine-tuning process is repeated until the model's performance reaches the expected level, that is, it can accurately distinguish and calculate similarity between positive and negative examples.

[0103] Through the aforementioned training and fine-tuning process, the initial speech coding model was gradually trained into a target speech coding model, which demonstrated better performance in terminology recognition and similarity calculation. This target speech coding model will be used for subsequent speech translation and terminology localization tasks, enabling it to more accurately encode the features of the input audio and locate and recognize specialized terms in speech.

[0104] Based on the above optional embodiments, by obtaining speech coding training samples, and then using the speech coding training samples to train the initial speech coding model, a similarity calculation loss is obtained. Finally, the model parameters of the initial speech coding model are fine-tuned based on the similarity calculation loss to generate the target speech coding model. The trained target speech coding model can process speech features more accurately, especially in the recognition of terms in professional fields. It can effectively distinguish and locate positive terms while ignoring or suppressing the influence of negative terms, thereby improving the translation quality and accuracy of speech translation in specific scenarios.

[0105] In one optional embodiment, the original terminology information includes: audio of the terminology's source language, text of the terminology's source language, and text of the terminology's target language. Based on the original terminology information, the speech to be translated is correlated and translated to obtain the following translation results:

[0106] The audio segments associated with the original term information in the speech to be translated are located to obtain the term location audio segments;

[0107] The terminology source language audio is replaced with terminology location audio segments, and the target term information is determined based on the terminology location audio segments, terminology source language text, and terminology target language text.

[0108] The translation result is obtained by associating the target term information with the speech to be translated.

[0109] Specifically, the audio segments associated with the original terminology information in the speech to be translated are located to obtain terminology-localized audio segments. These segments are audio fragments containing specialized terms or new words appearing in the speech to be translated. Using the target speech coding model and retrieval and localization techniques trained above, the input speech to be translated is analyzed. A sliding window similarity calculation is used to find audio segments that match the original terminology information; these are the terminology-localized audio segments. The terminology-localized audio segments need to contain clear phonetic expressions of the specialized terms and minimize interference from irrelevant information so that subsequent associative translation can accurately capture the meaning of the terms.

[0110] After obtaining the term localization audio segment, the original language audio of the term in the original term information is replaced with the term localization audio segment. This constructs a multimodal information containing the term localization audio segment, the original language text of the term, and the target language text of the term, thus obtaining the target term information.

[0111] The target term information and the speech to be translated are provided as input to the target speech translation model. The target speech translation model can perform related translation of terms based on the input data. That is, during the translation process, the target speech translation model will refer to the input target term information to ensure the accurate translation of professional terms.

[0112] Based on the above optional embodiments, by locating the audio segments associated with the original terminology information in the speech to be translated, terminology location audio segments are obtained. Then, the original language audio of the terminology is replaced with the terminology location audio segments. Based on the terminology location audio segments, the original language text of the terminology, and the target language text of the terminology, the target terminology information is determined. Finally, the speech to be translated is associated with the target terminology information to obtain the translation result. This enables the target speech translation model to reduce the influence of errors and irrelevant information when translating professional terms, and further improves the accuracy and reliability of terminology translation.

[0113] In one optional embodiment, determining target term information based on term-localized audio segments, term source language text, and term target language text includes:

[0114] Add start and end markers to the target language text of the terminology. The start and end markers are used to indicate that the terminology is translated based on the target language text of the terminology in the process of translating the speech to be translated.

[0115] The target term information is determined based on the term-localized audio segment, the term's original language text, the start marker, the term's target language text, and the end marker.

[0116] Specifically, specific start-of-term and end-of-term markers are added before and after the target language text for each candidate term. These markers act as special tokens, forming a clear boundary in the translated text and informing the target speech translation model of the start and end positions of the terms. These markers guide the model to switch to term translation mode promptly during translation generation, utilizing known target language terms for accurate translation, rather than relying on model guesswork or contextual inference.

[0117] Furthermore, the term location audio clip, the term's source language text, the start of term marker, the term's target language text, and the end of term marker are combined to form a structured multimodal term information, thereby constructing a complete target term information structure for guidance in the subsequent translation process.

[0118] Based on the above optional embodiments, by adding start and end markers to the target language text of the term, and then determining the target term information based on the term location audio segment, the term source language text, the start marker, the term target language text, and the end marker, it is possible not only to identify and locate professional terms in the audio, but also to accurately translate the term according to the guidance of the term target language text and markers, thereby further improving the accuracy and efficiency of translation.

[0119] In one optional embodiment, the speech to be translated is correlated based on the target term information, and the translation result includes:

[0120] A target speech translation model is used to associate and translate the target term information with the speech to be translated, and the translation result is obtained. The text corresponding to the term location audio segment is copied from the target language text of the term contained in the target term information. The text corresponding to the other audio segments in the speech to be translated, except for the term location audio segment, is generated based on audio translation.

[0121] Specifically, in the process of using a target speech translation model to associate target term information with the speech to be translated, term content can be processed by term copying, and ordinary text content other than terms can be processed by audio translation generation to generate the final translation result.

[0122] During terminology copying, for audio segments located in the speech to be translated (i.e., term-localized audio segments), the target speech translation model can directly translate the terms based on the target language text provided in the target term information. The target speech translation model copies the terminology translations determined beforehand through retrieval and matching, rather than regenerating the translations, thereby ensuring the accuracy and consistency of specialized terms or new words and reducing translation errors.

[0123] Besides locating the audio segments containing the terminology, other parts of the speech to be translated are generated based on their audio signals. During this process, the translation of ordinary text content other than terminology relies on the speech-to-text conversion capabilities learned by the target speech translation model during training. The target speech translation model generates the corresponding text translation based on the content of the audio signal. This ensures the natural fluency and semantic coherence of ordinary text content, except for specialized terminology, while also demonstrating the target speech translation model's ability to understand and process speech signals.

[0124] By combining terminology copying with audio translation generation, the associative translation process ensures both accurate translation of specialized terminology and the natural generation of ordinary text content, thereby improving the accuracy and professionalism of the overall text translation. In practical applications, this audio translation processing method is particularly important for audio translation of professional documents, conference speeches, court transcripts, and other documents containing a large number of specialized terms, significantly improving translation quality and user satisfaction, especially in scenarios requiring a high degree of professionalism and terminology consistency.

[0125] Based on the above optional embodiments, by using a target speech translation model to associate and translate the target term information with the speech to be translated, the translation result is obtained. This can maintain sensitivity and accuracy to professional terms when dealing with complex and varied speech input, while also flexibly handling the translation of ordinary text content, demonstrating the powerful capabilities and practicality of the target speech translation model in end-to-end speech translation scenarios.

[0126] In an optional embodiment, the speech processing method in this disclosure further includes:

[0127] Obtain speech translation training samples, which include: speech translation training sentences, given terminology information, and given audio.

[0128] The initial speech translation model is trained using speech translation training samples to obtain the token prediction loss;

[0129] The model parameters of the initial speech translation model are fine-tuned based on token prediction loss to generate the target speech translation model.

[0130] The aforementioned speech translation training statements can be the model's expected translation results, i.e., the speech signal translated into a text representation in the target language. The given terminology information can include information on specialized terms in both the source and target languages, including the text representations and correct translations of the terms. The given audio can be the original speech signal to be translated, containing the content that the initial speech translation model needs to recognize and translate.

[0131] For example, during model training, an initial speech translation model can first be used to predict the translation of a given audio. During prediction, the initial speech translation model generates a series of tokens and attempts to match them with the given speech translation training utterance. Then, by comparing the difference between the token sequence generated by the initial speech translation model and the expected translation, the token prediction loss is calculated, which is a measure of the model's error in predicting the next token. The token prediction loss function guides how the initial speech translation model adjusts its behavior to generate more accurate translations of the target language.

[0132] Furthermore, using the calculated token prediction loss, a fine-tuning algorithm can be used to adjust the model parameters to reduce the loss. The model parameter update process can be iterative; after each training round, the initial speech translation model can predict again based on the new model parameters, calculating a new loss, until the initial speech translation model's performance on the training set meets the target requirements. After training and fine-tuning, the initial speech translation model gradually becomes more accurate and reliable, ultimately generating a target speech translation model capable of effectively handling both technical terms and ordinary text content.

[0133] Based on the above optional embodiments, by obtaining speech translation training samples, and then using the speech translation training samples to train the initial speech translation model, a token prediction loss is obtained. Finally, the model parameters of the initial speech translation model are fine-tuned based on the token prediction loss to generate the target speech translation model. Thus, by focusing on reducing the token prediction loss, the initial speech translation model can continuously improve its translation performance, especially when dealing with professional terms, which can significantly improve the accuracy and efficiency of translation.

[0134] In one optional embodiment, the initial speech translation model is trained using speech translation training samples to obtain the token prediction loss, which includes:

[0135] During the initial training of the speech translation model, the speech translation training sentences are divided into multiple tokens;

[0136] For the current token among multiple tokens, the token prediction loss is obtained by performing a logarithmic summation of conditional probabilities based on given terminology information, given audio, and the preceding tokens of the current token.

[0137] Specifically, during the initial training of the speech translation model, the speech translation training sentences are divided into multiple tokens. These tokens can be words, subwords, or specific character sequences, depending on the model's word segmentation strategy and vocabulary. Each token represents a basic unit in the sentence.

[0138] For the current token among multiple tokens, the token prediction loss is obtained by performing a logarithmic summation of the conditional probabilities based on the given terminology information, the given audio, and the preceding tokens of the current token. The token prediction loss can be calculated using the following formula 3:

[0139] Where N is the length of the token sequence, i.e., the total number of tokens in the speech translation training sentences, w i Given the current token among multiple tokens, K′ represents the given terminology information, u represents the given audio, and w represents the current token among multiple tokens. <i L is the preceding token of the current token. LLM The token prediction loss measures the deviation between the tokens predicted by the model and the actual tokens.

[0140] Based on the above optional embodiments, during the training process of the initial speech translation model, the speech translation training sentences are divided into multiple tokens. Then, for the current token among these tokens, a logarithmic summation of conditional probabilities is performed based on given terminology information, given audio, and the preceding tokens of the current token to obtain the token prediction loss. This loss is used to fine-tune the model parameters of the initial speech translation model, generating the target speech translation model. This allows the initial speech translation model to gradually learn during training how to more accurately predict each token and how to generate the entire translated sentence given audio and terminology information. The conditional probability-based training method ensures that the initial speech translation model can effectively utilize contextual information and terminology knowledge, improving translation quality and the accuracy of terminology translation.

[0141] Figure 5 is a schematic diagram of a speech processing method according to an embodiment of the present disclosure. As shown in Figure 5, firstly, the speech to be translated is acquired. The speech to be translated can originate from a speech input device, such as a microphone, and contains the content that the user wants to translate. Then, term localization is performed. A large number of pre-processed term triples are stored in the retrieval pool, including the term's source language audio, the term's source language text, and the term's target language text. The speech to be translated is encoded using a target speech encoding model to obtain a first feature vector. The speech audio corresponding to multiple candidate terms is also encoded using the same target speech encoding model to obtain a second feature vector. The similarity between the first and second feature vectors is calculated, and the original term information is obtained from the retrieval pool. The audio segments associated with the original term information in the speech to be translated are located to obtain term localization audio segments. Then, the term's source language audio is replaced with the term localization audio segments. Based on the term localization audio segments, the term's source language text, and the term's target language text, target term information is determined. Finally, the speech to be translated is associated with the target term information to obtain the translation result.

[0142] Figure 6 is a schematic diagram of another speech processing method according to an embodiment of the present disclosure. As shown in Figure 6, after acquiring the speech to be translated, a target speech coding model is used to encode the speech to be translated to obtain a first feature vector. The target speech coding model is also used to encode the audio of multiple candidate terms corresponding to the terms to obtain a second feature vector. A preset storage area is used to store the multimodal information of multiple candidate terms according to a preset data structure. The length of the speech feature corresponding to the first feature vector is 6, and the length of the term feature corresponding to the second feature vector is 3. The sliding window size is determined to be 3 based on the term feature length, and the number of similarity calculations is determined to be 4 based on the length of the speech feature and the length of the term feature. The sliding window is moved along the speech feature corresponding to the first feature vector according to the sliding window size 3 and the preset step size 1, and the similarity is calculated 4 times according to the number of similarity calculations, resulting in multiple calculation results of 0.2, 0.3, 0.8, and 0.4. The original term information is obtained based on the comparison results between the multiple calculation results. The original term information includes: the original language audio of the term, the original language text of the term, and the target language text of the term. In the speech to be translated, the audio segments associated with the original terminology information are located to obtain term-localized audio segments. The original language audio of the terminology is then replaced with the term-localized audio segments. Based on the term-localized audio segments, the original language text of the terminology, and the target language text of the terminology, the target terminology information is determined. A target speech translation model is then used to associate the target terminology information with the speech to be translated, yielding the translation result.

[0143] The speech processing method of this disclosure improves the recall rate of retrieving original terminology information through sliding window-based similarity calculation, thereby greatly reducing the introduction of irrelevant terminology information. Simultaneously, the retrieved knowledge is at the terminology level, rather than sentence level, avoiding the introduction of other information unrelated to the terminology within the sentence. By replacing the original language audio of the terminology with terminology-localized audio segments, a terminology co-occurrence association between terminology knowledge and the original sentence audio is explicitly established. Furthermore, special tags can explicitly link the switching between the two behavioral modes of terminology knowledge and translation decoding, further improving translation efficiency and accuracy.

[0144] Figure 7 is a flowchart of another speech processing method according to an embodiment of the present disclosure. As shown in Figure 7, the method may include the following steps:

[0145] Step S71: Obtain the court hearing recording to be translated;

[0146] Step S72: Retrieve the original legal terminology information corresponding to the court hearing recording to be translated from the preset storage area. The preset storage area is used to store multimodal information of multiple candidate legal terms according to a preset data structure.

[0147] Step S73: Based on the original legal terminology information, perform associated translation on the court hearing recording to be translated to obtain the translation result.

[0148] Based on steps S71 to S73 above, the court hearing recording to be translated is acquired, and then the original legal terminology information corresponding to the recording is retrieved from a preset storage area. This preset storage area stores multimodal information of multiple candidate legal terms according to a preset data structure. Finally, the court hearing recording is translated based on the original legal terminology information to obtain the translation result. This method allows for the retrieval of original terminology information related to the court hearing recording from the preset storage area, effectively avoiding the introduction of a large amount of irrelevant information, thereby reducing errors and misleading information during the speech translation process, and further improving translation efficiency and quality. The multimodal information of multiple candidate terms stored in the preset storage area further enriches the translation basis during the associative translation process, helping to more accurately understand the context and meaning of professional terms in the court hearing recording, further improving the relevance and efficiency of the translation process. This solves the technical problems of low translation efficiency and poor accuracy in related technologies when performing speech translation.

[0149] Figure 8 is a flowchart of another speech processing method according to an embodiment of the present disclosure. As shown in Figure 8, the method may include the following steps:

[0150] Step S81: Obtain a voice processing request through the first application programming interface, wherein the request data carried in the voice processing request includes: the voice to be translated;

[0151] Step S82: Return a voice processing response through the second application programming interface, wherein the response data carried in the voice processing response includes: translation result, which is generated according to any one of the voice processing methods in the embodiments of this disclosure.

[0152] Based on steps S81 to S82 above, a voice processing request is obtained through a first application programming interface (API). The request data carried in the voice processing request includes the voice to be translated. Then, a voice processing response is returned through a second API. The response data carried in the voice processing response includes the translation result, which is generated according to any one of the voice processing methods in this embodiment. By obtaining the voice to be translated, the original terminology information corresponding to the voice to be translated is retrieved from a preset storage area. The preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure. Finally, the voice to be translated is correlated based on the original terminology information to obtain the translation result. This allows for the retrieval of original terminology information related to the voice to be translated from the preset storage area, effectively avoiding the introduction of a large amount of irrelevant information, thereby reducing errors and misleading information in the voice translation process, and further improving translation efficiency and quality. The multimodal information of multiple candidate terms stored in the preset storage area further enriches the translation basis in the correlated translation process, helping to more accurately understand the context and meaning of professional terms in the voice to be translated, further improving the relevance and efficiency of the translation process. This solves the technical problems of low translation efficiency and poor accuracy in related technologies when performing voice translation.

[0153] Figure 9 is a flowchart of another speech processing method according to an embodiment of the present disclosure. As shown in Figure 9, the method may include the following steps:

[0154] Step S91: Obtain the currently input voice processing dialogue request, wherein the request data carried in the voice processing dialogue request includes: the voice to be translated;

[0155] Step S92: In response to the voice processing dialogue request, a voice processing dialogue response is returned, wherein the information carried in the voice processing dialogue response includes: translation result, which is generated according to any one of the voice processing methods in the embodiments of this disclosure;

[0156] Step S93: Display the translation results within the graphical user interface.

[0157] Based on steps S91 to S93 above, by acquiring the currently input voice processing dialogue request, the request data carried in the voice processing dialogue request includes: the voice to be translated. Then, in response to the voice processing dialogue request, a voice processing dialogue reply is returned. The information carried in the voice processing dialogue reply includes: the translation result. The translation result is generated according to any one of the voice processing methods in this embodiment of the present disclosure. Finally, the translation result is displayed in the graphical user interface. By acquiring the voice to be translated, the original terminology information corresponding to the voice to be translated is retrieved from a preset storage area. The preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure. Finally, the voice to be translated is correlated and translated based on the original terminology information to obtain the translation result. This allows for the retrieval of original terminology information related to the voice to be translated from the preset storage area, effectively avoiding the introduction of a large amount of irrelevant information, thereby reducing errors and misleading information in the voice translation process, and further improving translation efficiency and quality. The multimodal information of multiple candidate terms stored in the preset storage area further enriches the translation basis in the associative translation process, which helps to more accurately understand the context and meaning of professional terms in the speech to be translated, and further improves the pertinence and efficiency of the translation process. This solves the technical problems of low translation efficiency and poor accuracy in speech translation in related technologies.

[0158] Figure 10 is a flowchart of another speech processing method according to an embodiment of the present disclosure. As shown in Figure 10, the method may include the following steps:

[0159] Step S101: In response to the input command applied to the operation interface, select the speech to be translated on the operation interface;

[0160] Step S102: In response to the processing command applied to the operation interface, the translation result is displayed on the operation interface; wherein the translation result is generated according to any of the speech processing methods in the embodiments of this disclosure.

[0161] Based on steps S101 and S102 above, by responding to input commands applied to the operation interface, the speech to be translated is selected on the operation interface, and then, in response to processing commands applied to the operation interface, the translation result is displayed on the operation interface. The translation result is generated according to any one of the speech processing methods in this embodiment. By acquiring the speech to be translated, the original terminology information corresponding to the speech to be translated is retrieved from a preset storage area. The preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure. Finally, the speech to be translated is correlated based on the original terminology information to obtain the translation result. This allows the retrieval of original terminology information related to the speech to be translated from the preset storage area, effectively avoiding the introduction of a large amount of irrelevant information, thereby reducing errors and misleading information in the speech translation process, and further improving translation efficiency and quality. The multimodal information of multiple candidate terms stored in the preset storage area further enriches the translation basis in the correlated translation process, helping to more accurately understand the context and meaning of professional terms in the speech to be translated, further improving the relevance and efficiency of the translation process, thereby solving the technical problems of low translation efficiency and accuracy in speech translation in the relevant technology.

[0162] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0163] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, because according to this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure.

[0164] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solutions of this disclosure, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.

[0165] According to an embodiment of this disclosure, a speech processing apparatus for implementing the above-described speech processing method is also provided. FIG11 is a structural block diagram of a speech processing apparatus according to an embodiment of this disclosure. As shown in FIG11, the apparatus includes:

[0166] Module 1101 is configured to acquire the speech to be translated.

[0167] The retrieval module 1102 is configured to retrieve the original terminology information corresponding to the speech to be translated from a preset storage area, wherein the preset storage area is used to store the multimodal information of multiple candidate terms according to a preset data structure;

[0168] Translation module 1103 is configured to perform associative translation on the speech to be translated based on the original terminology information to obtain the translation result.

[0169] Optionally, the retrieval module 1102 is further configured to: use a target speech coding model to perform speech coding on the speech to be translated to obtain a first feature vector, and use the target speech coding model to perform speech coding on the audio of the terms corresponding to multiple candidate terms to obtain a second feature vector; calculate the similarity between the first feature vector and the second feature vector to obtain the original term information.

[0170] Optionally, the retrieval module 1102 is further configured to: obtain the length of the speech feature to be translated corresponding to the first feature vector and the length of the term feature corresponding to the second feature vector; and perform a sliding window similarity calculation on the first feature vector and the second feature vector based on the length of the speech feature to be translated and the length of the term feature to obtain the original term information.

[0171] Optionally, the retrieval module 1102 is further configured to: determine the sliding window size based on the term feature length, and determine the number of similarity calculations based on the length of the speech feature to be translated and the length of the term feature; slide the sliding window on the speech feature to be translated corresponding to the first feature vector according to the sliding window size and the preset step size, and perform multiple similarity calculations according to the number of similarity calculations to obtain multiple calculation results; and obtain the original term information based on the comparison results between the multiple calculation results.

[0172] Optionally, the acquisition module 1101 is further configured to acquire speech coding training samples, wherein the speech coding training samples include: speech coding training statements, positive example terms, and negative example terms, where positive example terms are terms actually contained in the speech coding training statements, and negative example terms are irrelevant terms not contained in the speech coding training statements; the speech processing device further includes: a training module 1104, configured to train an initial speech coding model using the speech coding training samples to obtain a similarity calculation loss; and a fine-tuning module 1105, configured to fine-tune the model parameters of the initial speech coding model based on the similarity calculation loss to generate a target speech coding model.

[0173] Optionally, the translation module 1103 is further configured to: locate the audio segment associated with the original term information in the speech to be translated to obtain the term location audio segment; replace the term source language audio with the term location audio segment, and determine the target term information based on the term location audio segment, the term source language text and the term target language text; and perform associated translation on the speech to be translated based on the target term information to obtain the translation result.

[0174] Optionally, the translation module 1103 is further configured to: add start and end markers to the target language text of the term, wherein the start and end markers are used to indicate that, during the translation of the speech to be translated, the term location audio segment is translated based on the term target language text; and the target term information is determined based on the term location audio segment, the term source language text, the start marker, the term target language text, and the end marker.

[0175] Optionally, the translation module 1103 is further configured to: use a target speech translation model to associate and translate the target term information with the speech to be translated, and obtain the translation result, wherein the text corresponding to the term location audio segment is copied based on the term target language text contained in the target term information, and the text corresponding to the other audio segments in the speech to be translated, excluding the term location audio segment, is generated based on audio translation.

[0176] Optionally, the acquisition module 1101 is further configured to acquire speech translation training samples, wherein the speech translation training samples include: speech translation training sentences, given terminology information, and given audio; the training module 1104 is further configured to train the initial speech translation model using the speech translation training samples to obtain the token prediction loss; the fine-tuning module 1105 is further configured to fine-tune the model parameters of the initial speech translation model based on the token prediction loss to generate the target speech translation model.

[0177] Optionally, the training module 1104 is further configured to: during the training of the initial speech translation model, divide the speech translation training sentences into multiple tokens; for the current token among the multiple tokens, perform a logarithmic summation of conditional probabilities based on the given term information, the given audio, and the preceding tokens of the current token to obtain the token prediction loss.

[0178] It should be noted that the acquisition module 1101, retrieval module 1102, and translation module 1103 correspond to steps S41 to S43 in the above embodiments. The three modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the server provided in the above embodiments.

[0179] Figure 12 is a structural block diagram of another voice processing device according to an embodiment of the present disclosure. As shown in Figure 12, the device includes:

[0180] Module 1201 is configured to acquire court hearing recordings to be translated.

[0181] The retrieval module 1202 is configured to retrieve the original legal terminology information corresponding to the court hearing recording to be translated from a preset storage area. The preset storage area is used to store multimodal information of multiple alternative legal terms according to a preset data structure.

[0182] Translation module 1203 is configured to perform associative translation of the court hearing recording to be translated based on the original legal terminology information, and obtain the translation result.

[0183] It should be noted that the acquisition module 1201, retrieval module 1202, and translation module 1203 correspond to steps S71 to S73 in the above embodiments. The three modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also run as part of the device in the server provided in the above embodiments.

[0184] Figure 13 is a structural block diagram of another voice processing apparatus according to an embodiment of the present disclosure. As shown in Figure 13, the apparatus includes:

[0185] The acquisition module 1301 is configured to acquire a speech processing request through a first application programming interface, wherein the request data carried in the speech processing request includes: speech to be translated;

[0186] The return module 1302 is configured to return a speech processing response via a second application programming interface, wherein the response data carried in the speech processing response includes: a translation result, which is generated according to any one of the speech processing methods in the embodiments of this disclosure.

[0187] It should be noted that the acquisition module 1301 and return module 1302 correspond to steps S81 to S82 in the above embodiments. The two modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of a device and run in the server provided in the above embodiments.

[0188] Figure 14 is a structural block diagram of another voice processing apparatus according to an embodiment of the present disclosure. As shown in Figure 14, the apparatus includes:

[0189] The acquisition module 1401 is configured to acquire the currently input speech processing dialogue request, wherein the request data carried in the speech processing dialogue request includes: speech to be translated;

[0190] Return module 1402 is configured to return a voice processing dialogue response in response to a voice processing dialogue request, wherein the information carried in the voice processing dialogue response includes: translation result, which is generated according to any one of the voice processing methods in the embodiments of this disclosure;

[0191] Display module 1403 is configured to display the translation results within a graphical user interface.

[0192] It should be noted that the acquisition module 1401, return module 1402, and display module 1403 correspond to steps S91 to S93 in the above embodiments. The three modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also run as part of the device in the server provided in the above embodiments.

[0193] Figure 15 is a structural block diagram of another speech processing apparatus according to an embodiment of the present disclosure. As shown in Figure 15, the apparatus includes:

[0194] Select module 1501 is configured to respond to input commands applied to the user interface and select the speech to be translated on the user interface.

[0195] The display module 1502 is configured to respond to processing instructions applied to the operation interface and display the translation result on the operation interface; wherein the translation result is generated according to any one of the speech processing methods in the embodiments of this disclosure.

[0196] It should be noted that the selection module 1501 and display module 1502 correspond to steps S101 to S102 in the above embodiments. The two modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of a device and run in the server provided in the above embodiments.

[0197] It should be noted that the preferred implementation schemes involved in the above embodiments of this disclosure are the same as the schemes, application scenarios and implementation processes provided in the above embodiments, but are not limited to the schemes provided in the above embodiments.

[0198] Embodiments of this disclosure also provide a speech processing system, including: a client for sending speech to be translated; a server connected to the client for retrieving original terminology information corresponding to the speech to be translated from a preset storage area, and performing associative translation on the speech to be translated based on the original terminology information to obtain a translation result, wherein the preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure; the client is also used to output the translation result.

[0199] Embodiments of this disclosure may provide a computing device. FIG16 is a structural block diagram of a computing device according to an embodiment of the present disclosure. As shown in FIG16, the computing device may include: one or more (only one is shown in the figure) processors 162, memory 164, memory controller, and peripheral interfaces.

[0200] The aforementioned computing device can be understood as an integrated smart terminal, including but not limited to servers, desktop computers, PCs (Personal Computers), all-in-one model machines, etc., and the computing device may have the model described in the above embodiments of this disclosure pre-installed.

[0201] Specifically, this computing device can pre-install various types of models, including but not limited to models in natural language processing, visual processing, speech processing, code processing, and multimodal task processing, thus providing diverse model selection. In different product forms, this computing device can support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference, and application. In some product forms, this computing device also supports model management, including but not limited to multi-type model management (supporting the management of discriminative, generative, and other model types), model version control (supporting the control of different model versions), and model evaluation (evaluating model performance and effectiveness based on model evaluation tools). In other product forms, this computing device can also create applications based on models, providing API calling capabilities, allowing models to be called into created applications through API interfaces, and providing application management tools for application management and monitoring.

[0202] Furthermore, the computing device may also include data management (supporting the creation and management of model tuning datasets), a training center (providing abundant training resources to help users learn and master AI technology), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, it provides a comprehensive and integrated device for AI development, training, deployment, and application.

[0203] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0204] The processor can invoke an executable program stored in memory via a transmission device to execute the method described in any of the above embodiments.

[0205] Embodiments of this disclosure can provide an electronic device. FIG17 is a structural block diagram of an electronic device according to an embodiment of this disclosure. As shown in FIG17, the electronic device may include: an input / output device 172; a memory 174; and a processor 176, wherein the processor 176 is connected to the input / output device 172 and the memory 174 via a bus 178.

[0206] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0207] The processor can invoke an executable program stored in memory via a transmission device to execute the method described in any of the above embodiments.

[0208] It will be understood by those skilled in the art that the structure shown in the figure is merely illustrative, and the computing device may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. This figure does not limit the structure of the aforementioned computing device. For example, the computing device 100 may also include more or fewer components (such as a network interface, a display device, etc.) than shown in the figure, or may have a different configuration than that shown in the figure.

[0209] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0210] Embodiments of this disclosure also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.

[0211] Optionally, in this embodiment, the storage medium may be located in a computing device.

[0212] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, which, when the executable program is running, controls the device where the computer-readable storage medium is located to execute the method described in any of the above embodiments.

[0213] Embodiments of this disclosure also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.

[0214] Embodiments of this disclosure also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which can be used to store a computer program that, when executed by a processor, implements the methods provided in the embodiments described above.

[0215] Embodiments of this disclosure also provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the method provided in the above embodiments.

[0216] In the above embodiments of this disclosure, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0217] In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0218] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0219] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0220] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0221] The above description is only a preferred embodiment of this disclosure. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure.

Claims

1. A speech processing method comprising: obtaining a speech to be translated; retrieving original term information corresponding to the speech to be translated from a preset storage area, wherein the preset storage area is configured to store multimodal information of a plurality of alternative terms in a preset data structure; performing associated translation on the speech to be translated based on the original term information to obtain a translation result.

2. The voice processing method of claim 1, wherein, Retrieving the original term information corresponding to the speech to be translated from the preset storage area comprises: performing speech coding on the speech to be translated using a target speech coding model to obtain a first feature vector, and performing speech coding on term audio corresponding to the plurality of alternative terms using the target speech coding model to obtain a second feature vector; performing similarity calculation on the first feature vector and the second feature vector to obtain the original term information.

3. The voice processing method of claim 2, wherein, Performing similarity calculation on the first feature vector and the second feature vector to obtain the original term information comprises: obtaining a speech feature length corresponding to the first feature vector and a term feature length corresponding to the second feature vector; based on the speech feature length and the term feature length, performing sliding window similarity calculation on the first feature vector and the second feature vector to obtain the original term information.

4. The voice processing method of claim 3, wherein, Based on the speech feature length and the term feature length, performing sliding window similarity calculation on the first feature vector and the second feature vector to obtain the original term information comprises: determining a sliding window size based on the term feature length, and determining a number of similarity calculations based on the speech feature length and the term feature length; performing sliding on the speech feature corresponding to the first feature vector according to the sliding window size and a preset step length, and performing multiple similarity calculations according to the number of similarity calculations to obtain a plurality of calculation results; obtaining the original term information according to a comparison result between the plurality of calculation results.

5. The voice processing method of claim 2, wherein, The speech processing method further comprises: obtaining a speech coding training sample, wherein the speech coding training sample comprises a speech coding training sentence, a positive example term, and a negative example term, the positive example term being a term actually contained in the speech coding training sentence, and the negative example term being an irrelevant term not contained in the speech coding training sentence; training an initial speech coding model using the speech coding training sample to obtain a similarity calculation loss; fine-tuning model parameters of the initial speech coding model based on the similarity calculation loss to generate the target speech coding model.

6. The voice processing method of claim 1, wherein, The original term information comprises term original language audio, term original language text, and term target language text, and performing associated translation on the speech to be translated based on the original term information to obtain the translation result comprises: locating an audio segment associated with the original term information in the speech to be translated to obtain a term located audio segment; The term's original language audio is replaced with the term's location audio segment, and the target term information is determined based on the term's location audio segment, the term's original language text, and the term's target language text. The target term information is used to perform associative translation on the speech to be translated, and the translation result is obtained.

7. The voice processing method of claim 6, wherein, Based on the terminology-localized audio segment, the determination of the target term information by the terminology source language text and the terminology target language text includes: Add start and end markers to the target language text of the term, wherein the start and end markers are used to indicate that, during the translation of the speech to be translated, the term is translated based on the target language text of the term; Based on the terminology-located audio segment, the terminology's original language text, the start marker symbol, the terminology's target language text, and the end marker symbol, the target term information is determined.

8. The voice processing method of claim 6, wherein, Based on the target terminology information, the speech to be translated is correlated and translated to obtain the translation results, including: A target speech translation model is used to associate the target term information with the speech to be translated to obtain the translation result. The text corresponding to the term location audio segment is copied based on the term target language text contained in the target term information. The text corresponding to the other audio segments in the speech to be translated, excluding the term location audio segment, is generated based on audio translation.

9. The voice processing method of claim 8, wherein, The speech processing method further includes: Obtain speech translation training samples, wherein the speech translation training samples include: speech translation training sentences, given terminology information, and given audio; The initial speech translation model is trained using the aforementioned speech translation training samples to obtain the token prediction loss; The model parameters of the initial speech translation model are fine-tuned based on the token prediction loss to generate the target speech translation model.

10. The speech processing method of claim 9, wherein, The token prediction loss is obtained by training the initial speech translation model using the speech translation training samples, and includes: During the training process of the initial speech translation model, the speech translation training sentences are divided into multiple tokens; For the current token among the plurality of tokens, the conditional probability is summed logarithmically based on the given term information, the given audio, and the preceding token of the current token to obtain the token prediction loss.

11. The voice processing method of claim 1, wherein, The preset data structure includes one of the following: term triplet structure, term graph structure, term embedding space structure, term knowledge base structure, term hash table structure, and term relation network structure.

12. The voice processing method of claim 1, wherein, The method of obtaining the speech to be translated includes one of the following: direct user input, pre-recorded audio files, real-time audio streams, and integration with mobile applications or devices.

13. The voice processing method of claim 2, wherein, Calculating the similarity between the first feature vector and the second feature vector includes: Calculate the cosine similarity between the first feature vector and the second feature vector; or, calculate the Euclidean distance between the first feature vector and the second feature vector.

14. A speech processing method, comprising: Obtain the court hearing recordings to be translated; Retrieve the original legal terminology information corresponding to the court hearing recording to be translated from the preset storage area, wherein the preset storage area is used to store multimodal information of multiple alternative legal terms according to a preset data structure; The court hearing recording to be translated is translated by associating it with the original legal terminology information.

15. A speech processing method, comprising: A voice processing request is obtained through a first application programming interface, wherein the request data carried in the voice processing request includes: the voice to be translated; A speech processing response is returned via a second application programming interface, wherein the speech processing response carries response data including a translation result, which is generated according to the speech processing method of any one of claims 1 to 13.

16. A speech processing method, comprising: Obtain the currently input voice processing dialogue request, wherein the request data carried in the voice processing dialogue request includes: the voice to be translated; In response to the voice processing dialogue request, a voice processing dialogue response is returned, wherein the information carried in the voice processing dialogue response includes: a translation result, which is generated according to the voice processing method according to any one of claims 1 to 13; The translation results are displayed within a graphical user interface.

17. A speech processing method, comprising: In response to an input command applied to the user interface, select the speech to be translated on the user interface; In response to processing instructions applied to the operation interface, the translation result is displayed on the operation interface; The translation result is generated according to the speech processing method described in any one of claims 1 to 13.

18. A speech processing system, comprising: The client is used to send the voice to be translated; The server, connected to the client, is used to retrieve the original terminology information corresponding to the speech to be translated from a preset storage area, and to perform associative translation on the speech to be translated based on the original terminology information to obtain a translation result. The preset storage area is used to store multimodal information of multiple candidate terms according to a preset data structure. The client is also used to output the translation results.

19. An electronic device comprising: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, generates the speech processing method according to any one of claims 1 to 17.

20. A computer readable storage medium comprising a stored executable program, wherein, When the executable program is executed, it controls the device containing the computer-readable storage medium to perform the speech processing method according to any one of claims 1 to 17 to generate speech.