Audio and video processing method supporting multi-language automatic translation
By introducing intelligent translation gateways and components into the audio and video system, combining multi-dimensional array cache and slice block alignment technology, the problem of automatic translation in multiple languages in the audio and video system is solved, and efficient and accurate web page and audio and video stream translation is achieved.
Patent Information
- Application Number
- CN202510217826.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-13
AI Technical Summary
The current audio and video system does not support automatic translation in multiple languages, and the existing multilingual translation technology is costly and inefficient, making it difficult to achieve real-time/quasi-real-time translation of web pages and audio and video streams.
By introducing intelligent translation gateway and intelligent translation components into the audio and video system, multi-dimensional array cache and slice block aligned with dual paths can be used to realize multi-language automatic translation of web page text and audio and video streams.
It realizes multi-language translation of audio and video system web pages and real-time/quasi-real-time translation of audio and video streams, improving the compatibility and accuracy of the translation system and reducing manual translation costs.
Smart Images

Figure CN119993160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio and video data processing, and in particular to an audio and video processing method supporting multi-language automatic translation, an audio and video processing device supporting multi-language automatic translation, an electronic device and a storage medium. Background Art
[0002] NLP (Natural Language Processing) is a branch of artificial intelligence and linguistics that is dedicated to enabling computers to understand, interpret, and generate content in human language. The goal of NLP is to narrow the gap between human language and computers, enabling computers to perform functions such as speech recognition, natural language understanding, machine translation, text generation, speech synthesis, and dialogue systems. The current key technologies of NLP mainly include machine learning, deep learning, semantic analysis, syntactic analysis, and semantic understanding.
[0003] With the development of globalization, the demand for cross-language communication is growing. Current audio and video systems usually do not support automatic translation of multiple languages. To perform translation, it means that a lot of modifications need to be made to the source code of the HTML file, which is labor-intensive and inefficient. At the same time, this modification method is also likely to affect the original structure and layout of the HTML file, bringing a series of compatibility issues. In addition, the current multilingual translation technology mainly adopts a method based on machine translation. That is, automatic translation is achieved by training large-scale language models. A large amount of data and computing resources are required, and the cost is high. Therefore, how to realize the multi-language translation of web pages in audio and video systems and efficiently realize real-time / quasi-real-time translation of audio and video streams is an urgent problem to be solved in the current audio and video data processing. Summary of the invention
[0004] The present invention provides an audio and video processing method supporting multi-language automatic translation, an audio and video processing device supporting multi-language automatic translation, an electronic device and a storage medium, which are used to solve or partially solve the technical problem of how to realize multi-language translation of web pages in an audio and video system and efficiently realize real-time / quasi-real-time translation of audio and video streams.
[0005] The present invention provides an audio and video processing method supporting multi-language automatic translation, which is applied to a multi-language client, wherein the multi-language client is connected to a server provided with an intelligent translation gateway through a cloud network; the method comprises:
[0006] When the html source file returned by the intelligent translation gateway is obtained through the browser and a web page is rendered based on the html source file, the intelligent translation component is loaded and the web page translation language is set;
[0007] Searching for content to be translated on a web page based on the intelligent translation component, wherein the content to be translated includes text information to be translated and real-time audio and video streams;
[0008] In conjunction with the intelligent translation gateway, the text information to be translated is translated based on a multidimensional array cache call to obtain a text translation result in the webpage translation language;
[0009] Audio information is extracted from the real-time audio and video stream in real time, and audio translation based on slice blocks and dual-path alignment is performed on the audio information to obtain an audio translation result of the audio information in the webpage translation language.
[0010] Optionally, the process of setting the webpage translation language includes:
[0011] When the intelligent translation component is loaded, the language used by the browser is obtained, and the language used by the browser is set as the default webpage translation language;
[0012] or,
[0013] When the intelligent translation component is loaded, a language selection box pops up on the display interface of the web page, and based on the selection operation on the language selection box, the selected language is used as the web page translation language.
[0014] Optionally, the local storage of the browser is provided with a first-level cache area, and the first-level cache area is used to cache historical web page translation information stored in a multidimensional array form; the text translation based on the multidimensional array cache call on the text information to be translated in combination with the intelligent translation gateway to obtain the text translation result in the web page translation language includes:
[0015] Matching the text information to be translated with the historical webpage translation information through the intelligent translation component;
[0016] If a first text translation result of the text information to be translated in the webpage translation language is successfully matched from the historical webpage translation information, the intelligent translation component directly calls the first text translation result from the first-level cache area, and displays the first text translation result at a corresponding position of the text information to be translated in the display interface of the webpage;
[0017] If the text translation result of the text information to be translated in the webpage translation language cannot be matched from the historical webpage translation information, the text information to be translated is sent to the intelligent translation gateway, so that the intelligent translation gateway dynamically translates the text information to be translated based on multi-dimensional array feature optimization to obtain a second text translation result;
[0018] The second text translation result returned by the intelligent translation gateway is received, and the second text translation result is displayed at a corresponding position of the text information to be translated in the display interface of the webpage through the intelligent translation component.
[0019] Optionally, a secondary cache area is provided on one side of the intelligent translation gateway; the process in which the intelligent translation gateway dynamically translates the text information to be translated based on multi-dimensional array feature optimization to obtain a second text translation result includes:
[0020] After receiving the text information to be translated, the intelligent translation gateway forwards the text information to be translated to the secondary cache area;
[0021] In the secondary cache area, the intelligent translation gateway simultaneously considers multi-dimensional parameters and constructs a queue of nodes to be translated of the text information to be translated in the form of a multi-dimensional array;
[0022] According to the number of node queues in the node queue to be translated and the size of the multidimensional array, constructing a cache misjudgment rate objective function of the node queue to be translated based on a hash function;
[0023] Taking minimizing the cache misjudgment rate objective function as the optimization solution goal, continuously adjusting the number of hash function values in the cache misjudgment rate objective function, and in the parameter adjustment process, taking the webpage translation language as the target translation language, calling the intelligent translator to dynamically translate the text information to be translated, and outputting the second text translation result when the cache misjudgment rate is minimized;
[0024] The second text translation result is saved, and the second text translation result is returned to the multilingual client.
[0025] Optionally, performing audio translation on the audio information based on slice blocks and dual-path alignment to obtain an audio translation result of the audio information in the webpage translation language includes:
[0026] Extracting a plurality of audio short-time frames in sequence from the audio information with a preset frame length as an extraction interval;
[0027] For each of the audio short-time frames, extract key audio features from the audio short-time frame;
[0028] Continuously dividing the key audio features into a plurality of slice blocks;
[0029] Constructing a speech-to-text alignment path loss and a text-to-text alignment path loss of the key audio features;
[0030] According to the speech-text alignment path loss and the text-text alignment path loss, combined with the cross entropy loss, a total loss function is constructed;
[0031] By using a pre-trained synchronous speech-to-text translation model, minimizing the total loss function is taken as the optimization solution target, and the webpage translation language is taken as the target translation language, the plurality of slice blocks are translated in real time to obtain an audio frame translation result of the audio short-time frame in the webpage translation language;
[0032] The audio frame translation results of all the audio short-time frames are integrated to obtain the audio translation result of the audio information in the webpage translation language.
[0033] Optionally, in the process of real-time translation of the several slice blocks based on the pre-trained synchronous speech-to-text translation model, each slice block is based on one-dimensional convolution and adopts a bidirectional self-attention mechanism to perform intra-block bidirectional processing, and any two adjacent slice blocks perform inter-block unidirectional transmission based on a masked self-attention mechanism.
[0034] Optionally, the method further comprises:
[0035] After obtaining the audio frame translation result, calculating the real-time translation decision value of the audio short-time frame through a real-time translation decision function;
[0036] When the real-time translation decision value is greater than a preset decision threshold, the audio frame translation result is immediately outputted at the audio and video subtitle display position of the webpage;
[0037] When the real-time translation decision value is less than or equal to a preset decision threshold, the audio frame translation result is not outputted temporarily, and the audio translation process of the next audio short-time frame is directly entered.
[0038] Optionally, the intelligent translation component is formed by the intelligent translation gateway combining preset screening conditions, matching the request path of the multilingual client through regular expressions, and performing MIME type analysis on the HTTP request header information and response content of the request path, and injecting intelligent translation code into the HTML source file when confirming that the HTML source file is returned.
[0039] Optionally, the smart translation component introduces a Mutation Observer by creating an observer instance and passing in a callback function to monitor regular DOM changes of a web page; the smart translation component also introduces an event bubbling mechanism to monitor specified DOM changes of a web page.
[0040] The present invention also provides an audio and video processing device supporting multi-language automatic translation, which is applied to a multi-language client, wherein the multi-language client is connected to a server provided with an intelligent translation gateway through a cloud network; the audio and video processing device comprises:
[0041] An intelligent translation component loading unit, used to load the intelligent translation component and set the webpage translation language when the html source file returned by the intelligent translation gateway is obtained through the browser and the webpage is rendered based on the html source file;
[0042] A content-to-be-translated searching unit, configured to search for content-to-be-translated of a web page based on the intelligent translation component, wherein the content-to-be-translated includes text information to be translated and real-time audio and video streams;
[0043] A text translation unit, used for performing text translation based on multi-dimensional array cache call on the text information to be translated in combination with the intelligent translation gateway, to obtain a text translation result in the webpage translation language;
[0044] The audio translation unit is used to extract audio information from the real-time audio and video stream in real time, perform audio translation on the audio information based on slice blocks and dual-path alignment, and obtain an audio translation result of the audio information in the webpage translation language.
[0045] The present invention also provides an electronic device, the device comprising a processor and a memory:
[0046] The memory is used to store program code and transmit the program code to the processor;
[0047] The processor is used to execute the audio and video processing method supporting multi-language automatic translation as described in any one of the above items according to the instructions in the program code.
[0048] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program codes, and the program codes are used to execute the audio and video processing method supporting multi-language automatic translation as described in any one of the above items.
[0049] It can be seen from the above technical solutions that the present invention has the following advantages:
[0050] Provided is an audio and video processing method that supports multi-language automatic translation. A multi-language client is connected to a server provided with an intelligent translation gateway through a cloud network. When the multi-language client obtains the html source file returned by the intelligent translation gateway through a browser and performs web page rendering, the intelligent translation component is loaded, the web page translation language is set, and then the text information to be translated and the real-time audio and video stream of the web page are searched based on the intelligent translation component. For text translation, the text information to be translated is translated based on a multi-dimensional array cache call in combination with the intelligent translation gateway, and the text translation result under the web page translation language is obtained, so that the intelligent translation component formed based on the intelligent translation gateway, combined with the text translation based on the multi-dimensional array cache call, can support texts in a variety of different formats and real-time audio and video stream translation, greatly improving the compatibility of the translation system. And based on the intelligent translation component, the multi-language translation function of the web page can be automatically realized according to the client browser language. For audio and video translation, audio information is extracted from real-time audio and video streams in real time, and audio translation is performed on the audio information based on slicing blocks and dual-path alignment to obtain the audio translation result of the audio information in the web page translation language. Based on audio slicing blocks and the introduction of dual-path alignment, a composite model of multiple concurrent learning methods can be obtained to achieve low-latency audio and video subtitle translation, thereby improving the accuracy and efficiency of translation. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0052] Figure 1 It is a structural schematic diagram of a multi-language audio and video system;
[0053] Figure 2 A flowchart of the steps of an audio and video processing method supporting multi-language automatic translation;
[0054] Figure 3 The overall flow diagram of an audio and video processing method supporting multi-language automatic translation is shown;
[0055] Figure 4 The present invention is a structural block diagram of an audio and video processing device that supports automatic translation of multiple languages. DETAILED DESCRIPTION
[0056] The embodiments of the present invention provide an audio and video processing method supporting multi-language automatic translation, an audio and video processing device supporting multi-language automatic translation, an electronic device and a storage medium, which are used to solve or partially solve the technical problem of how to realize multi-language translation of web pages in an audio and video system and efficiently realize real-time / quasi-real-time translation of audio and video streams.
[0057] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] As an example, with the development of globalization, the demand for cross-language communication is growing. Current audio and video systems generally do not support automatic translation functions in multiple languages. Performing translation means that a large amount of modification of the HTML file source code is required, which is labor-intensive and inefficient. At the same time, this modification method is also likely to affect the original structure and layout of the HTML file, causing a series of compatibility issues. In addition, the current multilingual translation technology mainly adopts a method based on machine translation. That is, automatic translation is achieved by training large-scale language models. A large amount of data and computing resources are required, and the cost is high. Therefore, how to realize the multilingual translation of web pages in audio and video systems and efficiently realize real-time / quasi-real-time translation of audio and video streams is an urgent problem to be solved in current audio and video data processing.
[0059] Therefore, one of the core invention points of the embodiment of the present invention is: in view of the shortcomings of the current technology, an audio and video processing system and a corresponding processing method that support automatic translation functions in multiple languages are proposed. A multilingual intelligent translation gateway is set up in the system, and an intelligent translation component is injected into the web page without perception through web page injection and function callback. Combined with the optimized XPath algorithm and the text translation model based on the multidimensional array, it can support text translation in a variety of different formats, greatly improving the compatibility of the translation system. And the multilingual translation function of the web page can be automatically realized according to the client browser language. Through the intelligent translation gateway, loose coupling (Loose Coupling) with the system where the source web page is located can be achieved. There is no need to modify the page source file, and there is no need to modify the back-end interface API (Application Programming Interface) source code, and the dynamic translation of the front-end and back-end interface API text during the web page interaction process can also be realized, greatly reducing the cost of manual translation. For audio and video translation, based on the audio slice block, multiple CTC (Connectionist Temporal Classification) classifiers are introduced at the same time to construct dual-path constraints, and a composite model of multiple concurrent learning methods is obtained to achieve low-latency audio and video subtitle translation. Through refined recognition of audio features and multimodal display, the accuracy and efficiency of the entire translation process are greatly improved.
[0060] Figure 1 A schematic structural diagram of a multi-language audio and video system provided by an embodiment of the present invention is shown.
[0061] Combination Figure 1 The multi-language audio and video system may mainly include a multi-language client 101 and a server 103 connected to the multi-language client 101 via a cloud network 102 .
[0062] The multi-language client 101 may mainly include a browser, an intelligent translation component, a video player, and a localStorage (local storage) cache, and the localStorage cache corresponds to a first-level cache area.
[0063] The server 103 side mainly includes an intelligent translation gateway, an intelligent translator, a cache (corresponding to the secondary cache area) and a web server that are respectively connected to the intelligent translation gateway. The web server and the cache are connected through a database. The server 103 side also includes a video server.
[0064] The multilingual audio and video system provided by the embodiment of the present invention builds an intelligent translation gateway on the server side 103, and all http / https requests / responses initiated by users will be intercepted by the intelligent translation gateway first. In actual applications, the intelligent translation gateway implements the proxy function of requests and responses, and provides a unified entrance for the multilingual audio and video system. Since various data such as html / css / js / json / xml are all passed through the intelligent translation gateway, the intelligent translation gateway innovatively injects intelligent translation components into the web page without perception. When the customer obtains the html source file and renders the web page through the browser (including but not limited to chrome / webkit / firefox / edge / opera / IE / mobile browser, etc.), the intelligent translation component is loaded by the client and executed in the browser's built-in JS engine (JavaScript Engine). The intelligent translation component determines the language currently used by the customer through the built-in function of the browser, automatically executes the translation function, and displays the translated web page on the user interface.
[0065] Reference Figure 2 , shows a flowchart of the steps of an audio and video processing method supporting multi-language automatic translation provided by an embodiment of the present invention, which is applied to a multi-language client, and the multi-language client is connected to a server provided with an intelligent translation gateway through a cloud network; the audio and video processing method may specifically include the following steps:
[0066] Step 201, when the html source file returned by the intelligent translation gateway is obtained through the browser and a web page is rendered based on the html source file, the intelligent translation component is loaded and the web page translation language is set;
[0067] In actual applications, the intelligent translation component is formed by the intelligent translation gateway combined with preset filtering conditions, matching the request path of multilingual clients through regular expressions, and performing MIME type analysis based on the HTTP request header information and response content of the request path. When it is confirmed that the html source file is returned, the intelligent translation code is injected into the html source file.
[0068] Specifically, the intelligent translation gateway first matches the http / https request path of the multilingual client through regular expressions. During the proxy process, it ignores requests with suffix formats such as .css / .js / .png / .ico / .ttf / .woff / .svg, and includes responses to requests with suffixes such as .htm / .html / .php / .aspx / .jsp in the matching considerations. Then, it combines the HTTP (HyperText Transfer Protocol) header information (for example, Content-Type type is text / html) and the response content (for example,<!DOCTYPE html> The declaration begins, conforming to the DTD (Document Type Definition) file syntax standard, including,, and other tags) for MIME (Multipurpose Internet Mail Extensions) type analysis, which is used to define the file type and encoding method to ensure that the client correctly parses and displays the data sent by the server. When it is confirmed that the return is an html source file, the smart translation gateway will unconsciously add a section of smart translation code to the web page. The specific injection location is at the end of the web page and before. The reference injection code is as follows:
[0069] <script src=" / translation-gateway-plugin.js">< / script>
[0070] In another optional embodiment, the javascript translation plug-in of the intelligent translation gateway can be made into an extension of a browser (including but not limited to chrome / webkit / firefox / edge / opera / IE / mobile browser, etc.). That is, it is developed and compiled and packaged according to the browser extension specification, and provided for download and installation. The client installs the browser plug-in, and after accessing the corresponding html website through the browser, the browser plug-in implements the steps of path interception, format matching, analysis, etc., and performs the automatic translation function.
[0071] When customers operate the web pages of multi-language audio and video systems, they usually call the back-end interface for information exchange, such as searching for certain device names, setting historical recording tags, retrieving group information, etc. In the multi-language translation scenario, in addition to the automatic translation of web page text or real-time audio and video streams, the client also needs to combine the intelligent translation gateway to quickly and accurately translate the input and output information generated by the information interaction operation.
[0072] Therefore, in the smart translation component, the Mutation Observer technology is introduced by creating an observer instance and passing in a callback function to monitor the changes of the common DOM (Document Object Model) of the web page and update it efficiently. You can refer to the following code:
[0073] const mobs = new MutationObserver(Callback);
[0074] For some specific types (specified types) of DOM changes, such as paragraph elements, monitoring and updating , inline elements , Text Area <textarea> , reference element< / textarea> <blockquote>and <q>etc., you can introduce the event bubbling mechanism to handle it, and use event delegation combined with native events to implement it. You can refer to the following code:
[0075] document.addEventListener('click', (event) => {});
[0076] In some practical applications, the intelligent translation gateway can intercept all user inputs and analyze the input text using the n-gram model (an important tool in natural language processing). The Transformer model is applied in the translation process. At the same time, the intelligent translation gateway can also use a variety of deep learning optimization techniques, such as gradient clipping, learning rate decay, regularization, etc., to improve the performance and generalization ability of the model. Furthermore, evaluation indicators can also be used to measure the quality of translation. For example, indicators such as BLEU (Bilingual Evaluation Understudy) and METEOR (Metric for Evaluation of Translation with Explicit Ordering). In the subsequent use process, after algorithm optimization and collection of user feedback, the multi-language translation accuracy of the multi-language audio and video system can be greatly improved to meet the needs of multi-language users.
[0077] After the smart translation component is injected into the html webpage, it is loaded by the browser's JS engine (including but not limited to V8 / SpiderMonkey / JavaScriptCore / Chakra / Rhino, etc.). The smart translation component is injected into the html source file and is on the same page as other js variables and html elements. It can obtain and set all data on the page and avoid cross-domain operations. After the overall html rendering is completed, the browser language is obtained by executing the following js code to set the default translation language of the webpage:
[0078] transGwPlugin.autoSetLanguage(navigator.language);
[0079] It should be noted that the smart translation component can also pop up a selection box / input box at the border of the web page to support users to manually select the translation language. After the user manually selects, the selection result is encrypted and saved in the browser's cookie. The subsequent language is based on the language selected by the user.
[0080] In the specific implementation, the process of setting the webpage translation language can be divided into the following two situations:
[0081] The first is an automatic setting scenario. When the intelligent translation component is loaded, the browser's language is obtained and the browser's language is set as the default web page translation language.
[0082] The second is a manual setting scenario. When the intelligent translation component is loaded, a language selection box pops up on the display interface of the web page, and based on the selection operation on the language selection box, the selected language is used as the web page translation language.
[0083] Step 202, searching for content to be translated on a web page based on the intelligent translation component, wherein the content to be translated includes text information to be translated and real-time audio and video streams;
[0084] The smart translation component waits for all HTML elements to be loaded and rendered, and uses the optimized XPath algorithm to find the content to be translated on the web page. This algorithm is better than depth-first or breadth-first algorithms such as NodeIterator and TreeWalker in writing complex query expressions to accurately locate DOM elements. For example, the following expression can be used to directly locate deeply nested elements with specific attributes:
[0085] / html / body / div[@class='ext'] / ul / li[5]
[0086] Ignore some tags and classes, such as 'style', 'script', 'link', 'pre', etc. Some attributes are included in the list to be translated, such as the alt attribute of the img tag, the placeholder attribute of the input tag, the title attribute of the tag, etc. The smart translation component uses NanoID (an algorithm for generating short and unique string IDs) technology to generate a unique Key value for each html element for accurate node identification. Through the unique identification of the Key value, even if the subsequent node order changes, the DOM can be accurately updated.
[0087] Step 203, combining the intelligent translation gateway to perform text translation based on multidimensional array cache call on the text information to be translated, and obtaining a text translation result in the webpage translation language;
[0088] In the embodiment of the present invention, a multidimensional array is designed to represent the node queue of the translation node. The multidimensional parameters involved in the multidimensional array mainly include the routing path (url), the unique serial number (nanoID), the XPath expression (xPathExp), the node attribute (attribute), the translation language (language), the original text (originText), the translation value (transText), the terminology library matching value (technicalTerm), the timeout time (expireTime), etc.
[0089] In one case, the browser localStorage is used to cache multidimensional arrays, and localStorage is used as a first-level cache. In other words, the browser's local storage is set with a first-level cache, which is used to cache historical web page translation information stored in the form of a multidimensional array. Since only some NanoIDs need to be translated, the translation nodes can be regarded as sparse matrices and compressed storage (Compressed Storage) can be used. Therefore, in the subsequent process, if a web page has a historical translation record, the translation result can be quickly restored from the first-level cache, thereby providing users with a fast experience without delay.
[0090] In another case, if there is no historical translation record for a web page (i.e., this is the first translation), the secondary cache area on the side of the intelligent translation gateway can be combined to implement text translation based on multidimensional array cache calls to obtain text translation results in the web page translation language.
[0091] In a specific implementation, the process of performing text translation based on multidimensional array cache call on the text information to be translated by the intelligent translation gateway to obtain the text translation result in the webpage translation language can be implemented by executing the following sub-steps S01 to S03:
[0092] Step S01: Matching the text information to be translated with the historical webpage translation information through the intelligent translation component;
[0093] Step S02-1: In the case where the translation has been performed before, that is, when the first text translation result of the text information to be translated in the webpage translation language can be successfully matched from the historical webpage translation information, the intelligent translation component directly calls the first text translation result from the first-level cache area, and displays the first text translation result at the corresponding position of the text information to be translated in the display interface of the webpage;
[0094] Step S02-2: in case that the translation has not been performed before, that is, when the text translation result of the text information to be translated in the webpage translation language cannot be matched from the historical webpage translation information, the text information to be translated is sent to the intelligent translation gateway, so that the intelligent translation gateway dynamically translates the text information to be translated based on multi-dimensional array feature optimization to obtain a second text translation result;
[0095] The embodiment of the present invention designs a probabilistic data structure based on Bloom Filter on the server side, and the translation results will be stored on the server side as a secondary cache. In other words, a secondary cache area is set on one side of the intelligent translation gateway. The cache misjudgment rate objective function can be constructed by the following formula:
[0096]
[0097] in, Indicates the cache misjudgment rate; Indicates the number of hash functions; Indicates the number of node queues for translation nodes; Indicates the size of a multidimensional array.
[0098] According to the characteristics of the multidimensional array of translation content, as As the value increases, the model and algorithm are continuously optimized and dynamically adjusted The size of the value, When , the misjudgment rate is the lowest (i.e., the optimal hash function calculation). Increase, need to adjust To ensure the false positive rate Below a threshold (e.g. 0.01%), optimal Values should balance the base and index The embodiment of the present invention reversely infers through the threshold: , the solution is (Because 0.5 14 ≈0.0061%<0.01%). Among them, dynamic The expression is:
[0099]
[0100] Among them, max means to find the maximum value; Indicates rounding up; ,make sure is an integer not less than 14. By combining the theoretical optimal Value and minimum Constraints, dynamic adjustment strategies can effectively balance the misjudgment rate and computational overhead. When increasing, the theoretical optimal , if insufficient (cannot obtain the theoretical optimal ) ensures Not less than 14, so that the misjudgment rate is always less than 0.01%. This approach takes into account both mathematical optimality and the actual limitations of the system, and is suitable for real-time dynamic adjustment scenarios. The introduction of the secondary cache greatly improves the response speed and reduces the translation cost on the server side.
[0101] In combination with the above content, the specific process of the intelligent translation gateway dynamically translating the text information to be translated based on the multi-dimensional array feature optimization and obtaining the second text translation result can be achieved by executing the following sub-steps S11 to S15:
[0102] Step S11: After receiving the text information to be translated, the intelligent translation gateway forwards the text information to be translated to the secondary cache area;
[0103] Step S12: In the secondary cache area, the intelligent translation gateway considers multi-dimensional parameters at the same time and constructs a queue of nodes to be translated of the text information to be translated in the form of a multi-dimensional array;
[0104] Step S13: constructing a cache misjudgment rate objective function of the node queue to be translated based on a hash function according to the number of node queues in the node queue to be translated and the size of the multidimensional array;
[0105] Step S14: taking the minimization of the cache misjudgment rate objective function as the optimization solution objective, continuously adjusting the number of hash function values in the cache misjudgment rate objective function, and in the parameter adjustment process, taking the webpage translation language as the target translation language, calling the intelligent translator to dynamically translate the text information to be translated, and outputting the second text translation result when the cache misjudgment rate is minimized;
[0106] Step S15: Save the second text translation result, and return the second text translation result to the multilingual client.
[0107] Step S03: receiving the second text translation result returned by the intelligent translation gateway, and displaying the second text translation result at a corresponding position of the text information to be translated in the display interface of the web page through the intelligent translation component.
[0108] Step 204: extract audio information from the real-time audio and video stream in real time, perform audio translation on the audio information based on slice blocks and dual-path alignment, and obtain an audio translation result of the audio information in the webpage translation language.
[0109] In a specific implementation, the process of performing audio translation based on the slice block and dual-path alignment on the audio information to obtain the audio translation result of the audio information in the webpage translation language can be implemented by executing the following sub-steps S21 to S27:
[0110] Step S21: extracting a number of short-time audio frames from the audio information in sequence with a preset frame length as the extraction interval;
[0111] The real-time audio and video stream is parsed in real time, and after obtaining the audio information therefrom, a frame length of 25ms is taken to divide the audio signal of the audio information into several audio short-time frames.
[0112] Step S22: for each audio short-time frame, extract key audio features from the audio short-time frame;
[0113] Extract FBank features (Filterbank features) from short-time audio frames, that is, extract key features of the audio, such as Mel-spectrogram, Mel Frequency Cepstral Coefficients, and time domain features of the audio signal.
[0114] Step S23: continuously dividing the key audio features into a plurality of slice blocks;
[0115] Due to the continuity of speech and the uncertainty of duration, it is difficult for the traditional S2TT (Speech Translation to Text) architecture to reasonably encode continuous speech input. Therefore, the embodiment of the present invention proposes a real-time translation model based on slicing blocks. First, the key audio features are converted into Continuously divided into several slice blocks. Among them, represents the feature matrix, represents the number of audio frames, and is the feature dimension corresponding to each frame. Specifically, it is divided into of Slice block sequence The convolution operation is performed in each slice block using the following bidirectional processing method:
[0116]
[0117] The following one-way transmission mechanism is used between any two adjacent slice blocks:
[0118]
[0119] in, Indicates The hidden layer output of the slice block; Indicates The batch normalization layer output of the slice block; Represents a one-dimensional convolution operation; represents bidirectional self-attention; Indicates The convolutional layer output of the slice block; Indicates The convolutional layer output of the slice block; It is a masked self-attention mechanism. That is, bidirectional self-attention and convolution operations are performed within the slice block, while unidirectional transfer operations are performed between slice blocks.
[0120] It can be understood that through framing and key feature extraction operations, physical signals can be converted into feature representations to generate fine-grained time features. By dividing the slice blocks, the feature sequence can be organized into coarse-grained units that can be processed by the model, that is, the feature representation is converted into model input. Through this hierarchical design, local details of speech can be retained based on framing and feature extraction, and long-term context modeling can be supported based on the division of slice blocks, thus providing an efficient solution for streaming speech processing.
[0121] Step S24: constructing a speech-text alignment path loss and a text-text alignment path loss of key audio features;
[0122] Construct a dual-path CTC constraint. The speech-text alignment path loss can be constructed as follows:
[0123]
[0124] in, Indicates a given condition Lower path probability.
[0125] The text-text alignment path loss is constructed as follows:
[0126]
[0127] in, , , Respectively represent the source speech input (i.e., the key audio features as input), the source speech text, and the target speech text; A feature dataset representing the alignment of source speech input and source speech text path; A feature dataset representing the path alignment between source speech text and target speech text; Indicates a given condition Down Path probability.
[0128] Step S25: constructing a total loss function according to the speech-text alignment path loss and the text-text alignment path loss combined with the cross entropy loss;
[0129] The overall objective function is:
[0130]
[0131] in, is the cross entropy loss, , Both are balance coefficients.
[0132] Through the above dual-path constraints and the overall objective function, a multi-task multi-concurrency learning framework is constructed. By introducing multiple CTC classifiers, they are placed on the source speech input With source speech text , and the source speech text and the target speech text It can solve the problem of different lengths and inability to align input and output sequences, and guide the text generation strategy accordingly.
[0133] Step S26: using a pre-trained synchronous speech-to-text translation model (Speech Translation toText, S2TT), with minimizing the total loss function as the optimization solution target, and with the webpage translation language as the target translation language, a number of slice blocks are translated in real time to obtain an audio frame translation result of the audio short-time frame in the webpage translation language;
[0134] In the process of real-time translation of several slice blocks based on the pre-trained synchronous speech-to-text translation model, each slice block is based on one-dimensional convolution and uses a bidirectional self-attention mechanism to perform intra-block bidirectional processing, and any two adjacent slice blocks perform inter-block unidirectional transmission based on a masked self-attention mechanism.
[0135] Furthermore, after obtaining the audio frame translation result, the real-time translation decision value of the audio short-time frame can be calculated through the real-time translation decision function; when the real-time translation decision value is greater than a preset decision threshold, the audio frame translation result is immediately output at the audio and video subtitle display position of the web page; when the real-time translation decision value is less than or equal to the preset decision threshold, the audio frame translation result is temporarily not output, and the audio translation process of the next audio short-time frame is directly entered.
[0136] Specifically, the real-time translation decision function is defined as follows:
[0137]
[0138] in, To translate decision values in real time; represents the activation function, ensuring that the output is in the range [0,1]; is the learnable weight vector; represents the learnable bias term; represents the encoder at time step Set a hidden state. As the decision threshold, Output immediately if , otherwise wait for new input.
[0139] Therefore, through a reasonable delayed output strategy, a controllable delay mechanism can be implemented to control the model to wait for a delay in time to recognize the source speech and then generate the corresponding target text. By presenting instant results in the real-time translation process, a more comprehensive low-latency audio translation communication experience is provided.
[0140] Step S27: Integrate the audio frame translation results of all audio short-time frames to obtain the audio translation result of the audio information in the webpage translation language.
[0141] Thus, for real-time audio and video streams (including but not limited to rtsp / rtmp / HTTP-FLV / HLS / M3U8 / etc. protocols), by executing steps S21 to S27, an audio and video player plug-in supporting multi-language translation functions based on a composite model is implemented, supporting low-latency synchronous speech to text translation (S2TT) technology. At the same time, by providing a translation decision output optimization strategy, the control model generates the corresponding target text at the appropriate time of speech input, which can present high-quality translation results in the synchronous translation process.
[0142] Furthermore, during the audio translation process, the multi-dimensional features of the audio signal can be recorded and analyzed, such as pitch, voiceprint, timbre, loudness, intonation, speed, tone, etc., and the audio features can be accurately identified and classified. After massive training and regression analysis of a large-scale corpus, the real-time translation model can identify instrumental music (violin, guitar, etc.), environmental sounds (wind, rain, waves, etc.), machine operation (factory machines, car engines, electrical appliances), sound effect signals (fighting sounds, footsteps, etc.), etc. At the same time, the model can accurately define labels for voice information, such as personnel classification (age, gender, region, etc.), and further perform the following refined identification on voice information:
[0143] Occupational characteristics: The occupation of the voice user is inferred based on the audio content and speaking style. For example, if the audio contains a lot of medical terms and discusses cases, the voice user may be a doctor. Another example is that the audio mostly uses education-related vocabulary and explains knowledge, the voice user may be a teacher, etc.
[0144] Social status: For example, if the voice user uses an imperative tone and has a rich vocabulary and involves decision-making matters, he may be a senior executive or a superior. If the voice user's tone is relatively calm and polite and the content revolves around execution-level tasks, he may be a grassroots employee, and so on.
[0145] Social relationship: For example, if the voice user speaks in an intimate tone and uses intimate names, the voice user and the person they are talking to may be a couple / couple / close friends. For another example, if the voice user speaks in a more formal and respectful manner, the voice user and the person they are talking to may be colleagues, partners, superiors, etc.
[0146] Personality labels: For example, people who speak fast, loud, and passionately tend to be extroverted, cheerful, and impulsive. People who speak slowly, with a steady tone and a moderate volume tend to be introverted, calm, and cool. People who speak in a strict tone and use formal words may be more serious, and so on.
[0147] Emotional characteristics: Through voice emotion recognition technology and combined with the speaking content, the emotional labels of voice users can be divided into happiness, sadness, anger, anxiety, calmness, etc.
[0148] Based on the audio features, we can deeply explore different industry scenarios and conduct efficient, detailed, and comprehensive personalized portraits of voice users. At the same time, we can distinguish different scenarios by combining the temporal and spatial features of the audio, such as whether it overlaps, and display the translated multimodal content on the web page's audio and video player interface. For example, before displaying the automatically translated subtitles, we can add the audio features:
[0149] Example 1: (Teenager - Happy): Hey, this time I tell you, I got a good score in English again.
[0150] Example 2: (Old man - anxious - excited): Xiaoli, please ride your bike carefully, there is a car coming on the road ahead.
[0151] Example 3: (Doctor - serious - calm): Judging from the test results now, the condition is not serious, but you must receive treatment in time and take your condition seriously.
[0152] It should be noted that the audio features in the brackets in the above examples will be gradually refined as the real-time audio stream is gradually parsed, and the content displayed in the audio and video player interface will also change in real time according to the translation situation. For example, at the beginning, the brackets only show (male-smooth tone), and as the audio features evolve, they may be changed to (doctor-serious) and so on. Therefore, in the real-time translation process of the audio stream, through the refined identification of audio features and their display, it can bring great accuracy and efficiency to the entire translation work. And it plays an important auxiliary role in the content of the real-time video.
[0153] In actual application scenarios, in video conferencing, online live broadcasts and other scenarios, real-time translation of audio streams often needs to be combined with video content to provide a more complete communication experience. Through the refined recognition of audio features, we can better understand the conversation situation in the video and provide more accurate contextual information for translation. For example, when a specific scene or action appears in the video, the multilingual audio and video system can infer the speaker's intentions and emotions based on the changes in audio features, thereby providing a more appropriate translation.
[0154] By extracting audio features in real time, as well as fine-tuning and multi-modal display of audio features, the automatic translation subtitles are gradually refined before they are displayed, greatly improving the accuracy and efficiency of the entire translation process. It has broad application prospects in video surveillance, live interaction, distance education and other scenarios.
[0155] In some embodiments, in order to further improve the accuracy of translation, when the translated content is displayed on the audio and video player interface of the web page, a feedback scenario for information interaction with the user can be set. For example, a query box "Is the current translation accurate?" pops up at a location on the web page that does not affect the user's reading of the translated content; the user selects accurate or inaccurate based on the judgment; the background regularly conducts further training and optimization of the real-time translation model based on the user's accuracy feedback results to improve the accuracy of recognition translation. When the user chooses to cross out the query box, it can be determined that the user does not like this information interaction, and the query box will no longer pop up during the next translation to avoid affecting the user experience.
[0156] In an embodiment of the present invention, an audio and video processing method supporting multi-language automatic translation function is proposed. Based on a multi-language intelligent translation gateway, an intelligent translation component is injected into the web page without perception through web page injection and function callback. Combined with the optimized XPath algorithm and the text translation model based on the multidimensional array, it can support text translation in a variety of different formats, greatly improving the compatibility of the translation system. And the multi-language translation function of the web page can be automatically realized according to the client browser language. Through the intelligent translation gateway, loose coupling with the system where the source web page is located can be achieved. There is no need to modify the page source file, and there is no need to modify the back-end interface API source code, and the dynamic translation of the front-end and back-end interface API texts during the web page interaction process can also be realized, greatly reducing the cost of manual translation. For audio and video translation, based on the audio slice block, multiple CTC classifiers are introduced at the same time to construct dual-path constraints, and a composite model of multi-concurrent learning methods is obtained to achieve low-latency audio and video subtitle translation. Through the refined recognition of audio features and multi-modal display, the accuracy and efficiency of the entire translation process are greatly improved.
[0157] For better explanation, refer to Figure 3 , showing an overall flow diagram of an audio and video processing method supporting multi-language automatic translation provided by an embodiment of the present invention. It should be pointed out that this embodiment only briefly describes the general flow of audio and video processing supporting multi-language automatic translation. The specific implementation process of each step can be understood by referring to the relevant content in the aforementioned embodiment, which will not be described here. It can be understood that the present invention is not limited to this.
[0158] The intelligent translation gateway combines preset screening conditions, matches the request path of multilingual clients through regular expressions, and performs MIME type analysis based on the HTTP request header information and response content of the request path. When it is confirmed that the html source file is returned, the intelligent translation code is injected into the html source file to form an intelligent translation component.
[0159] The smart translation gateway introduces Mutation Observer in the smart translation component by creating an observer instance and passing in a callback function. It also introduces an event bubbling mechanism to monitor DOM changes on the web page. It returns the HTML source file to the multilingual client.
[0160] The multi-language client obtains the HTML source file through the browser, and when rendering the web page based on the HTML source file, loads the intelligent translation component, and sets the web page translation language based on automatic or manual methods;
[0161] The multilingual client searches for the text information to be translated and the real-time audio and video streams of the web page through the intelligent translation component, and then combines with the intelligent translation gateway to perform text translation based on multidimensional array cache calls on the text information to be translated, to obtain the text translation result in the web page translation language, and extracts audio information in real time from the real-time audio and video stream, and performs audio translation based on slicing blocks and dual-path alignment on the audio information, to obtain the audio translation result of the audio information in the web page translation language.
[0162] Reference Figure 4 , shows a structural block diagram of an audio and video processing device supporting multi-language automatic translation provided by an embodiment of the present invention, which is applied to a multi-language client, and the multi-language client is connected to a server provided with an intelligent translation gateway through a cloud network; the audio and video processing device may specifically include:
[0163] The intelligent translation component loading unit 401 is used to load the intelligent translation component and set the webpage translation language when the html source file returned by the intelligent translation gateway is obtained through the browser and the webpage is rendered based on the html source file;
[0164] A content-to-be-translated searching unit 402 is used to search for content-to-be-translated of a web page based on the intelligent translation component, wherein the content-to-be-translated includes text information to be translated and real-time audio and video streams;
[0165] A text translation unit 403 is used to perform text translation based on multi-dimensional array cache call on the text information to be translated in combination with the intelligent translation gateway to obtain a text translation result in the webpage translation language;
[0166] The audio translation unit 404 is used to extract audio information from the real-time audio and video stream in real time, perform audio translation on the audio information based on slicing blocks and dual-path alignment, and obtain an audio translation result of the audio information in the webpage translation language.
[0167] In an optional embodiment, the intelligent translation component loading unit 401 includes:
[0168] A first setting unit for webpage translation language, configured to obtain the language used by the browser after the intelligent translation component is loaded, and set the language used by the browser as a default webpage translation language;
[0169] The second setting unit for webpage translation language is used to pop up a language selection box on the display interface of the webpage after the intelligent translation component is loaded, and based on the selection operation on the language selection box, use the selected language as the webpage translation language.
[0170] In an optional embodiment, the local storage of the browser is provided with a first-level cache area, and the first-level cache area is used to cache historical web page translation information stored in a multidimensional array form; the text translation unit 403 includes:
[0171] An information matching unit, used for matching the text information to be translated with the historical webpage translation information through the intelligent translation component;
[0172] a text translation result calling unit, configured to, when a first text translation result of the text information to be translated in the webpage translation language is successfully matched from the historical webpage translation information, directly call the first text translation result from the first-level cache area by the intelligent translation component, and display the first text translation result at a corresponding position of the text information to be translated in the display interface of the webpage;
[0173] a text information sending unit to be translated, configured to send the text information to be translated to the intelligent translation gateway when a text translation result of the text information to be translated in the webpage translation language cannot be matched from the historical webpage translation information, so that the intelligent translation gateway dynamically translates the text information to be translated based on multidimensional array feature optimization to obtain a second text translation result;
[0174] The second text translation result receiving unit is used to receive the second text translation result returned by the intelligent translation gateway, and display the second text translation result at a corresponding position of the text information to be translated in the display interface of the web page through the intelligent translation component.
[0175] In an optional embodiment, a secondary cache area is provided on one side of the intelligent translation gateway; the process of the intelligent translation gateway dynamically translating the text information to be translated based on multi-dimensional array feature optimization to obtain a second text translation result includes:
[0176] After receiving the text information to be translated, the intelligent translation gateway forwards the text information to be translated to the secondary cache area;
[0177] In the secondary cache area, the intelligent translation gateway simultaneously considers multi-dimensional parameters and constructs a queue of nodes to be translated of the text information to be translated in the form of a multi-dimensional array;
[0178] According to the number of node queues in the node queue to be translated and the size of the multidimensional array, constructing a cache misjudgment rate objective function of the node queue to be translated based on a hash function;
[0179] Taking minimizing the cache misjudgment rate objective function as the optimization solution goal, continuously adjusting the number of hash function values in the cache misjudgment rate objective function, and in the parameter adjustment process, taking the webpage translation language as the target translation language, calling the intelligent translator to dynamically translate the text information to be translated, and outputting the second text translation result when the cache misjudgment rate is minimized;
[0180] The second text translation result is saved, and the second text translation result is returned to the multilingual client.
[0181] In an optional embodiment, the audio translation unit 404 includes:
[0182] An audio short-time frame extraction unit, used to extract a plurality of audio short-time frames in sequence from the audio information with a preset frame length as an extraction interval;
[0183] A key audio feature extraction unit, configured to extract key audio features from each of the audio short-time frames;
[0184] A slice block division unit, used for continuously dividing the key audio feature into a plurality of slice blocks;
[0185] A dual path loss construction unit, used to construct a speech-text alignment path loss and a text-text alignment path loss of the key audio feature;
[0186] A total loss function construction unit, used to construct a total loss function according to the speech-text alignment path loss and the text-text alignment path loss in combination with the cross entropy loss;
[0187] An audio information real-time translation unit is used to perform real-time translation of the plurality of slice blocks using a pre-trained synchronous speech-to-text translation model, with minimization of the total loss function as an optimization solution target and the webpage translation language as a target translation language, to obtain an audio frame translation result of the audio short-time frame in the webpage translation language;
[0188] The audio frame translation result integration unit is used to integrate the audio frame translation results of all the audio short-time frames to obtain the audio translation result of the audio information in the webpage translation language.
[0189] In an optional embodiment, in the process of real-time translation of the plurality of slice blocks based on a pre-trained synchronous speech-to-text translation model, each of the slice blocks is based on one-dimensional convolution and adopts a bidirectional self-attention mechanism to perform intra-block bidirectional processing, and any two adjacent slice blocks perform inter-block unidirectional transmission based on a masked self-attention mechanism.
[0190] In an optional embodiment, the device further includes a real-time translation decision unit, and the real-time translation decision unit is specifically used to:
[0191] After obtaining the audio frame translation result, calculating the real-time translation decision value of the audio short-time frame through a real-time translation decision function;
[0192] When the real-time translation decision value is greater than a preset decision threshold, the audio frame translation result is immediately outputted at the audio and video subtitle display position of the webpage;
[0193] When the real-time translation decision value is less than or equal to a preset decision threshold, the audio frame translation result is not outputted temporarily, and the audio translation process of the next audio short-time frame is directly entered.
[0194] In an optional embodiment, the intelligent translation component is formed by the intelligent translation gateway combining preset screening conditions, matching the request path of the multilingual client through regular expressions, and performing MIME type analysis on the HTTP request header information and response content of the request path. When it is confirmed that an html source file is returned, the intelligent translation code is injected into the html source file.
[0195] In an optional embodiment, the intelligent translation component introduces a Mutation Observer by creating an observer instance and passing in a callback function to monitor regular DOM changes of a web page; the intelligent translation component also introduces an event bubbling mechanism to monitor specified DOM changes of a web page.
[0196] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the aforementioned method embodiment.
[0197] It should be noted that, in order to enable those skilled in the art to better distinguish data of the same type but with different actual meanings, some technical features are distinguished by the first and the second in the embodiments of the present invention. The first and the second are only used for data distinction and have no other special meanings. It can be understood that the present invention is not limited to this.
[0198] An embodiment of the present invention further provides an electronic device, the device comprising a processor and a memory:
[0199] The memory is used to store the program code and transmit the program code to the processor;
[0200] The processor is used to execute the audio and video processing method supporting multi-language automatic translation according to the instructions in the program code.
[0201] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the audio and video processing method supporting multi-language automatic translation of any embodiment of the present invention.
[0202] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0203] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0204] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0205] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0206] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0207] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.< / q> < / blockquote>
Claims
1. An audio and video processing method supporting multi-language automatic translation, characterized in that: Applied to a multi-language client, the multi-language client is connected to a server provided with an intelligent translation gateway through a cloud network; the audio and video processing method includes: When the html source file returned by the intelligent translation gateway is obtained through the browser and a web page is rendered based on the html source file, the intelligent translation component is loaded and the web page translation language is set; Searching for content to be translated on a web page based on the intelligent translation component, wherein the content to be translated includes text information to be translated and real-time audio and video streams; In conjunction with the intelligent translation gateway, the text information to be translated is translated based on a multidimensional array cache call to obtain a text translation result in the webpage translation language; Audio information is extracted from the real-time audio and video stream in real time, and audio translation based on slice blocks and dual-path alignment is performed on the audio information to obtain an audio translation result of the audio information in the webpage translation language.
2. The audio and video processing method according to claim 1, characterized in that: The process of setting the webpage translation language includes: When the intelligent translation component is loaded, the language used by the browser is obtained, and the language used by the browser is set as the default webpage translation language; or, When the intelligent translation component is loaded, a language selection box pops up on the display interface of the web page, and based on the selection operation on the language selection box, the selected language is used as the web page translation language.
3. The audio and video processing method according to claim 1, characterized in that: The local storage of the browser is provided with a first-level cache area, and the first-level cache area is used to cache historical web page translation information stored in a multidimensional array form; the text translation based on the multidimensional array cache call of the text information to be translated in combination with the intelligent translation gateway is performed to obtain the text translation result in the web page translation language, including: Matching the text information to be translated with the historical webpage translation information through the intelligent translation component; If a first text translation result of the text information to be translated in the webpage translation language is successfully matched from the historical webpage translation information, the intelligent translation component directly calls the first text translation result from the first-level cache area, and displays the first text translation result at a corresponding position of the text information to be translated in the display interface of the webpage; If the text translation result of the text information to be translated in the webpage translation language cannot be matched from the historical webpage translation information, the text information to be translated is sent to the intelligent translation gateway, so that the intelligent translation gateway dynamically translates the text information to be translated based on multi-dimensional array feature optimization to obtain a second text translation result; The second text translation result returned by the intelligent translation gateway is received, and the second text translation result is displayed at a corresponding position of the text information to be translated in the display interface of the webpage through the intelligent translation component.
4. The audio and video processing method according to claim 3, characterized in that: The intelligent translation gateway is provided with a secondary cache area on one side; the intelligent translation gateway dynamically translates the text information to be translated based on multi-dimensional array feature optimization to obtain a second text translation result, including: After receiving the text information to be translated, the intelligent translation gateway forwards the text information to be translated to the secondary cache area; In the secondary cache area, the intelligent translation gateway simultaneously considers multi-dimensional parameters and constructs a queue of nodes to be translated of the text information to be translated in the form of a multi-dimensional array; According to the number of node queues in the node queue to be translated and the size of the multidimensional array, constructing a cache misjudgment rate objective function of the node queue to be translated based on a hash function; Taking minimizing the cache misjudgment rate objective function as the optimization solution goal, continuously adjusting the number of hash function values in the cache misjudgment rate objective function, and in the parameter adjustment process, taking the webpage translation language as the target translation language, calling the intelligent translator to dynamically translate the text information to be translated, and outputting the second text translation result when the cache misjudgment rate is minimized; The second text translation result is saved, and the second text translation result is returned to the multilingual client.
5. The audio and video processing method according to claim 1, characterized in that: The step of performing an audio translation on the audio information based on the slice block and the dual-path alignment to obtain an audio translation result of the audio information in the webpage translation language includes: Extracting a plurality of audio short-time frames in sequence from the audio information with a preset frame length as an extraction interval; For each of the audio short-time frames, extract key audio features from the audio short-time frame; Continuously dividing the key audio features into a plurality of slice blocks; Constructing a speech-to-text alignment path loss and a text-to-text alignment path loss of the key audio features; According to the speech-text alignment path loss and the text-text alignment path loss, combined with the cross entropy loss, a total loss function is constructed; By using a pre-trained synchronous speech-to-text translation model, minimizing the total loss function is taken as the optimization solution target, and the webpage translation language is taken as the target translation language, the plurality of slice blocks are translated in real time to obtain an audio frame translation result of the audio short-time frame in the webpage translation language; The audio frame translation results of all the audio short-time frames are integrated to obtain the audio translation result of the audio information in the webpage translation language.
6. The audio and video processing method according to claim 5, characterized in that: In the process of real-time translation of the plurality of slice blocks based on the pre-trained synchronous speech-to-text translation model, each slice block is based on one-dimensional convolution and adopts a bidirectional self-attention mechanism to perform intra-block bidirectional processing, and any two adjacent slice blocks perform inter-block unidirectional transmission based on a masked self-attention mechanism.
7. The audio and video processing method according to claim 5, characterized in that: Also includes: After obtaining the audio frame translation result, calculating the real-time translation decision value of the audio short-time frame through a real-time translation decision function; When the real-time translation decision value is greater than a preset decision threshold, the audio frame translation result is immediately outputted at the audio and video subtitle display position of the webpage; When the real-time translation decision value is less than or equal to a preset decision threshold, the audio frame translation result is not outputted temporarily, and the audio translation process of the next audio short-time frame is directly entered.
8. The audio and video processing method according to any one of claims 1 to 7, characterized in that: The intelligent translation component is formed by the intelligent translation gateway combining preset screening conditions, matching the request path of the multilingual client through regular expressions, and performing MIME type analysis in combination with the HTTP request header information and response content of the request path. When it is confirmed that the html source file is returned, the intelligent translation code is injected into the html source file.
9. The audio and video processing method according to claim 8, characterized in that: The smart translation component introduces Mutation Observer by creating an observer instance and passing in a callback function to monitor regular DOM changes of web pages; the smart translation component also introduces an event bubbling mechanism to monitor specified DOM changes of web pages.
10. An audio and video processing device supporting multi-language automatic translation, characterized in that: Applied to a multi-language client, the multi-language client is connected to a server provided with an intelligent translation gateway via a cloud network; The audio and video processing device comprises: An intelligent translation component loading unit, used to load the intelligent translation component and set the webpage translation language when the html source file returned by the intelligent translation gateway is obtained through the browser and the webpage is rendered based on the html source file; A content-to-be-translated searching unit, configured to search for content-to-be-translated of a web page based on the intelligent translation component, wherein the content-to-be-translated includes text information to be translated and real-time audio and video streams; A text translation unit, used for performing text translation based on multi-dimensional array cache call on the text information to be translated in combination with the intelligent translation gateway, to obtain a text translation result in the webpage translation language; The audio translation unit is used to extract audio information from the real-time audio and video stream in real time, perform audio translation on the audio information based on slice blocks and dual-path alignment, and obtain an audio translation result of the audio information in the webpage translation language.