Multi-language matching method of cloud mobile phone and related equipment

By obtaining multimodal data and combining edge node distributed processing and blockchain verification mechanisms, dynamically update the translation model and adaptive interface layout, the lag and display confusion of cloud mobile phone multilingual support is solved, and efficient and real-time multilingual matching and interface adaptation are achieved.

CN120337948APending Publication Date: 2025-07-18启朔(深圳)科技有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510492952.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In terms of multi-language support, existing cloud mobile phones have problems such as lagging language library updates, insufficient support for low resources and small languages, and confusing interface display, which affects user experience and operation efficiency.

Method used

By acquiring multimodal data, dynamically update the translation model based on edge node distributed processing and blockchain verification mechanism, and combining context-aware translation and adaptive interface layout to achieve multilingual matching.

Benefits of technology

It improves the accuracy and real-time nature of language recognition, ensures the real-time and security of language library updates, generates translation results that meet the needs of the scenario, and reduces the overlap rate of interface elements, and improves user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337948A_ABST
    Figure CN120337948A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-language matching method of a cloud mobile phone and related equipment, and relates to the technical field of cloud mobile phones, and the method comprises the following steps: obtaining multi-modal data input by a user; based on the multi-modal data, determining a target language type and to-be-translated content; obtaining a dynamically updated translation model through a block chain verification mechanism based on edge node distributed processing; based on the translation model, performing context-aware translation on the to-be-translated content to generate a multi-language matching result; and dynamically adjusting a user interface layout to adapt to display of a multi-language matching result based on the character density characteristics of the target language type. According to the method, the vertical roll steel clamping accident is effectively prevented when the plate blank is wide, the production interruption risk is reduced, and the rolling efficiency and the equipment safety are improved. According to the method, translation accuracy, real-time performance and interface adaptability are improved, the problems of language updating lagging and poor display compatibility of a traditional scheme are solved, and efficient support is provided for cross-border interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cloud mobile phones, and particularly to a multi-language matching method for cloud mobile phones and related devices. Background Art

[0002] With the acceleration of the globalization process, cloud mobile phones, as a new type of mobile terminal technology, are increasingly widely used in the international market. However, existing cloud mobile phones have significant deficiencies in multi-language support. On the one hand, traditional translation technologies rely on a centralized language library update mechanism, resulting in a lag in language library updates, being unable to adapt to emerging vocabulary and dynamic language environments, especially insufficient support for low-resource minority languages, and translation results often having semantic deviations or not conforming to language habits, seriously affecting the user experience. On the other hand, in multi-language interaction scenarios, the user interface layout often uses a fixed template, making it difficult to adapt to differences in character density of different languages, resulting in a high overlap rate of interface elements and a chaotic display, further reducing the user operation efficiency. Therefore, there is an urgent need for a multi-language matching method for cloud mobile phones to solve the above-mentioned technical problems. Summary of the Invention

[0003] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further elaborated in detail in the Detailed Description section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0004] In a first aspect, this application provides a multi-language matching method for a cloud mobile phone, the method comprising:

[0005] Obtain multi-modal data input by a user, wherein the multi-modal data includes voice data, text data, and image data;

[0006] Based on the multi-modal data, determine the target language type and the content to be translated;

[0007] Based on edge node distributed processing, obtain a dynamically updated translation model through a blockchain verification mechanism;

[0008] Based on the translation model, perform context-aware translation on the content to be translated to generate a multi-language matching result;

[0009] Based on the character density characteristics of the target language type, dynamically adjust the user interface layout to adapt to the display of the multi-language matching result.

[0010] In some embodiments, obtaining the multi-modal data input by the user includes:

[0011] Collect the user's voice signal, and preprocess the voice signal through a preset noise reduction algorithm to generate voice data;

[0012] Receive the text input of the user, and standardize the text input with a unified character set through the encoding conversion module to generate text data;

[0013] Capture the image information of the user, and use optical character recognition technology to extract the foreign language text content in the image information to generate image data;

[0014] Perform timestamp synchronization processing on the voice data, text data and image data to generate the fused multi-modal data.

[0015] In some embodiments, based on the multi-modal data, determine the target language type and the content to be translated, including:

[0016] Parse the voice data through a voice recognition engine to generate the first text content;

[0017] Perform semantic boundary division on the text data through a syntax parsing module to extract independent semantic units;

[0018] Perform regional text detection on the image data through an optical character recognition algorithm to generate the second text content;

[0019] Based on the first text content, independent semantic units and the second text content, output the initial language type probability distribution through a pre-trained language classification model;

[0020] Based on the blockchain consensus protocol and the edge node cluster, obtain the latest language library version hash value matching the initial language type probability distribution;

[0021] According to the latest language library version hash value, verify the validity of the initial language type probability distribution and determine the target language type;

[0022] Based on the grammar rules of the target language type, segment the continuous data segments that meet the translation conditions from the first text content, independent semantic units and the second text content as the content to be translated.

[0023] In some embodiments, based on edge node distributed processing, obtain a dynamically updated translation model through a blockchain verification mechanism, including:

[0024] Generate a candidate block containing the translation model update parameters based on the master node, where the candidate block includes a model version identifier and an update data hash value;

[0025] Broadcast the candidate block to multiple slave nodes in the edge node cluster to trigger each slave node to perform hash consistency verification;

[0026] Based on the Byzantine fault tolerance algorithm, count the number of slave nodes in the edge node cluster that pass the verification;

[0027] When the number of slave nodes is greater than the first preset ratio, determine that the translation model update parameters pass blockchain consensus verification;

[0028] Based on the translation model update parameters that have passed verification, load a dynamically updated translation model corresponding to the model version identifier from the distributed language library.

[0029] In some embodiments, based on the translation model, perform context-aware translation on the content to be translated to generate multilingual matching results, including:

[0030] Obtain historical interaction data in the user behavior log, where the historical interaction data includes the secondary editing frequency of the translation result and the interface dwell time;

[0031] Based on the historical interaction data, calculate the scenario translation preference parameters through a preset weight assignment rule, where the scenario translation preference parameters include business term weights and colloquial expression weights;

[0032] Based on the scenario translation preference parameters, input the content to be translated into the translation model to generate an initial translation result;

[0033] Based on the translation model, extract the context semantic features of the content to be translated;

[0034] Based on the context semantic features, dynamically correct the initial translation result to generate multilingual matching results.

[0035] In some embodiments, based on the character density feature of the target language type, dynamically adjust the user interface layout to adapt to the display of multilingual matching results, including:

[0036] Based on the character density feature of the target language type, determine the median character width;

[0037] Count the number of characters displayed in a single line in the current user interface to obtain the line character count;

[0038] Based on the median character width and the line character count, calculate the control spacing through a preset non-linear formula;

[0039] When the line character count is greater than the preset character threshold, based on the control spacing, adjust the control positions in the user interface;

[0040] According to the adjusted control positions, re-render the display content of the multilingual matching results to reduce the overlap rate of interface elements.

[0041] In some embodiments, the preset non-linear formula is determined based on the following formula, expressed as:

[0042] Control spacing = (base value × median character width) / (1 + e^(-k × line character count))

[0043] Where k is a preset attenuation coefficient, and the reference value is determined based on the character density characteristics of the target language type.

[0044] In a second aspect, the present application provides a multi-language matching device for a cloud phone, including:

[0045] A multi-modal data acquisition unit, configured to acquire multi-modal data input by a user, where the multi-modal data includes voice data, text data, and image data;

[0046] A language translation content determination unit, configured to determine the target language type and the content to be translated based on the multi-modal data;

[0047] A dynamic translation model acquisition unit, configured to acquire a dynamically updated translation model through a blockchain verification mechanism based on edge node distributed processing;

[0048] A multi-language matching result generation unit, configured to perform context-aware translation on the content to be translated based on the translation model to generate a multi-language matching result;

[0049] A language page layout adjustment unit, configured to dynamically adjust the user interface layout based on the character density characteristics of the target language type to adapt to the display of the multi-language matching result.

[0050] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program stored in the memory, the steps of the multi-language matching method for a cloud phone according to any one of the first aspects are implemented.

[0051] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the multi-language matching method for a cloud phone according to any one of the first aspects is implemented.

[0052] In summary, the present application improves the multilingual matching ability of the cloud phone by obtaining user multimodal data and dynamically updating the translation model based on the blockchain verification mechanism, combining the context-aware translation algorithm with the adaptive interface layout adjustment. Specifically, the target language type and the content to be translated are determined based on the multimodal data, effectively improving the accuracy of language recognition and the integrity of the input information; the dynamically updated translation model is obtained through edge node distributed processing and the blockchain consensus mechanism, ensuring the real-time and security of the language library update; the context-aware translation algorithm combines the user behavior log to optimize the translation preference and generate a translation result that better meets the scenario requirements; the interface layout is dynamically adjusted based on the target language character density feature, reducing the interface element overlap rate and improving the multilingual display adaptability. The overall solution realizes collaborative optimization in terms of translation accuracy, real-time performance and user interaction experience, providing efficient and reliable technical support for the international application of cloud phones.

[0053] The multilingual matching method of the cloud phone proposed in this application. Other advantages, objectives and features of this application will be partially reflected by the following description, and will also be understood by those skilled in the art through the research and practice of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this specification. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0055] Figure 1 It is a schematic flowchart of a multilingual matching method for a cloud phone provided by an embodiment of the present application;

[0056] Figure 2 It is a schematic structural diagram of a multilingual matching device for a cloud phone provided by an embodiment of the present application;

[0057] Figure 3 It is a schematic structural diagram of an electronic device for multilingual matching of a cloud phone provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] In the description, claims, and above-mentioned drawings of this application, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments.

[0059] Please refer to Figure 1 , which is a schematic flowchart of a multi-language matching method for a cloud mobile phone provided by an embodiment of this application, and specifically may include:

[0060] S110. Obtain multi-modal data input by a user, where the multi-modal data includes voice data, text data, and image data;

[0061] Exemplarily, in the multi-language matching method of a cloud mobile phone, obtaining multi-modal data input by a user is the basis for the system to achieve real-time language interaction. The multi-modal data includes three input forms: voice, text, and image. By integrating information of different modalities, the interaction intention of the user can be captured more comprehensively. For example, voice data can reflect the user's oral expression habits, text data provides structured language input, and image data supplements the text recognition ability through visual information. The three work together to improve the coverage and accuracy of language recognition.

[0062] The acquisition of multi-modal data aims to solve the limitations of a single input form. For example, voice input may be interfered by environmental noise, and text input cannot cover foreign language content in images. By synchronously collecting and preliminarily processing multi-source data, the system constructs a fused multi-modal input framework, providing unified underlying data support for subsequent language type recognition, dynamic translation, and interface adaptation. The core of this step lies in enhancing data completeness through multi-modal fusion, laying a foundation for the efficient cooperation of subsequent modules.

[0063] S120. Based on the multi-modal data, determine the target language type and the content to be translated;

[0064] Exemplarily, when determining the target language type and the content to be translated, the system realizes the accurate recognition of the language type and content segmentation based on the collaborative analysis of multi-modal data. The voice data is parsed by the voiceprint feature extraction and speech recognition engine to generate the preliminary text content; the text data is divided into semantic boundaries by the grammar parsing module to extract independent semantic units; the image data extracts visual text information through the OCR technology. By fusing the features of the three types of data, the system constructs a multi-dimensional language feature vector, inputs it into the pre-trained language classification model, and outputs the probability distribution of the initial language type, so as to comprehensively determine the target language type and ensure that the recognition result covers the complexity of different input scenarios.

[0065] Furthermore, the system combines the blockchain consensus mechanism with the collaborative verification of the edge node cluster to ensure the consistency between the language library version and the initial language type. Through hash value matching and the distributed verification process, a reliable language library version dynamically adapted to the current language environment is screened out. On this basis, according to the grammar rules of the target language type, continuous semantic segments that meet the translation conditions are segmented from the multi-modal data, providing structured input for subsequent context-aware translation and avoiding the problems of fragmented translation content or semantic fragmentation.

[0066] S130. Based on the distributed processing of edge nodes, obtain a dynamically updated translation model through the blockchain verification mechanism;

[0067] Exemplarily, in the process of dynamically updating the translation model, the system relies on the distributed computing power of the edge node cluster to achieve the efficient collaborative update of the language library and the translation model. The master node is responsible for generating candidate blocks containing the update parameters of the translation model and broadcasting them to the slave nodes through the blockchain network; each edge node performs consistency verification on the block hash value based on the improved Byzantine Fault Tolerance algorithm (PBFT) to ensure the integrity and legality of the updated data. When the number of verified nodes reaches the preset ratio, the system determines that the update parameters are valid and triggers the dynamic loading of the language library, thus ensuring that the translation model can be adapted to emerging vocabulary, minority languages, and user personalization requirements in real time.

[0068] The blockchain verification mechanism plays a key role in this process. By hash value matching and the distributed consensus rule, it solves the single point of failure and data tampering risks existing in the traditional centralized update mechanism. The distributed characteristics of the edge nodes further reduce the communication overhead and improve the real-time performance of model updates. Finally, the verified translation model is loaded from the distributed language library to the cloud mobile phone terminal, forming a closed loop of dynamic update to ensure the continuous optimization of multi-language matching capabilities and scenario adaptability.

[0069] S140. Based on the translation model, perform context-aware translation on the content to be translated to generate multi-language matching results;

[0070] Exemplarily, the core of context-aware translation lies in combining user behavior data with semantic features, dynamically optimizing translation strategies to adapt to different scenario requirements. The translation model is trained and generated based on the federated learning framework, integrating global language rules and the personalized preferences of local users. During the translation process, the model analyzes the user's historical interaction data, extracts context semantic features, generates a preliminary translation result, and dynamically adjusts the word selection and sentence structure to ensure that the output not only conforms to the grammar norms of the target language but also is close to the semantic logic of the actual application scenario.

[0071] Through lightweight model deployment and real-time semantic analysis capabilities, the initial translation result is dynamically corrected. For example, for polysemous words or ambiguous sentences, the model performs probability weighting based on context semantic features to optimize the accuracy and naturalness of the translation result. This process effectively solves the problem of rigid expression caused by semantic fragmentation in traditional translation technologies, and finally generates multi-language results that highly match the user's needs, improving the practicality and fluency of cross-language interaction.

[0072] S150 dynamically adjusts the user interface layout based on the character density features of the target language type to adapt to the display of multi-language matching results.

[0073] Exemplarily, the core of dynamically adjusting the user interface layout lies in dynamically calculating the control spacing through a preset non-linear formula based on the character density features of the target language. The system calculates the line character count and character width distribution of the current displayed content, and combines the layout rules of the target language to generate control position adjustment parameters, thus avoiding overlapping or misaligned display of interface elements. This mechanism solves the compatibility problem of traditional fixed layouts in multi-language scenarios, ensuring that the translation results can be clearly and smoothly presented in different language environments.

[0074] Furthermore, based on the real-time rendering engine and adaptive layout algorithm, the system rearranges the interface elements according to the dynamically calculated control spacing. For example, when the line character count exceeds the preset threshold, the non-linear spacing adjustment logic is triggered, and by reducing the control density or optimizing the line break strategy, the overlapping rate of the interface elements is reduced. At the same time, for dynamic adaptation to different writing directions, it ensures the visual natural coherence of multi-language content, ultimately improving the operation efficiency and experience fluency of users in multi-language interaction scenarios.

[0075] In summary, in the embodiments of the present application, by integrating multi-modal data input, blockchain dynamic verification, and context-aware translation technologies, the comprehensive performance of cloud phones in multi-language interaction scenarios is improved. The multi-modal data integrates voice, text, and image information to ensure the comprehensiveness and accuracy of language type recognition and the extraction of content to be translated. Based on edge node distributed processing and blockchain consensus mechanism, real-time and secure updates of the translation model are realized, effectively supporting the rapid adaptation of low-resource minority languages. Context-aware translation dynamically optimizes the output in combination with user behavior logs to generate matching results that meet personalized needs and scene semantics. The lightweight model deployment further reduces latency to meet the requirements of millisecond-level real-time interaction. At the same time, the adaptive layout engine based on character density features dynamically adjusts the spacing and arrangement logic of interface elements to solve the display overlap problem caused by the width difference of different language characters, improving visual clarity and operation fluency in multi-language mixed scenarios. Through privacy protection mechanisms and weak network optimization strategies, while ensuring data security, high robustness and low resource consumption are taken into account, providing efficient and reliable multi-language support for global applications.

[0076] In some examples, obtaining the multi-modal data input by the user includes:

[0077] Collecting the user's voice signal, preprocessing the voice signal through a preset noise reduction algorithm to generate voice data;

[0078] Receiving the user's text input, and standardizing the unified character set of the text input through an encoding conversion module to generate text data;

[0079] Capturing the user's image information, and using optical character recognition technology to extract the foreign language text content in the image information to generate image data;

[0080] Performing timestamp synchronization processing on the voice data, text data, and image data to generate the fused multi-modal data.

[0081] Exemplarily, the acquisition of voice data is achieved by collecting the user's original voice signal and preprocessing the signal using a preset noise reduction algorithm. The noise reduction algorithm is based on frequency domain filtering and time domain waveform analysis technologies to filter out environmental background noise and device interference signals and extract pure voice features. The preprocessed voice data is converted into a standard audio format (such as PCM encoding) to ensure compatibility with the input of subsequent speech recognition engines. This step lays the foundation for the accurate conversion of speech to text by improving the signal-to-noise ratio of the voice signal and avoiding semantic parsing errors caused by noise interference.

[0082] The text data input by the user is standardized to a unified character set through an encoding conversion module. This module identifies the encoding format of the original text (such as UTF-8, GBK, etc.), converts it to a preset unified encoding standard (such as Unicode), and eliminates the garbled characters caused by character set incompatibility in multi-language scenarios. The standardized text data is further divided into semantic boundaries and segmented into independent semantic units through natural language processing techniques to ensure structured input for subsequent language classification and translation processing, improving the efficiency and accuracy of semantic parsing.

[0083] The foreign language text content in the image data is extracted through optical character recognition technology (OCR). The OCR algorithm performs regional text detection on the image based on a convolutional neural network. After locating the text region, a sequence recognition model (such as CRNN) is used to parse the character content line by line. The extracted text is initially classified by language type and associated with the speech and text data to generate structured image text data. This technology can effectively identify text information in complex backgrounds or low-resolution images, supplement the visual text source in multi-modal input, and enhance the coverage of language recognition.

[0084] The speech, text, and image data are fused into a unified multi-modal input stream through a timestamp synchronization mechanism. The system generates timestamp tags accurate to the millisecond level for each type of data, and performs temporal matching on heterogeneous data based on a time window alignment strategy to ensure the consistency of multi-modal input in the time dimension in the same interaction scenario. The fused data is transmitted to the language recognition module through an encrypted channel, providing a complete and temporally associated input source for subsequent target language type determination and translation content extraction, avoiding semantic breaks or scenario misjudgments caused by data asynchrony.

[0085] In some instances, based on the multi-modal data, the target language type and the content to be translated are determined, including:

[0086] Parse the speech data through a speech recognition engine to generate the first text content;

[0087] Perform semantic boundary division on the text data through a syntax parsing module to extract independent semantic units;

[0088] Perform regional text detection on the image data through an optical character recognition algorithm to generate the second text content;

[0089] Based on the first text content, independent semantic units, and the second text content, output the initial language type probability distribution through a pre-trained language classification model;

[0090] Based on the blockchain consensus protocol and the edge node cluster, obtain the hash value of the latest language library version that matches the initial language type probability distribution;

[0091] Verify the validity of the initial language type probability distribution based on the hash value of the latest language library version, and determine the target language type;

[0092] Based on the grammar rules of the target language type, split the continuous data segments that meet the translation conditions from the first text content, independent semantic units, and the second text content as the content to be translated.

[0093] Exemplarily, the speech data is parsed into the first text content by a speech recognition engine. Based on a deep neural network model (such as the Transformer architecture), the speech recognition engine converts the preprocessed speech signal into a corresponding text sequence. During the parsing process, a joint decoding strategy of an acoustic model and a language model is adopted, and the recognition result is optimized by combining context information to ensure the accuracy of speech-to-text conversion. The generated text content retains the semantic integrity and intonation features of the original speech, providing the basic input for subsequent language classification.

[0094] The text data is subjected to semantic boundary division by a grammar parsing module to extract independent semantic units. The grammar parsing module adopts dependency parsing technology to identify the subject-predicate-object structure and modification relationships in the sentence, and divides the text into independent semantic units (such as phrases, clauses) according to grammar markers such as punctuation marks and conjunctions. The extracted semantic units retain the context relevance and avoid semantic fragmentation caused by over-segmentation, providing a structured input for multimodal data fusion and language type determination.

[0095] The image data is subjected to regional text detection by an optical character recognition algorithm to generate the second text content. The algorithm locates the text regions in the image based on a target detection model (such as YOLO or Faster R-CNN), and identifies the character sequence through a recurrent neural network (RNN) and a connectionist temporal classification (CTC) decoder. For complex backgrounds or skewed text, perspective transformation and binarization preprocessing are adopted to enhance the recognition robustness. The generated second text content annotates the text position information in the original image, supplementing the visual semantic information of the multimodal input.

[0096] The first text content, independent semantic units, and the second text content are input into a pre-trained language classification model, and the initial language type probability distribution is output. The model is fine-tuned based on the multilingual BERT architecture, captures cross-modal semantic correlations through a multi-head attention mechanism, calculates the probability scores of each language type, and the model outputs the probability distribution of each candidate language type, reflecting the matching degrees of different languages at the semantic, grammatical, and lexical levels, providing an initial judgment basis for subsequent verification.

[0097] Based on the blockchain consensus protocol, the edge node cluster validates the effectiveness of the initial language type probability distribution. The primary node generates a candidate block containing a hash value according to the latest language library version. The secondary nodes verify the hash consistency through an improved PBFT algorithm (the total number of nodes N≥3f + 1, where f is the number of Byzantine nodes). When more than 75% of the nodes confirm, the system obtains the hash value of the latest language library version that matches the initial probability distribution, ensuring the real-time nature and data authority of language type determination.

[0098] According to the hash value of the latest language library version, verify the effectiveness of the initial language type probability distribution. The hash value calculates the integrity of the language library data through the SHA-256 algorithm. If the initial probability distribution is consistent with the language library version corresponding to the hash value, the target language type is determined to be the language with the highest score in the probability distribution. If there is a version conflict, trigger the incremental update mechanism to reload the language library until consistency verification is achieved.

[0099] Based on the grammar rules of the target language type, split the continuous data segments that meet the translation conditions from the first text content, independent semantic units, and the second text content. The system matches the syntactic structure features of the target language through a rule engine, identifies the core semantic units to be translated, and eliminates redundant or repetitive content. The segmented data segments retain context coherence and serve as the input for context-aware translation, ensuring that the translation results conform to the expression habits and grammar norms of the target language.

[0100] In some instances, based on edge node distributed processing, obtain a dynamically updated translation model through the blockchain verification mechanism, including:

[0101] Based on the primary node generating a candidate block containing translation model update parameters, where the candidate block includes a model version identifier and an update data hash value;

[0102] Broadcast the candidate block to multiple secondary nodes in the edge node cluster, triggering each secondary node to perform hash consistency verification;

[0103] Based on the Byzantine fault tolerance algorithm, count the number of secondary nodes in the edge node cluster that pass the verification;

[0104] When the number of secondary nodes is greater than the first preset ratio, determine that the translation model update parameters pass the blockchain consensus verification;

[0105] Based on the verified translation model update parameters, load the dynamically updated translation model corresponding to the model version identifier from the distributed language library.

[0106] Exemplarily, the master node generates a candidate block according to the translation model update requirements. The candidate block contains a model version identifier and an update data hash value. The model version identifier is used to uniquely identify the current update version. The hash value is generated by one-way encrypting the update data (such as new words and grammar rules) through the SHA-256 algorithm to ensure the integrity and immutability of the data. The master node synchronously records the timestamp and the edge node cluster status information when generating the block, providing metadata support for subsequent distributed verification. The core of this step is to encapsulate the language library update requirements into a blockchain data structure, realizing the decentralization and traceability of the update process.

[0107] The master node broadcasts the candidate block to multiple slave nodes in the edge node cluster, triggering each slave node to perform a hash consistency check. After receiving the candidate block, the slave node independently calculates the hash value of the update data and compares it with the hash value recorded in the block. If the hash values are consistent, the check passes; if not, it is marked as an abnormal node. This process is implemented through the PBFT consensus algorithm, where the master node and the slave nodes interact through an encrypted communication channel (such as the TLS1.3 protocol) to ensure the security of data transmission. This mechanism ensures the global synchronization of update requests through distributed broadcasting, avoiding update interruptions caused by single-point failures.

[0108] The slave node performs a hash consistency check on the candidate block based on the improved PBFT algorithm. Specifically, the slave node calculates the hash value of the received update data and compares it with the hash value in the candidate block. If they are consistent, it is marked as passing the check. The PBFT algorithm requires the total number of nodes N to satisfy N≥3f + 1 (f is the number of Byzantine nodes), ensuring that consensus can still be reached under the condition of partial node failures or malicious behaviors. During the check process, the slave node sequentially executes message passing and signature confirmation in three stages: pre-prepare, prepare, and commit, and finally generates a check result and feedbacks it to the master node.

[0109] The master node counts the number of slave nodes in the edge node cluster that pass the check and determines whether to pass the blockchain consensus verification according to a preset ratio (such as 75%). If the number of slave nodes that pass the verification is greater than the first preset ratio (for example, 3 / 4 node confirmation), it is determined that the translation model update parameters are legal and valid; otherwise, a retransmission or rollback mechanism is triggered. This threshold design is based on the fault tolerance optimization scheme of the PBFT algorithm. Compared with the 66% confirmation threshold of the traditional PBFT, the fault tolerance rate of this application is increased to 25% in a 4-node cluster, while reducing the communication overhead by more than 50%, improving the verification efficiency and system robustness.

[0110] Based on the verified model version identifier, load the corresponding dynamically updated translation model from the distributed language library. During the loading process, the local cache of the edge node is preferentially selected to reduce network transmission latency. The updated translation model adapts to the resource limitations of the cloud mobile phone through lightweight deployment technologies (such as the Transformer model compressed by knowledge distillation), ensuring translation real-time performance (latency less than ±35ms). After the model is loaded, update logs are recorded to the blockchain, and the updated translation model takes effect in real-time, supporting the dynamic adaptation of low-resource minority languages and user personalized needs. The edge node cluster synchronously updates the local language library copy to ensure that subsequent translation requests can be executed based on the latest model, forming a closed-loop update mechanism, achieving millisecond-level language support response and the traceability and anti-tampering of the language library change trajectory, providing a data basis for subsequent incremental updates.

[0111] In some instances, based on the translation model, perform context-aware translation on the content to be translated to generate multilingual matching results, including:

[0112] Obtain historical interaction data in the user behavior log, where the historical interaction data includes the secondary editing frequency of the translation result and the interface stay duration;

[0113] Based on the historical interaction data, calculate the scenario translation preference parameters through a preset weight assignment rule, where the scenario translation preference parameters include business term weights and colloquial expression weights;

[0114] Based on the scenario translation preference parameters, input the content to be translated into the translation model to generate an initial translation result;

[0115] Based on the translation model, extract the context semantic features of the content to be translated;

[0116] Based on the context semantic features, dynamically correct the initial translation result to generate multilingual matching results.

[0117] Exemplarily, the user behavior log data is obtained by recording the historical operations of the user interacting with the cloud mobile phone, including the secondary editing frequency of the translation result and the interface stay duration. The secondary editing frequency reflects the user's dissatisfaction with the translation result, and the interface stay duration is used to evaluate the user's attention to specific translation content. The system collects and stores these data through an encrypted channel to ensure that user privacy meets the requirements of local differential privacy (ε = 0.5). After the log data is cleaned and normalized, it is input into a preset weight assignment module to provide basic input for subsequent calculation of scenario translation preference parameters.

[0118] Based on historical interaction data, calculate the scenario translation preference parameters through a preset weight allocation rule. The weight rule is dynamically adjusted according to the statistical distribution of user editing frequency and staying duration: if the editing frequency of terms in the business scenario is high, increase the weight of business terms; if the staying duration in the tourism scenario is long and the editing of colloquial expressions is frequent, increase the weight of colloquial expressions. The parameter calculation uses a weighted average algorithm and combines a time decay factor (such as exponential smoothing method) to optimize the timeliness of historical data, ensuring that the preference parameters adapt to the latest needs of users.

[0119] When inputting the content to be translated into the translation model, the system adjusts the model inference strategy based on the scenario translation preference parameters. The translation model uses a lightweight architecture compressed by knowledge distillation technology (such as a 6-layer Transformer), with the number of model parameters less than 200M and the end-to-end latency reduced by 60%. During the generation of the initial translation result, the model first matches the business term library or colloquial expression library of the target language, and combines the dynamically loaded language library data in real time to output the preliminary translation text. This step achieves millisecond-level response through edge node distributed computing, meeting the real-time interaction requirements under 5G / 6G networks.

[0120] Input the content to be translated and the scenario translation preference parameters into the translation model together to generate the initial translation result. The translation model is trained based on the federated learning framework, integrating global language rules and local user personalized data (such as business term libraries, colloquial expression habits). The model uses a lightweight Transformer architecture (such as DistilBERT compressed to 6 layers), captures the semantic relevance of the content to be translated through the attention mechanism, and combines the preference parameters to adjust the priority of vocabulary selection (such as preferentially matching high-weight terms) to generate the preliminary translation text.

[0121] The translation model extracts the context semantic features of the content to be translated through the multi-head attention mechanism, including grammatical structure, sentiment tendency, and cross-sentence relevance. For example, for polysemous words or ambiguous phrases, the model calculates the semantic weight according to the context before and after, and identifies the most likely meaning. At the same time, combined with the translation preferences recorded in the user behavior log (such as emphasizing formal expressions in the business scenario), the semantic features are dynamically weighted to generate a context-enhanced semantic vector, providing a basis for correcting the translation result.

[0122] Based on the context semantic features, the system dynamically corrects the initial translation result. The correction process includes: for polysemous words or ambiguous sentences, selecting the most suitable translation option according to the context features (e.g., "bank" is translated as "bank" in the financial scenario and "riverbank" in the natural scenario); adjusting the sentence structure to conform to the grammar habits of the target language (e.g., changing an active Chinese sentence to a passive English sentence); optimizing the term consistency (e.g., unifying the translation of repeated terms in the same conversation). The corrected result is output in real time through a lightweight model, with an end-to-end delay of less than ±35ms, meeting the real-time interaction requirements, and finally generating an accurate, natural and scenario-adapted multi-language matching result.

[0123] In some instances, based on the character density features of the target language type, the user interface layout is dynamically adjusted to adapt to the display of the multi-language matching result, including:

[0124] Based on the character density features of the target language type, determine the median character width;

[0125] Count the number of characters displayed in a single line in the current user interface to obtain the line character count;

[0126] Based on the median character width and the line character count, calculate the control spacing through a preset non-linear formula;

[0127] When the line character count is greater than the preset character threshold, based on the control spacing, adjust the positions of the controls in the user interface;

[0128] According to the adjusted control positions, re-render the display content of the multi-language matching result to reduce the overlap rate of interface elements.

[0129] Exemplarily, according to the character density features of the target language type, calculate the median character width. The character density features are obtained by analyzing the standard font specifications of the target language (such as the ligature width of Arabic and the square character width of Chinese), and are dynamically calibrated in combination with actual rendering test data (such as Monte Carlo simulation results). For example, the median Arabic character width is determined by statistically averaging the widths of common ligature combinations, and for Chinese, the average width of a single character is calculated based on the Unicode standard Chinese character library. This median serves as a reference parameter, reflecting the typical placeholder requirements of the target language in the interface layout.

[0130] Real-time statistics of the number of characters displayed in a single line (line character count) in the current user interface, and the statistics process is based on the text flow analysis of the rendering engine. For right-to-left (RTL) languages (such as Arabic), the line character count is calculated starting from the right end; for left-to-right (LTR) languages (such as English), it is calculated starting from the left end. When the line character count exceeds the preset character threshold (such as 15 characters), the system triggers the adaptive layout adjustment logic to avoid control overlap or display overflow caused by high-density characters.

[0131] The control spacing is dynamically calculated through a preset non - linear formula, which is expressed as:

[0132] Control spacing=(Base value×Median character width) / (1 + e^(-k×Characters per line))

[0133] Among them, the base value is preset based on the character density characteristics of the target language (for example, the base value for Arabic is 0.8, and for Chinese is 1.2), and k is a preset attenuation coefficient (the default k = 0.5). The formula dynamically adjusts the attenuation rate of the spacing as the number of characters per line increases through an exponential function: when the number of characters per line is small, the spacing is large to adapt to wide - character languages; after the number of characters per line exceeds the threshold, the spacing shrinks non - linearly to balance the display density and readability. This formula is verified through Monte Carlo simulation, which can reduce the overlap rate of Arabic interface elements from 17.3% in the traditional scheme to 2.1%.

[0134] When the number of characters per line exceeds the preset threshold, the system adjusts the positions of the controls in the user interface based on the calculated control spacing. The adjustment process includes horizontal spacing redistribution, vertical line - spacing optimization, and control size scaling. For example, for languages written from right to left, the controls are rearranged in a right - aligned manner, and blank areas are dynamically inserted in combination with the spacing value. The adaptive layout engine updates the control coordinate data through a real - time rendering interface to ensure a smooth and flicker - free adjustment process. At the same time, the virtual input method sandbox isolates the running environment of multi - language input methods to avoid the loss of input focus or context interruption caused by layout adjustment.

[0135] The adjusted control positions redraw the display content through a GPU - accelerated rendering engine. The system converts the new layout parameters into vertex data and texture coordinates based on OpenGL or Vulkan graphics interfaces to refresh the interface elements in real - time. After rendering, the overlap rate of the elements is verified through an overlap detection algorithm to ensure that it meets the preset standard (such as ≤2.1%). If the detection fails, it is fed back to the layout engine for iterative optimization. Finally, the multi - language matching result is presented with a clear and coherent visual effect. For example, in the Chinese wide - character scenario, the control spacing automatically expands to avoid truncation or extrusion, enhancing the user interaction fluency and visual comfort.

[0136] Please refer to Figure 2 , which is a schematic structural diagram of a multi - language matching device for a cloud phone provided by an embodiment of this application, including:

[0137] A multi - modal data acquisition unit 21, which is used to acquire multi - modal data input by the user. Among them, the multi - modal data includes voice data, text data, and image data;

[0138] A language translation content determination unit 22, which determines the target language type and the content to be translated based on the multi - modal data;

[0139] The dynamic translation model acquisition unit 23 obtains a dynamically updated translation model based on edge node distributed processing through a blockchain verification mechanism;

[0140] The multilingual matching result generation unit 24 performs context-aware translation on the content to be translated based on the translation model to generate multilingual matching results;

[0141] The language page layout adjustment unit 25 dynamically adjusts the user interface layout based on the character density characteristics of the target language type to adapt to the display of multilingual matching results.

[0142] Please refer to Figure 3 , this embodiment of the present application also provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any method for multilingual matching of a cloud phone.

[0143] Since the electronic device introduced in this embodiment is the device used for implementing a multilingual matching device of a cloud phone in this embodiment of the present application, based on the method introduced in this embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in this embodiment of the present application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in this embodiment of the present application belongs to the scope protected by this application.

[0144] In the specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.

[0145] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0146] Those skilled in the art should understand that the embodiments of the present application can provide a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media containing computer-readable program code.

[0147] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a device for implementing the functions specified in one or more blocks.

[0148] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a device for implementing the functions specified in one or more blocks.

[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a device for implementing the functions specified in one or more blocks.

[0150] Embodiments of the present application also provide a computer program product, which includes computer software instructions. When the computer software instructions run on a processing device, the processing device is caused to execute Figure 1 the process of a multi-language matching method for a cloud mobile phone in a corresponding embodiment.

[0151] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium, an optical medium, or a semiconductor medium, etc.

[0152] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0153] In several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may be in electrical, mechanical, or other forms.

[0154] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0155] In addition, the functional units in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware and / or software functional units.

[0156] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods of the various embodiments of the present application.

[0157] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

[0158] Although the preferred embodiments of this specification have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments and all changes and modifications that fall within the scope of this specification.

[0159] Obviously, those skilled in the art can make various changes and deformations to this specification without departing from the spirit and scope of this specification. Thus, if these modifications and deformations of this specification fall within the scope of the claims of this specification and their equivalent technologies, this specification also intends to include these changes and deformations.

Claims

1. A multi-language matching method for cloud mobile phones, characterized in that, Including: Obtain multi-modal data input by the user, where the multi-modal data includes speech data, text data, and image data; Based on the multi-modal data, determine the target language type and the content to be translated; Based on edge node distributed processing, obtain a dynamically updated translation model through a blockchain verification mechanism; Based on the translation model, perform context-aware translation on the content to be translated to generate a multi-language matching result; Based on the character density characteristics of the target language type, dynamically adjust the user interface layout to adapt to the display of the multi-language matching result.

2. The method according to claim 1, wherein The obtaining of the multi-modal data input by the user includes: Collect the user's speech signal, preprocess the speech signal through a preset noise reduction algorithm to generate speech data; Receive the user's text input, and perform unified character set standardization on the text input through an encoding conversion module to generate text data; Capture the user's image information, and use optical character recognition technology to extract the foreign language text content in the image information to generate image data; Perform timestamp synchronization processing on the speech data, the text data, and the image data to generate fused multi-modal data.

3. The method according to claim 1, wherein The determining of the target language type and the content to be translated based on the multi-modal data includes: Parse the speech data through a speech recognition engine to generate a first text content; Perform semantic boundary division on the text data through a syntax parsing module to extract independent semantic units; Perform regional text detection on the image data through an optical character recognition algorithm to generate a second text content; Based on the first text content, the independent semantic units, and the second text content, output an initial language type probability distribution through a pre-trained language classification model; Based on the blockchain consensus protocol and the edge node cluster, obtain the hash value of the latest language library version that matches the initial language type probability distribution; According to the hash value of the latest language library version, verify the validity of the initial language type probability distribution and determine the target language type; Based on the grammar rules of the target language type, segment continuous data segments that meet the translation conditions from the first text content, the independent semantic units, and the second text content as the content to be translated.

4. The method according to claim 1, characterized in that, The obtaining of the dynamically updated translation model through the blockchain verification mechanism based on edge node distributed processing includes: Generate a candidate block containing translation model update parameters based on the master node, where the candidate block includes a model version identifier and an update data hash value; Broadcast the candidate block to multiple slave nodes in the edge node cluster to trigger each slave node to perform hash consistency verification; Based on the Byzantine fault tolerance algorithm, count the number of slave nodes that pass the verification in the edge node cluster; When the number of slave nodes is greater than the first preset ratio, determine that the translation model update parameters pass the blockchain consensus verification; Based on the verified translation model update parameters, load a dynamically updated translation model corresponding to the model version identifier from the distributed language library.

5. The method according to claim 1, characterized in that, The performing of context-aware translation on the content to be translated based on the translation model to generate a multi-language matching result includes: Obtain historical interaction data in the user behavior log, where the historical interaction data includes the secondary editing frequency of the translation result and the interface stay duration; Based on the historical interaction data, calculate scene translation preference parameters through a preset weight allocation rule, where the scene translation preference parameters include business term weights and colloquial expression weights; Based on the scene translation preference parameters, input the content to be translated into the translation model to generate an initial translation result; Based on the translation model, extract the context semantic features of the content to be translated; Based on the context semantic features, dynamically correct the initial translation result to generate a multilingual matching result.

6. The method according to claim 1, wherein The dynamically adjusting the user interface layout based on the character density features of the target language type to adapt to the display of the multilingual matching result includes: Based on the character density features of the target language type, determine the median character width; Count the number of characters displayed in a single row in the current user interface to obtain the row character count; Based on the median character width and the row character count, calculate the control spacing through a preset non-linear formula; When the row character count is greater than the preset character threshold, based on the control spacing, adjust the positions of the controls in the user interface; According to the adjusted control positions, re-render the display content of the multilingual matching result to reduce the overlap rate of interface elements.

7. The method according to claim 6, wherein The preset non-linear formula is determined based on the following formula, expressed as: Control spacing = (reference value × median character width) / (1 + e^(-k × row character count)) where k is a preset attenuation coefficient, and the reference value is determined based on the character density features of the target language type.

8. A multi-language matching device for cloud mobile phones, characterized in that, Including: A multimodal data acquisition unit for acquiring multimodal data input by the user, where the multimodal data includes voice data, text data, and image data; A language translation content determination unit for determining the target language type and the content to be translated based on the multimodal data; A dynamic translation model acquisition unit for acquiring a dynamically updated translation model through a blockchain verification mechanism based on edge node distributed processing; A multilingual matching result generation unit for performing context-aware translation on the content to be translated based on the translation model to generate a multilingual matching result; A language page layout adjustment unit for dynamically adjusting the user interface layout based on the character density features of the target language type to adapt to the display of the multilingual matching result.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that when the processor executes the computer program stored in the memory, it implements the steps of the multilingual matching method of the cloud phone according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program, when executed by the processor, implements the multilingual matching method of the cloud phone according to any one of claims 1 to 7.

Citation Information

Cited By

  • Translation method based on government affair window double-screen intelligent translator and translator

    CN120874862A

  • Multi-language intelligent translation method and system based on short video

    CN121078267A

  • A short video-based multilingual intelligent translation method and system

    CN121078267B