Integrating scene-aware machine translation methods, storage media, and electronic devices
By integrating scene labels generated from scene-aware data with the text to be translated during the encoding stage of the Transformer network, the problem of low accuracy in contextualized short text translation by machine translation devices is solved, achieving higher translation accuracy and user experience.
Patent Information
- Application Number
- CN202011079936.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-10
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2040-10-10
AI Technical Summary
Existing machine translation devices based on Transformer neural networks lack key contextual information when processing short, contextualized texts, resulting in low translation accuracy.
By fusing scene labels generated from scene-aware data with the text to be translated during the encoding stage of the Transformer network, and then encoding and decoding them, the translation accuracy is improved.
It significantly improves the translation accuracy of contextualized short texts, solves the problem of inaccurate translation due to lack of context, and enhances the user experience.
Smart Images

Figure CN114330374B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of neural network machine translation technology, specifically to a fusion scene-aware machine translation method, storage medium, and electronic device. Background Technology
[0002] From early dictionary matching to rule-based translation combining dictionaries with linguistic expert knowledge, and then to corpus-based statistical machine translation, with the improvement of computing power and the explosive growth of data, translation based on deep neural networks, namely Neural Machine Translation (NMT), has become increasingly widely used. The development of NMT can be divided into two stages: the first stage (2014-2017) was NMT based on Recurrent Neural Networks (RNNs), whose core network architecture was an RNN; the second stage (2017 to present) is NMT based on Transformer neural networks (hereinafter referred to as NMT-Transformer), whose core network architecture is a Transformer model. In the field of machine translation, RNNs have been gradually replaced by Transformers.
[0003] However, current mainstream translation devices or products based on NMT-Transformer face the same problem: low translation accuracy when processing contextualized short texts. This is mainly because contextualized short texts are usually composed of several words or characters, and their meaning is closely related to the context. However, contextualized short texts lack key contextual information, which leads to inaccurate NMT-Transformer translation. Summary of the Invention
[0004] This application provides a method for fusion scene-aware machine translation, a storage medium, and an electronic device. Scene tags are generated based on scene-aware data collected by the electronic device. These scene tags and the text to be translated are then fused and encoded together as a source language sequence in the encoding stage of a Transformer network, and information from the source language sequence is extracted. Finally, in the decoding stage of the Transformer network, the information from the source language sequence is converted into the target language, and a translation result conforming to the scene of the text to be translated is obtained, thereby greatly improving the translation accuracy of contextualized short texts.
[0005] In a first aspect, embodiments of this application provide a fusion scene-aware machine translation method for an electronic device with machine translation capabilities, comprising: acquiring text to be translated and scene-aware data, wherein the scene-aware data is collected by the electronic device and used to determine the scene in which the electronic device is located; determining the scene in which the electronic device is located based on the scene-aware data; generating scene tags corresponding to the scene based on the scene in which the electronic device is located; inputting the scene tags and the text to be translated together as a source language sequence into an encoder for translation for encoding to obtain fusion scene-aware encoded data; and decoding the fusion scene-aware encoded data through a decoder for translation and converting it into a target language to obtain a fusion scene-aware translation result.
[0006] For example, a mobile phone with machine translation capabilities can collect scene-aware data. The collected scene-aware data is used to determine the current scene of the mobile phone, and then a scene label corresponding to the current scene of the mobile phone can be generated. When using the mobile phone for translation, the mobile phone can integrate the generated scene label for machine translation encoding and decoding, and then obtain a translation result that integrates scene awareness.
[0007] In one possible implementation of the first aspect above, the method further includes: determining the characteristics of the scene in which the electronic device is located based on the scene perception data to obtain scene state data, the scene state data being used to characterize the scene in which the electronic device is located; and classifying and statistically analyzing the scene state data to determine the scene in which the electronic device is located.
[0008] For example, a mobile phone with machine translation capabilities can determine the characteristics of the current scene based on collected scene perception data, and record these characteristics as scene state information. By classifying and statistically analyzing the scene state data obtained above, scene types containing one or more scene state information can be obtained, and each scene type corresponds to a scene in which the mobile phone is located.
[0009] In one possible implementation of the first aspect described above, the method further includes: the scene perception data is acquired by a detection element disposed in the electronic device, the detection element including at least one of a GPS element, a camera, a microphone, and a sensor. The scene perception data includes one or more of location data, image data, sound data, acceleration data, ambient temperature data, ambient light intensity data, and angular motion data.
[0010] For example, a mobile phone can continuously collect its current location data through its GPS component, ambient sound data through its microphone, temperature and light intensity data of the current scene through its temperature and ambient light sensors, and angular motion data through its gyroscope. The phone can also collect image data of features in the surrounding environment and image data of the text to be translated through its camera. Some of this scene perception data collected by the phone may be invalid when used further, but most of the scene perception data is valid data when determining the scene in which the phone is located.
[0011] In one possible implementation of the first aspect above, determining the characteristics of the scene in which the electronic device is located based on the scene perception data includes one or more of the following: determining the location name of the scene based on the location data; determining characteristic text or characteristic objects in the scene based on one or more of text and target objects in the image data, and determining the environmental characteristics of the scene; determining the noise type or noise level in the scene based on one or more of frequency, voiceprint, and amplitude in the sound data, and determining whether the scene is indoors or outdoors; determining the motion state of the electronic device in the scene based on the acceleration data and the angular motion data; determining the temperature level and light intensity level of the scene based on the ambient temperature data and the ambient light intensity data, and determining whether the scene is indoors or outdoors.
[0012] For example, a mobile phone can determine the location of its surroundings using location data, such as a shopping mall or airport. It can also identify features in the scene using image data captured by its camera. For instance, if the captured image data includes subway seats and station information, it can determine that the scene is a subway ride. If the text to be translated is subway station information, the phone can integrate the subway ride scenario for contextualized short text translation. Furthermore, a mobile phone can determine whether it is indoors or outdoors using collected sound data. Some sound data can even provide a preliminary assessment of the scene; for example, construction noise indicates an outdoor location, while the sound of mahjong tiles clattering indicates an indoor location, possibly a mahjong game. Finally, acceleration data collected by the accelerometer and angular motion data collected by the gyroscope can be used to determine the phone's current motion state. For example, the acceleration data differs between subway and bus rides, allowing the phone to determine the mode of transportation. The mobile phone can also use ambient temperature sensor and ambient light sensor to collect ambient temperature data or ambient light intensity data to determine whether the scene is indoors or outdoors. Generally, the indoor temperature is lower than the outdoor temperature in summer and higher than the outdoor temperature in winter. During the day, the indoor light intensity is lower than the outdoor light intensity, and at night, the indoor light intensity is higher than the outdoor light intensity when the lights are on.
[0013] In one possible implementation of the first aspect above, the method further includes: determining the user's motion state based on the scene perception data, wherein the user's motion state is used to determine the characteristics of the scene; wherein the scene perception data includes one or more of heart rate data and blood oxygen data.
[0014] For example, mobile phones can acquire users' heart rate and blood oxygen data by connecting to wearable devices, such as smartwatches or fitness trackers. This data helps determine if the user is exercising; during exercise or with increased intensity, heart rate and blood oxygen levels will rise significantly. The user's activity data can further determine the phone's context. For instance, a user's heart rate fluctuates considerably while exercising in a gym. When using translation services within the gym, the phone can use these heart rate changes to determine the fitness environment and enable context-aware translation of short texts. Heart rate and blood oxygen data can also determine the user's current altitude. For example, during a hike, altitude and location data can determine the user's location, allowing the phone to translate short texts in a context-aware manner.
[0015] In one possible implementation of the first aspect above, the method further includes: the order in which the scene label and the text to be translated are input into the encoder is determined based on the relevance between the text content in the text to be translated and the scene label; the greater the relevance between the text content in the text to be translated and the scene label, the closer the input distance between the text content in the text to be translated and the scene label.
[0016] The fused scene-aware encoded data includes scene feature information from the scene tags extracted by the encoder during the encoding process and text content information from the text to be translated. The encoder extracts the scene feature information and the text content information in the order in which the scene tags and the text to be translated are input into the encoder.
[0017] For example, in a dining scenario, since the scenario label is more relevant to the dish name, the scenario label "restaurant" can be placed before the dish name in the menu text to be translated when inputting it into the encoder. The scenario feature information extracted by the encoder during encoding is closer to the dish name, so the scenario feature information has a greater impact on the translation of the dish name.
[0018] In one possible implementation of the first aspect above, the method further includes: the decoder selecting words corresponding to the text content information in the target language based on the scene feature information, and generating the fused scene-aware translation result.
[0019] For example, in a dining scenario, the decoder selects words corresponding to the dish names in the target language based on the characteristics of the dining scenario to form the translation, and finally obtains an accurate translation of the dish names.
[0020] In one possible implementation of the first aspect described above, the method further includes: the generation of scene labels based on the scene state data is implemented through a classifier, and the encoder and the decoder are implemented through a neural network model. The classifier performs classification calculations on the scene state data through a classification algorithm, which includes any one of gradient boosting tree classification algorithm, support vector machine algorithm, logistic regression algorithm, and iterative binary tree three-generation decision tree algorithm; the neural network model includes a recurrent neural network machine translation model based on the Transformer network.
[0021] For example, a classifier within the phone can be used to classify and statistically analyze scene state data, determining the scene the phone is in and generating scene labels. When translating using the phone, the scene labels are fused and the trained NMT-Transformer translation model completes the translation of the contextualized short text.
[0022] Secondly, embodiments of this application provide a readable medium storing instructions that, when executed on an interactive device, cause the electronic device to perform the aforementioned fusion scene-aware machine translation method.
[0023] Thirdly, embodiments of this application provide an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the processors of the electronic device, for executing the above-described fusion scene-aware machine translation method. Attached Figure Description
[0024] Figure 1 This is a schematic diagram illustrating the application scenario of the fusion scene-aware machine translation method of this application;
[0025] Figure 2 This is a diagram illustrating examples of incorrect translations by current translation devices when translating short, contextualized texts.
[0026] Figure 3 This is a schematic diagram illustrating the steps of the fusion scene-aware machine translation method of this application;
[0027] Figure 4 This is a schematic diagram of the data transformation process in the fusion scene-aware machine translation method of this application;
[0028] Figure 5 This is a schematic diagram illustrating the process of embedding scene tags in the text to be translated, as described in this application.
[0029] Figure 6 This is a schematic diagram comparing the interfaces of a scenario-based short text translation result according to this application;
[0030] Figure 7 This is a comparative diagram of the interface for another scenario-based short text translation result in this application;
[0031] Figure 8 This is a schematic diagram of the structure of a mobile phone 100 according to an embodiment of this application;
[0032] Figure 9 This is a software structure block diagram of a mobile phone 100 according to an embodiment of this application. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of the embodiments of this application will be further described in detail below with reference to the accompanying drawings and implementation schemes.
[0034] Figure 1 The diagram shows an application scenario of the scene-aware machine translation method of this application.
[0035] like Figure 1As shown, the scenario includes electronic device 100 and electronic device 200, which are connected via a network and interact with each other. Electronic device 100 or electronic device 200 has machine translation capabilities. Users can translate text by taking photos or short videos using electronic device 100, or by directly inputting the text to be translated. The photos or videos taken by electronic device 100, as well as the voice commands input by the user, need to be converted into text data to be translated using the image-to-text conversion function or speech recognition function of electronic device 100 before translation.
[0036] Electronic device 200 can be used to train an NMT-Transformer to enable it to fuse scene labels for translation encoding and decoding. Electronic device 200 can also be used to train a classifier to enable it to generate scene labels based on scene-aware data. The NMT-Transformer and classifier trained by electronic device 200 can be ported to electronic device 100 for use.
[0037] Electronic device 100 can translate the text by using its built-in translation function, or by opening locally installed translation software or by opening an online translation webpage to interact with electronic device 200 and complete the translation of the text to be translated.
[0038] Electronic device 100 is a terminal device that interacts with users. It is equipped with application software or an application system capable of executing NMT-Transformer based on Neural Machine Translation (NMT). Electronic device 100 may also be equipped with a human-computer dialogue system to recognize user voice commands requesting the translation function, and electronic device 100 may have the function of recognizing text on images or videos to convert images or videos into text data for further translation.
[0039] Neural Machine Translation (NMT) is a machine translation method proposed in recent years. Compared to traditional Statistical Machine Translation (SMT), NMT can train a neural network that can map from one sequence to another, outputting a variable-length sequence. This results in excellent performance in translation, dialogue, and text summarization. NMT is essentially an encoder-decoder system. The encoder encodes the source language sequence and extracts information from it. The decoder then converts this information into the target language, thus completing the translation. Currently, the mainstream machine translation method is based on NMT-Transformer, whose core network architecture is a Transformer network. This means that the encoder and decoder functions in NMT-Transformer are implemented through a Transformer network. The core of the Transformer network is the self-attention layer, which calculates self-attention in the vector space. Self-attention can be understood as relevance; two vectors with high relevance have high self-attention, while two vectors with low or no relevance have low or zero self-attention.
[0040] As mentioned above, current NMT-Transformer translation methods suffer from inaccurate translations of short, contextualized texts due to a lack of context. For example, ... Figure 2 As shown, in a restaurant dining scenario, current translation devices might incorrectly translate "battered whiting" as "severely damaged cod," which is irrelevant to the actual dish name (the correct translation should be: fried cod). Therefore, in contextualized short text translation scenarios, existing machine translation methods suffer from low accuracy, leading to a poor user experience. Contextualized short texts include, but are not limited to, menu names, store names in shopping malls, and specialized terminology on immigration cards. To address this technical issue, the traditional approach is to add additional information from other dimensions during the decoding or post-processing stages of machine translation to replace the text's contextual information and improve translation accuracy. However, this approach is essentially a secondary correction of the translation result. Because the added additional information is relatively simple and its matching accuracy with the text to be translated is not high, the translation accuracy of short texts remains relatively low. Furthermore, this method of adding additional information results in significant noise during decoding and information extraction, making it unsuitable for use with Transformer networks and therefore not applicable to the current mainstream NMT-Transformer.
[0041] To address the aforementioned technical problems, this application provides a scene-aware machine translation method. Based on NMT-Transformer, scene labels are generated using scene-aware data collected by electronic devices (e.g., location data of the scene, noise data of the scene, and image data to be translated). These generated scene labels are then fused and encoded together with the text to be translated as a source language sequence in the encoding stage of the Transformer network, and information from this source language sequence is extracted. Finally, in the decoding stage of the Transformer network, the information from the source language sequence is converted into the target language, resulting in a translation that matches the scene of the text to be translated. This application significantly improves the translation accuracy of contextualized short texts by fusing scene labels into the text to be translated.
[0042] The electronic device 100 or electronic device 200 that applies the fusion scene-aware machine translation method of this application can generate corresponding scene tags based on the scene-aware data collected in real time by the electronic device 100, and embed the scene tags as part of the source language into the text to be translated for translation encoding and decoding, and finally obtain a short text translation result that conforms to the scene in which the electronic device 100 is located.
[0043] Compared to traditional methods that improve short text translation accuracy by adding extra information to replace contextual information during the decoding stage of machine translation, this application's solution comprehensively integrates scene-aware data, enabling the generated scene tags to participate in both the encoding and decoding stages of machine translation. This results in a stable improvement in translation accuracy. The scene tags integrated in this solution are more diverse than those added in existing technologies, leading to higher accuracy. Furthermore, the missing scene-aware data (e.g., a sensor element on the electronic device 100 not being turned on) does not affect the judgment of scene tags. In some embodiments, if a scene-aware data point is missing, other scene-aware data can be promptly added to generate accurate scene tags. Therefore, the scene tags integrated in this solution contain higher-quality scene feature information and allow for more flexible embedding of scene feature information. This application also solves the problem of high noise in the decoded information obtained by directly adding extra information during the decoding stage, thus improving the user experience.
[0044] It is understood that, in this application, electronic device 100 includes, but is not limited to, laptop computers, tablet computers, mobile phones, wearable devices, head-mounted displays, servers, mobile email devices, portable game consoles, portable music players, e-reader devices, televisions with one or more processors embedded or coupled thereto, or other interrupting electronic devices capable of accessing networks. Electronic device 100 can collect scene perception data through its own sensors, Global Positioning System (GPS), and cameras, and can also be used to train a classifier to generate scene labels based on the scene perception data.
[0045] It is understood that electronic device 200 includes, but is not limited to, cloud, server, laptop computer, desktop computer, tablet computer, and other electronic devices that have one or more processors embedded therein and are capable of accessing a network.
[0046] For ease of explanation, the following text uses electronic device 100 as a mobile phone and electronic device 200 as a server as examples to explain the technical solution of this application in detail.
[0047] The following is combined Figure 3 This document details the specific procedures outlined in this application. Figure 3 As shown, the fusion scene-aware machine translation method of this application includes the following steps:
[0048] 301: Mobile phone 100 acquires the text to be translated and scene-aware data, and obtains scene state data based on the scene-aware data.
[0049] The text to be translated can be obtained in various ways, including directly inputting the text through the input interface of the mobile phone 100, or taking a photo or video with the mobile phone 100. The text to be translated can also be text data obtained by recognizing and converting user voice commands, and there are no restrictions on this.
[0050] For example, after the mobile phone 100 takes a photo or video, it uses its own image recognition system to extract text information from the photo or video and convert it into text to be translated.
[0051] Mobile phone 100 acquires voice commands issued by the user. For example, the user can send voice commands to mobile phone 100 by waking up the voice assistant. Mobile phone 100 recognizes the text information in the user's voice commands through its own human-computer dialogue system and converts it into text to be translated.
[0052] The methods for acquiring scene perception data include image and sound data collected by various detection elements such as the camera, microphone, infrared sensor or depth sensor of the mobile phone.
[0053] Figure 4 This is a schematic diagram of the data transformation process in the fusion scene-aware machine translation method of this application. Figure 4 As shown, the mobile phone 100 can acquire scene perception data through microphone, gyroscope, accelerometer, GPS, and computer vision (CV). The mobile phone 100 can also acquire health status data (e.g., heart rate data acquired by PPG sensor, blood oxygen data acquired by blood oxygen sensor, etc.) or step count data acquired by wristband or watch as one of the scene perception data, which is not limited here.
[0054] In addition, there can be one or more ways to obtain the same scene perception data. For example, location information can be obtained through GPS as mentioned above, or it can be obtained through Wi-Fi signal. There are no restrictions here.
[0055] Furthermore, such as Figure 4 As shown, mobile phone 100 can analyze and obtain scene state data based on the scene perception data collected above. In the prior art, one or more scene state data determination rules can be preset in mobile phone 100.
[0056] As an example of a judgment rule, mobile phone 100 can determine whether it is indoors or outdoors based on noise type or noise level. For instance, mobile phone 100 can set noise levels in the range of 2 to 4 as indoor noise and levels 4 to 6 as outdoor noise. When the sound picked up by the microphone of mobile phone 100 is identified as level 2 noise, the corresponding scene status data obtained by mobile phone 100 is indoor noise. When the sound picked up by the microphone of mobile phone 100 is identified as level 5 noise, the corresponding scene status data obtained by mobile phone 100 is outdoor noise. Mobile phone 100 can define outdoor noise as construction noise (such as construction noise from road construction machinery) and traffic noise (such as car horns, car engine sounds, and tire friction sounds). It can also define indoor noise as sounds played in public places such as shopping malls, airports, and train stations (such as announcements and music). Furthermore, it can define some everyday noises (such as the sound of playing mahjong or noise from other entertainment venues). These noise types are primarily identified based on the different frequencies and voiceprints of the various sounds. In some embodiments, mobile phone 100 can also simply use the frequency range of the sound as the criterion for noise type determination; this is not a limitation. Therefore, mobile phone 100 can analyze the noise type or noise level of the sound (noise) collected by the microphone to obtain scene status data, determining whether it is an indoor or outdoor scene.
[0057] As an example of the judgment rule, mobile phone 100 can determine whether the current scene is a shopping mall or a station based on GPS location information and online map data; mobile phone 100 can also analyze the scene status data obtained from GPS location data and CV-captured target to be translated (text or image, such as a menu) to determine whether it is a restaurant or a restaurant in a shopping mall.
[0058] In other embodiments, as examples of determination rules, the mobile phone 100 can also analyze and obtain scenario state data such as walking, running, riding status, step count data, and movement trajectory based on its own gyroscope, accelerometer, GPS-measured location data, and heart rate data measured by wearable devices connected to the mobile phone 100, such as watches.
[0059] In other embodiments, as an example of the determination rule, the mobile phone 100 can also analyze the transportation scenario that the user is currently riding based on the images collected by CV. For example, if the images collected by CV are seats on the subway, subway station announcement display screen, bus seats, station information icons posted inside the bus, etc., the mobile phone 100 can determine that the user is in a scenario of riding the subway or taking the bus.
[0060] Understandably, when analyzing scene status data, if certain scene perception data is missing, mobile phone 100 can use other scene perception data to replace the missing scene perception data to determine the scene status data. For example, when GPS is not enabled on mobile phone 100, it cannot collect location data. Mobile phone 100 can obtain scene status data by analyzing data such as sound collected by the microphone and environmental features collected by the infrared sensor.
[0061] Scene state data can be obtained based on the analysis of one or more scene perception data. Generally, for simple and easily distinguishable scenes, the mobile phone can determine its scene state data based on a relatively small amount of scene perception data. For example, for a station or airport scene, the mobile phone may only need to integrate location information collected by GPS or sound types collected by the microphone within the station to determine the basic scene state. For more complex and difficult-to-distinguish scenes, the mobile phone may need to integrate multiple perception data to comprehensively determine the basic scene state, and no restrictions are imposed here.
[0062] It is understandable that the different methods of obtaining the text to be translated mentioned above can correspond to different scene-aware data and different ways of obtaining scene-aware data, and no restrictions are imposed here.
[0063] For example, when mobile phone 100 acquires text to be translated by taking a photo or recording a video, the scene perception data acquired by mobile phone 100 may include scene feature image data (e.g., image data acquired through CV), location data (e.g., angular motion data acquired through mobile phone gyroscope, location data acquired through mobile phone GPS), sound data (e.g., sound data acquired through microphone), and so on.
[0064] When manually inputting text to be translated, the camera on phone 100 does not need to be turned on and is therefore off. Phone 100 can acquire image data without using CV (Continuous Vision) to obtain scene perception data. The methods phone 100 acquires include location data (e.g., angular motion data collected via gyroscope, location data collected via GPS, etc.); sound data (e.g., sound data collected via microphone, etc.); motion data (e.g., heart rate data collected via smartwatches or smart bands, acceleration data collected via accelerometers, etc.); environmental data (e.g., ambient temperature data collected via temperature sensors, ambient light intensity data collected via ambient light sensors, etc.), and so on. It is understandable that in some scenarios, even without a camera interface on the phone 100 screen, the camera can still work in the background to collect CV signals.
[0065] When acquiring text to be translated via voice commands, to prevent interference between audio data and to obtain scene-aware data, the camera of phone 100 does not need to be turned on in this scenario. Phone 100 can acquire image data without using voice recognition (CV). The scene-aware data acquired by phone 100 can include location data (e.g., angular motion data measured by a gyroscope, location data acquired via GPS, etc.); user motion status data (e.g., heart rate data, blood oxygen data, etc. acquired via smartwatches, smart bracelets, etc.); environmental data (e.g., ambient temperature data acquired via a temperature sensor, ambient light intensity data acquired via an ambient light sensor, etc.), and so on. It is understandable that in some scenarios, even without a camera interface on the phone 100 screen, the phone 100 camera can still work in the background to acquire CV signals.
[0066] It is understandable that the device acquiring scene perception data and the device determining the basic scene state can be the same electronic device (for example, mobile phone 100 can both collect scene perception data and directly analyze scene state data), or they can be different electronic devices. For example, the scene perception data collected by mobile phone 100 can be sent to server 200 for further analysis to obtain scene state data; or scene perception data can be collected through smart wearable devices such as watches and bracelets and sent to mobile phone 100 for further analysis to obtain scene state data. There are no restrictions here.
[0067] 302: Mobile phone 100 generates scene tags based on the obtained scene state data.
[0068] Specifically, the mobile phone 100 classifies and labels the scene state data obtained from the above analysis, and different scene state data can correspond to the same scene label. Therefore, it can be understood that the correspondence between scene state data and scene labels is a many-to-one or one-to-one correspondence.
[0069] The scene labels generated by mobile phone 100 based on scene state data can be accomplished using a pre-trained scene classifier. For example, mobile phone 100 can train a Gradient Boosting Decision Tree (GBDT) classifier by inputting the scene state data into it, and labeling scene state data in the same or similar categories with the same scene label. The scene state data can be either sample scene state data specifically collected for training the classifier, or scene state data analyzed in actual machine translation applications. Furthermore, this scene state data can accumulate over time to form a scene state database, and the corresponding scene labels can also accumulate over time to form a scene label library. Since the classifier algorithm occupies relatively little storage space, training the classifier can be performed on electronic devices such as mobile phone 100, or on server 200; there are no restrictions on this.
[0070] The GBDT classifier described above applies the GBDT algorithm. GBDT is one of the best-fitting algorithms for the true distribution among traditional machine learning algorithms. It can be used for both classification and regression, and can also filter features. The principle of the GBDT algorithm is through multiple iterations, each iteration generating a weak classifier. Each classifier is trained based on the residuals of the previous classifier. The weak classifiers are generally required to be simple enough, with low variance and high bias. This is because the training process continuously improves the accuracy of the final classifier by reducing bias. The decision tree used in the GBDT algorithm is the CART regression tree. During training, the GBDT classifier can classify scene state data. Then, scene state data belonging to the same or similar categories can be manually labeled with the same scene label by 100 mobile phones or 200 servers. After training with a large amount of scene state sample data, a scene label database is formed. For example, the scene status data that mobile phone 100 can obtain based on GPS data is a well-known restaurant or a shopping mall; the scene status data that can be obtained based on the noise type of the sound collected by the microphone (such as indoor noise); and the scene status data that can be obtained by analyzing the photos of the target object taken by CV and the photos of the surrounding environment is a menu. Combining the above scene status data, it can be determined that the current scene is a menu translation scene in a restaurant, so "restaurant" or "restaurant" can be labeled as scene tag.
[0071] It is understandable that mobile phone 100 or server 200 can also train other classification algorithm models to generate scene labels corresponding to scene state data. Other classification algorithms include, but are not limited to, Support Vector Machine (SVM) algorithm, Logistic Regression (LR) algorithm, Iterative Dichotomiser 3 (ID3) decision tree algorithm, etc., and are not limited here.
[0072] 303: Mobile phone 100 encodes the above scene labels and the text to be translated together as a source language sequence through the encoder in NMT-Transformer, and extracts the scene and text information to be translated from the source language sequence.
[0073] Specifically, the scene label and the text to be translated are input into the encoder (encoding layer in the Transformer network) of the NMT-Transformer for encoding. The encoding layer in the Transformer network is implemented by a multi-layer self-attention network, where the attention vector output by each self-attention network layer is used as the input of the next self-attention network layer.
[0074] When inputting scene tags and the text to be translated into the encoder, two aspects need to be considered regarding how to embed scene tags into the text to be translated: first, considering which type of text in the text to be translated is more relevant to the content of the scene tags; second, considering that the position of the scene tags embedded in the text to be translated should be closer to the text content with higher relevance.
[0075] Figure 5 This diagram illustrates a process of embedding scene tags into the text to be translated during the encoding process. For example... Figure 5 As shown, for example in a menu translation scenario, X5, X6, X7, and X8 represent the words that make up the scenario labels, and X9, X... 10 X 11 X 12This refers to the words that make up the text to be translated. The text to be translated is the text data obtained from the photographed menu, and the scene label is "restaurant". The dish names in the menu have a higher relevance to the scene label, while the prices have a lower or no relevance. Therefore, when inputting the scene label and the text to be translated into the Transformer network, the scene label should be input before the dish name text, making the scene label closer to the dish name text. For example, each line of text on the menu consists of dish name + price + dish description, so the scene label "restaurant" can be input before each line of text on the menu is input into the Transformer network. It can be understood that the closer the text to the scene label is to the text to be translated, the more scene information the Transformer network extracts from the scene label while encoding and extracting textual information from it; that is, the greater the attention the scene label pays to the text. Figure 5 As shown, a strong correlation curve represents the text to be translated; conversely, the farther away the text is from the scene label, the less scene information the Transformer network extracts from the scene label while encoding and extracting text information, that is, the less attention the scene label pays to the text, which is represented by a weak correlation curve.
[0076] For example, given the menu text BATTERED WHITING 16.0M|19.0NM, if restaurant is entered before it, the source language sequence input to the Transformer network will be: Restaurant BATTERED WHITING 16.0M|19.0NM. Here, the closer the distance between Restaurant and BATTERED WHITING, the greater attention Restaurant pays to BATTERED WHITING, and the greater its influence on the translation result. Under the influence of Restaurant, BATTERED WHITING is translated as fried cod, instead of the previously incorrect translation of severely damaged cod. Conversely, the greater distance between Restaurant and 16.0M|19.0NM results in less attention Restaurant has and a smaller influence on the translation result, which remains unchanged.
[0077] 304: Mobile Phone 100 uses the decoder in NMT-Transformer to decode the scene information and text information extracted from the scene tags and text to be translated during the encoding stage into a translation expressed in the target language, and outputs the translation result.
[0078] Specifically, the decoder (the decoding layer in the Transformer network) selects the target language for decoding based on the scene information extracted by the encoder during the encoding process to obtain the translated text. Among them, the decoding layer in the Transformer network is also implemented by multiple self-attention networks.
[0079] It can be understood that since in the encoding stage of NMT-Transformer, the Transformer network has already extracted the scene information in the embedded scene tags while extracting the text information to be translated, in the decoding stage of NMT-Transformer, the Transformer network can directly select the target language for decoding based on the scene information in the scene tags, so as to obtain a translated text that is more in line with the scene.
[0080] Figure 6 This is a schematic diagram for comparing the interfaces of the translation results of a scenario-based short text of the present application. Among them, as Figure 6 (a) shows, it is the translation result interface of a traditional translation device or translation equipment. For the dish name Fisherman's Basket to be translated, the finally decoded translation is: Fisherman's Basket, which is obviously an incorrect translation result; as Figure 6 (b) shows, it is the translation result interface of a translation equipment applying the fusion scenario-aware machine translation method of the present application. Based on the scene information extracted from the scene tag (restaurant), the decoded translation for the dish Fisherman's Basket is Seafood Platter, and the translation result is correct.
[0081] Next, in combination with Figure 7 Another implementation scenario is introduced.
[0082] Figure 7 This is another schematic diagram for comparing the interfaces of the translation results of a scenario-based short text of the present application. As Figure 7 shown, it is the display of the translation result of the entry card filled in the entry and exit scenario. Among them, Figure 7 (a) shows the source language text of the entry card (customs declaration card) to be translated; Figure 7 (b) shows the translation result of a traditional translation equipment; Figure 7 (c) shows the translation result of a translation equipment applying the fusion scenario-aware machine translation method of the present application.
[0083] Combined with Figure 3 and its related description, the process of obtaining the translation result shown in Figure 7 (c) by applying the fusion scenario-aware machine translation method of the present application includes the following steps:
[0084] S1: Obtain the text to be translated and scene-aware data, and obtain scene state data based on the scene-aware data.
[0085] Specifically, on the one hand, by opening the camera of the mobile phone 100 to take a picture of the immigration card page, the mobile phone 100 uses its own image recognition system to extract the text to be translated on the photographed immigration card.
[0086] On the other hand, the mobile phone 100 collects location data via GPS, sound data via microphone, and environmental feature image data via CV as scene perception data. Based on the collected scene perception data, scene status data is obtained. For example, location data collected via GPS determines whether the current location or nearby marked geographical locations are an airport or customs; sound data collected via microphone determines whether the environment is indoors or outdoors; and environmental feature image data collected via CV determines whether there are registration windows, registration forms, etc. If some entry / exit scenarios restrict the use of CV, then scene status data can be obtained from other scene perception data instead of collecting image data via CV. The quantity and type of scene perception data collected by the mobile phone 100 are not limited here. Refer to step 301 and related descriptions above for details, which will not be repeated here.
[0087] S2: Generate a scene label "entry / exit" based on the scene state data obtained in S1 above. The pre-trained classifier in mobile phone 100 can directly and quickly generate entry / exit scene labels based on the scene state data obtained above. For details, please refer to step 302 and related descriptions above, which will not be repeated here.
[0088] S3: Mobile phone 100 uses the encoder in NMT-Transformer to encode the generated entry / exit scene tags and the text to be translated together as a source language sequence, and extracts the entry / exit scene information and the text to be translated information from the source language sequence. The specific encoding process is described in step 303 above and will not be repeated here.
[0089] S4: Mobile Phone 100 uses the decoder in the NMT-Transformer to decode the entry / exit scene information and the text to be translated extracted from the entry / exit scene tags and the text to be translated during the encoding stage, word by word, to produce a translation expressed in the target language, and outputs the translation result, such as... Figure 7 As shown in (c). The specific decoding process is described in step 304 above and will not be repeated here.
[0090] In such Figure 7In the translation results shown in (c), there are some professional terms on the immigration card, such as "Please print in capital letters", which is accurately translated as "Please fill in with capital letters". Here, "print" is correctly translated as "fill in".
[0091] In contrast, Figure 7 (b) shows the traditional translation result, which renders the sentence "Please print in capital letters" as "Please print in capital letters," which is clearly incorrect. Therefore, the translation result incorporating the scene tag (entry / exit scene) is more accurate and provides a better user experience.
[0092] As described above, in practical applications, mobile phone 100 can embed an application that uses the fusion scene-aware machine translation method of this application to achieve accurate translation of contextualized short texts; mobile phone 100 can also install application software and send the text to be translated to server 200 through interaction with server 200, and server 200 will complete the translation based on the fusion scene-aware machine translation method and then feed back the translation result to mobile phone 100; mobile phone 100 can also access the web version of the translation engine through its own browser and send the text to be translated to server 200 through interaction with server 200, and server 200 will complete the translation based on the fusion scene-aware machine translation method and then feed back the translation result to mobile phone 100, without any restrictions.
[0093] An exemplary structure of an electronic device 100 is given below with reference to embodiments of this application.
[0094] Figure 8 A schematic diagram of the structure of a mobile phone 100 according to an embodiment of this application is shown.
[0095] Mobile phone 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, buttons 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0096] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0097] The processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0098] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0099] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0100] In the embodiments of this application, the mobile phone 100 can train a scene classifier and train the encoder and decoder of NMT-Transformer through the processor 110. In the actual contextualized short text translation process, the processor 110 processes the scene-aware data and the text to be translated acquired by the mobile phone 100 and executes the fusion scene-aware machine translation method described in steps 301 to 304 above.
[0101] In some embodiments, processor 110 may include one or more interfaces.
[0102] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may also adopt different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0103] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the mobile phone 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.
[0104] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0105] The wireless communication function of mobile phone 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor. Mobile phone 100 communicates and transmits data with server 200 through the above wireless communication function.
[0106] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.
[0107] Mobile communication module 150 can provide wireless communication solutions including 2G / 3G / 4G / 5G for use on mobile phone 100. Wireless communication module 160 can provide wireless communication solutions for use on mobile phone 100 including wireless local area networks (WLAN), such as wireless fidelity (Wi-Fi), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), and other wireless communication technologies.
[0108] In some embodiments, antenna 1 of mobile phone 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling mobile phone 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0109] Mobile phone 100 implements display functions through a GPU, display screen 194, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information. The image or text captured by mobile phone 100 for the text to be translated in a contextualized short text is displayed on the aforementioned display screen 194, and the translation result of the text to be translated by mobile phone 100 is also displayed on the aforementioned display screen 194 to provide feedback to the user.
[0110] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. Mobile phone 100 may include one or N displays screens 194, where N is a positive integer greater than 1.
[0111] The SIM card interface 195 is used to connect the SIM card.
[0112] Mobile phone 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194, and application processor. Mobile phone 100 can also acquire CV signals through the above shooting functions, that is, it can acquire images of the surrounding scene or images of the text to be translated.
[0113] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, mobile phone 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0114] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The internal memory 121 can be used to store computer executable program code, including instructions. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area may store data created during the use of the mobile phone 100 (such as audio data, phonebook, etc.). Furthermore, the internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the mobile phone 100 by running instructions stored in the internal memory 121 and / or instructions stored in memory located within the processor.
[0115] In embodiments of this application, the processor 110 executes the fusion scene-aware machine translation method of this application by running instructions stored in the internal memory 121 and / or instructions stored in the memory disposed in the processor.
[0116] The mobile phone 100 can achieve audio functions such as music playback and recording through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0117] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0118] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Mobile phone 100 can be equipped with at least one microphone 170C. In some embodiments, mobile phone 100 can be equipped with two microphones 170C, which, in addition to collecting sound signals, can also achieve noise reduction. In other embodiments, mobile phone 100 can also be equipped with three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and achieve directional recording functions, etc. In the implementation of this application, sound signals can be collected through microphone 170C, and the noise level or noise type of the collected sound signals can be determined to further analyze scene state data, such as whether it is indoors or outdoors.
[0119] The 170D headphone jack is used to connect wired headphones.
[0120] The gyroscope sensor 180B can be used to determine the motion attitude of the mobile phone 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the mobile phone 100 around three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the mobile phone 100's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the mobile phone 100 through reverse movement, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing gaming scenarios.
[0121] Accelerometer 180E can detect the magnitude of acceleration of mobile phone 100 in various directions (generally three axes). When mobile phone 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the phone's posture and applied to applications such as screen orientation switching and pedometers. In the implementation of this application, certain scenario state data, such as the user's walking, running, or riding states, can be obtained by analyzing the jitter state data measured by gyroscope sensor 180B and acceleration data measured by accelerometer 180E.
[0122] A distance sensor 180F is used to measure distance. The mobile phone 100 can measure distance via infrared or laser. In some embodiments, during a shooting scenario, the mobile phone 100 can utilize the distance sensor 180F to measure distance for fast focusing.
[0123] An ambient light sensor 180L is used to sense ambient light intensity. The mobile phone 100 can adaptively adjust the brightness of its display screen 194 based on the sensed ambient light intensity. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking photos. The ambient light sensor 180L can also work in conjunction with a proximity sensor 180G to detect whether the mobile phone 100 is in a pocket, preventing accidental touches. In the implementation of this application, scene state data can be analyzed based on the ambient light intensity sensed by the ambient light sensor 180L, for example, determining whether the current scene is indoors or outdoors.
[0124] Keypad 190 includes a power button, volume buttons, etc. Keypad 190 can be a mechanical keypad or a touch keypad. Mobile phone 100 can receive keypad input and generate key signal inputs related to user settings and function control of mobile phone 100.
[0125] The software system of mobile phone 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to illustrate the software structure of mobile phone 100.
[0126] Figure 9 This is a software structure block diagram of the mobile phone 100 according to an embodiment of the present invention.
[0127] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0128] The application layer can include a series of application packages.
[0129] like Figure 9As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.
[0130] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0131] like Figure 9 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0132] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0133] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0134] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0135] The phone manager is used to provide communication functions for the mobile phone 100. For example, it manages call status (including connection, hang-up, etc.).
[0136] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0137] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, playing alert sounds, vibrating the phone at 100Hz, and flashing indicator lights.
[0138] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0139] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0140] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0141] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0142] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0143] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0144] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0145] A 2D graphics engine is a graphics engine for 2D drawing.
[0146] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0147] The following example, using a menu translation scenario, illustrates the workflow of the mobile phone's software and hardware.
[0148] When touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including operations such as opening translation software or opening camera 193). The raw input event is stored in the kernel layer. The application framework layer retrieves the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking a touch click operation as an example, where the control corresponding to the click operation is the camera application icon, the camera application calls the interface of the application framework layer to launch the camera application, and then calls the kernel layer to launch the camera driver, capturing a static image or video of the menu to be translated through camera 193.
[0149] In this specification, the reference to "an embodiment" or "an embodiment" means that a specific feature, structure, or characteristic described in connection with the embodiment is included in at least one exemplary implementation or technology disclosed in this application. The phrase "in an embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment.
[0150] This application also discloses means for performing operations in text. Such means may be specifically constructed for the claimed purpose or may include a general-purpose computer selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored on a computer-readable medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, application-specific integrated circuits (ASICs), or any type of medium suitable for storing electronic instructions, and each may be coupled to a computer system bus. Furthermore, the computer mentioned in the specification may include a single processor or may employ an architecture involving multiple processors for increased computing power.
[0151] The processes and displays presented herein do not inherently relate to any specific computer or other device. Various general-purpose systems may also be used with the programs based on the teachings herein, or it may prove convenient to construct more specialized devices to perform one or more method steps. Structures for various such systems are discussed in the following description. Furthermore, any specific programming language sufficient to implement the techniques and embodiments disclosed herein can be used. Various programming languages may be used to implement this disclosure, as discussed herein.
[0152] Furthermore, the language used in this specification has been primarily chosen for readability and instructional purposes and may not have been chosen to depict or limit the subject matter disclosed. Therefore, this application disclosure is intended to illustrate, rather than limit, the scope of the concepts discussed herein.
Claims
1. A scene-aware machine translation method for use in electronic devices with machine translation capabilities, characterized in that, The method includes: The text to be translated and scene-aware data are acquired. The scene-aware data is collected by the electronic device and used to determine the scene in which the electronic device is located. Based on the scene perception data, the scene in which the electronic device is located is determined; Based on the scene in which the electronic device is located, generate scene tags corresponding to the scene; The scene labels and the text to be translated are input together as source language sequences into an encoder for translation to obtain scene-aware encoded data. The order in which the scene labels and the text to be translated are input into the encoder is determined based on the relevance between the text content in the text to be translated and the scene labels. The greater the relevance between the text content in the text to be translated and the scene labels, the closer the text content in the text to be translated is to the input distance of the scene labels. When the text to be translated is encoded, the more scene information is extracted from the scene labels. The fused scene-aware encoded data is decoded and converted to the target language using a decoder for translation to obtain the fused scene-aware translation result.
2. The method according to claim 1, characterized in that, Based on the scene perception data, the scene in which the electronic device is located is determined, including: Based on the scene perception data, the characteristics of the scene in which the electronic device is located are determined to obtain scene state data, which is used to characterize the scene in which the electronic device is located; and The scene status data is classified and statistically analyzed to determine the scene in which the electronic device is located.
3. The method according to claim 2, characterized in that, The scene perception data is collected by a detection element installed in the electronic device, which includes at least one of a GPS element, a camera, a microphone, and a sensor.
4. The method according to claim 3, characterized in that, The scene perception data includes one or more of the following: location data, image data, sound data, acceleration data, ambient temperature data, ambient light intensity data, and angular motion data.
5. The method according to claim 4, characterized in that, The characteristics of the scene in which the electronic device is located are determined based on the scene perception data, including one or more of the following: Based on the location data, determine the location name of the scene; Based on one or more of the text and target objects in the image data, determine the characteristic text or object in the scene, and determine the environmental features of the scene; Based on one or more of the frequency, voiceprint, and amplitude in the sound data, determine the noise type or noise level in the scene, and determine whether the scene is indoors or outdoors; Based on the acceleration data and the angular motion data, the motion state of the electronic device in the scene is determined; Based on the ambient temperature data and the ambient light intensity data, the temperature level and light intensity level of the scene are determined, and it is determined whether the scene is indoors or outdoors.
6. The method according to claim 5, characterized in that, Determining the scene in which the electronic device is located based on the scene perception data further includes: The user's motion state is determined based on the scene perception data, and the user's motion state is used to determine the characteristics of the scene. The scene perception data includes one or more of heart rate data and blood oxygen data.
7. The method according to claim 1, characterized in that, The fused scene-aware encoded data includes scene feature information extracted by the encoder from the scene tags during the encoding process and text content information from the text to be translated. The encoder extracts the scene feature information and the text content information in the order in which the scene label and the text to be translated are input into the encoder.
8. The method according to claim 7, characterized in that, The decoder selects words in the target language that correspond to the text content information based on the scene characteristics information, and generates the fused scene-aware translation result.
9. The method according to claim 8, characterized in that, The generation of scene labels based on the scene state data is achieved through a classifier, and the encoder and decoder are achieved through a neural network model.
10. The method according to claim 9, characterized in that, The classifier performs classification calculations on the scene state data through a classification algorithm, which includes any one of the following: gradient boosting tree classification algorithm, support vector machine algorithm, logistic regression algorithm, and iterative binary tree 3rd generation decision tree algorithm; The neural network model includes a recurrent neural network machine translation model based on the Transformer network.
11. A readable medium, characterized in that, The readable medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the method of any one of claims 1 to 10.
12. An electronic device, characterized in that, include: Memory, used to store instructions executed by one or more processors of an electronic device, and A processor is one of the processors in an electronic device, used to perform the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Preprocessing method, equipment and storage medium of human-machine interaction system
CN108803879A
Instant translation method and device, computer equipment and storage medium
CN111709431A