Game real-time intelligent translation method and device based on interactive reinforcement learning
Through interactive augmented learning and augmented reality technology, foreign language text in the game screen can be recognized and translated in real time, solving the problem of difficult foreign language text in overseas games, realizing real-time, automatic and immersive translation effects, and improving user experience.
Patent Information
- Application Number
- CN202510560068.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In the prior art, foreign language texts of overseas games are difficult to be understood by domestic users in real time, resulting in cumbersome operations, lack of real-time and immersion, and cannot meet users' translation needs.
Through an interactive enhanced learning method, game screen information and user operation habits are obtained, user attention areas are identified, context-enhanced image recognition model is used for text recognition, and target language text is generated by combining multimodal fusion translation model, real-time translation is used for augmented reality, and translation parameters are dynamically adjusted to adapt to game scene changes.
Real-time, automatic and immersive translation of game screen content, can process a large number of continuous text, and the translation results are in line with the game scene, improving the user experience and translation accuracy.
Smart Images

Figure CN120087378B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of game screen translation, and in particular to a method and device for real-time intelligent translation of games based on interactive reinforcement learning. Background Art
[0002] With the globalization of the gaming industry, overseas games are becoming increasingly popular in the domestic market. However, many overseas games have not yet been localized, and the foreign language text contained in the game interface (such as plot dialogue, item descriptions, and operating instructions) is difficult for domestic users to understand, seriously affecting the user experience.
[0003] Existing technologies often require users to manually capture game footage and switch to third-party translation software for text recognition and translation. This approach is not only cumbersome but also lacks real-time translation results, leading to frequent interruptions during gameplay and reducing user immersion and efficiency. Therefore, achieving efficient, accurate, and real-time translation of foreign language text within gaming environments has become an urgent technical need. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a method and device for real-time intelligent translation of games based on interactive reinforcement learning to solve the problems of the existing technology, such as cumbersome operation, lack of real-time performance, difficulty in processing large amounts of continuous text, and translation results that do not conform to the game scenario.
[0005] In a first aspect of an embodiment of the present application, a method for real-time intelligent translation of games based on interactive reinforcement learning is provided, comprising: when detecting that a user has turned on a translation function, obtaining current game screen information and user operation habit information, and determining a corresponding screen capture mode based on the current game screen information and user operation habit information; dynamically sensing the current game screen according to the screen capture mode, identifying the user's focus area and performing local capture to obtain image data within the user's focus area; using a context-enhanced image recognition model to perform text recognition on the image data in the user's focus area, and obtaining recognized text based on a contextual knowledge base specific to the game scene and user operation history information; inputting the recognized text, corresponding image data, and user operation history information into a multimodal fusion translation model to generate target language text; based on the correspondence between the recognized text and the target language text, rendering the target language text on the original foreign language area using augmented reality, and dynamically adjusting the presentation parameters of the target language text according to changes in the game screen.
[0006] According to a second aspect of an embodiment of the present application, a real-time intelligent translation device for games based on interactive reinforcement learning is provided, comprising: an acquisition module for acquiring current game screen information and user operation habit information when detecting that a user has turned on a translation function, and determining a corresponding screen capture mode based on the current game screen information and user operation habit information; a perception module for dynamically perceiving the current game screen according to the screen capture mode, identifying the user's focus area, and performing local capture to obtain image data within the user's focus area; a recognition module for performing text recognition on the image data within the user's focus area using a context-enhanced image recognition model, and obtaining recognized text based on a contextual knowledge base specific to the game scene and user operation history information; a generation module for inputting the recognized text, corresponding image data, and user operation history information into a multimodal fusion translation model to generate target language text; and a rendering module for rendering the target language text on the original foreign language area using an augmented reality method based on the correspondence between the recognized text and the target language text, and dynamically adjusting the presentation parameters of the target language text according to changes in the game screen.
[0007] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.
[0008] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0009] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:
[0010] When it is detected that the user has turned on the translation function, the current game screen information and user operation habit information are obtained, and the corresponding screen capture mode is determined according to the current game screen information and user operation habit information; according to the screen capture mode, the current game screen is dynamically perceived, the user's focus area is identified and local capture is performed to obtain image data in the user's focus area; the context-enhanced image recognition model is used to perform text recognition on the image data in the user's focus area, and the recognized text is obtained based on the context knowledge base exclusive to the game scene and the user's operation history information; the recognized text and the corresponding image data and the user's operation history information are input into the multimodal fusion translation model to generate the target language text; based on the correspondence between the recognized text and the target language text, the target language text is rendered on the original foreign language area using augmented reality, and the presentation parameters of the target language text are dynamically adjusted according to the changes in the game screen. This application realizes real-time, automatic, and immersive translation of game screen content, which can not only process large amounts of continuous text, but also make the translation results more in line with the game scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 1 is a flow chart of a method for real-time intelligent translation of games based on interactive reinforcement learning provided by an embodiment of the present application;
[0013] Figure 2 1 is a schematic diagram of the structure of a real-time intelligent game translation device based on interactive enhanced learning provided by an embodiment of the present application;
[0014] Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0015] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0016] With the rapid development of the global gaming market, many overseas games have not been localized (i.e., translating or localizing in-game text, prompts, and plots). This has resulted in Chinese users experiencing difficulty understanding foreign language text within the game interface. In-game text often includes game instructions, plot dialogue, and item descriptions, directly impacting players' understanding of the game content and overall gameplay experience.
[0017] The common practice in current technology is that when players need translation, they often manually take screenshots of the game screen and then switch to third-party translation software for text recognition and translation. This process has the following major shortcomings:
[0018] The operation is cumbersome: manually take a screenshot, switch applications, wait for translation, and then switch back to the game.
[0019] Lack of real-time performance: Players cannot see the translation results directly and instantly in the game, which affects the game's continuity and immersion.
[0020] Difficulty processing large amounts of continuous text: For game text that appears frequently or changes rapidly, manual screenshots cannot keep up with the progress of the game, nor can they meet the needs of continuous scene translation.
[0021] Therefore, the existing translation methods have obvious shortcomings in user experience and efficiency, and there is an urgent need for a more automated, real-time and gaming-friendly translation solution.
[0022] The contents of the technical solution of this application are described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] Figure 1 Schematic diagram of the process of the real-time intelligent translation method of the game based on interactive enhanced learning provided by the embodiment of the present application. Figure 1 As shown, the real-time intelligent translation method for games based on interactive reinforcement learning may specifically include:
[0024] S101, when it is detected that the user turns on the translation function, obtaining current game screen information and user operation habit information, and determining a corresponding screen capture mode according to the current game screen information and user operation habit information;
[0025] S102, dynamically sensing the current game screen according to the screen capture mode, identifying the user's focus area and performing partial capture to obtain image data within the user's focus area;
[0026] S103, using a context-enhanced image recognition model to perform text recognition on the image data within the user's focus area, and obtaining recognized text based on a contextual knowledge base specific to the game scene and user operation history information;
[0027] S104, inputting the recognized text, corresponding image data, and user operation history information into a multimodal fusion translation model to generate a target language text;
[0028] S105 , based on the correspondence between the recognized text and the target language text, the target language text is rendered on the original foreign language area using augmented reality, and the rendering parameters of the target language text are dynamically adjusted according to changes in the game screen.
[0029] In some embodiments, obtaining current game screen information and user operation habit information, and determining a corresponding screen capture mode according to the current game screen information and user operation habit information includes:
[0030] Analyze the current game screen information to obtain the screen change frequency and text density;
[0031] Collect statistics on user operation habits to obtain user interaction frequency and historical usage behavior;
[0032] Determine the screen capture mode based on the frequency of screen changes, text density, user interaction frequency, and historical usage behavior;
[0033] Set the corresponding floating display permission according to the screen capture mode, and apply the screen capture mode and floating display permission to the subsequent local capture and translation rendering process.
[0034] Specifically, in this embodiment, when a user clicks or activates the translation function in an overseas game (for example, by clicking the floating ball on the game interface), the system first initiates the collection and analysis of current game screen information and user operation habits. This process specifically includes the following steps:
[0035] First, the system performs real-time or periodic sampling statistics on the current game screen in the background to determine whether there are a large number of continuous text prompts in the game scene, or whether the screen is in a high-frequency update state.
[0036] For example, the system can identify the screen switching speed by reading the image frame rate, and analyze the text density by detecting the number and distribution of characters in the screenshot.
[0037] If it is detected that the game is in a plot dialogue scene or a scene where operation prompts appear frequently, the text density is determined to be high; if it is detected that the game screen is mainly static scenes, or there is only a small amount of text in the scene, the text density is determined to be low.
[0038] Furthermore, the system also collects statistics and analyzes the user's operating habits to obtain the user's interactive operation frequency and historical usage behavior.
[0039] For example, the system determines whether the user is in a high-frequency operation state based on the number of mouse clicks, keyboard input frequency, or touch operation records in the game.
[0040] Through historical user usage behavior, we can understand the frequency of users' translation needs in previous game processes. For example, some users tend to frequently enable the translation function in plot scenes, but rarely use translation in ordinary level battles.
[0041] Furthermore, the system performs correlation analysis on the picture change frequency, text density and user operation habit information to comprehensively determine the screen capture mode.
[0042] If the game screen updates quickly, the text appears densely, and the user operates frequently, the system will intelligently recommend and adaptively switch to high-frequency capture mode.
[0043] If the game screen is relatively static and the user's translation needs are limited, switch to low-frequency capture mode or on-demand capture mode to reduce the use of system resources.
[0044] Furthermore, after the corresponding capture mode is determined, the system will set corresponding floating display permissions according to different capture modes.
[0045] In high-frequency capture mode, the floating display permission allows frequent screenshots of the game in short time intervals and displays translation prompts on the screen in real time.
[0046] In low-frequency capture mode, floating display permissions are more limited, and screenshots and translations are only performed when possible text content is detected or manually triggered by the user.
[0047] At the same time, the system will prompt the user which capture mode has been entered through a pop-up window or a light notification, so that the user knows the balance between the current performance overhead and translation frequency.
[0048] Ultimately, the determined screen capture mode and floating display permissions are applied to subsequent steps of the method described in this embodiment, such as partial image capture, text recognition, and translation rendering. The system continuously monitors changes in the game scene and user operations. If significant changes are found, the capture mode can be updated again and the floating display permissions adjusted accordingly.
[0049] Based on the above-mentioned embodiments of this application, it is possible to understand in which circumstances the system will switch to which screen capture mode, and how to allocate floating display permissions in combination with user operation habit information. This example only provides a detailed description of the portion of the technical solution of this application that "obtains current game screen information and user operation habit information, and determines the corresponding screen capture mode and floating display permission accordingly." For other technical details not covered, please refer to other embodiments or related descriptions of this application.
[0050] In some embodiments, a context-enhanced image recognition model is used to perform text recognition on image data within the user's focus area, and recognized text is obtained based on a contextual knowledge base specific to the game scene and user operation history information, including:
[0051] The image data within the user's focus area is fed into a context-enhanced image recognition model to extract text information;
[0052] Obtain terms and expressions related to the current game environment from the game scene-specific contextual knowledge base;
[0053] Based on the user operation history information, the extracted text information is matched with the context, and the text information is recognized and corrected in combination with terms and expressions to generate recognized text.
[0054] Specifically, in this embodiment, to improve text recognition accuracy and save system resources, when a user activates the floating ball by clicking or touching it, the system first detects the user's focus area in the game interface (such as the location of the mouse cursor or the touch hotspot). For the detected user focus area, the system performs the following processing steps:
[0055] The system compares the game screen of the previous frame and the current frame, determines the changes in the local area, and performs differential capture on the high-value area (that is, the area of user attention) instead of taking a screenshot of the entire screen.
[0056] For example, when the user's cursor is detected hovering over a certain item description or subtitle, the system captures only the image within that focus area. This localized capture method significantly reduces the amount of image data to be processed and improves recognition efficiency.
[0057] Furthermore, the cropped image data of the region of interest is input into an enhanced OCR model (such as a Transformer-OCR model) to extract possible text information.
[0058] This OCR model has the ability to recognize irregular fonts, small amounts of blurred or overlapping text, and can extract text content from relatively complex game screens.
[0059] Compared with traditional OCR, this model can learn richer character structure features and contextual associations during training.
[0060] Furthermore, to improve the accuracy and scene adaptability of text recognition, this embodiment introduces a contextual knowledge base dedicated to game scenes during the OCR recognition process. This knowledge base contains:
[0061] Specialized terms, item names, character nicknames, plot dialogue templates, etc. in a specific gaming field; vocabulary collections pre-imported by developers or automatically accumulated after long-term use.
[0062] When the OCR engine recognizes text, it will compare and match the preliminary recognition results with the knowledge base to check whether the key words appear in the game-specific dictionary or are similar to it, so as to make corrections or completions.
[0063] At the same time, the system combines the user's historical operation information in the game to further improve the accuracy of text recognition, for example:
[0064] If the user has translated certain specific nouns or dialogues multiple times in the same or similar scenarios, the proper nouns or dialogue expressions will have a higher priority matching weight;
[0065] When new professional vocabulary or dialogue forms are detected, the system records them in the knowledge base for reference in subsequent recognition.
[0066] By combining user operation history with the knowledge base, the OCR model can more accurately correct recognition errors, distinguish similar words or phrases, and form recognized text that is more consistent with the actual content of the game.
[0067] Finally, the system completes the correction and confirmation of the recognized text based on the OCR extraction results, knowledge base comparison and user historical feedback, and provides the generated recognized text to the subsequent translation unit for translation processing.
[0068] If the recognized text is a description of a game item, the system may also associate it with the corresponding item icon, item attributes, and other information so that the translation module can obtain a more complete context when generating the translation.
[0069] According to the above-mentioned embodiments of the present application, it can be seen that on the basis of only differential local capture of the user's attention area, the context-enhanced OCR engine is combined with the game scene-specific knowledge base and the user operation history, which can achieve high-precision recognition and real-time processing of complex game texts without increasing too much system burden.
[0070] In some embodiments, the recognized text, corresponding image data, and user operation history information are input into a multimodal fusion translation model to generate target language text, including:
[0071] Preprocess the recognized text, the corresponding image data, and the user operation history information to obtain input features that can be used for multimodal fusion;
[0072] The input features are input into the multimodal fusion translation model so that the multimodal fusion translation model translates the recognized text based on the comprehensive semantic information of the text, image and user operation history to obtain the target language text.
[0073] Specifically, in this embodiment, to obtain translation results that are more relevant to the game scene, after the system completes the recognition of foreign text in the game screen, it will input the recognized text, corresponding image data, and user operation history information into the multimodal fusion translation model to generate the target language text that conforms to the game context. The main process includes the following steps:
[0074] First, text preprocessing is performed: the text obtained through OCR recognition is segmented or segmented, and basic character normalization is performed. For example, special characters, symbols, and common delimiters in games (such as "—" and "·") can be unified or cleaned to ensure consistency in subsequent translation.
[0075] Next, image data extraction takes place: the image area or local image information corresponding to the recognized text is formatted and adjusted in resolution so that the multimodal model can extract visual features related to the text content. Specifically, the system uses image cropping, scaling, and feature vectorization to represent the image data as a model-recognizable graph vector or embedding vector.
[0076] Next, user operation history information is summarized: the system extracts the user's recent interactive behaviors during the game, such as mouse click locations, touch hotspots, frequency of opening the translation function, and the user's previous correction records for similar text. To enable this information to be fully utilized by the multimodal model, this embodiment adds timestamps or operation category labels to these interactive behaviors and converts them into vectorized representations or feature codes that can be processed by the model.
[0077] Furthermore, after completing the above preprocessing, the system will perform a preliminary combination of text features, image features, and user operation history features before multimodal fusion. In this embodiment, the system first embeds or vectorizes the text features, image features, and operation history features separately, and then matches and aligns them based on temporal or content relevance, thereby obtaining an input feature matrix that can be directly processed by the multimodal model.
[0078] For example, the position of the text on the screen is identified through index OCR and paired with the image features at the same position; at the same time, the user's operation records in that time period or scene are annotated to the same feature index to construct a multimodal information unit for that spatiotemporal position.
[0079] Furthermore, the constructed multimodal input features are passed to the multimodal fusion translation model, which can translate the recognized text based on the comprehensive semantics of text, images and user operation information.
[0080] In this embodiment, the multimodal fusion translation model is capable of simultaneously processing and correlating multiple input sources. For example, the model uses image features to determine the likely game context of the text (e.g., item descriptions, plot dialogues, or system prompts). It also considers the user's previous translations or corrections of text in similar contexts to determine which vocabulary or sentence structures are most likely to align with the player's expectations.
[0081] By applying joint attention or other fusion mechanisms to text, images, and action history, the model can integrate various contextual information to generate more accurate target language text. For example, if a player repeatedly changes "Stage" to "Level" in similar scenarios, the model will be more inclined to translate "Stage" directly to "Level" in subsequent translations.
[0082] Based on the comprehensive analysis of the above multimodal information, the model finally outputs the target language text and returns this result to the subsequent rendering or display module.
[0083] In some cases, if the text contains specific in-game terminology (such as item names or NPC character names), the model will automatically perform a more refined translation based on the user's historical actions and the image context. This might include retaining the item name in the foreign language and adding a bracketed annotation, or maintaining the phonetic transcription of a person's name. This differentiated translation is based on the user's previous feedback on the game and the model's pre-learned context.
[0084] According to the aforementioned embodiments of this application, before integrating the recognized text, image data, and user operation history into the multimodal fusion translation model, the system preprocesses and aligns the features of each data type. This then leverages the contextual fusion capabilities of the multimodal model to generate target language text that better aligns with the actual game scene. This process not only fully utilizes visual cues within the game screen but also incorporates the user's past translation habits and operational behaviors, resulting in a more coherent and accurate translation result.
[0085] In some embodiments, based on the correspondence between the recognized text and the target language text, the target language text is rendered on the original foreign language area using augmented reality, and the rendering parameters of the target language text are dynamically adjusted according to changes in the game screen, including:
[0086] Obtain the position information of the recognized text in the game screen and establish a mapping relationship between the recognized text and the target language text;
[0087] The target language text is rendered and superimposed using augmented reality at the location corresponding to the original foreign language area;
[0088] Monitor changes in perspective, scenes, or objects in the game screen, adjust presentation parameters of the rendered overlaid target language text based on the changes, and output the adjusted target language text to the game screen interface.
[0089] Specifically, in this embodiment, to ensure that the translation results are perfectly aligned with the original foreign language and adapt to dynamic changes in the game screen, after obtaining the recognized text and the corresponding target language text, the system uses augmented reality (AR) technology to overlay the target language text onto the corresponding foreign language position, thereby presenting an immersive game translation effect. The main process includes the following steps:
[0090] When the OCR engine recognizes foreign text in the game screen, it will also record the location information of the text in screen coordinates or scene coordinates, such as the coordinate points of the upper left and lower right corners, the text rectangular area, or text key points (such as character boundaries, text center, etc.).
[0091] The system will establish a one-to-one or one-to-many mapping table based on the correspondence between the recognized text and the target language text, ensuring that every sentence, every paragraph and even every word can be correctly mapped to the foreign language area during rendering.
[0092] The mapping relationship may also include other metadata, such as the text area size, the estimated height of the original font, the visual depth, etc., for reference by the subsequent rendering engine.
[0093] After the mapping relationship is established, the system calls the augmented reality rendering module to "paste" the target language text to the corresponding position in the foreign language area, for example:
[0094] If the game is a 2D scene, the system can perform simple text overlay based on the screen pixel coordinates; if the game has a 3D space or layered perspective, the system will combine the camera perspective information provided by the game engine to perform three-dimensional projection and fit of the text to present a depth effect consistent with the background.
[0095] In some examples, the system can set a default text style, including font, size, color, transparency, and stroke effect, so that users can clearly and naturally see the translated text in the game interface.
[0096] Before rendering, if it is detected that the length of the original foreign language text is significantly different from that of the target language text, the system can reduce occlusion and ensure readability through automatic line wrapping, dynamic scaling, and line spacing adjustment.
[0097] Furthermore, during the game, the user may use the mouse or controller to rotate the view, or control the character to move, which may cause changes in the game screen, scene, or object position. To ensure that the translated text maintains the accurate position, size, and direction after the scene changes, the system can monitor the following information in real time:
[0098] Scene object changes: Detect whether the game objects corresponding to the translated text (such as props, dialog boxes, etc.) have been moved or replaced;
[0099] Perspective change: Read the rotation and translation information of the camera or character's perspective and determine the coordinate transformation of the AR overlay point under the new perspective;
[0100] Scene lighting and special effects: If there are significant lighting changes or special effects obstruction in the game, the system can appropriately adjust the brightness or transparency of the translated text to make it more integrated with the picture.
[0101] Based on the above detection results, the system calculates new target language text rendering parameters, which may include the following parameters:
[0102] Position: offset or alignment relative to the foreign text area;
[0103] Size: Maintain a similar visual proportion to the original text area;
[0104] Transparency: Increase translucency to reduce distractions in complex background scenes;
[0105] Direction: In a 3D scene, consider that the orientation of the translated text will rotate or tilt accordingly with the player's perspective.
[0106] After adjusting the parameters, the system outputs the updated rendering results to the game screen. If the user continues to move or switch scenes, the system will continue to repeat this process, thus achieving real-time dynamic updates of the translated text.
[0107] Ultimately, through the aforementioned AR rendering overlay and dynamic adjustment process, the target language text can be precisely aligned with the original text and automatically adapt as the scene changes, allowing users to enjoy an immersive, real-time translation experience during the game with almost no manual intervention.
[0108] If the user closes or exits the game, the system can record the current translation status and rendering parameters to quickly restore the user's personalized translation experience the next time it is opened.
[0109] According to the above-mentioned embodiment of the present application, after obtaining the correspondence between the recognized text and the target language text, the system uses augmented reality technology to superimpose the Chinese translation on the foreign language part of the game screen, and automatically dynamically corrects its position, size, orientation and other parameters according to the perspective, scene and object changes, so that the translated text is highly matched with the game screen, thereby enhancing the user's immersive experience.
[0110] In some embodiments, after generating the target language text, the method further includes:
[0111] Determine the confidence level of the target language text and, if the confidence level is below a preset threshold, provide the user with an interactive translation option and collect user corrections or confirmations of the target language text;
[0112] The user's correction or confirmation information on the target language text is used as interactive feedback to update the parameters of the multimodal fusion translation model in real time, and the updated model status is saved when the user exits or runs in the background.
[0113] Specifically, in this embodiment, after the system completes the recognition of the foreign language text and generates the corresponding target language text through the multimodal fusion translation model, in order to further improve the translation accuracy and user experience, the system will perform a confidence assessment on the translation results and provide the user with interactive translation options when necessary. The specific process is as follows:
[0114] While outputting the target language text, the multimodal fusion translation model calculates attention distribution, generates probability distribution and other data based on its internal network to comprehensively derive a confidence score.
[0115] This score reflects the reliability of the translation results in the model's prediction. For example, when the translation content involves highly professional terms, custom character names in games, or relatively ambiguous image areas, the model may not be clear about its meaning and will output a lower confidence value.
[0116] Furthermore, the system pre-sets an adjustable confidence threshold, such as 0.7 or other suitable ranges.
[0117] If the confidence score of the translation result is higher than the threshold, the system determines that the translation result has a high reliability and directly displays the translated text in AR rendering or provides it to the user in other ways.
[0118] If the confidence score of the translation result is lower than the threshold, the system determines that the accuracy of the current translation result needs to be verified and needs to provide the user with an interactive translation option.
[0119] Furthermore, when a low confidence level is detected, the system will display a prompt or pop-up window on the translated text portion in the game screen or floating translation interface to guide the user to quickly confirm or correct the translation result.
[0120] In some interface designs, this prompt can be used to remind users by highlighting the border, marking "may be inaccurate" or appearing an "edit" icon next to the translated text.
[0121] Users can click or touch the prompt to view detailed information about the current translation and perform the following operations:
[0122] Confirm: If the user believes that there are no obvious errors in the translation result, he or she can click "Confirm" or "Accept";
[0123] Correction: If the user finds that a specific word, phrase or sentence structure in the translation is inaccurate, they can directly edit and replace it;
[0124] Ignore: If the user does not care about the accuracy of the text for the time being, they can choose Ignore and the system will temporarily retain the translation.
[0125] Furthermore, the system records the user's confirmation information or correction content for the translation result as interactive feedback, forming a "translation correction label" or "user approval label".
[0126] This interactive feedback is fed into the online learning or incremental learning module of the multimodal fusion translation model in real time to update the model parameters. For example, if a user makes the same replacement for a specific term multiple times, the system will increase the priority matching of that term in subsequent scenarios. If the user continues to choose the same translation in the same type of scenario, the model will automatically learn and memorize this preference so that it can be automatically applied in similar situations.
[0127] When the user exits the game or puts the game back to the background, the system triggers a round of quick save logic to persist the latest model parameters or learning status locally or in the cloud to ensure that the previous learning results are retained when the game is started next time.
[0128] Through this interactive feedback mechanism, the system can continuously accumulate data during each user's use, gradually improving its understanding of game scene terminology, user personal preferences, and even contextual information, thereby gradually improving translation quality.
[0129] If the system detects similar low-confidence content again in the future, the model will refer to previously obtained user feedback to perform priority corrections to minimize repeated errors.
[0130] According to the above-mentioned embodiments of the present application, it can be seen that after generating the target language text, the system provides an interactive correction method for translations with low confidence, and combines user feedback to iterate and update the multimodal fusion translation model in real time, thereby effectively improving translation accuracy and user satisfaction.
[0131] In some embodiments, the user's correction or confirmation information of the target language text is used as interactive feedback to update the parameters of the multimodal fusion translation model in real time, including:
[0132] Obtaining interactive information generated by the user's correction or confirmation operations on the target language text;
[0133] Associate the interaction information with the recognized text and the target language text to form feedback tags for the translation results;
[0134] The feedback label is input into the parameter update process of the multimodal fusion translation model to iteratively optimize the parameters of the multimodal fusion translation model.
[0135] Specifically, in this embodiment, to enable the translation system to continuously adapt to user needs and gradually optimize its performance, after the user obtains the translation result, the system collects the user's corrections or confirmations of the target language text through an interactive feedback mechanism and updates the multimodal fusion translation model in real time. The specific process includes the following steps:
[0136] When the system displays the target language text to the user in the game interface or floating translation window, it provides the user with interactive options such as "Satisfied / Correct / Replace";
[0137] If the user chooses to modify or replace, the system records the modified vocabulary or sentence structure and automatically generates a modification field;
[0138] If the user confirms the translation result, such as clicking "Satisfied", the system will record the user's approval information;
[0139] These interactive operations will be accompanied by a timestamp, scene identifier, and unique identifier of the translated text, so as to facilitate traceability during subsequent model updates.
[0140] Furthermore, after the user makes an interactive operation, the system matches the information with the current recognized text and the target language text to form a feedback tag for the translation result. For example, the feedback tag may include:
[0141] Original recognized text: foreign text recognized by OCR;
[0142] Corrected text or confirmation mark: If the user makes corrections, the specific fields of the correction can be recorded; if it is just confirmation, the "approved" label can be recorded;
[0143] Scene information: such as the location of the game character, the game stage, or the context of the dialogue;
[0144] User identity or historical behavior characteristics: This facilitates subsequent statistics on the translation preferences of the same user.
[0145] Through this labeling method, the original translation process and user feedback information are closely linked, providing a traceable database for subsequent iterative optimization.
[0146] Furthermore, the system inputs the feedback labels into the reinforcement learning or online learning module of the multimodal fusion translation model to iteratively optimize the model parameters;
[0147] When a user makes the same correction to the same proper noun multiple times, the model will internally increase the priority translation weight of the word or automatically remember the user's preference;
[0148] If users tend to use more colloquial expressions in certain scenarios (such as plot dialogues), the model will gradually learn this style and automatically adjust the translation output the next time it encounters a similar context;
[0149] Through regular or real-time parameter updates, the system can reflect improvements in the short term and further enhance translation accuracy and user satisfaction in the long term.
[0150] Furthermore, when the user turns off the translation function (e.g., by clicking the floating ball) or switches the game to the background, the system will automatically pause the real-time capture of the game screen to reduce resource usage;
[0151] At the same time, the system will save the latest parameters or weights of the current model locally or in the cloud, ensuring that the translation model can continue to use the previous learning results when the user opens the game or translation function again;
[0152] If the system is used by multiple people or across devices, user feedback and model update results can also be shared on the server or cloud side, so that different terminals can obtain real-time iterative translation capabilities.
[0153] According to the above-mentioned embodiments of the present application, it can be seen that by providing users with feedback channels such as correction or confirmation after the translation results are generated, and promptly applying this feedback to the online learning process of the multimodal fusion translation model, the adaptability and accuracy of the translation system in various game scenarios can be effectively improved.
[0154] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0155] Figure 2 This is a schematic diagram of the structure of the game real-time intelligent translation device based on interactive enhanced learning provided by the embodiment of the present application. Figure 2 As shown, the game real-time intelligent translation device based on interactive enhanced learning includes:
[0156] Acquisition module 201, for acquiring current game screen information and user operation habit information when detecting that the user has turned on the translation function, and determining a corresponding screen capture mode based on the current game screen information and user operation habit information;
[0157] The perception module 202 is used to dynamically perceive the current game screen according to the screen capture mode, identify the user's focus area and perform partial capture to obtain image data within the user's focus area;
[0158] Recognition module 203, configured to perform text recognition on image data within the user's focus area using a context-enhanced image recognition model, and obtain recognized text based on a game scene-specific contextual knowledge base and user operation history information;
[0159] A generation module 204 is configured to input the recognized text, corresponding image data, and user operation history information into a multimodal fusion translation model to generate a target language text;
[0160] The rendering module 205 is used to render the target language text on the original foreign language area using augmented reality based on the correspondence between the recognized text and the target language text, and dynamically adjust the rendering parameters of the target language text according to changes in the game screen.
[0161] In some embodiments, Figure 2 The acquisition module 201 analyzes the current game screen information to obtain the screen change frequency and text density; collects statistics on user operation habit information to obtain the user interaction operation frequency and historical usage behavior; determines the screen capture mode based on the screen change frequency, text density, user interaction operation frequency and historical usage behavior; sets corresponding floating display permissions based on the screen capture mode, and applies the screen capture mode and floating display permissions to the subsequent local capture and translation rendering process.
[0162] In some embodiments, Figure 2 The recognition module 203 inputs the image data within the user's focus area into the context-enhanced image recognition model to extract text information; obtains terms and expressions related to the current game environment from the context knowledge base exclusive to the game scene; performs context-related matching on the extracted text information based on the user's operation history information, and recognizes and corrects the text information in combination with the terms and expressions to generate recognized text.
[0163] In some embodiments, Figure 2 The generation module 204 pre-processes the recognized text and the corresponding image data and user operation history information to obtain input features that can be used for multimodal fusion; the input features are input into the multimodal fusion translation model so that the multimodal fusion translation model translates the recognized text based on the comprehensive semantic information of the text, image and user operation history to obtain the target language text.
[0164] In some embodiments, Figure 2 The rendering module 205 obtains the position information of the recognized text in the game screen and establishes a mapping relationship between the recognized text and the target language text; renders and overlays the target language text at the position corresponding to the original foreign language area using augmented reality; monitors the changes in the perspective, scene or object in the game screen, and adjusts the presentation parameters of the rendered and overlaid target language text according to the changes, and outputs the adjusted target language text to the game screen interface.
[0165] In some embodiments, Figure 2After generating the target language text, the update module 206 determines the confidence of the target language text. When the confidence is lower than a preset threshold, the update module 206 provides the user with an interactive translation option and collects the user's correction or confirmation information on the target language text; the user's correction or confirmation information on the target language text is used as interactive feedback to update the parameters of the multimodal fusion translation model in real time, and saves the updated model state when the user exits or runs in the background.
[0166] In some embodiments, Figure 2 The update module 206 obtains the interactive information generated by the user's correction or confirmation operation on the target language text; associates the interactive information with the recognized text and the target language text to form a feedback label for the translation result; and inputs the feedback label into the parameter update process of the multimodal fusion translation model to iteratively optimize the parameters of the multimodal fusion translation model.
[0167] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0168] Figure 3 Schematic diagram of the structure of the electronic device 3 provided in the embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, the steps of the above-mentioned method embodiments are implemented. Alternatively, when the processor 301 executes the computer program 303, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0169] For example, computer program 303 may be divided into one or more modules / units, which are stored in memory 302 and executed by processor 301 to implement the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of computer program 303 in electronic device 3.
[0170] The electronic device 3 may be a desktop computer, a notebook, a PDA, a cloud server or other electronic device. The electronic device 3 may include but is not limited to a processor 301 and a memory 302. Those skilled in the art will understand that Figure 3 It is only an example of electronic device 3 and does not constitute a limitation of electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0171] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0172] Memory 302 can be an internal storage unit of electronic device 3, such as a hard drive or memory of electronic device 3. Memory 302 can also be an external storage device of electronic device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, memory 302 can include both an internal storage unit of electronic device 3 and an external storage device. Memory 302 is used to store computer programs and other programs and data required by the electronic device. Memory 302 can also be used to temporarily store data that has been output or is about to be output.
[0173] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0174] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0175] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0176] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which may be electrical, mechanical or other forms.
[0177] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0178] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0179] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium.
[0180] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the technical solutions of the present application are described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A real-time intelligent translation method for games based on interactive reinforcement learning, characterized in that: include: When detecting that the user has turned on the translation function, obtaining current game screen information and user operation habit information, and determining a corresponding screen capture mode according to the current game screen information and user operation habit information; According to the screen capture mode, the current game screen is dynamically sensed, the user's focus area is identified and local capture is performed to obtain image data in the user's focus area; Performing text recognition on the image data within the user's focus area using a context-enhanced image recognition model, and obtaining recognized text based on a game scene-specific contextual knowledge base and user operation history information; Inputting the recognized text, corresponding image data, and user operation history information into a multimodal fusion translation model to generate a target language text; Based on the correspondence between the recognized text and the target language text, the target language text is rendered on the original foreign language area using augmented reality, and the rendering parameters of the target language text are dynamically adjusted according to changes in the game screen; The obtaining of current game screen information and user operation habit information, and determining a corresponding screen capture mode according to the current game screen information and user operation habit information, includes: Analyze the current game screen information to obtain screen change frequency and text density; Collect statistics on the user's operating habits to obtain the user's interactive operation frequency and historical usage behavior; Determining the screen capture mode according to the screen change frequency, text density, user interaction frequency, and historical usage behavior; A corresponding floating display permission is set according to the screen capture mode, and the screen capture mode and the floating display permission are applied to subsequent partial capture and translation rendering processes.
2. The method according to claim 1, characterized in that The context-enhanced image recognition model performs text recognition on the image data within the user's focus area, and obtains recognized text based on the game scene-specific context knowledge base and user operation history information, including: Inputting the image data within the user's focus area into the context-enhanced image recognition model to extract text information; Obtain terms and expressions related to the current game environment from the game scene-specific contextual knowledge base; Based on the user operation history information, context association matching is performed on the extracted text information, and recognition correction is performed on the text information in combination with the terms and expressions to generate the recognized text.
3. The method according to claim 1, characterized in that The step of inputting the recognized text, corresponding image data, and user operation history information into a multimodal fusion translation model to generate a target language text includes: Preprocessing the recognized text, the corresponding image data, and the user operation history information to obtain input features that can be used for multimodal fusion; The input features are input into a multimodal fusion translation model, so that the multimodal fusion translation model translates the recognized text based on the comprehensive semantic information of the text, image and user operation history to obtain the target language text.
4. The method according to claim 1, wherein The method of rendering the target language text on the original foreign language area using augmented reality based on the correspondence between the recognized text and the target language text, and dynamically adjusting the rendering parameters of the target language text according to changes in the game screen, includes: Obtaining position information of the recognized text in the game screen, and establishing a mapping relationship between the recognized text and the target language text; Rendering and superimposing the target language text at a position corresponding to the original foreign language area using augmented reality; Monitor changes in perspective, scene, or object in the game screen, adjust presentation parameters of the rendered overlaid target language text according to the changes, and output the adjusted target language text to the game screen interface.
5. The method according to claim 3, characterized in that After generating the target language text, the method further includes: Determining the confidence level of the target language text, and when the confidence level is lower than a preset threshold, providing the user with an interactive translation option and collecting user correction or confirmation information on the target language text; The user's correction or confirmation information on the target language text is used as interactive feedback to update the parameters of the multimodal fusion translation model in real time, and the updated model state is saved when the user exits or the model is run in the background.
6. The method according to claim 5, characterized in that The method of using the user's correction or confirmation information of the target language text as interactive feedback to update the parameters of the multimodal fusion translation model in real time includes: Acquiring interactive information generated by a user performing a correction or confirmation operation on the target language text; Associating the interaction information with the recognized text and the target language text to form a feedback tag for the translation result; The feedback tag is input into the parameter updating process of the multimodal fusion translation model to iteratively optimize the parameters of the multimodal fusion translation model.
7. A real-time intelligent translation device for games based on interactive enhanced learning, characterized in that: include: an acquisition module, configured to acquire current game screen information and user operation habit information when detecting that the user has turned on the translation function, and determine a corresponding screen capture mode based on the current game screen information and user operation habit information; A perception module is used to dynamically perceive the current game screen according to the screen capture mode, identify the user's focus area and perform local capture to obtain image data within the user's focus area; A recognition module, configured to perform text recognition on the image data within the user's focus area using a context-enhanced image recognition model, and obtain recognized text based on a contextual knowledge base specific to the game scene and user operation history information; A generation module, configured to input the recognized text, corresponding image data, and user operation history information into a multimodal fusion translation model to generate a target language text; a rendering module for rendering the target language text on the original foreign language area using augmented reality based on the correspondence between the recognized text and the target language text, and dynamically adjusting the rendering parameters of the target language text according to changes in the game screen; The acquisition module is configured to analyze the current game screen information to obtain screen change frequency and text density; collect statistics on the user operation habit information to obtain user interaction operation frequency and historical usage behavior; and determine the screen capture mode based on the screen change frequency, text density, user interaction operation frequency, and historical usage behavior. A corresponding floating display permission is set according to the screen capture mode, and the screen capture mode and the floating display permission are applied to subsequent partial capture and translation rendering processes.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Translation display method and device based on augmented reality, computing equipment and medium
CN108681393A
AR translation processing method and electronic equipment
CN118230203A