Game real-time intelligent translation method and device based on interactive reinforcement learning

By adopting a real-time intelligent translation method based on interactive augmented learning in the game, dynamically perceive and process game screen information, use a multimodal fusion translation model to generate target language text, and use augmented reality to render the translation results, the problems of cumbersome translation operations, lack of real-timeness and inconsistent with the game scene in the existing technology are solved, and efficient, real-time and immersive game translation effects are achieved.

CN120087378AActive Publication Date: 2025-06-03QINGFENG (BEIJING) TECH CO LTD

Patent Information

Application Number
CN202510560068.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is complicated to translate foreign language texts in games, lack real-time, and it is difficult to process a large number of continuous texts, and the translation results do not conform to the game scenario.

Method used

The real-time intelligent translation method of game based on interactive augmented learning is adopted. By obtaining game screen information and user operation habits, the screen capture mode is determined, the image data of the user's attention area is dynamically perceived and partially captured, the context-enhanced image recognition model is used for text recognition, and the target language text is generated in combination with the multimodal fusion translation model, the translation results are rendered using augmented reality, and the presentation parameters are dynamically adjusted.

Benefits of technology

Real-time, automatic and immersive translation of game screen content, can process a large number of continuous text, and the translation results are more in line with the game scene, improving the user's play experience and usage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087378A_ABST
    Figure CN120087378A_ABST
Patent Text Reader

Abstract

The invention provides a game real-time intelligent translation method and device based on interactive reinforcement learning. The method comprises the following steps: determining a corresponding screen capture mode according to current game picture information and user operation habit information; dynamically sensing the current game picture according to the screen capture mode; performing text recognition on the image data in the user attention area by using a context-enhanced image recognition model, and obtaining a recognition text according to a context knowledge base exclusive to the game scene and user operation historical information; inputting the recognition text, the corresponding image data and user operation historical information into a multi-modal fusion translation model to generate a target language text; and based on the corresponding relationship between the recognition text and the target language text, rendering the target language text on the original foreign language region in an augmented reality mode. According to the method, real-time, automatic and immersive translation of the game picture content is realized, a large number of continuous texts can be processed, and the translation result better conforms to a game scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of game screen translation, and particularly to a real-time intelligent translation method and device for games based on interactive reinforcement learning. Background Art

[0002] With the globalization development of the game industry, overseas games are increasingly widely used in the domestic market. However, many overseas games have not been localized, and the foreign language texts (such as plot dialogues, item descriptions, and operation guides) contained in the game interface are difficult for domestic users to understand, seriously affecting the user's gaming experience.

[0003] In the prior art, users usually need to manually capture the game screen and switch to a third-party translation software for text recognition and translation. This method is not only cumbersome to operate, but also unable to present the translation results in real time, resulting in frequent interruptions during the game process, reducing the user's immersion and usage efficiency. Therefore, how to achieve efficient, accurate, and real-time translation of foreign language texts in the game environment has become a technical need to be urgently solved. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a real-time intelligent translation method and device for games based on interactive reinforcement learning to solve the problems of cumbersome operation, lack of real-time performance, difficulty in processing a large amount of continuous text, and translation results not conforming to the game scenario existing in the prior art.

[0005] In the first aspect of the embodiments of this application, a real-time intelligent translation method for games based on interactive reinforcement learning is provided, including: when it is detected that the user enables the translation function, obtaining the current game screen information and the user operation habit information, and determining the corresponding screen capture mode according to the current game screen information and the user operation habit information; according to the screen capture mode, dynamically perceiving the current game screen, identifying the user's attention area and performing local capture to obtain the image data within the user's attention area; using a context-enhanced image recognition model to perform text recognition on the image data within the user's attention area, and obtaining the recognized text according to the context knowledge base exclusive to the game scenario and the user operation history information; inputting the recognized text, the corresponding image data, and the user operation history information into a multi-modal fusion translation model to generate the target language text; based on the correspondence between the recognized text and the target language text, using the augmented reality method to render the target language text on the original foreign language area, and dynamically adjusting the presentation parameters of the target language text according to the changes in the game screen.

[0006] In the second aspect of the embodiments of the present application, a real-time intelligent translation device for games based on interactive reinforcement learning is provided, including: an acquisition module, configured to obtain current game screen information and user operation habit information when detecting that the user enables the translation function, and determine a corresponding screen capture mode according to the current game screen information and user operation habit information; a perception module, configured to dynamically perceive the current game screen according to the screen capture mode, identify the user's attention area and perform local capture, so as to obtain image data within the user's attention area; an identification module, configured to perform text identification on the image data within the user's attention area by using a context-enhanced image recognition model, and obtain the identified text according to the context knowledge base specific to the game scenario and the user operation history information; a generation module, configured to input the identified text, the corresponding image data and the user operation history information into a multimodal fusion translation model to generate target language text; a rendering module, configured to render the target language text on the original foreign language area in an augmented reality manner based on the correspondence between the identified text and the target language text, and dynamically adjust the presentation parameters of the target language text according to the changes in the game screen.

[0007] In the third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0008] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0009] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects: By obtaining current game screen information and user operation habit information when detecting that the user enables the translation function, determining a corresponding screen capture mode according to the current game screen information and user operation habit information; dynamically perceiving the current game screen according to the screen capture mode, identifying the user's attention area and performing local capture to obtain image data within the user's attention area; performing text identification on the image data within the user's attention area by using a context-enhanced image recognition model, and obtaining the identified text according to the context knowledge base specific to the game scenario and the user operation history information; inputting the identified text, the corresponding image data and the user operation history information into a multimodal fusion translation model to generate target language text; rendering the target language text on the original foreign language area in an augmented reality manner based on the correspondence between the identified text and the target language text, and dynamically adjusting the presentation parameters of the target language text according to the changes in the game screen. The present application realizes real-time, automatic and immersive translation of game screen content, can not only process a large amount of continuous text, but also the translation result is more in line with the game scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 is a schematic flowchart of a game real-time intelligent translation method based on interactive reinforcement learning provided by an embodiment of the present application; Figure 2 is a schematic structural diagram of a game real-time intelligent translation device based on interactive reinforcement learning provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0013] With the rapid development of the global game market, many overseas games have not been localized in a timely manner (that is, the text, prompts, plot, etc. in the game are translated or localized), resulting in domestic users facing the problem of being unable to understand the foreign language text included in the interface when playing games. The text in games often involves game guides, plot dialogues, item descriptions, etc., which directly affects players' understanding of the game content and gaming experience.

[0014] The common practice of the current prior art is that when players need translation, they often manually take screenshots of the game screen and then switch to a third-party translation software for text recognition and translation. The following are the main deficiencies of this process: Complicated operation: Manually taking screenshots, switching applications, waiting for translation, and then switching back to the game.

[0015] Lack of real-time performance: Players cannot directly and immediately see the translation results in the game, which affects the coherence and immersion of the game.

[0016] Difficulty in processing a large amount of continuous text: For frequently appearing or rapidly changing game text, manual screenshots cannot keep up with the game progress and cannot meet the continuous scene translation requirements.

[0017] Therefore, the existing translation methods have obvious deficiencies in terms of user experience and usage efficiency, and there is an urgent need for a more automated, real-time translation solution that conforms to the game scenario.

[0018] The content of the technical solution of this application will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0019] Figure 1 It is a schematic flowchart of a real-time intelligent translation method for games based on interactive reinforcement learning provided by an embodiment of this application. As Figure 1 shown, the real-time intelligent translation method for games based on interactive reinforcement learning may specifically include: S101, when it is detected that the user enables the translation function, obtain the current game screen information and user operation habit information, and determine the corresponding screen capture mode according to the current game screen information and user operation habit information; S102, according to the screen capture mode, dynamically sense the current game screen, identify the user's attention area and perform local capture to obtain the image data within the user's attention area; S103, use a context-enhanced image recognition model to perform text recognition on the image data within the user's attention area, and obtain the recognized text according to the context knowledge base exclusive to the game scenario and the user operation history information; S104, input the recognized text, the corresponding image data, and the user operation history information into a multi-modal fusion translation model to generate the target language text; S105, based on the correspondence between the recognized text and the target language text, use the augmented reality method to render the target language text on the original foreign language area, and dynamically adjust the presentation parameters of the target language text according to the changes in the game screen.

[0020] In some embodiments, obtaining the current game screen information and user operation habit information, and determining the corresponding screen capture mode according to the current game screen information and user operation habit information includes: Analyze the current game screen information to obtain the screen change frequency and text density; Statistically analyze the user operation habit information to obtain the user interaction operation frequency and historical usage behavior; Determine the screen capture mode according to the screen change frequency, text density, user interaction operation frequency, and historical usage behavior; Set the corresponding floating display permission according to the screen capture mode, and apply the screen capture mode and the floating display permission to the subsequent local capture and translation rendering processes.

[0021] Specifically, in this embodiment, when it is detected that the user clicks or activates the translation function in an overseas game (for example, clicks on the floating ball on the game interface), the system first starts the process of collecting and analyzing the current game screen information and the user's operation habits. The specific steps are as follows: First, the system performs real-time or periodic sampling and statistics on the current game screen in the background to determine whether there are a large number of continuous text prompts in the game scene or whether the screen is in a high-frequency update state.

[0022] For example, the system can identify the screen switching speed by reading the frame rate and analyze the text density by detecting the number of characters and the character distribution in the screenshot.

[0023] If it is detected that the game is in a plot dialogue scene or a scene where operation prompts frequently appear, it is determined that the text density is high; if it is detected that the game screen is mainly static or there is only a small amount of text in the scene, it is determined that the text density is low.

[0024] Furthermore, the system simultaneously conducts statistics and analysis on the user's operation habits to obtain the user's interaction operation frequency and historical usage behavior.

[0025] For example, the system determines whether the user is in a high-frequency operation state based on the number of mouse clicks, keyboard input frequency, or touch operation records of the user in the game.

[0026] Through the user's historical usage behavior, understand the frequency of the user's translation needs during previous game processes. For example, some users are more inclined to frequently activate the translation function in the plot scene and use it less during ordinary level battles.

[0027] Furthermore, the system performs correlation analysis on the screen change frequency, text density, and user operation habit information to comprehensively determine the screen capture mode.

[0028] If the game screen is updated quickly, the text appears densely, and the user operates frequently, the system will intelligently recommend and adaptively switch to the high-frequency capture mode.

[0029] If the game screen is relatively static and the user's translation needs are limited, it will switch to the low-frequency capture mode or the on-demand capture mode, thereby reducing the occupation of system resources.

[0030] Furthermore, when the corresponding capture mode is determined, the system will set the corresponding floating display permission according to different capture modes.

[0031] In the high-frequency capture mode, the floating display permission allows frequent screenshots of the game screen at short time intervals and displays translation prompts on the screen in real time.

[0032] In the low - frequency capture mode, the floating display permission is more limited, and screen capture and translation are only performed when text content may appear or when the user manually triggers.

[0033] At the same time, the system will prompt the user in the form of a pop - up window or a lightweight notification to indicate which capture mode has been entered, so that the user can be aware of the current balance between performance overhead and translation frequency.

[0034] Finally, the determined screen capture mode and floating display permission will be applied to the subsequent steps of the method described in this embodiment, such as local image capture, text recognition, and translation rendering processes. The system continuously monitors changes in the game scene and user operations. If there are significant changes, the above - mentioned capture mode can be updated again and the floating display permission can be adjusted accordingly.

[0035] According to the above - mentioned embodiments of the present application, it can be understood under what circumstances the system will switch to which screen capture mode and how to allocate the floating display permission in combination with user operation habit information. This example only details the part of the technical solution of the present application "obtaining the current game screen information and user operation habit information, and determining the corresponding screen capture mode and floating display permission". Other technical details not involved can be referred to other embodiments of the present application or related descriptions.

[0036] In some embodiments, a context - enhanced image recognition model is used to perform text recognition on the image data within the user - focused area, and the recognized text is obtained based on the context knowledge base exclusive to the game scene and the user operation history information, including: Input the image data within the user - focused area into the context - enhanced image recognition model to extract text information; Obtain the terms and expressions related to the current game environment from the context knowledge base exclusive to the game scene; Based on the user operation history information, perform context - related matching on the extracted text information, and combine the terms and expressions to correct the recognition of the text information to generate the recognized text.

[0037] Specifically, in this embodiment, to improve the accuracy of text recognition and save system resources, when the user activates the floating ball by clicking or touching, the system will first detect the user - focused area in the game interface (such as the position pointed to by the mouse cursor or the touch hot spot). For the detected user - focused area, the system performs the following processing steps: The system compares the game screens of the previous frame and the current frame, judges the change situation in the local area, and performs differential capture on the high - value area (i.e., the user - focused area), rather than taking a screenshot of the full screen.

[0038] For example, when it is detected that the user's cursor is staying at a certain item description or subtitle position, the system only captures the image within the focus area. Through this local capture method, the amount of image data to be processed can be significantly reduced, and the recognition efficiency can be improved.

[0039] Furthermore, the image data of the concerned area obtained by capture is input into an enhanced OCR model (such as the Transformer-OCR model) to extract the possible text information contained.

[0040] This OCR model has the ability to recognize irregular fonts, small amounts of blurred or overlapping text, and can extract text content in relatively complex game scenes.

[0041] Compared with traditional OCR, this model can learn richer character structure features and context associations during the training process.

[0042] Furthermore, to improve the accuracy and scene adaptability of text recognition, in this embodiment, a context knowledge base exclusive to the game scene is introduced during the OCR recognition process. This knowledge base includes: Special terms, item names, character nicknames, plot dialogue templates, etc. in specific game fields; a vocabulary set pre-imported by developers or automatically accumulated after long-term use.

[0043] When the OCR engine recognizes text, it will compare and match the preliminary recognition result with the said knowledge base to check whether the keyword appears in or is similar to the game-specific dictionary, so as to make corrections or supplements.

[0044] At the same time, the system combines the user's operation history information in this game to further improve the accuracy of text recognition. For example: If the user has translated certain specific proper nouns or dialogue contents multiple times in the same or similar scenes before, then this proper noun or dialogue expression will have a higher priority matching weight; When a new professional vocabulary or dialogue form is detected, the system records it in the knowledge base for reference during subsequent recognition.

[0045] By combining the user operation history with the knowledge base, the OCR model can more accurately correct recognition errors, distinguish approximate words or phrases, and form recognition text that is more in line with the actual content of the game.

[0046] Finally, the system completes the correction and confirmation of the recognition text based on the OCR extraction result, knowledge base comparison, and user historical feedback, and provides the generated recognition text to the subsequent translation unit for translation processing.

[0047] If the recognized text is a game item description, the system may also associate it with corresponding item icons, item attributes, and other information, so that the translation module can obtain more complete context when generating translations.

[0048] According to the embodiments of the present application described above, it can be seen that on the basis of only performing differential local capture in the user's attention area, by combining the context-enhanced OCR engine with the game scene-specific knowledge base and the user's operation history, high-precision recognition and real-time processing of complex game texts can be achieved without adding too much system burden.

[0049] In some embodiments, the recognized text, corresponding image data, and user operation history information are input into a multimodal fusion translation model to generate target language text, including: Preprocess the recognized text, corresponding image data, and user operation history information to obtain input features that can be used for multimodal fusion; Input the input features into the multimodal fusion translation model, so that the multimodal fusion translation model can translate the recognized text based on the comprehensive semantic information of the text, image, and user operation history to obtain the target language text.

[0050] Specifically, in this embodiment, in order to obtain a translation result that better fits the game scene, after the system completes the recognition of foreign language text in the game screen, it will input the recognized text, corresponding image data, and user operation history information into the multimodal fusion translation model together, so as to generate the target language text that conforms to the game context. Its main process includes the following steps: First, text information preprocessing: Segment or break the text obtained by OCR recognition, and perform basic character regularization operations on it. For example, for relatively special characters, symbols, and common delimiters in the game (such as “—”, “·”, etc.), unified or cleaning processing can be performed to ensure consistency in the subsequent translation process.

[0051] Next, image data extraction: Convert the format and adjust the resolution of the screen area or local image information corresponding to the recognized text, so that the multimodal model can extract visual features related to the text content. Specifically, the system can represent the image data as a graphical vector or embedding vector that can be recognized by the model through image cropping, scaling, and feature vectorization.

[0052] Then, summarize the user operation history information: The system will extract the user's interaction behaviors during the recent game process, such as the mouse click position, touch hotspots, the frequency of opening the translation function, the user's correction records for similar texts, etc. In order to make this information fully utilized by the multimodal model, in this embodiment, time stamps or operation category labels will be added to these interaction behaviors and converted into a vectorized representation or feature encoding that can be processed by the model.

[0053] Furthermore, after completing the above preprocessing, the system will perform a preliminary combination of text features, image features, and user operation history features before multimodal fusion. In this embodiment, the system first embeds or vectorizes the text features, image features, and operation history features, and then matches and aligns them according to time association or content association, thereby obtaining an input feature matrix that can be directly processed by the multimodal model.

[0054] For example, the screen location of the text is identified by index OCR and paired with the image features at the same location; at the same time, the user's operation records in the time period or scene are annotated to the same feature index to construct a multimodal information unit of the spatiotemporal location.

[0055] Furthermore, the constructed multimodal input features are passed to the multimodal fusion translation model, and the model can translate the recognized text based on the comprehensive semantics of text, image and user operation information.

[0056] In this embodiment, the multimodal fusion translation model has the ability to process multiple input sources simultaneously and correlate them. For example, the model will determine what kind of game scene the text may be in based on image features (such as prop descriptions, plot dialogues, or system prompts), and combine the user's previous translation or correction operations on text in similar scenes to determine which words or sentences are more in line with the player's expectations.

[0057] By performing joint attention or other fusion mechanisms on text, images, and operation history, the model can integrate various contextual information to generate more accurate target language text. For example, if the player corrects "Stage" to "Level" in similar scenarios many times, the model will be more inclined to translate "Stage" directly into "Level" in subsequent translations.

[0058] Based on the comprehensive analysis of the above multimodal information, the model finally outputs the target language text and returns this result to the subsequent rendering or display module.

[0059] In some cases, if the text content is detected to involve special terms in the game (names of props, NPC characters, etc.), the model will automatically perform a more sophisticated translation based on the user's operation history and image background information, such as keeping the prop name in foreign language and adding bracketed annotations at the end, or keeping the name of the person in transliteration, etc. This differentiated translation depends on the user's previous operation feedback in the game and the contextual experience learned by the model in advance.

[0060] According to the above-mentioned embodiment of the present application, before the recognition text and image data and user operation history information are comprehensively input into the multimodal fusion translation model, the system will first pre-process and feature-align each type of data, and then generate a target language text that is more in line with the actual game scene through the context fusion capability of the multimodal model. This process not only makes full use of the visual clues in the game screen, but also combines the user's past translation habits and operation behaviors, making the final translation result more coherent and accurate.

[0061] In some embodiments, based on the correspondence between the recognized text and the target language text, the target language text is rendered on the original foreign language area using an augmented reality method, and the rendering parameters of the target language text are dynamically adjusted according to the changes in the game screen, including: Obtaining the position information of the recognized text in the game screen, and establishing a mapping relationship between the recognized text and the target language text; At the location corresponding to the original foreign language area, the target language text is rendered and superimposed using augmented reality; Monitor changes in the perspective, scene or object in the game screen, adjust the presentation parameters of the rendered overlay target language text according to the changes, and output the adjusted target language text to the game screen interface.

[0062] Specifically, in this embodiment, in order to perfectly align the translation result with the original foreign language and adapt it to the dynamic changes of the game screen, after obtaining the recognized text and the corresponding target language text, the system superimposes the target language text to the corresponding foreign language position based on augmented reality (AR) technology, thereby presenting an immersive game translation effect. Its main process includes the following steps: When the OCR engine recognizes foreign text in the game screen, it will also record the location information of the text in screen coordinates or scene coordinates, such as the upper left and lower right corner coordinates, text rectangular area or text key points (such as character boundaries, text center, etc.).

[0063] The system will establish a one-to-one or one-to-many mapping table based on the correspondence between the recognized text and the target language text, ensuring that each sentence, each paragraph and even each word can be correctly mapped to the foreign language area during rendering.

[0064] The mapping relationship may also include other metadata, such as text area size, estimated original font height, visual depth, etc., for reference by subsequent rendering engines.

[0065] After the mapping relationship is established, the system calls the augmented reality rendering module to "paste" the target language text to the corresponding position of the foreign language area, for example: If the game is a 2D scene, the system can perform simple text overlay based on the screen pixel coordinates; if the game has a 3D space or layered perspective, the system will combine the camera perspective information provided by the game engine to perform three-dimensional projection and fit of the text in order to present a depth effect consistent with the background.

[0066] In some examples, the system can set a default text style, including font, size, color, transparency, and stroke effect, so that users can clearly and naturally see the translated text in the game interface.

[0067] Before rendering, if it is detected that there is a large difference in the length of the original foreign text and the target language text, the system can reduce occlusion and ensure readability through automatic line wrapping, dynamic scaling, line spacing adjustment, etc.

[0068] Furthermore, during the game, the user may use the mouse or controller to rotate the view, or control the character to move, which may cause changes in the game screen, scene or object position. To ensure that the translated text can maintain the accurate position, size and direction after the scene is switched, the system can monitor the following information in real time: Scene object changes: detect whether the game objects (such as props, dialog boxes, etc.) corresponding to the translated text have been moved or replaced; Perspective change: read the rotation and translation information of the camera or character's perspective, and determine the coordinate transformation of the AR overlay point under the new perspective; Scene lighting and special effects: If there are significant lighting changes or special effects occlusion in the game, the system can appropriately adjust the brightness or transparency of the translated text to make it more integrated with the picture.

[0069] Based on the above detection results, the system will calculate new target language text rendering parameters, which may include the following parameters: Position: offset or alignment relative to the foreign text area; Size: Keep a similar visual proportion to the original text area; Transparency: In complex background scenes, the translucency can be increased to reduce interference; Direction: In a 3D scene, the orientation of the translated text needs to be rotated or tilted accordingly with the player's perspective.

[0070] After the system completes the parameter adjustment, it will output the updated rendering results to the game screen interface. If the user continues to move or switch scenes, the system will continue to repeat the process, thus achieving real-time dynamic update of the translated text.

[0071] Finally, through the above AR rendering overlay and dynamic adjustment process, the target language text can be accurately aligned with the original text and automatically adapted to changes in the scene. During the game, users can obtain an immersive real-time translation experience with little manual intervention.

[0072] If the user closes or exits the game, the system can record the current translation status and rendering parameters to quickly restore the user's personalized translation experience when the game is launched next time.

[0073] According to the embodiments of the present application above, after obtaining the correspondence between the recognized text and the target language text, the system uses augmented reality technology to overlay the Chinese translation on the foreign language part of the game screen and automatically dynamically corrects parameters such as its position, size, and orientation according to changes in the perspective, scene, and object, making the translated text highly match the game screen, thereby enhancing the user's immersive experience.

[0074] In some embodiments, after generating the target language text, the method further includes: Determining the confidence level of the target language text. When the confidence level is lower than a preset threshold, providing the user with interactive translation options and collecting the user's correction or confirmation information for the target language text; Taking the user's correction or confirmation information for the target language text as an interactive feedback, updating the parameters of the multi-modal fusion translation model in real time, and saving the updated model state when the user exits or runs in the background.

[0075] Specifically, in this embodiment, when the system completes the recognition of the foreign language text and generates the corresponding target language text through the multi-modal fusion translation model, to further improve the translation accuracy and user experience, the system will perform a confidence level determination on the translation result and provide interactive translation options to the user when necessary. The specific process is as follows: When the multi-modal fusion translation model outputs the target language text, it will comprehensively obtain a confidence score based on data such as calculating the attention distribution and generating the probability distribution within its internal network.

[0076] This score can reflect the reliability of the translation result in the model prediction. For example, when the translation content involves highly specialized terms, custom character names in the game, or relatively ambiguous image regions, the model may not be able to clearly judge its meaning, resulting in a lower confidence value being output.

[0077] Furthermore, the system pre-sets an adjustable confidence threshold, such as 0.7 or other appropriate ranges.

[0078] If the confidence score of the translation result is higher than this threshold, the system determines that the translation result has a high reliability and directly renders and displays the translated text in AR or provides it to the user in other ways.

[0079] If the confidence score of the translation result is lower than this threshold, the system determines that the accuracy of the current translation result needs to be verified, and it is necessary to provide the user with interactive translation options.

[0080] Furthermore, when low confidence is detected, the system presents a prompt or pop-up window in the game screen or floating translation interface for the translated text part to guide the user to quickly confirm or correct the translation result.

[0081] In some interface designs, this prompt can be used to remind the user by means of a highlighted border, marking "May be inaccurate", or presenting an "Edit" icon next to the translated text.

[0082] The user can click or touch this prompt to view the detailed information of the current translation content and perform the following operations: Confirm: If the user believes that there are no obvious errors in the translation result, they can click "Confirm" or "Accept"; Correct: If the user finds that a specific word, phrase, or sentence pattern in the translation is inaccurate, they can directly edit and replace it; Ignore: If the user is not currently concerned about the accuracy of this text, they can choose to ignore it, and the system will temporarily retain this translation.

[0083] Furthermore, the system records the user's confirmation information or correction content for the translation result as interactive feedback, forming a "translation correction label" or "user approval label".

[0084] This interactive feedback is sent in real time to the online learning or incremental learning module of the multi-modal fusion translation model for updating model parameters; for example, if the user makes the same replacement for a certain specific term multiple times, the system will increase the priority matching degree of this term in subsequent scenarios; if the user continuously selects the same translation in the same type of scenario, the model will automatically learn and remember this preference for automatic application in similar situations.

[0085] When the user exits the game or minimizes the game to the background, the system will trigger a round of quick save logic to persist the latest model parameters or learning status to local or cloud to ensure that the previous learning results can be retained when starting next time.

[0086] Through the above interactive feedback mechanism, this system can continuously accumulate data during each user usage process, gradually improving its understanding of game scenario terms, user personal preferences, and even context information, thereby achieving a gradual improvement in translation quality.

[0087] If the system detects similar low-confidence content again later, the model will refer to the previously obtained user feedback to perform priority corrections to minimize repeated errors.

[0088] According to the embodiments of the present application described above, it can be seen that after generating the target language text, the system provides an interactive correction approach for translations with low confidence, and iteratively updates the multi-modal fusion translation model in combination with user feedback, thereby effectively improving translation accuracy and user satisfaction.

[0089] In some embodiments, the correction or confirmation information of the user for the target language text is used as interactive feedback to update the parameters of the multi-modal fusion translation model in real time, including: Obtain the interactive information generated by the user's correction or confirmation operation on the target language text; Associate the interactive information with the recognized text and the target language text to form a feedback label for the translation result; Input the feedback label into the parameter update process of the multi-modal fusion translation model to iteratively optimize the parameters of the multi-modal fusion translation model.

[0090] Specifically, in this embodiment, to enable the translation system to continuously adapt to user needs and gradually optimize its own performance, after the user obtains the translation result, the system collects the correction or confirmation information of the user for the target language text through an interactive feedback mechanism and updates the multi-modal fusion translation model in real time. The specific process includes the following steps: When the system displays the target language text to the user in the game interface or the floating translation window, it will provide the user with interactive options such as "satisfied / correct / replace", etc.; If the user selects to correct or replace, the system records the modified words or sentences by the user and automatically generates a correction field; If the user performs a confirmation operation on the translation result, such as clicking "satisfied", the system records the user's approval information; These interactive operations will carry a timestamp, the scene identifier, and the unique identifier of the translation text, etc., for subsequent traceability processing when the model is updated.

[0091] Further, after the user makes an interactive operation, the system associates this information with the current recognized text and the target language text to form a feedback label for this translation result. For example, the feedback label may include: Original recognized text: The foreign language obtained by OCR recognition; Corrected text or confirmation identifier: If the user makes a correction, the specific corrected field can be recorded; if only a confirmation is made, the "approval" label can be recorded; Scene information: Such as the location of the game character, the game stage, or the dialogue context; User identity or historical behavior characteristics: Facilitate subsequent statistics of the translation preferences of the same user.

[0092] In this tagging method, the original translation process is closely associated with user feedback information, providing a traceable database for subsequent iterative optimization.

[0093] Furthermore, the system inputs the feedback tags into the reinforcement learning or online learning module of the multi-modal fusion translation model to iteratively optimize the model parameters; When the user makes the same correction to the same proper noun multiple times, the model will internally increase the priority translation weight of this word or automatically remember the user's preference; If in certain scenarios (such as plot dialogues) the user tends to adopt a more colloquial expression, the model will gradually learn this style and automatically adjust the translation output when encountering a similar context next time; Through regular or real-time parameter updates, the system can demonstrate improvement effects in the short term and further enhance translation accuracy and user satisfaction during long-term use.

[0094] Furthermore, when the user closes the translation function (such as clicking the floating ball to close) or switches the game to the background, the system will automatically pause the real-time capture of the game screen to reduce resource occupancy; At the same time, the system will save the latest parameters or weights of the current model locally or in the cloud to ensure that when the user re-opens the game or the translation function, the translation model can continue the previous learning results; If the system is in a multi-user or cross-device scenario, it can also share user feedback and model update results on the server side or in the cloud, enabling different terminals to obtain real-time iterative translation capabilities.

[0095] According to the embodiments of the present application described above, it can be seen that by providing feedback channels such as correction or confirmation for users after the translation result is generated and applying this feedback in a timely manner to the online learning process of the multi-modal fusion translation model, the adaptability and accuracy of the translation system in various game scenarios can be effectively improved.

[0096] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0097] Figure 2 It is a schematic structural diagram of a game real-time intelligent translation device based on interactive reinforcement learning provided by an embodiment of the present application. As Figure 2 shown, the game real-time intelligent translation device based on interactive reinforcement learning includes: An acquisition module 201, configured to obtain current game screen information and user operation habit information when detecting that the user enables the translation function, and determine a corresponding screen capture mode according to the current game screen information and user operation habit information; The perception module 202 is configured to dynamically perceive the current game screen according to the screen capture mode, identify the user's attention area and perform local capture, so as to obtain the image data within the user's attention area; The recognition module 203 is configured to perform text recognition on the image data within the user's attention area by using a context-enhanced image recognition model, and obtain the recognized text according to the context knowledge base exclusive to the game scenario and the user operation history information; The generation module 204 is configured to input the recognized text, the corresponding image data and the user operation history information into a multi-modal fusion translation model to generate the target language text; The rendering module 205 is configured to render the target language text on the original foreign language area in an augmented reality manner based on the correspondence between the recognized text and the target language text, and dynamically adjust the presentation parameters of the target language text according to the changes in the game screen.

[0098] In some embodiments, Figure 2 the acquisition module 201 analyzes the current game screen information to obtain the screen change frequency and text density; statistically analyzes the user operation habit information to obtain the user interaction operation frequency and historical usage behavior; determines the screen capture mode according to the screen change frequency, text density, user interaction operation frequency and historical usage behavior; sets the corresponding floating display permission according to the screen capture mode, and applies the screen capture mode and the floating display permission to the subsequent local capture and translation rendering processes.

[0099] In some embodiments, Figure 2 the recognition module 203 inputs the image data within the user's attention area into a context-enhanced image recognition model to extract text information; obtains the terms and expressions related to the current game environment from the context knowledge base exclusive to the game scenario; performs context association matching on the extracted text information based on the user operation history information, and corrects the recognition of the text information by combining the terms and expressions to generate the recognized text.

[0100] In some embodiments, Figure 2 the generation module 204 preprocesses the recognized text, the corresponding image data and the user operation history information to obtain input features available for multi-modal fusion; inputs the input features into the multi-modal fusion translation model, so that the multi-modal fusion translation model translates the recognized text based on the comprehensive semantic information of the text, image and user operation history to obtain the target language text.

[0101] In some embodiments, Figure 2The rendering module 205 obtains the position information of the recognized text in the game screen and establishes a mapping relationship between the recognized text and the target language text; at the position corresponding to the original foreign language area, the target language text is rendered and superimposed in an augmented reality manner; monitors the changes in the perspective, scene, or objects in the game screen, adjusts the presentation parameters of the rendered and superimposed target language text according to the changes, and outputs the adjusted target language text to the game screen interface.

[0102] In some embodiments, Figure 2 The update module 206 determines the confidence level of the target language text after generating the target language text. When the confidence level is lower than the preset threshold, it provides the user with an interactive translation option and collects the user's correction or confirmation information for the target language text; uses the user's correction or confirmation information for the target language text as interactive feedback to update the parameters of the multimodal fusion translation model in real time, and saves the updated model state when the user exits or runs in the background.

[0103] In some embodiments, Figure 2 The update module 206 obtains the interaction information generated by the user's correction or confirmation operation for the target language text; associates the interaction information with the recognized text and the target language text to form a feedback label for the translation result; inputs the feedback label into the parameter update process of the multimodal fusion translation model to iteratively optimize the parameters of the multimodal fusion translation model.

[0104] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0105] Figure 3 is a schematic structural diagram of the electronic device 3 provided by the embodiments of the present application. As Figure 3 shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the above various method embodiments. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the above various device embodiments.

[0106] Exemplarily, the computer program 303 can be divided into one or more modules / units. One or more modules / units are stored in the memory 302 and executed by the processor 301 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 303 in the electronic device 3.

[0107] The electronic device 3 can be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 3 can include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art can understand that Figure 3 These are merely examples of the electronic device 3 and do not constitute a limitation on the electronic device 3. It may include more or fewer components than shown in the figure, or combine certain components, or have different components. For example, the electronic device may also include input / output devices, network access devices, a bus, etc.

[0108] The processor 301 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0109] The memory 302 can be an internal storage unit of the electronic device 3. For example, the hard disk or memory of the electronic device 3. The memory 302 can also be an external storage device of the electronic device 3. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 3. Further, the memory 302 can also include both the internal storage unit and the external storage device of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device. The memory 302 can also be used to temporarily store data that has been output or will be output.

[0110] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0111] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0112] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0113] In the embodiments provided in this application, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there can be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0114] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0116] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0117] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the technical solutions of the present application have been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A method for real-time intelligent translation of games based on interactive enhanced learning, characterized in that: include: When it is detected that the user turns on the translation function, current game screen information and user operation habit information are obtained, and a corresponding screen capture mode is determined according to the current game screen information and user operation habit information; According to the screen capture mode, the current game screen is dynamically sensed, the user's focus area is identified and partial capture is performed to obtain image data in the user's focus area; Using a context-enhanced image recognition model to perform text recognition on the image data in the user's focus area, and obtaining recognized text based on a game scene-specific context knowledge base and user operation history information; Inputting the recognized text, corresponding image data and user operation history information into a multimodal fusion translation model to generate a target language text; Based on the correspondence between the recognized text and the target language text, the target language text is rendered on the original foreign language area using augmented reality, and the rendering parameters of the target language text are dynamically adjusted according to changes in the game screen.

2. The method according to claim 1, characterized in that The obtaining of current game screen information and user operation habit information, and determining a corresponding screen capture mode according to the current game screen information and user operation habit information, includes: Analyze the current game screen information to obtain screen change frequency and text density; Collect statistics on the user's operation habit information to obtain the user's interactive operation frequency and historical usage behavior; Determining the screen capture mode according to the picture change frequency, text density, user interaction operation frequency and historical usage behavior; The corresponding floating display permission is set according to the screen capture mode, and the screen capture mode and the floating display permission are applied to the subsequent local capture and translation rendering process.

3. The method according to claim 1, characterized in that The context-enhanced image recognition model performs text recognition on the image data in the user's focus area, and obtains the recognized text according to the game scene-specific context knowledge base and user operation history information, including: Inputting the image data within the user focus area into the context-enhanced image recognition model to extract text information; Obtain terms and expressions related to the current game environment from the game scene-specific contextual knowledge base; Based on the user operation history information, context association matching is performed on the extracted text information, and recognition correction is performed on the text information in combination with the terms and expressions to generate the recognized text.

4. The method according to claim 1, characterized in that The step of inputting the recognized text, the corresponding image data and the user operation history information into a multimodal fusion translation model to generate a target language text includes: Preprocessing the recognized text and the corresponding image data and user operation history information to obtain input features that can be used for multimodal fusion; The input features are input into a multimodal fusion translation model, so that the multimodal fusion translation model translates the recognized text based on the comprehensive semantic information of the text, image and user operation history to obtain the target language text.

5. The method according to claim 1, characterized in that Based on the correspondence between the recognized text and the target language text, the target language text is rendered on the original foreign language area by using an augmented reality method, and the rendering parameters of the target language text are dynamically adjusted according to the changes in the game screen, including: Acquire the position information of the recognized text in the game screen, and establish a mapping relationship between the recognized text and the target language text; At a position corresponding to the original foreign language area, the target language text is rendered and superimposed by using an augmented reality method; Monitor changes in the viewing angle, scene or object in the game screen, adjust the presentation parameters of the rendered superimposed target language text according to the changes, and output the adjusted target language text to the game screen interface.

6. The method according to claim 4, characterized in that After generating the target language text, the method further includes: Determining the confidence of the target language text, and when the confidence is lower than a preset threshold, providing the user with an interactive translation option, and collecting the user's correction or confirmation information on the target language text; The correction or confirmation information of the target language text by the user is used as interactive feedback to update the parameters of the multimodal fusion translation model in real time, and the updated model state is saved when the user exits or runs in the background.

7. The method according to claim 6, characterized in that The method of using the user's correction or confirmation information on the target language text as interactive feedback to update the parameters of the multimodal fusion translation model in real time includes: Acquiring interactive information generated by a user performing a correction or confirmation operation on the target language text; Associating the interaction information with the recognition text and the target language text to form a feedback tag for the translation result; The feedback label is input into the parameter updating process of the multimodal fusion translation model to iteratively optimize the parameters of the multimodal fusion translation model.

8. A real-time intelligent game translation device based on interactive enhanced learning, characterized in that: include: An acquisition module, for acquiring current game screen information and user operation habit information when detecting that the user turns on the translation function, and determining a corresponding screen capture mode according to the current game screen information and user operation habit information; A perception module, used to dynamically perceive the current game screen according to the screen capture mode, identify the user's focus area and perform local capture to obtain image data in the user's focus area; A recognition module, used to perform text recognition on the image data in the user's focus area using a context-enhanced image recognition model, and obtain recognized text based on a game scene-specific context knowledge base and user operation history information; A generation module, used for inputting the recognized text, corresponding image data and user operation history information into a multimodal fusion translation model to generate a target language text; The rendering module is used to render the target language text on the original foreign language area by using augmented reality based on the correspondence between the recognized text and the target language text, and dynamically adjust the rendering parameters of the target language text according to the changes in the game screen.

9. An electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Translation display method and device based on augmented reality, computing equipment and medium

    CN108681393A

  • AR translation processing method and electronic equipment

    CN118230203A

  • Text processing method and device in game, storage medium and electronic device

    CN118384494A

  • Apparatus and method of translating chatting data in online game

    KR1020100089673A

  • Method and apparatus for character selection based on character recognition, and terminal device

    US20230169785A1

Cited By

  • Intelligent equipment online language teaching translation system based on image recognition

    CN121189346A

  • Dynamic screen area monitoring and self-adaptive character extraction system based on visual positioning

    CN121937959A