Text display method and apparatus, wearable device, and storage medium
By obtaining and translating text in images on a wearable device and overlaying the translation results at the target position of the text, the problem that the display position of text translation results in the prior art cannot be dynamically adjusted, and users' understanding of translation results and application accuracy are improved.
Patent Information
- Application Number
- PCT/CN2024/117803
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-09-09
- Publication Date
- 2025-05-08
AI Technical Summary
The prior art cannot dynamically adjust the display position of the translation results when using wearable devices to translate text, making it difficult for users to associate text with translation results, which in turn affects the accuracy of understanding and application of translation results.
By obtaining images collected by the wearable device, identifying and translating the text in the image, and when the effectiveness of the translation result is greater than the preset threshold, the translation result is superimposed and displayed on the target position of the text to improve the degree of fit between the translation result and the text.
Real-time fit between text translation results and text is achieved, improving the user's ability to associate text with translation results, thereby enhancing the understanding and application accuracy of translation results.
Smart Images

Figure CN2024117803_08052025_PF_FP_ABST
Abstract
Description
Text display method, device, wearable device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 3, 2023, with application number 2023114631569 and invention name “Text display method, device, wearable device and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence technology, and in particular to a text display method based on a wearable device, a text display device, a wearable device, and a computer-readable storage medium. Background Art
[0003] Currently, there are several common methods for translating text: The first involves users manually inputting text into a terminal device for translation; the second involves users taking a photo of a real scene with a terminal device, translating the text in the photo, and then performing the translation. All of these methods utilize the terminal device to translate text, requiring users to manually input text into the terminal device or to hold the terminal device constantly. This can lead to user fatigue and low translation efficiency.
[0004] In light of the above reasons, some technicians have proposed technologies for text translation using wearable devices. However, while wearable devices can translate text in real time, improving translation efficiency and reducing user fatigue, they cannot dynamically adjust the display position of the translation results when the text moves or the wearable device shakes. This means that the translation results cannot be aligned with the text, resulting in users being unable to associate the text with the translation results, and thus being unable to accurately understand and apply the translation results.
[0005] Summary of the Invention
[0006] The present application provides a text display method based on a wearable device, a text display device, a wearable device, and a computer-readable storage medium, aiming to improve the degree of fit between the translation results displayed on the wearable device and the text, so that the user can better associate the text with the translation results, and then accurately understand and apply the translation results.
[0007] To achieve the above objectives, the present application provides a text display method based on a wearable device, the method comprising:
[0008] Acquire a first image captured by the wearable device, wherein the first image includes at least text;
[0009] Translating the text on the first image to obtain a first text translation result;
[0010] When the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, the text on the first image is obtained at a target position on the current display screen, and the first text translation result is superimposed and displayed on the target position.
[0011] In addition, to achieve the above-mentioned purpose, the present application also provides a text display device, which includes: an acquisition module for acquiring a first image captured by the wearable device, wherein the first image includes at least text; a translation module for translating the text on the first image to obtain a first text translation result; a display module for acquiring the text on the first image at a target position on the current display screen when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, and superimposing and displaying the first text translation result on the target position.
[0012] In addition, to achieve the above-mentioned purpose, the present application further provides a wearable device, the wearable device comprising a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the following steps when executing the computer program:
[0013] Acquire a first image captured by the wearable device, wherein the first image includes at least text; translate the text on the first image to obtain a first text translation result; when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, obtain the text on the first image at a target position on the current display screen, and superimpose the first text translation result on the target position.
[0014] In addition, to achieve the above-mentioned object, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the following steps:
[0015] Acquire a first image captured by the wearable device, wherein the first image includes at least text; translate the text on the first image to obtain a first text translation result; when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, obtain the text on the first image at a target position on the current display screen, and superimpose the first text translation result on the target position.
[0016] The text display method based on a wearable device, the wearable device and the computer-readable storage medium disclosed in the embodiment of the present application can obtain a first image captured by the wearable device, wherein the first image includes at least text. And the text on the first image is translated to obtain a first text translation result. Furthermore, when the validity of the first text translation result is greater than a first preset threshold value and the text of the first image is on the current display screen, the target position of the text on the first image is obtained in the current display screen, and the first text translation result is superimposed and displayed on the target position. Thus, the first text translation result can be superimposed and displayed on the target position. The present application aims to obtain the target position of the text in the display screen while translating the text, and then superimpose the text translation result on the target position in real time to improve the fit between the translation result and the text, so that the user can better associate the text with the translation result, and then accurately understand and apply the translation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] FIG1 is a flow chart of a text display method based on a wearable device provided in an embodiment of the present application;
[0019] FIG2 is a schematic diagram of a scenario of a text display method based on a wearable device provided in an embodiment of the present application;
[0020] FIG3 is a schematic diagram of a scenario in which a first image is captured by a wearable device according to an embodiment of the present application;
[0021] FIG4 is a schematic diagram of a scenario in which a first text translation result is superimposed and displayed on a real-time position, provided by an embodiment of the present application;
[0022] FIG5 is a schematic diagram of a process for refreshing text translation results provided by an embodiment of the present application;
[0023] FIG6 is a schematic block diagram of a text display device provided in an embodiment of the present application;
[0024] FIG7 is a schematic block diagram of a wearable device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] The flowcharts shown in the accompanying drawings are illustrative only and do not necessarily include all content and operations / steps, nor do they necessarily need to be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual order of execution may vary depending on the actual situation. In addition, although the functional modules are divided in the device schematics, in some cases, the module division may be different from that shown in the device schematics.
[0027] The term "and / or" as used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0028] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.
[0029] Please refer to Figures 1 to 3. Figure 1 is a flow chart of a text display method based on a wearable device provided in an embodiment of the present application; Figure 2 is a scene chart of a text display method based on a wearable device provided in an embodiment of the present application; Figure 3 is a scene chart of capturing a first image through a wearable device provided in an embodiment of the present application.
[0030] As shown in FIG1 , the text display method based on a wearable device includes steps S11 to S13 .
[0031] Step S11: Acquire a first image captured by the wearable device, wherein the first image includes at least text.
[0032] The first image is an image initially captured by an image acquisition unit of the wearable device, and includes at least some text. Therefore, the first image can be used to obtain and translate the text.
[0033] The wearable device may be smart glasses, such as AR glasses, VR glasses, etc., which is not limited in this application.
[0034] It should be noted that this application does not limit the manner in which the wearable device captures the first image. For example, the first image may be captured by an image capture unit on the wearable device. The image capture unit includes components such as a camera, a sensor, and a lens mounted on the wearable device. These components work together to capture images and information about the surrounding environment.
[0035] Furthermore, this application does not limit the type of the above-mentioned camera. For example, the camera type can be an RGB camera, a depth camera, a panoramic camera, etc. This application and the camera type are described as an RGB camera. RGB cameras are generally used to capture color images and can generate color images by simultaneously capturing information from three color channels: red, green, and blue. Therefore, RGB cameras are used in wearable devices to implement tasks such as image or video acquisition, image recognition, and processing.
[0036] Specifically, a first image may be captured by a component such as an RGB camera. Since the first image includes text, subsequent text translation steps may be performed based on the captured first image.
[0037] Optionally, the wearable device has no stored images before capturing the first image, or clears the originally stored images to prevent other images from conflicting with the first image, thereby preventing the accurate acquisition of the translated text required by the user.
[0038] Optionally, the initial display screen corresponding to the first image does not display any text before the image is captured.
[0039] The initial display screen is the display screen corresponding to the first image.
[0040] It should be noted that the first image is an image of the real world captured by the wearable device, and the display screen is the screen on the wearable device used to display virtual content. Therefore, the user can observe images and information in the virtual world through the display screen.
[0041] It is understandable that when a user wears a wearable device, he or she can not only obtain images of the real world through the wearable device, but also obtain the display content of the display screen at the same time. Therefore, if text is displayed in the corresponding display screen when the first image is collected, it may cause the final translation result to conflict with the displayed text, and then the user will not be able to associate the translation result with the text to be translated, and will not be able to understand and apply the translation result. In addition, the displayed text in the display screen may also cause the algorithm to misjudge or confuse. Therefore, ensuring that the initial display screen corresponding to the first image has no displayed text before the image is collected will help reduce interference factors in the subsequent text translation process.
[0042] Optionally, before the image acquisition unit of the wearable device acquires the first image, the process includes: determining first state information of the wearable device through a preset sensor; when the first state information meets a first preset condition, acquiring the first image through the wearable device; or, in response to a user operation instruction, controlling the wearable device to acquire the first image.
[0043] It should be noted that this application does not limit the conditions for triggering the wearable device to capture the first image. For example, a preset sensor may be used to trigger the image capture unit of the wearable device to capture the first image, or the image capture unit of the wearable device may be directly controlled to capture the first image based on a user operation instruction.
[0044] Among them, the first state information is the motion state information of the wearable device before collecting the image; the preset sensor can be an IMU (Inertial Measurement Unit) sensor, a visual sensor, and a positioning sensor, etc. This application takes the preset sensor as an IMU sensor as an example for explanation.
[0045] It's important to note that IMU sensors typically include multiple built-in components, such as gyroscopes, accelerometers, and magnetometers, to measure and monitor a device's motion and orientation without the need for an external reference point. IMU sensors can provide data about a device's three-dimensional motion and orientation, such as acceleration and three-degree-of-freedom information, thereby determining the device's motion state.
[0046] Specifically, after the motion state information of the current wearable device is detected by the IMU sensor, it can be further determined whether it meets the first preset condition, and then when the motion state information of the wearable device meets the first preset condition, the image acquisition unit of the wearable device is controlled to capture the first image.
[0047] It should be noted that this application does not limit the first preset condition. For example, the first preset condition is that the motion state of the wearable device is relatively stable, which can be specifically determined based on the acceleration information or three-degree-of-freedom information of the wearable device. This application will explain this in detail later. It is understandable that when it is determined that the motion state of the wearable device has reached a relatively stable motion state, the image acquisition unit of the wearable device can be controlled to capture a stable first image, and then the text on the first image can be clearly obtained to improve the quality of subsequent text translation.
[0048] It should be noted that the three-degree-of-freedom information can be used to represent the motion state of the device, which generally involves the device's three motion states of horizontal rotation, vertical rotation, and pitch rotation. Therefore, by detecting the three-degree-of-freedom information of the wearable device, it can be used to determine the current motion state of the wearable device. Then, when the motion state of the wearable device meets the first preset condition, the image acquisition unit is controlled to capture a stable first image.
[0049] In addition, the wearable device may be controlled to capture the first image in direct response to a user's operating instruction. The user's operating instruction may be triggered by pressing a button on the wearable device or by a terminal device connected to the wearable device, and this application does not limit this.
[0050] Based on the above embodiment, the state of the wearable device includes a first state, a second state and a third state. When the first state information meets the first preset condition, the first image is collected by the wearable device, including: when the first state information meets the transition from the first state or the second state to the third state, and the stay time in the third state exceeds the first preset time length, the first image is collected by the wearable device.
[0051] The first state is a violent fluctuation state; the second state is a medium-speed fluctuation state; and the third state is a stable state. The first preset time length can be 0.5s, 1s, or 2s. This application takes the first preset time length of 0.5s as an example for explanation.
[0052] It should be noted that the above-mentioned violent fluctuation state, medium-speed fluctuation state, and stable state can be determined based on the three-degree-of-freedom information or acceleration information of the wearable device. This application uses three-degree-of-freedom information as an example. For example, the three-degree-of-freedom information of the wearable device can be classified according to specific intervals, and each specific interval corresponds to a state. Therefore, after the three-degree-of-freedom information of the wearable device is detected by the IMU sensor, the state of the wearable device can be determined based on the interval in which the three-degree-of-freedom information of the wearable device is located.
[0053] Specifically, when the wearable device changes from a violent fluctuation state or a medium-speed fluctuation state to a stable state, and the stay time in the stable state exceeds 0.5s, it can be determined that the motion state of the wearable device has reached a relatively stable motion state. At this time, the image acquisition unit of the wearable device can be controlled to capture a stable first image, and then the text on the first image can be clearly obtained to improve the quality of subsequent text translation.
[0054] In an embodiment of the present application, the first image may be captured by an image capture unit of the wearable device to obtain text on the first image for subsequent text translation.
[0055] Step S12: Translate the text on the first image to obtain a first text translation result.
[0056] The first text translation result is a translation result of the text on the first image.
[0057] It should be noted that this application does not limit the method for translating text. For example, the wearable device can communicate with the cloud or other devices with translation functions (such as mobile phones, tablets, or computers), and then the wearable device can send the text to the cloud or other devices with translation functions for text translation. Alternatively, optical character recognition technology and a translation engine can be integrated into the wearable device to achieve instant local translation on the wearable device.
[0058] In the embodiment of the present application, the text on the first image can be translated to obtain a corresponding first text translation result, so as to realize real-time text translation.
[0059] Step S13: When the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, obtain the target position of the text on the first image on the current display screen, and superimpose the first text translation result on the target position.
[0060] The validity of the first text translation result indicates the quality or accuracy of the first text translation result; and the target position is the real-time position of the text on the first image in the current display screen.
[0061] It can be understood that when the validity of the first text translation result is greater than the first preset threshold, it means that the quality or accuracy of the first text translation result is high, that is, the first text translation result is valid.
[0062] It should be noted that the present application does not limit the first preset threshold, for example, the first preset threshold is 80%, 90%, etc. In addition, the present application does not limit the method for obtaining the validity of the text translation result. For example, the accuracy of the text translation can be analyzed and evaluated using natural language processing tools, or the text translation quality can be automatically evaluated using BLEU (Bilingual Evaluation Understudy) or ROUGE to obtain the validity of the first text translation result.
[0063] Furthermore, when the validity of the first text translation result is greater than a first preset threshold and the text on the first image is in the current display screen, the real-time position of the text in the current display screen can be obtained, thereby superimposing the first text translation result on the real-time position of the text so that the user can conveniently view the translation result at the position corresponding to the text.
[0064] Optionally, the first text translation result may also be displayed superimposed near the real-time position of the text, for example, superimposed above or below the text, etc., which is not limited in this application.
[0065] It should be noted that this application does not limit the method for obtaining the real-time position of text on the display screen. For example, it can be obtained through visual odometry, particle filter or optical flow tracking. This application uses the optical flow tracking method to obtain the real-time position of text on the display screen as an example.
[0066] Optical flow tracking is a computer vision technique used to track the real-time position of a moving object across consecutive image frames based on the movement patterns of pixels within the image. Specifically, optical flow tracking acquires a series of consecutive image frames containing a specific target, analyzes the pixels in each frame, and estimates their positions in the next frame. This allows tracking of the target's motion by combining the movement information of the pixels within that target.
[0067] In an embodiment of the present application, the position of the text on the display page can be tracked by an optical flow tracking method and continuously updated to reflect the real-time position of the text on the display page.
[0068] Optionally, after translating the text on the first image to obtain a first text translation result, the method further includes: when the validity of the first text translation result is less than or equal to a first preset threshold, re-capturing a second image by the image acquisition unit.
[0069] It is understandable that when the validity of the first text translation result is less than or equal to the first preset threshold, it means that the text quality on the first image is poor or the accuracy of the translation result is not high enough, so it is necessary to obtain a better translation result by re-capturing the image.
[0070] In an embodiment of the present application, the first text translation result can be superimposed on the real-time position of the text, thereby improving the fit between the translation result and the text, allowing the user to better associate the text with the translation result, and then accurately understand and apply the translation result.
[0071] Optionally, after translating the text on the first image and obtaining a first text translation result, the method further includes: when the validity of the first text translation result is greater than a first preset threshold and the text on the first image is not on the current display screen, obtaining the end position of the text on the first image leaving the current display screen; and displaying the first text translation result at the end position.
[0072] Specifically, if the validity of the first text translation result is greater than a first preset threshold, and the text on the first image is not currently displayed, this indicates that the text on the first image may have left the current display screen due to movement of the wearable device or text. In this case, the end position where the text on the first image leaves the current display screen can be obtained, and the first text translation result can be displayed at the end position. This allows the user to associate the first text translation result with the text, allowing them to accurately understand and apply the translation result.
[0073] Optionally, when the text on the first image is not on the current display screen, the first text translation result can also be displayed at other fixed positions on the current display screen, such as the middle position or the top position, so that the user can associate the first text translation result with the text, and then accurately understand and apply the translation result.
[0074] In an embodiment of the present application, when the validity of the first text translation result is greater than a first preset threshold and the text on the first image is not currently displayed, the first text translation result may be displayed at the end position. This allows the user to associate the first text translation result with the text, thereby enabling the user to accurately understand and apply the translation result.
[0075] The text display method based on a wearable device, the wearable device and the computer-readable storage medium disclosed in the embodiment of the present application can capture a first image through a wearable device, wherein the first image includes at least text. And the text on the first image is translated to obtain a first text translation result. Furthermore, when the validity of the first text translation result is greater than a first preset threshold value and the text of the first image is on the current display screen, the target position of the text on the first image on the current display screen is obtained, and the first text translation result is superimposed and displayed on the target position. The present application aims to obtain the target position of the text in the display screen while translating the text, and then superimpose the text translation result on the target position in real time to improve the fit between the translation result and the text, so that the user can better associate the text with the translation result, and then accurately understand and apply the translation result.
[0076] Please continue to refer to Figure 5, which is a schematic diagram of a process for refreshing text translation results provided by an embodiment of the present application. As shown in Figure 5, refreshing text translation results can be achieved through steps S21 to S24.
[0077] Step S21: determining second state information of the wearable device through a preset sensor.
[0078] Step S22: When the second state information satisfies a second preset condition, a second image is captured by the wearable device, wherein the second image includes at least text.
[0079] Among them, the second state information is the motion state information of the wearable device after the text translation result is displayed; the second image is the image captured for the second time by the image acquisition unit of the wearable device; the preset sensor can be an IMU (Inertial Measurement Unit) sensor, a visual sensor, and a positioning sensor, etc. This application takes the preset sensor as an IMU sensor as an example for explanation. For its specific description, please refer to the above embodiment. To avoid repetition, it will not be repeated here.
[0080] As is understandable, since the accuracy of the first text translation result primarily depends on the clarity of the first image, that is, whether the text on the image can be clearly captured, to improve the quality of the text translation result, after displaying the text translation result, a second image displaying text can be further acquired to translate the text on the second image to obtain a second text translation result. When the second text translation result is more effective than the first text translation result, the text translation result can be replaced, thereby refreshing the text translation result.
[0081] Specifically, after the motion state information of the wearable device after displaying the text translation result is detected by the IMU sensor, it can be further determined whether it meets the second preset condition, and then when the motion state information of the wearable device meets the second preset condition, the image acquisition unit of the wearable device is controlled to capture the second image.
[0082] It should be noted that this application does not limit the second preset condition. For example, the second preset condition is that the motion state of the wearable device is relatively stable, which can be determined based on the acceleration information or three-degree-of-freedom information of the wearable device. For specific descriptions, please refer to the above embodiments. It is understandable that when it is determined that the motion state of the wearable device has reached a relatively stable motion state, the image acquisition unit can be controlled to capture a stable second image, and then the text on the second image can be clearly obtained to improve the quality of text translation.
[0083] Optionally, when the second state information satisfies a second preset condition, a second image is captured through the wearable device, including: when the second state information satisfies the transition from the first state or the second state to the third state, and the stay time in the third state exceeds a first preset time length, the second image is captured through the wearable device; or, when the second state information satisfies the stay time in the second state and in the third state exceeds a second preset time length, the second image is captured through the wearable device.
[0084] The first state is a violent fluctuation state; the second state is a medium-speed fluctuation state; and the third state is a stable state. The first preset duration can be 0.5s, 1s, or 2s, etc.; the second preset duration can be 3s, 5s, 6s, etc. This application uses the first preset duration of 0.5s and the second preset duration of 6s as an example for explanation.
[0085] It should be noted that the above-mentioned violent fluctuation state, medium-speed fluctuation state, and stable state can be determined based on the three-degree-of-freedom information or acceleration information of the wearable device. This application uses three-degree-of-freedom information as an example. For example, the three-degree-of-freedom information of the wearable device can be classified according to specific intervals, and each specific interval corresponds to a state. Therefore, after the three-degree-of-freedom information of the wearable device is detected by the IMU sensor, the state of the wearable device can be determined by determining the interval in which it is located.
[0086] Specifically, when the wearable device changes from a violent fluctuation state or a medium-speed fluctuation state to a stable state, and the stay time in the stable state exceeds 0.5s, it is determined that the motion state of the wearable device has reached a relatively stable motion state. At this time, the image acquisition unit of the wearable device can be controlled to capture a stable second image. Alternatively, when the wearable device stays in the medium-speed fluctuation state and the stable state for more than 6s, it is determined that the motion state of the wearable device has reached a relatively stable motion state. At this time, the image acquisition unit of the wearable device can be controlled to capture a stable second image. In this way, the text on the second image can be clearly obtained to improve the quality of subsequent text translation.
[0087] Step S23: Translate the text on the second image to obtain a second text translation result.
[0088] Step S24: When the validity of the second text translation result is greater than the validity of the first text translation result, the first text translation result is replaced with the second text translation result using the display unit of the wearable device.
[0089] The second text translation result is a translation result of the text on the second image.
[0090] Specifically, the text on the second image can be translated to obtain a corresponding second text translation result, thereby achieving real-time text translation. Furthermore, the validity of the second text translation result can be determined and compared with the validity of the first text translation result.
[0091] It should be noted that the methods for implementing text translation and obtaining the validity of text translation results can be found in the above embodiments, and are not described here in detail to avoid repetition.
[0092] It is understandable that if the validity of the second text translation result is greater than the validity of the first text translation result, it means that the quality of the second text translation result is higher than that of the first text translation result. In this case, the first text translation result can be replaced with the second text translation result. Conversely, it means that the quality of the second text translation result is not as high as that of the first text translation result. In this case, the second text translation result can be discarded and the first text translation result can be retained.
[0093] In an embodiment of the present application, a second image containing text can be further acquired based on the state of the wearable device to translate the text on the second image to obtain a second text translation result. Furthermore, when the quality of the second text translation result is higher than that of the first text translation result, the first text translation result is replaced with the second text translation result. This allows for the dynamic provision of more accurate, real-time text translation services, ensuring that users receive relatively accurate translation results while meeting their translation needs under specific conditions.
[0094] As shown in Figure 2, the scenario of the text display method based on wearable devices of the present application mainly includes the following three processes: input process (Input process), processing process (Process process) and output process (Output process). Among them, the IMU (Inertial Measurement Unit) + 3dof (degree of freedom) scheme image acquisition process corresponds to the present application's determination of the degree of freedom information of the wearable device based on the IMU sensor, and then determining the state of the wearable device based on the degree of freedom information of the wearable device, so that when the state of the wearable device reaches the preset condition, the image acquisition unit of the wearable device is controlled to capture a stable first image.
[0095] Specifically, during the input process, the image acquisition unit can be controlled to capture the first image in response to the user's operating instructions, that is, by pressing a button on the wearable device (clicking to take a photo). The image acquisition unit of the wearable device has no stored images before capturing the first image, or clears the original stored images to prevent other images from conflicting with the first image, thereby preventing the user from accurately obtaining the translation text required by the user. In addition, the initial display screen corresponding to the first image does not display text before the image is captured to reduce interference factors in the subsequent text translation process. 。
[0096] It should be noted that this application does not limit the above-mentioned method of implementing the image acquisition unit of the wearable device to have no stored image before acquiring the first image, and implementing the method of implementing the initial display screen corresponding to the first image to have no displayed text before acquiring the image. For example, the "forced start function" of the image acquisition unit can be turned on to clear the original text and status of the image acquisition unit (that is, the initial display screen corresponding to the first image has no displayed text before acquiring the image); or, the "initialization function" of the image acquisition unit can be turned on to implement the above method of implementing the image acquisition unit to have no stored image before acquiring the first image, and the initial display screen corresponding to the first image has no displayed text before acquiring the image.
[0097] Furthermore, before the image acquisition unit of the wearable device captures the first image, it can also obtain the degree of freedom information of the wearable device through a preset sensor (such as an IMU sensor) to determine the state information of the wearable device. When the wearable device transitions from a violent fluctuation state or a medium-speed fluctuation state to a stable state, and the duration of the stable state exceeds 0.5 seconds, it can be determined that the motion state of the wearable device has reached a relatively stable motion state. At this time, the image acquisition unit of the wearable device can be controlled to capture a stable first image.
[0098] During the processing, the text on the first image can be translated via the cloud. Thus, during the output process, the first text translation result can be output via the cloud and displayed superimposed on the target location; or the text displayed on the first image can be moved away from the end of the current display screen.
[0099] Optionally, the validity of the first text translation result can be determined, that is, whether a valid result is generated. If the first text translation result is valid (a valid result is generated), the first text translation result is maintained. If the first text translation result is invalid, the image acquisition unit of the wearable device is used to capture an image again, and the above text translation steps are repeated.
[0100] Furthermore, when the first text translation result is valid, the first text translation result may be determined as the original translation copy, and the following steps may be continued:
[0101] During the output process, the state information of the wearable device (corresponding to the state of the new text in Figure 2) continues to be determined through a preset sensor (such as an IMU sensor). When the wearable device stays in a medium-speed fluctuation state and a stable state for more than 6 seconds, or changes from a violent fluctuation state or a medium-speed fluctuation state to a stable state, and the stay time in the stable state exceeds 0.5 seconds, it is determined that the motion state of the wearable device has reached a relatively stable motion state. At this time, the image acquisition unit of the wearable device can be controlled to capture a stable second image.
[0102] Thus, during the processing, the text on the first image can be translated through the cloud, and during the output process, a second text translation result can be obtained through the cloud output.
[0103] Optionally, the validity of the second text translation result can be further compared with the validity of the first text translation result (original translation document), that is, whether a valid result is generated. If the validity of the second text translation result is greater than the validity of the first text translation result (original translation document) (a valid result is generated), the first text translation result (original translation document) is replaced with the second text translation result. If the validity of the second text translation result is less than or equal to the validity of the first text translation result (original translation document) (no valid result is generated), the first text translation result (original translation document) is maintained.
[0104] The methods of the present application can be used in a wide variety of general-purpose or specialized computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer devices, network PCs, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.
[0105] Please refer to Figure 6, which is a schematic block diagram of a text display device provided in an embodiment of the present application. The text display device can be configured in a wearable device to execute the aforementioned wearable device-based text display method.
[0106] As shown in FIG6 , the text display device 200 includes: an acquisition module 201 , a translation module 202 , and a display module 203 .
[0107] An acquisition module 201 is configured to acquire a first image captured by the wearable device, wherein the first image includes at least text;
[0108] A translation module 202 is configured to translate the text on the first image to obtain a first text translation result;
[0109] The display module 203 is configured to obtain the text on the first image at a target position on the current display screen, and to superimpose and display the first text translation result on the target position when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen.
[0110] The acquisition module 201 is further used to determine the first state information of the wearable device through a preset sensor; when the first state information meets a first preset condition, capture the first image through the wearable device; or, in response to a user operation instruction, control the wearable device to capture the first image.
[0111] The acquisition module 201 is further configured to capture the first image through the wearable device when the first state information satisfies the transition from the first state or the second state to the third state, and the stay time in the third state exceeds a first preset time length.
[0112] The display module 203 is further configured to, when the validity of the first text translation result is greater than a first preset threshold and the text on the first image is not on the current display screen, obtain an end position where the text on the first image leaves the current display screen; and display the first text translation result at the end position.
[0113] The display module 203 is further configured to determine second state information of the wearable device through the preset sensor; when the second state information satisfies a second preset condition, capture a second image through the wearable device, wherein the second image includes at least text; translate the text on the second image to obtain a second text translation result; and when the validity of the second text translation result is greater than the validity of the first text translation result, replace the first text translation result with the second text translation result.
[0114] The acquisition module 201 is also used to collect the second image through the wearable device when the second state information satisfies the transition from the first state or the second state to the third state, and the stay time in the third state exceeds the first preset time length; or, when the second state information satisfies the stay time in the second state and in the third state exceeds the second preset time length, collect the second image through the wearable device.
[0115] The acquisition module 204 is further configured to re-capture the second image through the wearable device when the validity of the first text translation result is less than or equal to the first preset threshold.
[0116] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0117] Exemplarily, the above method and apparatus may be implemented in the form of a computer program, which may be run on a wearable device as shown in FIG. 7 .
[0118] Please refer to FIG7 , which is a schematic diagram of a wearable device provided in an embodiment of the present application.
[0119] As shown in FIG7 , the wearable device 400 includes a processor 401 and a memory 402 connected via a system bus, wherein the memory 402 may include a volatile storage medium, a non-volatile storage medium, and an internal memory.
[0120] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can enable the processor 401 to perform any one of the text display methods based on the wearable device.
[0121] The processor 401 is used to provide computing and control capabilities to support the operation of the entire wearable device 400.
[0122] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 401, the processor 401 can execute any text display method based on the wearable device.
[0123] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that the structure of the wearable device 400 is merely a block diagram of a portion of the structure related to the present application solution, and does not constitute a limitation on the wearable device 400 to which the present application solution is applied. The specific wearable device 400 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0124] It should be understood that the processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0125] In some embodiments, the processor 401 is used to run a computer program stored in the memory 402 to implement the following steps: obtaining a first image captured by the wearable device, wherein the first image includes at least text; translating the text on the first image to obtain a first text translation result; when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, obtaining the text on the first image at a target position on the current display screen, and superimposing the first text translation result on the target position.
[0126] In some embodiments, the processor 401 is further used to determine first state information of the wearable device through a preset sensor; when the first state information meets a first preset condition, capture the first image through the wearable device; or, in response to a user operation instruction, control the wearable device to capture the first image.
[0127] In some embodiments, the processor 401 is further used to collect the first image through the wearable device when the first state information satisfies the transition from the first state or the second state to the third state, and the stay time in the third state exceeds a first preset time length.
[0128] In some embodiments, the processor 401 is further configured to, when the validity of the first text translation result is greater than the first preset threshold and the text on the first image is not on the current display screen, obtain the end position where the text on the first image leaves the current display screen; and display the first text translation result at the end position.
[0129] In some embodiments, the processor 401 is further used to determine second state information of the wearable device through the preset sensor; when the second state information meets a second preset condition, capture a second image through the wearable device, wherein the second image includes at least text; translate the text on the second image to obtain a second text translation result; when the validity of the second text translation result is greater than the validity of the first text translation result, replace the first text translation result with the second text translation result.
[0130] In some embodiments, the processor 401 is further used to collect the second image through the wearable device when the second state information satisfies the transition from the first state or the second state to the third state, and the stay time in the third state exceeds the first preset time length; or, collect the second image through the wearable device when the second state information satisfies the stay time in the second state and in the third state exceeds the second preset time length.
[0131] In some embodiments, the processor 401 is further configured to re-capture the second image through the wearable device when the validity of the first text translation result is less than or equal to the first preset threshold.
[0132] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions, and when the program instructions are executed, any one of the text display methods based on a wearable device provided in the embodiment of the present application is implemented.
[0133] In some embodiments, when the program instructions are executed, the following steps are implemented: obtaining a first image captured by the wearable device, wherein the first image includes at least text; translating the text on the first image to obtain a first text translation result; when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, obtaining the text on the first image at a target position on the current display screen, and superimposing the first text translation result on the target position.
[0134] The computer-readable storage medium may be an internal storage unit of the wearable device described in the aforementioned embodiment, such as a hard disk or memory of the wearable device. The computer-readable storage medium may also be an external storage device of the wearable device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the wearable device.
[0135] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application program required for at least one function, and the like.
[0136] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A text display method based on a wearable device, wherein: The method comprises: Acquire a first image captured by the wearable device, wherein the first image includes at least text; Translating the text on the first image to obtain a first text translation result; When the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, the text on the first image is acquired at a target position on the current display screen, and the first text translation result is superimposed and displayed on the target position.
2. The method according to claim 1, wherein: Before acquiring the first image captured by the wearable device, the method includes: Determining first state information of the wearable device through a preset sensor; When the first state information satisfies a first preset condition, collecting the first image through the wearable device; or, In response to a user operation instruction, the wearable device is controlled to capture the first image.
3. The method according to claim 2, wherein: The state of the wearable device includes a first state, a second state, and a third state, and when the first state information satisfies a first preset condition, collecting the first image through the wearable device includes: When the first state information satisfies the requirement of being transformed from the first state or the second state to the third state, and the stay time in the third state exceeds a first preset time length, the first image is collected through the wearable device.
4. The method according to claim 1, wherein: After translating the text on the first image to obtain a first text translation result, the method further includes: When the validity of the first text translation result is greater than a first preset threshold and the text on the first image is not on the current display screen, obtaining an end position of the text on the first image leaving the current display screen; The first text translation result is displayed at the end position.
5. The method according to claim 1, wherein: After translating the text on the first image to obtain a first text translation result, the method further includes: When the validity of the first text translation result is greater than a first preset threshold and the text on the first image is not in the current display screen, the first text translation result is displayed at a fixed position in the current display screen.
6. The method according to claim 5, wherein: The fixed position of the current display picture includes a middle position or a top position of the current display picture.
7. The method according to claim 1, wherein: After the first text translation result is superimposed and displayed on the target position, the method further includes: Determining second state information of the wearable device through the preset sensor; When the second state information satisfies a second preset condition, collecting a second image through the wearable device, wherein the second image includes at least text; Translating the text on the second image to obtain a second text translation result; When the validity of the second text translation result is greater than the validity of the first text translation result, the first text translation result is replaced with the second text translation result.
8. The method according to claim 7, wherein: The state of the wearable device includes a first state, a second state and a third state, and when the second state information satisfies a second preset condition, collecting a second image through the wearable device includes: When the second state information satisfies the transition from the first state or the second state to the third state, and the stay time in the third state exceeds the first preset time length, the second image is collected by the wearable device; or, When the second state information satisfies that the stay time in the second state and in the third state exceeds a second preset time length, the second image is collected through the wearable device.
9. The method according to claim 3 or 8, wherein: The first state is a violent fluctuation state, the second state is a medium-speed fluctuation state, and the third state is a stable state.
10. The method according to claim 9, wherein: The violent fluctuation state, the medium-speed fluctuation state and the stable state are determined according to the interval in which the three-degree-of-freedom information or acceleration information of the wearable device is located.
11. The method according to claim 1, wherein: The wearable device has no stored images before acquiring the first image, or clears the originally stored images.
12. The method according to claim 1, wherein: The initial display screen corresponding to the first image has no displayed text before the image is captured.
13. The method according to claim 1, wherein: The validity of the first text translation result indicates the quality or accuracy of the first text translation result, and the target position is the real-time position of the text on the first image in the current display screen.
14. The method according to claim 13, wherein: The first text translation result is displayed in a superimposed manner near the real-time position.
15. The method according to claim 1, wherein: After translating the text on the first image to obtain a first text translation result, the method further includes: When the validity of the first text translation result is less than or equal to the first preset threshold, the second image is collected again through the wearable device.
16. A text display device, wherein: The text display device comprises: An acquisition module, configured to acquire a first image captured by the wearable device, wherein the first image includes at least text; A translation module, used for translating the text on the first image to obtain a first text translation result; The display module is used for obtaining the text on the first image at a target position of the current display screen, and superimposing and displaying the first text translation result on the target position when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen.
17. A wearable device, wherein: include: A memory and a processor; wherein the memory is connected to the processor and is used to store a program; the processor is used to implement the following steps by running the program stored in the memory: Acquire a first image captured by the wearable device, wherein the first image includes at least text; translate the text on the first image to obtain a first text translation result; when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, acquire the text on the first image at a target position on the current display screen, and overlay and display the first text translation result on the target position.
18. The wearable device according to claim 17, wherein: The processor is also used to determine the first state information of the wearable device through a preset sensor; when the first state information meets a first preset condition, collect the first image through the wearable device; or, in response to a user operation instruction, control the wearable device to collect the first image.
19. The wearable device according to claim 18, wherein: The processor is further configured to capture the first image through the wearable device when the first state information satisfies the requirement of being transformed from the first state or the second state to the third state, and when the stay time in the third state exceeds a first preset time length.
20. A computer-readable storage medium, wherein: The computer readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the following steps: Acquire a first image captured by the wearable device, wherein the first image includes at least text; translate the text on the first image to obtain a first text translation result; when the validity of the first text translation result is greater than a first preset threshold and the text of the first image is on the current display screen, acquire the text on the first image at a target position on the current display screen, and overlay and display the first text translation result on the target position.
Citation Information
Patent Citations
Electronic device recognizing text in image
CN109840465A
Text shooting method, wearable equipment and storage medium
CN111970437A
Text translation method and device in page, computer equipment and storage medium
CN112507729A
Translation method and device, wearable equipment and storage medium
CN115131791A