Text translation method and device, computer equipment and storage medium
The acquisition and processing of images for text translation by smart wearable devices solves the problem of low translation efficiency in the prior art, real-time translation without operating a smart terminal is realized, and the user's vision is detected through sensors, improving the accuracy and efficiency of translation.
Patent Information
- Application Number
- CN202311463090.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
The existing text translation technology has shortcomings in real-time translation efficiency, especially when users cannot operate smart terminals, resulting in a decrease in translation efficiency.
The image to be processed is collected by the smart wearable device, text processing is performed to obtain the text to be translated, and translated based on the preset configuration parameters. Finally, the translated text is displayed in the display interface using the preset conversion matrix.
It realizes text translation directly through the smart wearable device without the user operating the smart terminal, which improves translation efficiency, and detects user head movement through preset sensors to determine the translated text corresponding to the text to be translated in the user's target field of view, further improving the accuracy and efficiency of translation.
Smart Images

Figure CN119940376A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text translation, and in particular to a text translation method, apparatus, computer equipment and storage medium. Background Art
[0002] Existing text translation technology is mainly used in applications of smart terminals, such as Youdao Dictionary, Google Lens, etc. These applications capture text through the camera of smart terminal devices and display the translation results on the screen of smart terminal devices. Users need to hold the terminal device continuously and cannot free their hands. In situations such as translating road signs when traveling, translating product information when shopping, and translating bills when paying, it is inconvenient to use smart terminals or the smart terminals are occupied, resulting in the inability to obtain translation results in time, reducing translation efficiency. Therefore, how to improve the efficiency of real-time translation of text has become an urgent problem to be solved. Summary of the invention
[0003] The present application provides a text translation method, apparatus, computer device and storage medium to improve text translation efficiency.
[0004] In a first aspect, the present application provides a text translation method, the method comprising:
[0005] Based on the smart wearable device, an image to be processed is collected, and original text in the image to be processed is processed to obtain a text to be translated;
[0006] Based on preset configuration parameters, the text to be translated is translated to obtain a translated text;
[0007] The translated text is mapped based on a preset conversion matrix, and the translated text is displayed in a display interface.
[0008] In a second aspect, the present application further provides a text translation device, the device comprising:
[0009] A module for obtaining text to be translated, which is used to collect images to be processed based on a smart wearable device, and perform text processing on the original text in the images to be processed to obtain text to be translated;
[0010] A translation text obtaining module, used to translate the text to be translated based on preset configuration parameters to obtain a translation text;
[0011] The translation text display module is used to map the translation text based on a preset conversion matrix and display the translation text in a display interface.
[0012] In a third aspect, the present application also provides a computer device, comprising a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the text translation method as described above when executing the computer program.
[0013] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the text translation method as described above.
[0014] The present application discloses a text translation method, device, computer equipment and storage medium. Based on a smart wearable device, an image to be processed is collected, and the original text in the image to be processed is processed to obtain a text to be translated; based on preset configuration parameters, the text to be translated is translated to obtain a translated text; based on a preset conversion matrix, the translated text is mapped and displayed in a display interface. On the one hand, the present invention can directly translate the text content in the image to be processed through a smart wearable device without the user having to operate the relevant application of the smart terminal, thereby improving the translation efficiency; on the other hand, the user's head movement is detected by a preset sensor, the user's target field of view at the second moment is determined, and the translated text corresponding to the text to be translated in the user's target field of view is displayed to the user, so that the user can quickly and accurately obtain the required translated text, thereby improving the translation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 is a schematic flow chart of a first embodiment of a text translation method provided by an embodiment of the present application;
[0017] Figure 2 It is a text display effect diagram of a text translation method provided in an embodiment of the present application;
[0018] Figure 3 is a schematic flow chart of a second embodiment of a text translation method provided in an embodiment of the present application;
[0019] Figure 4 It is a schematic diagram of a paragraph of a text to be corrected in a text translation method provided in an embodiment of the present application;
[0020] Figure 5is a schematic flow chart of a third embodiment of a text translation method provided by an embodiment of the present application;
[0021] Figure 6 A rendering of a translated text display of a target area of a text translation provided in an embodiment of the present application;
[0022] Figure 7 A text translation display effect diagram of another target area provided in an embodiment of the present application;
[0023] Figure 8 A schematic block diagram of a text translation device provided in an embodiment of the present application;
[0024] Fig. 9 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0026] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0027] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0028] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0029] The embodiments of the present application provide a text translation method, apparatus, computer equipment and storage medium. The text translation method can be applied to a server to implement text translation through a smart wearable device to improve text translation efficiency. The server can be an independent server or a server cluster.
[0030] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0031] See also Figure 1 , Figure 1 1 is a schematic flow chart of a text translation method provided by an embodiment of the present application. The text translation method can be applied to a smart wearable device to translate text content in an image to be processed through the smart wearable device, thereby improving text translation efficiency.
[0032] like Figure 1 As shown, the text translation method specifically includes steps S101 to S103.
[0033] S101. Based on a smart wearable device, an image to be processed is collected, and original text in the image to be processed is processed to obtain a text to be translated.
[0034] In one embodiment, the image to be processed is obtained by photographing the image in the user's field of view through the camera module of the smart wearable device (such as AR glasses, AR helmets, etc.). It is understandable that the camera module can be a camera built into the smart wearable device, or it can be a camera connected to the smart wearable device. Among them, the image to be processed may include the original text to be translated and the scene image. For example, the menu is translated using a smart wearable device, and the image to be processed includes the name of the dish on the menu, the image of the dish, the desktop, etc.
[0035] Furthermore, based on the smart wearable device, after collecting the image to be processed, it also includes: based on the conversion matrix, converting the image to be processed to display the image to be processed in the display interface.
[0036] In one embodiment, a conversion matrix pre-stored in the smart wearable device is obtained, wherein the conversion matrix is a matrix generated when the smart wearable device is calibrated for virtuality and reality before leaving the factory, and is used to obtain an image displayed on a display interface after multiplying an image captured by a camera module of the smart wearable device by the conversion matrix, that is, mapping objects in reality on the display interface of the smart wearable device for user viewing.
[0037] In one embodiment, text recognition is performed on the image to be processed by a layout detection algorithm, and when the original text in the image to be processed is recognized, the position of the original text in the image to be processed is obtained. The layout detection algorithm can automatically analyze, identify and understand the image, text, table information and position relationship in the layout.
[0038] In one embodiment, the position of the original text is represented by the upper left corner point and the lower right corner point of the original text. According to the upper left corner point and the lower right corner point, the size of the area to be cropped can be obtained. The image to be processed is cropped according to the size of the cropping area to obtain a target text area that only contains text.
[0039] In one embodiment, text processing includes text detection and text correction. Text detection is performed on the target text area to extract the text in the target text area. Due to problems such as environmental light, the brightness and darkness of each area of the image to be processed collected are different. Therefore, errors may occur in the text extracted by text detection. For example, word errors (such as, I “门”), grammar errors (可爱 “地”), etc. Therefore, the extracted text is used as the text to be corrected. In order to improve the translation accuracy, error correction and text correction need to be performed on the text to be corrected.
[0040] In one embodiment, a large model (large-scale deep learning model) is used to perform error correction and text correction on the text to be corrected. The error items in the text to be corrected are corrected to obtain the text to be translated after correction.
[0041] S102. Translate the text to be translated based on preset configuration parameters to obtain a translation text.
[0042] In one embodiment, the preset configuration parameters include the language type, and can also include the font, font size, spacing, line spacing, etc. of the translation text.
[0043] In one embodiment, the configuration parameters are parameters related to the translation text. They can be set by the user in advance in the system of the smart wearable device. For example, the language type set by the user for the current smart wearable device system is Chinese, and “Chinese” is obtained as the language type of the translation text. When the language type of the text to be translated obtained by the smart wearable device does not belong to Chinese text, the text to be translated is translated into Chinese; it can also ask the user after obtaining the text to be translated to obtain the language type set by the user. For example, after obtaining the text to be translated, a pop-up window, voice, etc. are initiated to ask the user what language type the text to be translated needs to be translated into, and the user operation is received to obtain the language type set by the user. Among them, the user operation can be the user's voice answer, or the smart wearable device gives voice type options and the user selects by nodding or shaking the head, or other operations that can determine the user's selection.
[0044] In one embodiment, the translation function of the smart wearable device can be enabled by the user in the function module before using the smart wearable device, and automatically translated according to the language type when the text to be translated is obtained;
[0045] In another embodiment, the translation function of the smart wearable device can also ask the user whether translation is needed when the smart wearable device obtains the text to be translated, and translate the text to be translated after receiving the user's instruction to confirm the translation.
[0046] In one embodiment, the smart wearable device can have a built-in translation module to complete the translation of the text to be translated, or it can complete the translation of the text to be translated through a cloud translation module connected to the smart wearable device. Among them, the cloud translation module can be a translation technology cloud platform, which is a cloud storage based on a corpus system and translation. Various corpus resources, various storage media and servers are stored together in the cloud to achieve translation resource sharing and form a mutually reinforcing translation ecosystem.
[0047] In one embodiment, when translating the text to be translated, a translation module is called to translate the text to be translated according to the configuration parameters to obtain a translated text.
[0048] In another embodiment, when translating the text to be translated, the cloud translation module is used to translate the text to be translated. The connection between the cloud translation module and the smart wearable device can be determined by the developer of the smart wearable device during the generation process, and the connection between the cloud translation module and the smart wearable device can be established, or the user can select the target cloud translation module for connection.
[0049] S103: Map the translation text based on a preset conversion matrix, and display the translation text in a display interface.
[0050] In one embodiment, the translated text is mapped to the display interface of the smart wearable device according to the conversion matrix, and is suspended in an overlay manner just above the original text in the display interface.
[0051] Furthermore, displaying the translated text in a display interface also includes: obtaining a paragraph text box of the text to be corrected; determining a font size of the translated text based on a size of the paragraph text box and the number of words in the translated text; and displaying the translated text in the display interface based on the font size.
[0052] In one embodiment, for the display of the translated text, the font size of the translated text is determined in combination with the number of words in the translated text and the text box size of the text to be translated, so that the translated text is just in the text box of the text to be translated, and it can be ensured that the translated text superimposed on the top of the text to be translated is the translated text corresponding to the text to be translated. For example, the text box corresponding to the text to be translated can store the text to be translated consisting of 100 words with a text font size of four, and the translated text corresponding to the translated text has 150 words, then the font size of the translated text needs to be adjusted down (such as five) to achieve that the display area of the translated text does not exceed the text box size of the text to be translated. It can be understood that if the font size of the translated text set by the user in the configuration parameters, the font size set by the user is displayed first when displayed, and when the display range of the translated text corresponding to the font size exceeds the text box size of the text to be translated, the font size is changed so that the display range of the translated text is in the corresponding text box of the text to be translated.
[0053] In a specific embodiment, a paragraph text box of the text to be revised is obtained, and the font size of the translated text is determined according to the size of the paragraph text box and the number of words in the translated text, so that the translated text can be displayed in the paragraph text box. For example, the paragraph text box corresponding to the text to be revised can store the text to be revised with a font size of 4 and 100 words, and the translated text corresponding to the text to be revised has 150 words, then the font size of the translated text needs to be reduced (such as 5).
[0054] In one embodiment, the translated text is displayed in an overlay manner just above the text to be revised. From the user's perspective, the translated text is suspended just above the text to be revised.
[0055] Furthermore, displaying the translated text in a display interface includes: detecting the position change and angle change of the image to be processed in three-dimensional space based on an optical flow algorithm; determining the three-dimensional change of the translated text based on the position change and angle change of the image to be processed, and displaying the translated text in a display interface based on the three-dimensional change.
[0056] In one embodiment, a three-dimensional reconstruction algorithm is used to perform three-dimensional reconstruction of the environment, and the three-dimensional information of the space is optimized from the spatially discrete depth information point cloud.
[0057] In one embodiment, the optical flow algorithm can find the position of a pixel on one image in another image. The two images are time series images, that is, two images at adjacent moments. The algorithm outputs the relative position change of the pixels in the two images. Through the optical flow algorithm, the three-dimensional position change of the image to be processed is tracked, such as the size change, position change and angle change of the image. Figure 2As shown, according to the three-dimensional information of the space and the three-dimensional position change of the image to be processed, the three-dimensional change of the translated text is obtained, and the translated text at different times is displayed in the display interface according to the three-dimensional change of the translated text, so that the translated text is kept superimposed and displayed at the position of the original text as the user moves. It can be understood that when the three-dimensional change of the translated text is the same as the three-dimensional change of the image to be processed, the translated text is always displayed in a superimposed floating manner directly above the text to be translated.
[0058] The above-mentioned embodiment provides a text translation method, device, computer equipment and storage medium. Based on the smart wearable device, the image to be processed is collected, and the original text in the image to be processed is processed to obtain the text to be translated; based on the preset configuration parameters, the text to be translated is translated to obtain the translated text; based on the preset conversion matrix, the translated text is mapped and the translated text is displayed in the display interface. On the one hand, the present invention can directly translate the text content in the image to be processed through the smart wearable device, thereby improving the translation efficiency; on the other hand, the user's head movement is detected by a preset sensor, the user's target field of view at the second moment is determined, and the translated text corresponding to the text to be translated in the user's target field of view is displayed to the user, so that the user can quickly and accurately obtain the required translated text, thereby improving the translation efficiency.
[0059] See also Figure 3 , Figure 3 This is a schematic flow chart of a text translation method provided by an embodiment of the present application. The text translation method can be applied to a smart wearable device to translate text content in an image to be processed through the smart wearable device, thereby improving text translation efficiency.
[0060] like Figure 3 As shown, the step S101 of the text translation method specifically includes steps S201 to S203.
[0061] S201, performing text recognition on the image to be processed, and determining the position of the original text in the image to be processed according to the recognition result;
[0062] S202, based on the position of the original text, cropping the image to be processed to obtain a target text area;
[0063] S203: Perform text detection and text correction on the target text area to obtain a text to be translated.
[0064] In one embodiment, text recognition is performed on the image to be processed by a layout detection algorithm to determine whether there is text in the image to be processed and obtain the position of the original text in the image to be processed. The layout detection algorithm can automatically analyze, identify and understand the image, text, table information and position relationship in the layout.
[0065] In one embodiment, the position of the original text can be represented by the upper left corner and the lower right corner of the original text. For example, the coordinates of the upper left corner are A and the coordinates of the lower right corner are B, then the position of the original text is represented as (A, B).
[0066] In one embodiment, the image to be processed is cropped according to the position of the original text to obtain a target text area containing only the text. The position of the original text is represented by the upper left corner point and the lower right corner point of the text, and the size of the area to be cropped can be obtained according to the upper left corner point and the lower right corner point. The image to be processed is cropped according to the size of the cropped area to obtain a target text area containing only the text.
[0067] Furthermore, the performing text detection and text correction on the target text area to obtain the text to be translated includes: based on a preset text detection algorithm, performing text detection on the target text area to obtain at least one single-line text; based on a paragraph recognition algorithm, performing paragraph division on at least one of the single-line texts to obtain at least one text to be corrected; based on a preset word library and a preset grammar library, performing text correction on at least one of the texts to be corrected to obtain the text to be translated.
[0068] In one embodiment, text detection is performed on the target text area using an OCR (Optical Character Recognition) algorithm, and the text in the target text area is divided into lines to obtain at least one single line of text, a text box, and a text box position in the target text area. The position of the text box is represented by the upper left corner point and the lower right corner point of the text box.
[0069] In one embodiment, through the paragraph recognition algorithm, the single-line texts belonging to the same paragraph are calculated as being divided into the same paragraph text. For example, the distance between two adjacent text boxes can be calculated, and when the distance is less than a preset value, the two adjacent single-line texts are considered to belong to the same paragraph text. Figure 4As shown, the single-line texts belonging to the same paragraph are merged to generate the text to be corrected, and the paragraph text box of the text to be corrected is obtained. Among them, the distance between two adjacent text boxes can be calculated from the positions of the text boxes. For example, if the position of the first text box is ((1, 2), (4, 3)) and the position of the second text box is ((1, 5), (4, 6)), it means that the first text box and the second text box are 3 unit lengths apart in the vertical direction. When the preset value is 4 units, the first single-line text and the second single-line text can be classified as the same paragraph text; when the preset value is 2 units, the first single-line text and the second single-line text do not belong to the same paragraph text. In another embodiment, the original text can also be segmented into paragraphs through algorithms such as machine learning.
[0070] In one embodiment, after all the single-line texts obtained are segmented, at least one text to be corrected is obtained. It can be understood that a single-line text can also be used as a text to be corrected.
[0071] It can be understood that due to problems such as ambient light, the brightness and darkness of each area of the image to be processed collected are different. Therefore, errors may occur in the text detected and extracted. For example, word errors (such as "我门" instead of "我们"), grammar errors (such as "可爱地" instead of "可爱的"), etc. Therefore, the extracted text is used as the text to be corrected. In order to improve the translation accuracy, it is necessary to correct and revise the text to be corrected.
[0072] In one embodiment, in one embodiment, a large model (large-scale deep learning model) is used to correct and revise the text to be corrected. The error items in the text to be corrected are corrected to obtain the text to be translated after correction. In another embodiment, the large model can also be used to revise the translated text obtained after translation to correct problems such as translation errors caused by language differences and improve the accuracy of the translated text.
[0073] In a specific embodiment, the large model is used to segment the text to be corrected, and the obtained words are compared with the words in the word library to determine the wrong words. The grammar of the text to be corrected is judged through each segmentation and the context semantics, and the wrong grammar is determined in combination with the grammar library. The word library and the grammar library can be set in the system of the smart wearable device, or they can be the cloud word library and the cloud grammar library. The large model is used to correct the identified wrong words and wrong grammar. For example, change "我门" to "我们" and "可爱地" to "可爱的".
[0074] In one embodiment, the translation can be performed directly based on the revised text to be translated to obtain the translated text, saving translation time; the revised text to be translated can also be displayed to the user, and the user can determine whether the revised text to be translated is the same as the original text by comparison. After the user confirms that the text to be translated is the same as the original text, the text to be translated will no longer be displayed, avoiding the overlapping display of the text to be translated and the translated text, which will cause reading difficulties, so that the user can view the translated text more intuitively. It is understandable that if the user finds that the text to be translated is not the same as the original text, a re-identification instruction can be issued to the smart wearable device, and the smart wearable device can regain the text to be translated; the text to be translated can also be directly modified to obtain the text to be translated that is the same as the original text, thereby improving the accuracy of the text to be translated and thus improving the accuracy of the translation.
[0075] The above-mentioned embodiment provides a text translation method, device, computer equipment and storage medium, which performs text recognition on the image to be processed, determines the position of the original text in the image to be processed according to the recognition result; based on the position of the original text, the image to be processed is cropped to obtain a target text area; the target text area is subjected to text detection and text correction to obtain the text to be translated. The method performs text detection on the collected image to be processed to determine the area of the text, and then crops the image to be processed to obtain the target text area, thereby avoiding unnecessary operations on areas where no text exists, thereby improving the translation efficiency. On the other hand, by performing text recognition and text processing on the text in the target text area, error correction can be performed on the identified text to be translated, thereby obtaining the paragraph text to be translated, improving the accuracy of the text to be translated, translating the paragraph text to be translated, obtaining an accurate translation text, and improving the accuracy of the translation.
[0076] See also Figure 5 , Figure 5 1 is a schematic flow chart of a text translation method provided by an embodiment of the present application. The text translation method can be applied to a smart wearable device to translate text content in an image to be processed through the smart wearable device, thereby improving text translation efficiency.
[0077] like Figure 5 As shown, the step S103 of the text translation method specifically includes steps S301 to S303.
[0078] S301, detecting a deflection value of the user's head based on a preset sensor;
[0079] S302, determining a movement direction and a movement distance of the user's head relative to a first visual field area at a first moment based on the deflection value, and determining a target visual field of the user at a second moment based on the movement direction and the movement distance;
[0080] S303: When the original text exists in the user's target field of view, a translation text corresponding to the original text is obtained, and the translation text is superimposed and displayed at the position of the original text in the display interface.
[0081] In one embodiment, the preset sensor may be an IMU (inertial measurement unit) sensor. The IMU sensor may detect the angle and direction of the user's head movement. For example, when the head moves downward, the gyroscope sensor in the IMU sensor detects that the head is turning downward, the accelerometer detects the acceleration of the movement, and the distance that the head is turning downward is calculated based on the acceleration. The size of the user's target visual field is certain, so the position of the user's target visual field at the second moment can be determined based on the angle direction of the head movement. For example, the user's current visual field is used as the visual field at the first moment, and the visual field at the first moment is used as a reference. When it is detected that the user's head moves relative to the visual field at the first moment at the second moment, the visual field at the second moment is determined based on the movement of the user's head. If the user turns left 10 degrees, the target visual field at the first moment moves left 10 degrees, and the left area of the target visual field at the first moment is used as the user's target visual field at the second moment. Among them, the IMU sensor is a combination of an accelerometer and a gyroscope sensor.
[0082] In one embodiment, depending on the target field of view of the user, when there is original text in the target field of view of the user, the translated text corresponding to the original text in different areas is displayed. The area of the target field of view of the user is smaller than the size of the image to be processed collected by the smart wearable device. The smart wearable device can translate all the text in the image to be processed, obtain the translated text and store it. After determining the target field of view area according to the user's head movement, it is determined whether there is original text in the target field of view area. When there is original text, the translated text corresponding to the original text in the current target field of view is retrieved and displayed directly above the original text in the display interface for the user to view.
[0083] In one embodiment, the smart wearable device relies on an IMU sensor to detect the movement of the user's head and thereby determine the target field of view. For example, when the head moves downward, the gyroscope sensor in the IMU sensor detects that the head is turning downward, and the accelerometer detects the acceleration of the movement. The distance the head turns downward is calculated based on the acceleration, thereby determining the target field of view and obtaining an image in the target field of view to display to the user.
[0084] In one embodiment, the accelerometer in the IMU sensor measures the acceleration on the three coordinate axes of the Cartesian coordinate system, and the gyroscope measures the angular velocity around the three coordinate axes. Through the data of these two sensors, the posture change of the IMU itself, that is, the posture change at the current moment relative to the previous moment, is calculated as the deflection value of the user's head on each coordinate axis.
[0085] In one embodiment, the movement direction and movement distance of the user's head relative to the first field of view at the first moment can be determined based on the deflection value of the user's head on each coordinate axis. For example, the right direction is the positive direction of the horizontal coordinate, and the upward direction is the positive direction of the vertical coordinate. The movement data of the user's head is that the movement distance on the horizontal axis is 3 units, and the movement distance on the vertical axis is 2 units, which means that the user's head is rotating diagonally to the upper right, and the area diagonally to the upper right is used as the user's target field of view. If there is also movement data on the third coordinate axis, it means that the user's head moves in the front-to-back direction, then the content in the target display area is enlarged when the user moves forward, and the content in the target display area is reduced when the user moves backward.
[0086] In one embodiment, Figure 6 and Figure 7 As shown, according to different target fields of vision of the user, the translation texts corresponding to the texts to be translated in different areas (rectangular areas in the figure) are displayed.
[0087] In one embodiment, for the display of the translated text, the font size of the translated text is determined in combination with the number of words in the translated text and the size of the text box of the text to be translated, so that the translated text is just within the text box of the text to be translated, and it can be ensured that the translated text superimposed on the top of the text to be translated is the translated text corresponding to the text to be translated. For example, the text box corresponding to the text to be translated can store the text to be translated consisting of 100 words with a text font size of four, and the translated text corresponding to the text to be translated has 150 words, then the font size of the translated text needs to be reduced (such as five) to achieve that the display area of the translated text does not exceed the text box size of the text to be translated.
[0088] In one embodiment, the translation can be performed directly based on the corrected text to be translated to obtain the translated text, thus saving translation time, and the text to be translated does not need to be displayed in the display interface; the corrected text to be translated can also be displayed in the display interface, and after the user determines that the text to be translated is the same as the original text on the image to be processed, the text to be translated is translated, and the text to be translated is no longer displayed, so as to avoid reading difficulties caused by the overlapping display of the text to be translated and the translated text, so that the user can view the translated text more intuitively.
[0089] The above embodiment provides a text translation method, device, computer equipment and storage medium, which detects the deflection value of the user's head based on a preset sensor; determines the movement direction and movement distance of the user's head relative to the first visual field area at the first moment based on the deflection value, and determines the user's target visual field at the second moment based on the movement direction and the movement distance; when the original text exists in the user's target visual field, obtains the translation text corresponding to the original text, and displays the translation text in a superimposed manner at the position of the original text in the display interface. The method detects the movement of the user's head through a preset sensor, determines the user's target visual field at the second moment, and displays the translation text corresponding to the original text in the user's target visual field to the user, so that the user can quickly and accurately obtain the required translation text, thereby improving translation efficiency.
[0090] See also Figure 8 , Figure 8 The embodiment of the present application provides a schematic block diagram of a text translation device, which is used to execute the above-mentioned text translation method. The text translation device can be configured in a smart wearable device.
[0091] like Figure 8 As shown, the text translation device 400 comprises:
[0092] The module 401 for obtaining text to be translated is used to collect the image to be processed based on the smart wearable device, and perform text processing on the original text in the image to be processed to obtain the text to be translated;
[0093] A translation text obtaining module 402 is used to translate the text to be translated based on preset configuration parameters to obtain a translation text;
[0094] The translation text display module 403 is used to map the translation text based on a preset conversion matrix and display the translation text in a display interface.
[0095] In one embodiment, the translation text display module 403 includes:
[0096] A deflection value obtaining module, used to detect the deflection value of the user's head based on a preset sensor;
[0097] A user target field of view determination module, configured to determine a movement direction and a movement distance of the user's head relative to a first field of view area at a first moment based on the deflection value, and determine a user target field of view at a second moment based on the movement direction and the movement distance;
[0098] The translation text display module is used to obtain the translation text corresponding to the original text when the original text exists in the user's target field of view, and to superimpose and display the translation text at the position of the original text in the display interface.
[0099] In one embodiment, the translation text display module 403 includes:
[0100] A three-dimensional change detection module, used to detect the position change and angle change of the image to be processed in three-dimensional space based on an optical flow algorithm;
[0101] The translated text display module is used to determine the three-dimensional change of the translated text based on the position change and angle change of the image to be processed, and based on the three-dimensional change, superimpose the translated text on the position of the original text in the display interface.
[0102] In one embodiment, the module 401 for obtaining text to be translated includes:
[0103] A text location determination module, used for performing text recognition on the image to be processed, and determining the location of the original text in the image to be processed according to the recognition result;
[0104] A target text area obtaining module, used for cropping the image to be processed based on the position of the original text to obtain a target text area;
[0105] The module for obtaining text to be translated is used to perform text detection and text correction on the target text area to obtain the text to be translated.
[0106] In one embodiment, the module for obtaining text to be translated includes:
[0107] A single-line text obtaining unit, configured to perform text detection on the target text area based on a preset text detection algorithm to obtain at least one single-line text;
[0108] A text-to-be-corrected obtaining unit, configured to divide at least one of the single-line texts into paragraphs based on a paragraph recognition algorithm to obtain at least one text-to-be-corrected;
[0109] The to-be-translated text obtaining unit is used to perform text correction on at least one of the to-be-corrected texts based on a preset word library and a preset grammar library to obtain the to-be-translated text.
[0110] In one embodiment, the translation text display module 403 further includes:
[0111] A paragraph text box obtaining unit is used to obtain the paragraph text box of the text to be corrected;
[0112] a font size determining unit, configured to determine the font size of the translated text based on the size of the paragraph text box and the number of words in the translated text;
[0113] The to-be-translated text display unit is used to display the translated text in the display interface based on the font size.
[0114] In one implementation, the text translation device 400 further includes:
[0115] The to-be-processed image conversion module is used to convert the to-be-processed image based on the conversion matrix so as to display the to-be-processed image in the display interface.
[0116] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device and each module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0117] The above-mentioned device can be implemented in the form of a computer program. Fig. 9 Runs on the computer device shown.
[0118] See also Fig. 9 , Fig. 9 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server.
[0119] See also Fig. 9 The computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0120] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any text translation method.
[0121] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0122] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any text translation method.
[0123] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Fig. 9The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0124] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0125] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:
[0126] Based on the smart wearable device, an image to be processed is collected, and original text in the image to be processed is processed to obtain a text to be translated;
[0127] Based on preset configuration parameters, the text to be translated is translated to obtain a translated text;
[0128] The translated text is mapped based on a preset conversion matrix, and the translated text is displayed in a display interface.
[0129] In one embodiment, when the processor implements mapping the translation text based on a preset conversion matrix and displays the translation text in a display interface, it is configured to implement:
[0130] Based on the preset sensor, the deflection value of the user's head is detected;
[0131] Based on the deflection value, determine the movement direction and movement distance of the user's head relative to the first visual field area at the first moment, and based on the movement direction and the movement distance, determine the user's target visual field at the second moment;
[0132] When the original text exists in the user's target field of view, a translation text corresponding to the original text is obtained, and the translation text is superimposed and displayed at the position of the original text in the display interface.
[0133] In one embodiment, when displaying the translated text in the user's target field of view, the processor is used to implement:
[0134] Based on the optical flow algorithm, detecting the position change and angle change of the image to be processed in the three-dimensional space;
[0135] Based on the position change and angle change of the image to be processed, a three-dimensional change of the translated text is determined, and based on the three-dimensional change, the translated text is superimposed and displayed at the position of the original text in the display interface.
[0136] In one embodiment, when the processor performs text processing on the original text in the image to be processed to obtain the text to be translated, the processor is used to implement:
[0137] Performing text recognition on the image to be processed, and determining the position of the original text in the image to be processed according to the recognition result;
[0138] Based on the position of the original text, the image to be processed is cropped to obtain a target text area;
[0139] Text detection and text correction are performed on the target text area to obtain the text to be translated.
[0140] In one embodiment, the processor implements text processing including text detection and text correction, and when performing text detection and text correction on the target text area to obtain the text to be translated, is used to implement:
[0141] Based on a preset text detection algorithm, perform text detection on the target text area to obtain at least one single line of text;
[0142] Based on a paragraph recognition algorithm, dividing at least one of the single-line texts into paragraphs to obtain at least one text to be corrected;
[0143] Based on a preset word library and a preset grammar library, text correction is performed on at least one of the texts to be corrected to obtain the text to be translated.
[0144] In one embodiment, after performing text detection and text correction on the target text area to obtain the text to be translated, the processor is further configured to implement:
[0145] Get the paragraph text box of the text to be modified;
[0146] Determining the font size of the translated text based on the size of the paragraph text box and the number of words in the translated text;
[0147] Based on the font size, the translated text is displayed in the display interface.
[0148] In one embodiment, after acquiring the image to be processed based on the smart wearable device, the processor is further used to implement:
[0149] Based on the conversion matrix, the image to be processed is converted to display the image to be processed in the display interface.
[0150] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program. The computer program includes program instructions. The processor executes the program instructions to implement any text translation method provided in the embodiment of the present application.
[0151] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.
[0152] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A text translation method, characterized in that: include: Based on the smart wearable device, an image to be processed is collected, and original text in the image to be processed is processed to obtain a text to be translated; Based on preset configuration parameters, the text to be translated is translated to obtain a translated text; The translated text is mapped based on a preset conversion matrix, and the translated text is displayed in a display interface.
2. The text translation method according to claim 1, characterized in that: The step of displaying the translated text in a display interface includes: Based on the preset sensor, the deflection value of the user's head is detected; Based on the deflection value, determine the movement direction and movement distance of the user's head relative to the first visual field area at the first moment, and based on the movement direction and the movement distance, determine the user's target visual field at the second moment; When the original text exists in the user's target field of view, a translation text corresponding to the original text is obtained, and the translation text is superimposed and displayed at the position of the original text in the display interface.
3. The text translation method according to claim 1, characterized in that: The step of displaying the translated text in a display interface further includes: Based on the optical flow algorithm, detecting the position change and angle change of the image to be processed in the three-dimensional space; Based on the position change and angle change of the image to be processed, a three-dimensional change of the translated text is determined, and based on the three-dimensional change, the translated text is superimposed and displayed at the position of the original text in the display interface.
4. The text translation method according to claim 1, characterized in that: The processing of the original text in the image to be processed to obtain the text to be translated includes: Performing text recognition on the image to be processed, and determining the position of the original text in the image to be processed according to the recognition result; Based on the position of the original text, the image to be processed is cropped to obtain a target text area; Text detection and text correction are performed on the target text area to obtain the text to be translated.
5. The text translation method according to claim 4, characterized in that: The performing text detection and text correction on the target text area to obtain the text to be translated includes: Based on a preset text detection algorithm, perform text detection on the target text area to obtain at least one single line of text; Based on a paragraph recognition algorithm, dividing at least one of the single-line texts into paragraphs to obtain at least one text to be corrected; Based on a preset word library and a preset grammar library, text correction is performed on at least one of the texts to be corrected to obtain the text to be translated.
6. The text translation method according to claim 1, characterized in that: The step of displaying the translated text in a display interface further includes: Get the paragraph text box of the text to be modified; Determining the font size of the translated text based on the size of the paragraph text box and the number of words in the translated text; Based on the font size, the translated text is displayed in the display interface.
7. The text translation method according to any one of claims 1 to 6, characterized in that: After collecting the image to be processed based on the smart wearable device, the method further includes: Based on the conversion matrix, the image to be processed is converted to display the image to be processed in the display interface.
8. A text translation device, characterized in that: include: A module for obtaining text to be translated, which is used to collect images to be processed based on a smart wearable device, and perform text processing on the original text in the images to be processed to obtain text to be translated; A translation text obtaining module, used to translate the text to be translated based on preset configuration parameters to obtain a translation text; The translation text display module is used to map the translation text based on a preset conversion matrix and display the translation text in a display interface.
9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the text translation method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the text translation method according to any one of claims 1 to 7.