A translation method and related equipment

By detecting the user's body movements to control the voice data acquisition and translation process of the translation device, the problem of frequent button operations required in existing technologies is solved, thus improving the user experience.

CN114912468BActive Publication Date: 2026-03-10IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-07
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing language translation products require users to frequently press buttons, resulting in a cumbersome translation process and a poor user experience.

Method used

The translation device's voice data acquisition and translation process can be controlled by detecting the user's body movements, such as moving the device to or away from the lips, to simplify the operation process.

Benefits of technology

It reduces button operations, improves user experience, and simplifies the language translation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114912468B_ABST
    Figure CN114912468B_ABST
Patent Text Reader

Abstract

This application discloses a translation method and related equipment. The method includes: when the translation equipment determines that the body movements of a first user meet a first action condition, the translation equipment starts to record the first user's voice to obtain first speech data, so that the first speech data can represent the content of the first user's speech. When the translation equipment determines that the first translation condition has been met, the translation equipment can determine that the first user has finished speaking. Therefore, the translation equipment can acquire and display the translation representation data of the first speech data, so that the first user or the first user's cross-language communication partner can understand the translation result of the first speech data from the translation equipment. This allows the language translation process to be controlled by simple body movements, so that the translation equipment can effectively overcome the adverse effects caused by controlling the language translation process by button operation. This can effectively simplify the language translation process and thus effectively improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a translation method and a related device thereof. BACKGROUND

[0002] With the development of internationalization trend, language translation scenarios (for example, language learning scenarios, cross-language communication scenarios, etc.) are increasingly rich, so that the user group with language translation needs is increasingly large, thereby making the application range of language translation products (for example, terminal devices with language translation functions, application programs with language translation functions, etc.) that can provide translation services for these user groups is increasingly wide.

[0003] However, due to defects of some language translation products, the language translation process based on these language translation products is relatively cumbersome, thereby causing relatively poor user experience. SUMMARY

[0004] The main purpose of the embodiments of the present application is to provide a translation method and a related device, which can effectively simplify the language translation process, thereby effectively improving the user experience.

[0005] The embodiments of the present application provide a translation method, which is applied to a translation device, and the method comprises the following steps:

[0006] The embodiments of the present application provide a translation device, which comprises:

[0007] A first obtaining unit is configured to obtain first voice data when it is determined that a body action of a first user meets a first action condition, wherein the first voice data carries speaking content of the first user.

[0008] A first display unit is configured to obtain and display translation representation data of the first voice data when it is determined that a first translation condition is met, wherein the translation representation data comprises translation text data and / or translation audio data.

[0009] The embodiments of the present application provide a translation device, which comprises a processor, a memory and a system bus. The processor and the memory are connected through the system bus. The memory is configured to store one or more programs, wherein the one or more programs comprise instructions which, when executed by the processor, cause the processor to execute any of the embodiments of the translation method provided by the embodiments of the present application.

[0010] The embodiment of the present application provides a computer readable storage medium, wherein instructions are stored in the computer readable storage medium, and when the instructions run on a terminal device, the terminal device executes any implementation manner of the translation method provided by the embodiment of the present application.

[0011] The embodiment of the present application also provides a computer program product, which, when running on a terminal device, causes the terminal device to execute any implementation manner of the translation method provided by the embodiment of the present application.

[0012] Based on the above technical solutions, the present application has the following beneficial effects:

[0013] For the translation device provided by the present application, the translation device can control the language translation process by means of some simple body movements. Moreover, the implementation process can be as follows: when the translation device determines that the body movement of the first user meets the first action condition (for example, the first user moves the translation device to the lips, etc.), the translation device starts to collect sound for the first user to obtain first voice data, so that the first voice data can represent the speaking content of the first user, so that when the translation device determines that the first translation condition is met (for example, the first user moves the translation device away from the lips, etc.), the translation device can determine that the first user ends speaking, and therefore the translation device can obtain and display the translated representation data (for example, translated text data and / or translated audio data) of the first voice data, so that the first user (or the cross-language communication object of the first user) can understand the translation result for the first voice data from the translation device.

[0014] It can be seen that the translation device provided by the present application can control the language translation process by means of some simple body movements (for example, moving the translation device to the user's lips, moving the translation device away from the user's lips, etc.), so that the user of the translation device does not need to perform too much or too long key operation in the language translation process, thereby enabling the translation device to effectively overcome the adverse effects (for example, complicated operation, key finger twitching, etc.) caused by controlling the language translation process by means of key operation, so as to effectively simplify the language translation process, thereby effectively improving the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0016] Figure 1 A flowchart of a translation method provided for an embodiment of the present application;

[0017] Figure 2 A schematic diagram of a triggering action of a voice translation flow provided for an embodiment of the present application;

[0018] Figure 3 A structural schematic diagram of a translation device provided with a distance sensor for an embodiment of the present application;

[0019] Figure 4 A schematic diagram of a gyroscope provided for an embodiment of the present application;

[0020] Figure 5 A schematic diagram of a triggering action of a language translation flow provided for an embodiment of the present application;

[0021] Figure 6 A schematic diagram of a voice translation flow for a first user provided for an embodiment of the present application;

[0022] Figure 7 A schematic diagram of a display manner of a language translation result displayed by a first user to a second user provided for an embodiment of the present application;

[0023] Figure 8 A schematic diagram of another display manner of a language translation result displayed by a first user to a second user provided for an embodiment of the present application;

[0024] Figure 9 A schematic diagram of still another display manner of a language translation result displayed by a first user to a second user provided for an embodiment of the present application;

[0025] Figure 10 A schematic diagram of a voice translation flow for a second user provided for an embodiment of the present application;

[0026] Figure 11 A structural schematic diagram of a translation device provided for an embodiment of the present application;

[0027] Figure 12 A structural schematic diagram of a translation device provided for an embodiment of the present application;

[0028] Figure 13 A structural schematic diagram of a translation device provided for an embodiment of the present application; DETAILED DESCRIPTION

[0029] The inventors found in the research on language translation products that some language translation products usually control the language translation process by means of key operation (for example, a user can hold a button to speak and release the button to translate; or the user can click a button to speak and click again to translate, etc. control process). For ease of understanding, the following is illustrated by taking a cross-language communication scenario as an example.

[0030] As an example, when A who is good at Chinese and B who is good at English communicate by means of a language translation product, A needs to hold a “Chinese key” on the language translation product and speak first, so that A releases the “Chinese key” when A finishes speaking, so that the language translation product presents an English translation result for the content A has spoken; then, A hands over the language translation product to B, so that B can view the English translation result on the language translation product; subsequently, when B holds an “English key” on the language translation product, B speaks, so that B releases the “English key” when it is determined that B finishes speaking, so that the language translation product presents a Chinese translation result for the content B has spoken, so that A can know what B has said on the language translation product, thus realizing a round of cross-language communication process between A and B.

[0031] Based on the above example, it can be seen that for some language translation products, because these language translation products usually control the language translation process by means of key operation, the user of the language translation product not only needs to speak, but also needs to constantly perform key operation in a round of cross-language communication process (or, in a language learning process), so that the language translation process based on these language translation products is relatively cumbersome, and the user of these language translation products is prone to physical discomfort (for example, a finger is prone to twitch when continuously used for key operation), so that the user experience is relatively poor.

[0032] Based on the above finding, in order to solve the technical problem shown in the background section, the embodiments of the present application provide a translation method, which comprises: when a translation device determines that the limb action of a first user meets a first action condition (for example, the first user moves the translation device to the lip, etc.), the translation device starts to collect sound for the first user, to obtain first speech data, so that the first speech data can represent the speaking content of the first user, so that when the translation device determines that a first translation condition is met (for example, the first user moves the translation device away from the lip, etc.), the translation device can determine that the first user finishes speaking, so that the translation device can obtain and display the translation representation data (for example, translation text data and / or translation audio data) of the first speech data, so that the first user (or the cross-language communication object of the first user) can understand the translation result for the first speech data from the translation device.

[0033] It can be seen that, by means of the translation device provided in the present application, the language translation process can be controlled by means of some simple body movements (for example, moving the translation device to the user's lips, moving the translation device away from the user's lips, etc.), so that the user of the translation device does not need to perform excessive or long key operations during the language translation process, thereby enabling the translation device to effectively overcome the adverse effects (for example, complicated operation, key finger twitching, etc.) caused by controlling the language translation process by means of key operations, thus effectively simplifying the language translation process, thereby effectively improving the user experience.

[0034] In addition, the present application embodiment does not limit the execution subject (that is, the translation device) of the translation method, for example, the translation method provided in the present application embodiment can be applied to any terminal device capable of realizing language translation function. Wherein, the terminal device can be a translation machine, a dictionary pen, a smart phone, a computer, a personal digital assistant (Personal Digital Assistant, PDA) or a tablet computer, etc.

[0035] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0036] Method embodiment one

[0037] Referring to Figure 1 , the figure is a flow chart of a translation method provided in the present application embodiment.

[0038] The translation method applied to the translation device provided in the present application embodiment comprises S1-S2:

[0039] S1: When the translation device determines that the body movement of the first user meets the first action condition, the translation device acquires the first voice data.

[0040] Wherein, the translation device refers to an electronic device with language translation function; and the present application embodiment does not limit the translation device, for example, it can be a translation special-purpose device similar to a translation machine, or a terminal device (for example, a smart phone, etc.) installed with an application program (Application, App) for realizing language translation function.

[0041] The first user is used to represent the user (for example, the holder) of the translation device, so that the first user has the right to use the translation device.

[0042] The above "limb action of the first user" is used to represent a body state (e.g., moving a mobile phone to the lips, etc.) of the first user.

[0043] The above "first action condition" is used to describe which limb actions the first user usually performs when the first user has a speech translation requirement, so that the first action condition can trigger a speech translation process (e.g., roughly going through a process of collecting speech -> speech recognition -> language translation, etc.) for the content spoken by the first user; and embodiments of the present application do not limit the first action condition, and for the convenience of understanding, the following will be described in combination with three cases.

[0044] Case 1, for the first user, when the first user wants to use the translation device to perform speech translation processing for the content spoken by the first user, the first user will usually subconsciously shorten the distance between the mouth and the translation device (e.g., as shown in Figure 2 , the first user may subconsciously move the translation device to the lips, etc.), so as to ensure that the content spoken by the first user can be collected as clearly as possible by the translation device.

[0045] Based on the above case 1, in order to better utilize the user behavior commonality shown in the last paragraph, embodiments of the present application provide a possible implementation of the above "first action condition", which can be specifically: the distance between the body articulatory part (e.g., head, face, mouth, lips, etc.) of the first user and the translation device satisfies a first distance condition (e.g., ≤10 cm). Wherein, the first distance condition can be pre-set.

[0046] In addition, in order to better measure the "distance between the body articulatory part of the first user and the translation device", a distance sensor can be deployed on the preset position of the translation device (e.g., as shown in Figure 3 , the distance sensor can be deployed on the front top area of the screen). Based on this, embodiments of the present application provide a possible implementation of the above "first action condition", in which implementation, when the translation device includes a distance sensor, the first action condition can be specifically: the distance description information collected by the distance sensor in the translation device satisfies a first information condition (e.g., the distance description information indicates that the distance between the distance sensor and the face of the first user is less than or equal to 10 cm).

[0047] Case 2, for the first user, when the first user wants to use the translation device to perform speech translation processing for the content spoken by the first user, as Figure 2As shown, the first user will usually subconsciously pick up the horizontally placed translation device (e.g., a translation device placed horizontally on a table) to make the translation device as vertical as possible, so that the translation device can better pick up the content spoken by the first user.

[0048] Based on the above situation 2, in order to better utilize the commonality of user behavior shown in the previous paragraph, this application embodiment provides a possible implementation of the "first action condition" mentioned above, which can be: the angle between the translation device and the reference horizontal plane (e.g., the ground, the desktop, etc.) satisfies the first angle condition (e.g., between 90°-α and 90°+α). The first angle condition can be preset.

[0049] It should be noted that the above α refers to a preset angle deviation value; and the embodiments of this application do not limit the α, for example, α can be 30°.

[0050] In addition, to better measure the angle between the translation device and the reference horizontal plane, a gyroscope can be deployed inside the translation device (e.g., Figure 4 (The gyroscope shown). Based on this, the embodiments of this application also provide a possible implementation of the "first operating condition" above. In this implementation, when the translation device includes a gyroscope, the first operating condition can specifically be: the angle description information collected by the gyroscope satisfies the second information condition (for example, the angle in the Y-axis direction of the angle description information can indicate that the angle between the translation device and the reference horizontal plane is between 90°-α and 90°+α).

[0051] Case 3: In order to better identify the user's translation needs, referring to the commonalities of user behavior shown in the two cases above, this application provides a possible implementation of the "first action condition" above, which can be: the distance between the first user's body part for vocalization and the translation device meets the first distance condition, or the angle between the translation device and the reference horizontal plane meets the first angle condition.

[0052] In other words, for the translation device, as long as it determines that the distance between the first user's body part for vocalization and the translation device satisfies the first distance condition, it can determine that the first user's body movement satisfies the first movement condition; moreover, as long as it determines that the angle between the translation device and the reference horizontal plane satisfies the first angle condition, it can determine that the first user's body movement satisfies the first movement condition.

[0053] It can be seen that, for the translation device, when the translation device comprises a distance sensor and a gyroscope, if it is determined that the distance description information collected by the distance sensor in the translation device satisfies the first information condition, the translation device can determine that the limb action of the first user satisfies the first action condition; and if it is determined that the angle description information collected by the gyroscope in the translation device satisfies the second information condition, the translation device can also determine that the limb action of the first user satisfies the first action condition.

[0054] Based on the related content of the above first action condition, it can be known that, for the first action condition, the first action condition can describe which limb actions the first user will usually perform when the first user has a speech translation requirement, so that the first action condition can be used to identify whether the first user has a speech translation requirement, so as to directly trigger the speech translation processing process for the content spoken by the first user based on the user's limb action, without the user performing additional key operations, so as to facilitate simplifying the triggering process of the speech translation processing process.

[0055] The above "first speech data" refers to the sound collection result of the content spoken by the first user, so that the first speech data carries the speaking content of the first user (for example, the speaking content "Hello, I want to ask, how to get to the Ocean Park from here?" and the like).

[0056] In addition, the embodiments of the present application do not limit the acquisition process of the first speech data, for example, it can be specifically: when the translation device determines that the limb action of the first user satisfies the first action condition, the translation device can start the sound collection mode, so that the translation device can collect the pronunciation data of the first user in real time to obtain the first speech data, so that the first speech data can represent the content spoken by the first user.

[0057] In addition, the embodiments of the present application do not limit the implementation of S1, in order to facilitate understanding, the following three implementation modes are described.

[0058] In the first possible implementation, if only the distance between the first user and the translation device is referred to to determine whether to trigger the speech translation processing process, S1 can be specifically: when it is determined that the distance between the body pronunciation part of the first user and the translation device satisfies the first distance condition, the first speech data is acquired.

[0059] It can be seen that, if only the distance between the first user and the translation device is referred to to determine whether to trigger the voice translation processing procedure, when it is determined that the distance between the body pronunciation part of the first user and the translation device meets the first distance condition, the translation device can determine that the limb action of the first user meets the first action condition, so that the translation device can further determine that the first user has a voice translation requirement, and therefore the translation device can start to collect pronunciation data of the first user to obtain first voice data, so that the first voice data can represent the speaking content of the first user, so as to subsequently perform language translation processing on the user semantics carried by the first voice data.

[0060] In addition, the embodiment of the present application does not limit the determination method of the distance between the body pronunciation part of the first user and the translation device, for example, a distance sensor can be used for implementation. Based on this, when the translation device includes a distance sensor, S1 can be specifically: when it is determined that the distance description information collected by the distance sensor in the translation device meets the first information condition, the first voice data is obtained.

[0061] It can be seen that, for the translation device deployed with a distance sensor, if only the distance between the first user and the translation device is referred to to determine whether to trigger the voice translation processing procedure, when it is determined that the distance description information collected by the distance sensor in the translation device meets the first information condition, the translation device can determine that the distance between the body pronunciation part of the first user and the translation device meets the first distance condition, so that the translation device can determine that the limb action of the first user meets the first action condition, so that the translation device can further determine that the first user has a voice translation requirement, and therefore the translation device can start to collect pronunciation data of the first user to obtain first voice data, so that the first voice data can represent the speaking content of the first user, so as to subsequently perform language translation processing on the user semantics carried by the first voice data.

[0062] In the second possible implementation, if only the angle between the translation device and the reference horizontal plane (for example, the ground, etc.) is referred to to determine whether to trigger the voice translation processing procedure, S1 can be specifically: when it is determined that the angle between the translation device and the reference horizontal plane meets the first angle condition, the first voice data is obtained.

[0063] It can be seen that if whether to trigger the speech translation processing procedure is determined only by referring to the included angle between the translation device and the reference horizontal plane (for example, the ground, etc.), when it is determined that the included angle between the translation device and the reference horizontal plane satisfies the first angle condition, the translation device can determine that the limb action of the first user satisfies the first action condition, so that the translation device can further determine that the first user has a speech translation demand, and thus the translation device can start to collect the pronunciation data of the first user to obtain the first speech data, so that the first speech data can represent the speaking content of the first user, so as to subsequently perform language translation processing on the user semantics carried by the first speech data.

[0064] In addition, the embodiments of the present application do not limit the determination method of the "included angle between the translation device and the reference horizontal plane", for example, a gyroscope can be used for implementation. Based on this, when the translation device includes a gyroscope, S1 can be specifically: when it is determined that the angle description information collected by the gyroscope in the translation device satisfies the second information condition, the first speech data is obtained.

[0065] It can be seen that for the translation device deployed with a gyroscope, if whether to trigger the speech translation processing procedure is determined only by referring to the included angle between the translation device and the reference horizontal plane (for example, the ground, etc.), when it is determined that the angle description information collected by the gyroscope in the translation device satisfies the second information condition, the translation device can determine that the included angle between the translation device and the reference horizontal plane satisfies the first angle condition, so that the translation device can determine that the limb action of the first user satisfies the first action condition, so that the translation device can further determine that the first user has a speech translation demand, and thus the translation device can start to collect the pronunciation data of the first user to obtain the first speech data, so that the first speech data can represent the speaking content of the first user, so as to subsequently perform language translation processing on the user semantics carried by the first speech data.

[0066] In a third possible implementation, if whether to trigger the speech translation processing procedure is determined by comprehensively referring to the distance between the first user and the translation device and the included angle between the translation device and the reference horizontal plane, S1 can be specifically: when it is determined that the distance between the body pronunciation part of the first user and the translation device satisfies the first distance condition, or the included angle between the translation device and the reference horizontal plane satisfies the first angle condition, the first speech data is obtained. In order to facilitate understanding, the following is described by taking an example.

[0067] As an example, when the translation device includes a distance sensor and a gyroscope, S1 can be specifically: when it is determined that the distance description information collected by the distance sensor satisfies the first information condition, or the angle description information collected by the gyroscope satisfies the second information condition, the first speech data is obtained.

[0068] It can be seen that, for the translation device deployed with the distance sensor and the gyroscope, the translation device can comprehensively refer to the distance description information collected by the distance sensor and the angle description information collected by the gyroscope to determine whether to trigger the voice translation processing procedure, and the determination process can be specifically as follows:

[0069] When it is determined that the distance description information collected by the distance sensor in the translation device satisfies the first information condition, the translation device can determine that the distance between the body pronunciation part of the first user and the translation device satisfies the first distance condition, so that the translation device can determine that the limb action of the first user satisfies the first action condition, so that the translation device can further determine that the first user has a voice translation demand, and therefore the translation device can start collecting pronunciation data of the first user to obtain first voice data, so that the first voice data can represent the speaking content of the first user, so as to subsequently perform language translation processing on the user semantics carried by the first voice data.

[0070] When it is determined that the angle description information collected by the gyroscope in the translation device satisfies the second information condition, the translation device can determine that the included angle between the translation device and the reference horizontal plane satisfies the first angle condition, so that the translation device can also determine that the limb action of the first user satisfies the first action condition, so that the translation device can further determine that the first user has a voice translation demand, and therefore the translation device can start collecting pronunciation data of the first user to obtain first voice data, so that the first voice data can represent the speaking content of the first user, so as to subsequently perform language translation processing on the user semantics carried by the first voice data.

[0071] When it is determined that the distance description information collected by the distance sensor in the translation device satisfies the first information condition, and it is determined that the angle description information collected by the gyroscope in the translation device satisfies the second information condition, the translation device can also determine that the limb action of the first user satisfies the first action condition, so that the translation device can further determine that the first user has a voice translation demand, and therefore the translation device can start collecting pronunciation data of the first user to obtain first voice data, so that the first voice data can represent the speaking content of the first user, so as to subsequently perform language translation processing on the user semantics carried by the first voice data.

[0072] Based on the related content of S1, for the first user holding the translation device, if the first user wants to use the translation device to realize the language translation demand (for example, language learning demand or cross-language communication demand, etc.), the first user can perform some limb actions that can meet the first action condition (for example, picking up the translation device placed on the horizontal desktop and moving to the lip, etc.), so that the limb action of the first user can represent that the first user has a language translation demand, and therefore the translation device can start real-time sound collection for the first user to obtain the first voice data, so that the first voice data can represent the speaking content of the first user, so as to subsequently perform language translation processing on the user semantics carried by the first voice data.

[0073] S2: When the translation device determines that the first translation condition is reached, the translation device obtains and displays the translation representation data of the first voice data.

[0074] Wherein, the first translation condition is used to trigger the language translation process for the content spoken by the first user; and the present application embodiment does not limit the first translation condition, for example, in order to improve the voice translation effect as much as possible, the first translation condition can be set according to the speaking end time of the first user, so that the triggering time of the language translation process for the content spoken by the first user is equal to or slightly later than the speaking end time of the first user. In order to facilitate understanding, the following will be explained in combination with three cases.

[0075] Case 1, for the first user, when the first user finishes recording the content spoken by the first user using the translation device, the first user will subconsciously move the translation device away from the lip (for example, as shown in the figure, the translation device is placed flat, so that the eyes of the first user can see the screen of the translation device, etc.), so that the first user can see the voice recognition result of the first voice data displayed on the translation device, to confirm whether the translation device correctly records the content spoken by the first user. Figure 5

[0076] It can be seen that in some cases, the speaking end of the first user can be determined by means of the limb action of the first user. Based on this, the present application embodiment provides a possible implementation of the above "first translation condition", which can be that the limb action of the first user meets the second action condition. That is, when the translation device determines that the limb action of the first user meets the second action condition, the translation device can determine that the first user has finished speaking, so that the translation device can determine that the first translation condition is reached, and therefore the translation device can obtain the language translation result for the voice information carried by the first voice data.

[0077] ​The second action condition is used to describe which body action the first user will perform after finishing speaking, so that the first action condition can trigger the language translation process for the content spoken by the first user. The second action condition is not limited in the embodiments of the present application. For the convenience of understanding, three examples are described below.

[0078] Example 1, for the first user, after the first user finishes recording the content spoken by the first user using the translation device, the first user will subconsciously pull the distance between the lips and the translation device to ensure that the translation device can not record the first user as much as possible.

[0079] Based on this, in order to better utilize the user behavior commonality shown in the previous paragraph, the embodiments of the present application provide a possible implementation of the second action condition described above, which can be that the distance between the body articulatory part of the first user and the translation device satisfies the second distance condition (for example, greater than 10 cm).

[0080] In fact, the measurement of the distance between the body articulatory part of the first user and the translation device can be realized by deploying a distance sensor on the translation device. Therefore, the embodiments of the present application also provide a possible implementation of the second action condition described above. In this implementation, when the translation device includes a distance sensor, the second action condition can be that the distance description information collected by the distance sensor satisfies the third information condition (for example, the distance description information indicates that the distance between the distance sensor and the face of the first user is greater than 10 cm).

[0081] It can be seen that for the translation device with a distance sensor, if it is determined that the distance description information collected by the distance sensor satisfies the third information condition, the translation device can determine that the distance between the body articulatory part of the first user and the translation device satisfies the second distance condition. Therefore, the translation device can determine that the first user actively pulls away the distance between the first user and the translation device, so that the translation device can determine that the body action of the first user satisfies the second action condition, and further determine that the first user has finished speaking. Therefore, the translation device can determine that the first translation condition has been reached, so the translation device can obtain the language translation result for the speech information carried by the first speech data.

[0082] Example 2, for the first user, after the first user finishes recording the content spoken by the first user using the translation device, the first user will subconsciously lay down the translation device in a vertical state (for example, the translation device is placed on a table or the ground) Figure 5The first user can clearly see the content (e.g., the speech recognition result for the first speech data, etc.) displayed on the display screen of the translation device, so that the first user can check whether the translation device correctly records the content said by the first user.

[0083] Based on this, in order to better utilize the user behavior commonality shown in the above paragraph, the embodiment of the present application provides a possible implementation of the "second action condition" in the above, which can be specifically: the angle between the translation device and the reference horizontal plane (e.g., the ground, etc.) satisfies the second angle condition (e.g., between 0° and 45°).

[0084] In fact, the measurement process for the "angle between the translation device and the reference horizontal plane" can be realized by deploying a gyroscope inside the translation device, so the embodiment of the present application also provides a possible implementation of the "second action condition" in the above, in which when the translation device includes a gyroscope, the second action condition can be specifically: the angle description information collected by the gyroscope satisfies the fourth information condition (e.g., the Y-axis direction angle in the angle description information can represent that the angle between the translation device and the reference horizontal plane is between 0° and 45°).

[0085] It can be seen that for the translation device deployed with a gyroscope, if it is determined that the angle description information collected by the gyroscope satisfies the fourth information condition, the translation device can determine that the angle between it and the reference horizontal plane (e.g., the ground, etc.) satisfies the second angle condition, so that the translation device can determine that the first user has performed the flattening operation on the translation device originally in the vertical state, so that the translation device can determine that the limb action of the first user satisfies the second action condition, and then the translation device can determine that the first user has finished speaking, so that the translation device can determine that the first translation condition has been reached, and therefore the translation device can obtain the language translation result for the speech information carried by the first speech data.

[0086] Example 3, for the first user, after the first user records the content said by the first user using the translation device, the first user may, subconsciously, pull the distance between the lips and the translation device, or may, subconsciously, flatten the translation device originally in the vertical state, or may perform both operations of pulling the distance between the lips and the translation device and flattening the translation device originally in the vertical state.

[0087] Based on this, in order to better utilize the user behavior commonality shown in the previous paragraph, the embodiments of the present application provide a possible implementation of the above "second action condition", which can be specifically: the distance between the body articulation part of the first user and the translation device satisfies a second distance condition, or the included angle between the translation device and the reference horizontal plane satisfies a second angle condition.

[0088] In fact, for the translation device deployed with the distance sensor and the gyroscope, the embodiments of the present application provide a possible implementation of the above "second action condition", which can be specifically: the distance description information collected by the distance sensor satisfies a third information condition, or the angle description information collected by the gyroscope satisfies a fourth information condition.

[0089] It can be seen that for the translation device deployed with the distance sensor and the gyroscope, the determination process of the translation device for the second action condition is as follows:

[0090] If it is determined that the distance description information collected by the distance sensor satisfies the third information condition, the translation device can determine that the distance between the body articulation part of the first user and the translation device satisfies the second distance condition, so that the translation device can determine that the first user actively pulls away the distance between the first user and the translation device, so that the translation device can determine that the body action of the first user satisfies the second action condition, and further so that the translation device can determine that the first user has finished speaking, so that the translation device can determine that the first translation condition has been reached, and therefore the translation device can obtain the language translation result for the speech information carried by the first speech data;

[0091] If it is determined that the angle description information collected by the gyroscope satisfies the fourth information condition, the translation device can determine that the included angle between the translation device and the reference horizontal plane (for example, the ground, etc.) satisfies the second angle condition, so that the translation device can determine that the first user has performed a flat operation on the translation device originally in a vertical state, so that the translation device can determine that the body action of the first user satisfies the second action condition, and further so that the translation device can determine that the first user has finished speaking, so that the translation device can determine that the first translation condition has been reached, and therefore the translation device can obtain the language translation result for the speech information carried by the first speech data;

[0092] If it is determined that the distance description information collected by the distance sensor satisfies the third information condition and the angle description information collected by the gyroscope satisfies the fourth information condition, the translation device can determine that the body movement of the first user satisfies the second movement condition, so that the translation device can determine that the first user has finished speaking, and thus the translation device can determine that the first translation condition is reached, and the translation device can obtain the language translation result of the speech information carried by the first speech data.

[0093] Based on the related content of the second movement condition, for the translation device in the real-time sound collection state for the first user, once the translation device determines that the body movement of the first user satisfies the second movement condition, the translation device can infer that the first user has finished speaking, so that the translation device can determine that the first translation condition is reached at this time, and the translation device can obtain the language translation result of the speech information carried by the first speech data.

[0094] Case 2, for the first user, if the first user has used the translation device for many times, the first user has believed that the translation device can accurately collect the content spoken by the first user, so that the first user usually does not need to check whether the translation device can accurately collect the content spoken by the first user, and thus when the first user finishes collecting the content spoken by the first user by using the translation device, the first user is very likely to not perform any operation (for example, an operation of checking the sound collection result by placing the translation device flat).

[0095] It can be seen that in some cases, it is possible that the body movement of the first user cannot be used to determine whether the first user has finished speaking, and thus a traditional speaking-finish-determination method can be directly used to determine whether the first user has finished speaking. Based on this, the present embodiment provides a possible implementation manner of the first translation condition, which can be that the duration of the silence state of the first user reaches a preset duration threshold (for example, 60 seconds).

[0096] That is, for the translation device in the real-time sound collection state for the first user, when the translation device determines that the duration of the silence state of the first user reaches the preset duration threshold, the translation device can determine that the first user has finished speaking, so that the translation device can determine that the first translation condition is reached, and the translation device can obtain the language translation result of the speech information carried by the first speech data.

[0097] It should be noted that the embodiments of the present application do not limit the above-mentioned method for identifying that the first user is in the silence state, for example, any method capable of identifying the user voice state (for example, Voice Activity Detection (VAD) and the like) can be used to implement the embodiments of the present application.

[0098] In case 3, in order to further improve the identification effect of the speaking end moment of the first user, the two cases described above can be combined to determine whether the first user ends speaking. Based on this, the embodiments of the present application provide a possible implementation of the above-mentioned "first translation condition", which can be specifically: the body movement of the first user meets the second movement condition, or the duration of the silence state of the first user reaches a preset duration threshold.

[0099] It can be seen that for the translation device in the real-time receiving state for the first user, the judgment process of the translation device for the first translation condition is as follows:

[0100] If it is determined that the body movement of the first user meets the second movement condition, the translation device can infer that the first user has ended speaking, so that the translation device can determine that the first translation condition has been reached at this time, and the translation device can obtain the language translation result of the voice information carried by the first voice data;

[0101] If it is determined that the duration of the silence state of the first user reaches a preset duration threshold, the translation device can infer that the first user has ended speaking, so that the translation device can determine that the first translation condition has been reached at this time, and the translation device can obtain the language translation result of the voice information carried by the first voice data;

[0102] If it is determined that not only the body movement of the first user meets the second movement condition, but also the duration of the silence state of the first user reaches a preset duration threshold, the translation device can infer that the first user has ended speaking, so that the translation device can determine that the first translation condition has been reached at this time, and the translation device can obtain the language translation result of the voice information carried by the first voice data.

[0103] The above-mentioned "translation representation data of the first voice data" is used to represent the translation result of the voice information carried by the first voice data; and the embodiments of the present application do not limit the "translation representation data of the first voice data", for example, it can include at least one of the translation text data of the first voice data and the translation audio data of the first voice data.

[0104] The "translated text data of the first voice data" is used to represent the translation result of the voice information carried by the first voice data in text form. The embodiment of the present application does not limit the obtaining process of the "translated text data of the first voice data". For example, the "translated text data of the first voice data" can be obtained by directly performing language translation processing on the voice recognition text of the first voice data by the translation device. For another example, the obtaining process of the "translated text data of the first voice data" can also be that the translation device sends the voice recognition text of the first voice data to the server, so that the server can perform language translation processing on the voice recognition text to obtain the translated text data of the first voice data, and the server feeds back the translated text data to the translation device, so that the translation device can use and / or display the translated text data.

[0105] The "voice recognition text of the first voice data" is used to represent the voice recognition result of the first voice data. The embodiment of the present application does not limit the obtaining process of the "voice recognition text of the first voice data". For example, the "voice recognition text of the first voice data" can be obtained by directly performing voice recognition processing on the first voice data by the translation device. For another example, the obtaining process of the "voice recognition text of the first voice data" can also be that the translation device sends the first voice data to the server, so that the server can perform voice recognition processing on the first voice data to obtain the voice recognition text of the first voice data, and the server feeds back the voice recognition text to the translation device, so that the translation device can use and / or display the voice recognition text.

[0106] It should be noted that the embodiment of the present application does not limit the implementation of the "sending the first voice data to the server" by the translation device. For example, any voice data sending method (for example, edge collection and edge sending) that can be performed between the terminal device and the server can be used for implementation. In addition, the embodiment of the present application does not limit the implementation of the "voice recognition processing on the first voice data". For example, any voice recognition processing method that can be performed on voice data can be used for implementation. In addition, the embodiment of the present application does not limit the implementation of the "language translation processing on the voice recognition text of the first voice data". For example, any language translation processing method that can be performed on text data can be used for implementation.

[0107] The "translated audio data of the first voice data" is used to represent the translation result of the voice information carried by the first voice data in the form of audio; and the embodiment of the present application does not limit the obtaining process of the "translated audio data of the first voice data", for example, it can be specifically: after the translation device obtains the translated text data of the first voice data, the translation device can directly perform audio synthesis processing on the translated text data to obtain the translated audio data of the first voice data, so that the semantic information carried by the translated audio data is the same as the semantic information expressed by the translated text data. For another example, it can also be specifically: after the translation device obtains the translated text data of the first voice data, the translation device can send the translated text data to the server, so that the server can perform audio synthesis processing on the translated text data to obtain the translated audio data of the first voice data, and the server feeds back the translated audio data to the translation device, so that the translation device can use and / or play the translated audio data.

[0108] It should be noted that the embodiment of the present application does not limit the implementation of the above "audio synthesis processing on the translated text data", for example, any method (for example, Text To Speech (TTS) and the like) that can convert text data into audio data can be used to implement it.

[0109] In addition, the embodiment of the present application does not limit the implementation of the above step "the translation device obtains and displays the translated representation data of the first voice data", for example, it can be implemented by the way of obtaining and displaying at the same time to improve the efficiency of voice translation.

[0110] In addition, the embodiment of the present application also does not limit the implementation of the above S2, in order to facilitate understanding, the following three implementation modes are described.

[0111] In the first possible implementation, the language translation processing flow can be triggered by means of the body action of the first user, so that S2 can be specifically: when it is determined that the body action of the first user meets the second action condition, the translation device obtains and displays the translated representation data of the first voice data.

[0112] It can be seen that, for the translation device in the real-time sound collecting state for the first user, if the translation device detects that the body movement of the first user meets the second movement condition, the translation device can determine that the first user has finished speaking, so that the translation device can determine that the first translation condition is reached, and thus the translation device can acquire and display the translation representation data of the first voice data, so that the first user (and / or the cross-language communication object of the first user, etc.) can view (and / or listen to) the language translation result of the content said by the first user on the translation device in the shortest possible time.

[0113] In addition, the embodiments of the present application do not limit the detection manner of the “body movement of the first user meeting the second movement condition”, for example, it can be specifically that: when it is determined that the distance between the body pronunciation part of the first user and the translation device meets the second distance condition, or it is determined that the included angle between the translation device and the reference horizontal plane meets the second angle condition, the translation device acquires and displays the translation representation data of the first voice data.

[0114] It can be seen that, for the translation device in the real-time sound collecting state for the first user, once the translation device can determine that the distance between the body pronunciation part of the first user and the translation device meets the second distance condition (for example, the distance description information collected by the distance sensor in the translation device meets the third information condition, etc.), the translation device can determine that the body movement of the first user meets the second movement condition, so that the translation device can determine that the first user has finished speaking, so that the translation device can determine that the first translation condition is reached, and thus the translation device can acquire and display the translation representation data of the first voice data, so that the first user (and / or the cross-language communication object of the first user, etc.) can view (and / or listen to) the language translation result of the content said by the first user on the translation device in the shortest possible time; in addition, once the translation device can determine that the included angle between the translation device and the reference horizontal plane meets the second angle condition (for example, the angle description information collected by the gyroscope in the translation device meets the fourth information condition, etc.), the translation device can determine that the body movement of the first user meets the second movement condition, so that the translation device can determine that the first user has finished speaking, so that the translation device can determine that the first translation condition is reached, and thus the translation device can acquire and display the translation representation data of the first voice data, so that the first user (and / or the cross-language communication object of the first user, etc.) can view (and / or listen to) the language translation result of the content said by the first user on the translation device in the shortest possible time.

[0115] Based on the related content of S2, for the translation device in the real-time sound collecting state for the first user, if the translation device determines that the first translation condition is met, the translation device can obtain and display the translated representation data of the first voice data, so that the first user (and / or the cross-language communication object of the first user, etc.) can view (and / or listen to) the language translation result of the content said by the first user on the translation device in the shortest possible time.

[0116] Based on the related content of S1 to S2, for the translation method provided by the embodiments of the present application, when the translation device determines that the first user's body action meets the first action condition (for example, the first user moves the translation device to the lips, etc.), the translation device starts to collect sound for the first user to obtain the first voice data, so that the first voice data can represent the content of the first user's speech, so that when the translation device determines that the first translation condition is met (for example, the first user moves the translation device away from the lips, etc.), the translation device can determine that the first user has finished speaking, and therefore the translation device can obtain and display the translated representation data (for example, translated text data and / or translated audio data) of the first voice data, so that the first user (or the cross-language communication object of the first user) can understand the translation result for the first voice data from the translation device.

[0117] It can be seen that the translation device provided by the present application can control the language translation process by means of some simple body actions (for example, moving the translation device to the user's lips, moving the translation device away from the user's lips, etc.), so that the user of the translation device does not need to perform excessive or long key operations during the language translation process, thereby enabling the translation device to effectively overcome the adverse effects caused by controlling the language translation process by means of key operations (for example, complicated operation, key finger twitch, etc.), which can effectively simplify the language translation process and thereby effectively improve the user experience.

[0118] Method embodiment two

[0119] In fact, in order to further improve the user experience, some prompts (for example, text prompts, voice prompts, or vibration prompts, etc.) can be given to the first user when the translation device starts to collect sound for the first user.

[0120] Based on this, the embodiments of the present application also provide another possible implementation of S1, which can be specifically: when the translation device determines that the first user's body action meets the first action condition, the translation device sends a first prompt to the first user and obtains the first voice data.

[0121] The first prompt is used to prompt that the first user can start speaking. The first prompt is not limited in the embodiment of the application, for example, the first prompt can include at least one of prompt text (for example, "I am listening, please speak Chinese" as shown in the figure), prompt voice (for example, audio data carrying voice information "I am listening, please speak Chinese" as shown in the figure), and prompt vibration. Figure 6 Figure 6

[0122] It can be seen that for the translation device, when the translation device detects that the limb action of the first user meets the first action condition, the translation device can determine that the first user has a language translation requirement, so the translation device not only sends the first prompt to the first user to make the first user know from the first prompt that he or she can start speaking, but also collects the first voice data in real time for the first user to make the first voice data represent the speaking content of the first user, so that subsequent language translation processing can be performed on the user semantics carried by the first voice data, which can effectively avoid the adverse effects caused by the first user not knowing when to start speaking (for example, wasting of audio resources caused by starting to speak too late), thereby improving the user experience.

[0123] Method embodiment three

[0124] In fact, the translation device provided by the embodiment of the application is not only suitable for the application scenario of single person using the translation device for language learning, but also suitable for the scenario of at least two people using the translation device for cross-language communication. In order to facilitate understanding, the translation process in each scenario is described below.

[0125] Scenario one: single person uses the translation device for language learning.

[0126] In scenario one, each user only wants to query the language translation result of some sentences by using the translation device, so that the translation device only needs to directly feed back the language translation result of the content spoken by the user. Based on this, the translation method suitable for scenario one provided by the embodiment of the application can specifically include steps 11-12:

[0127] Step 11: When the translation device determines that the limb action of the first user meets the first action condition, the translation device acquires the first voice data, so that the first voice data can represent the sentence content that the first user wants to learn when learning a language.

[0128] It should be noted that the related content of step 11 can be referred to the related content of S1 in the above, and for the sake of brevity, it will not be repeated here.

[0129] ​​Step 12: When the translation device determines that the first translation condition is reached, the translation device acquires and displays the translated representation data of the first voice data, so that the translated representation data can represent the language translation result of the sentence content that the first user wants to learn in language learning.

[0130] It should be noted that the related content of step 12 can be referred to the related content of S2, which will not be repeated here for brevity.

[0131] Based on the above related content of steps 11-12, for the first user with language learning needs, if the first user wants to learn the language translation result of some sentence content using the translation device, the first user can first trigger the translation device to collect the content of the first user by some simple body movements (such as picking up the translation device placed on the horizontal desktop and moving it to the lips, etc.) that can meet the first action condition, so that the translation device can acquire the first voice data, so that the first voice data can represent the sentence content that the first user wants to learn, so that when the translation device determines that the first translation condition is reached, the translation device acquires and displays the translated representation data of the first voice data, so that the translated representation data can represent the language translation result of the sentence content that the first user wants to learn in language learning, so that the first user can learn the translation result of the first voice data from the translation device. Thus, it can be achieved by triggering the voice translation process with simple body movements, so as to effectively improve the voice translation effect, and further improve the user experience.

[0132] Scenario two: at least two people communicate across languages by using the translation device.

[0133] In scenario two, a user not only needs to query the language translation result of some sentence by using the translation device, but also needs to display the language translation result to the cross-language communication object (for example, the second user shown below) of the user, so that the cross-language communication object can know what the user wants to say from the language translation result.

[0134] Based on this, the translation method provided by the embodiments of the present application suitable for scenario one can specifically include steps 21-22:

[0135] Step 21: When the translation device determines that the body movement of the first user meets the first action condition, the translation device acquires the first voice data, so that the first voice data can represent the sentence content that the first user wants to express when communicating across languages with at least one person (for example, Figure 6 The "Hello, I want to ask, how to get to the Ocean Park from here?" shown in the figure).

[0136] It should be noted that the relevant content of step 21 can be found in the relevant content of S1 above, and will not be repeated here for the sake of brevity.

[0137] Step 22: When the translation device determines that the first translation condition has been met, the translation device acquires and displays translation representation data of the first speech data so that the translation representation data can represent the language translation result of the sentence content that the first user wants to express when communicating with at least one other person in a cross-language manner, so that the cross-language communication object of the first user (that is, the "second user" referred to below) can understand what the first user wants to express from the translation device.

[0138] It should be noted that the relevant content of step 22 can be found in the relevant content of S2 above, and will not be repeated here for the sake of brevity.

[0139] In addition, to help the first user's cross-language communication partner better understand what the first user wants to express, the first user can adjust the screen of the translation device in a certain way (e.g., Figure 7 The screen is shown being rotated 180° horizontally from vertical to horizontal. Figure 8 The screen is displayed horizontally and vertically, or Figure 9 The screen is rotated 180° horizontally (as shown) towards the cross-language communication recipient so that the recipient can better understand what the first user wants to express (e.g., Figure 10 ("Excuse me, how can I get to Ocean Park from here?").

[0140] Based on this, the present application embodiment also provides a possible implementation of the above step "displaying the translation representation data of the first voice data", which can be: when it is determined that the body movement of the first user meets the third action condition, the translation representation data of the first voice data is displayed according to the screen content display method corresponding to the third action condition.

[0141] The third action condition describes what physical actions the first user typically performs when the first user faces the screen of the translation device toward the first user's cross-language communication partner; and the embodiments of this application do not limit the third action condition. For ease of understanding, three examples are used for illustration below.

[0142] Example 1, such as Figure 7As shown, the first user can horizontally flip the translation device by 180° in a portrait mode to show the translation device to the cross-language communication object of the first user. Based on this, the present embodiment provides a possible implementation of the above-mentioned "third action condition", which can be specifically: the first user shows the screen of the translation device in a portrait mode.

[0143] In addition, in order to better detect the above-mentioned "showing the screen of the translation device in a portrait mode", a gyroscope and a gravity sensor can be deployed in the translation device. Based on this, the present embodiment also provides a possible implementation of the above-mentioned "third action condition", in which when the translation device includes a gyroscope and a gravity sensor, the third action condition can be specifically: the angle description information collected by the gyroscope satisfies a fifth information condition (for example, the angle description information indicates that the translation device has changed by approximately 180° (for example, between 170° and 190°) in the X-axis horizontal angle, etc.), or the gravity acceleration description information collected by the gravity sensor satisfies a sixth information condition (for example, the gravity acceleration description information indicates that the gravity acceleration is displaced in the Z-axis, etc.).

[0144] Example 2, as Figure 8 shown, the first user can show the translation device to the cross-language communication object of the first user in a horizontal mode. Based on this, the present embodiment provides a possible implementation of the above-mentioned "third action condition", which can be specifically: the first user shows the screen of the translation device in a horizontal mode.

[0145] In fact, the detection process for the above-mentioned "showing the screen of the translation device in a horizontal mode" can be realized by deploying a gyroscope and a gravity sensor in the translation device, so the present embodiment also provides a possible implementation of the above-mentioned "third action condition", in which when the translation device includes a gyroscope and a gravity sensor, the third action condition can be specifically: the angle description information collected by the gyroscope satisfies a seventh information condition (for example, the angle description information indicates that the translation device has changed by approximately 0°-45° in the Y-axis horizontal angle, etc.), or the gravity acceleration description information collected by the gravity sensor satisfies an eighth information condition (for example, the gravity acceleration description information indicates that the gravity acceleration is displaced in the Z-axis, etc.).

[0146] Example 3, as Figure 9As shown, the first user can show the translation device to the cross-language communication object of the first user by turning the wrist to horizontally flip the translation device by 180°. Based on this, the application embodiment provides a possible implementation of the above-mentioned "third action condition", which can be specifically: the first user shows the screen of the translation device in a horizontal flip manner.

[0147] In fact, the detection process of "showing the screen of the translation device in a horizontal flip manner" can be realized by deploying a gyroscope and a gravity sensor in the translation device, so the application embodiment also provides a possible implementation of the above-mentioned "third action condition". In this implementation, when the translation device includes a gyroscope and a gravity sensor, the third action condition can be specifically: the angle description information collected by the gyroscope satisfies the ninth information condition (for example, the angle description information indicates that the translation device has changed approximately 180° (for example, between 170° and 190°) in the X-axis horizontal angle, and has changed approximately 90° (for example, between 80° and 100°) in the Y-axis vertical angle, etc.), or the gravity acceleration description information collected by the gravity sensor satisfies the tenth information condition (for example, the gravity acceleration description information indicates that the gravity acceleration is displaced in the Z-axis, etc.).

[0148] Based on the above-mentioned related content of the "third action condition", for the third action condition, the third action condition can describe how the first user shows the screen of the translation device to the cross-language communication object of the first user, so that the translation device can subsequently display the translated representation data of the first voice data in a screen content display manner corresponding to the manner, so that the cross-language communication object can better understand what the first user wants to express from the translation device.

[0149] The "screen content display manner corresponding to the third action condition" mentioned above refers to a more suitable screen content display manner (for example, full-screen pop-up manner, etc.) for the translation device under the third action condition. Moreover, the application embodiment does not limit the "screen content display manner corresponding to the third action condition", for example, for ease of understanding, the following is described in conjunction with examples.

[0150] As an example, the above-mentioned step "when it is determined that the limb action of the first user satisfies the third action condition, showing the translated representation data of the first voice data in a screen content display manner corresponding to the third action condition" can specifically include at least one of the following steps 31-33:

[0151] Step 31: when it is determined that the first user shows the screen of the translation device to the second user in a horizontal longitudinal manner, the translated representation data of the first voice data is flipped and displayed.

[0152] In the embodiments of the present application, for the translation device (for example, the translation device which has not yet displayed the translation representation data of the first voice data or the translation device which is displaying the translation representation data of the first voice data), if the translation device detects that the first user displays the screen of the translation device to the second user in a horizontal longitudinal manner, the translation device can determine that the second user will look at the screen of the translation device in reverse, so that the translation device can determine that the second user has inconvenience in viewing the content displayed on the screen, and therefore, in order to overcome the problem, the translation device can flip and display the translation representation data of the first voice data (especially, the translation device can flip and display the translation representation data by means of a full-screen pop-up window), so that the second user can view the translation representation data in a forward direction on the translation device, so that the second user can better understand what the first user wants to express.

[0153] Step 32: when it is determined that the first user displays the screen of the translation device to the second user in a longitudinal horizontal flip manner, the translation representation data of the first voice data is displayed in an enlarged manner.

[0154] In the embodiments of the present application, for the translation device (for example, the translation device which has not yet displayed the translation representation data of the first voice data or the translation device which is displaying the translation representation data of the first voice data), if the translation device detects that the first user displays the screen of the translation device to the second user in a longitudinal horizontal flip manner, the translation device presumes that the second user may have a phenomenon of viewing the screen of the translation device from a distance, and therefore, in order to overcome the adverse effects caused by the phenomenon, the translation device can display the translation representation data of the first voice data in an enlarged manner (especially, the translation device can display the translation representation data in an enlarged manner by means of a full-screen pop-up window), so that the second user can clearly view the translation representation data on the translation device, so that the second user can better understand what the first user wants to express.

[0155] Step 33: when it is determined that the first user displays the screen of the translation device to the second user in a horizontal horizontal flip manner, the translation representation data of the first voice data is displayed in an enlarged manner in a horizontal screen.

[0156] In the embodiments of the present application, for the translation device (for example, the translation device that has not yet displayed the translation representation data of the first voice data or the translation device that is displaying the translation representation data of the first voice data), if the translation device detects that the first user displays the screen of the translation device to the second user in a horizontal flip manner, the translation device presumes that the second user not only finds it inconvenient to view the screen content due to looking at the screen of the translation device horizontally, but also may view the screen of the translation device from a distance. Therefore, in order to overcome these problems, the translation device can display the translation representation data of the first voice data in a zoomed-in horizontal screen manner (especially, the translation device can display the translation representation data in a full-screen pop-up manner), so that the second user can better view the translation representation data on the translation device, and thus the second user can better understand what the first user wants to express.

[0157] Based on the related content of steps 31 to 33, for the translation device (for example, the translation device that has not yet displayed the translation representation data of the first voice data or the translation device that is displaying the translation representation data of the first voice data), when the translation device detects that the body movement of the first user meets the third movement condition, the translation device can determine that the first user displays the screen of the translation device to the second user in a certain manner. Therefore, in order to facilitate the second user to better view the display content of the translation device, the translation device can display the translation representation data of the first voice data in the screen content display manner corresponding to the “certain manner”, so that the translation device can display the translation representation data in a manner as convenient as possible for the second user to view the screen content, thereby enabling the second user to better understand what the first user wants to express, and thus facilitating to improve the user experience.

[0158] Based on the above-mentioned related content of steps 21-22, for the first user and the second user with cross-language communication needs, if the first user wants to use the translation device to convey certain semantic information to the second user, the first user can first trigger the translation device to collect the content of the first user by some simple body movements (for example, picking up the translation device placed on the horizontal desktop and moving it to the lips, etc.) that can meet the first action condition, so that the translation device can collect the content of the first user to obtain first voice data, so that the first voice data can represent the semantic information that the first user wants to convey to the second user, so that when the translation device determines that the first translation condition is met, the translation device obtains and displays the translation representation data of the first voice data, so that the translation representation data can represent the semantic information that the first user wants to convey to the second user, so that the second user can understand what the first user wants to express from the translation device. This can realize triggering voice translation process by simple body movements, thereby effectively reducing voice translation effect, and further improving user experience.

[0159] Method embodiment four

[0160] In fact, for the cross-language communication scene, after the first user directs the screen of the translation device towards the second user, it is difficult for the first user to understand the current display state of the translation device, which may lead to the first user being unable to determine whether the second user can normally view the translation representation data of the first voice data.

[0161] Based on this, the embodiment of the present application also provides another possible implementation manner of the translation method, in which the translation method not only includes the above-mentioned steps 21-22, but also includes step 23:

[0162] Step 23: When the translation device determines to display the translation representation data of the first voice data in the screen content display manner corresponding to the third action condition, the translation device controls the translation device to vibrate, so as to prompt the first user holding the translation device to understand that the translation representation data is being displayed in the screen content display manner corresponding to the third action condition.

[0163] Therefore, when the translation device detects that its screen begins to display the translation representation data of the first voice data in accordance with the screen content display method corresponding to the third action condition (e.g., flipping to display the translation representation data, zooming in to display the translation representation data, or zooming in to display the translation representation data in landscape mode), the translation device can determine that its screen can display the translation representation data in a way that is convenient for the second user to view the screen content. Therefore, the translation device can control itself to vibrate to inform the first user who is holding the translation device that the translation representation data is being displayed in a way that is convenient for the second user to view the screen content. This allows the first user to infer that the second user can normally view the translation representation data, which is beneficial to improving the user experience.

[0164] It should be noted that the embodiments of this application are not limited to the above-described implementation of "controlling the translation device to vibrate". For example, it can be implemented by means of a vibration motor deployed in the translation device.

[0165] Method embodiment five

[0166] In practice, in cross-language communication scenarios, after the second user understands what the first user has said from the translation device, the second user usually begins to speak (e.g., Figure 10 (As shown). Based on this, the embodiments of this application also provide another possible implementation of the translation method. In this implementation, the translation method may include not only all or some of the above steps (e.g., steps 21-22, or steps 21-23, etc.), but may also include steps 24-25:

[0167] Step 24: When the preset recording conditions are met, the translation device acquires the second voice data.

[0168] The preset recording conditions refer to the conditions that need to be met when the voice translation processing for the second user begins during cross-language communication between the first user and the second user. Moreover, the embodiments of this application do not limit the preset recording conditions. For example, when the above-mentioned "translation representation data" includes translation audio data, the preset recording conditions may specifically be: the playback of the translation audio data of the first voice data ends.

[0169] The aforementioned "second speech data" refers to the audio recording results of what the second user said, so that the second speech data can represent the content of the second user's speech (e.g., Figure 10 The content shown is "Go straight along...". Here, the second user represents the first user's cross-language communication counterpart.

[0170] In addition, the language category to which the second voice data belongs (for example, Chinese shown in the figure) is different from the language category to which the first voice data belongs (for example, English shown in the figure). Figure 10 Figure 6

[0171] Based on the above-mentioned related content of step 24, for the translation device that is playing the translated audio data of the first voice data, when the translation device detects that the playing of the translated audio data ends, the translation device can infer that the second user may speak, and therefore the translation device can determine that the preset recording condition is reached, so that the translation device starts the sound collecting process for the second user to obtain the second voice data, so that the second voice data can represent the content said by the second user, so as to subsequently perform the voice translation process on the second voice data.

[0172] Step 25: When it is determined that the second translation condition is reached, the translation device obtains and displays the translated representation data of the second voice data.

[0173] The second translation condition is used to trigger the language translation process for the content said by the second user, and the embodiments of the present application do not limit the second translation condition. For example, in order to improve the voice translation effect as much as possible, the second translation condition can be set according to the end time of the speech of the first user (for example, the second translation condition can be specifically that the duration of the second user in the silent state reaches a preset duration threshold, etc.), so that the triggering time of the language translation process for the content said by the second user is equal to or slightly later than the end time of the speech of the second user.

[0174] The above-mentioned "translated representation data of the second voice data" is used to represent the translation result of the voice information carried by the second voice data, and the embodiments of the present application do not limit the "translated representation data of the second voice data". For example, it can include at least one of the translated text data of the second voice data and the translated audio data of the second voice data.

[0175] It should be noted that the above-mentioned related content of the "translated representation data of the second voice data" is similar to the above-mentioned related content of the "translated representation data of the first voice data", and for the sake of brevity, it will not be repeated here.

[0176] In addition, the embodiments of the present application do not limit the implementation mode of the above-mentioned step "the translation device obtains and displays the translated representation data of the second voice data", for example, it can be implemented in the way of obtaining and displaying at the same time, so as to improve the voice translation efficiency.

[0177] ​​Based on the above-mentioned related content of step 24 to step 25, for the translation device that is displaying the translated representation data of the first voice data to the second user, if the translation device determines that the preset recording condition has been reached (for example, the playing of the translated audio data of the first voice data ends, etc.), the translation device can collect the second voice data for the second user, so that the second voice data can represent the content said by the second user, so that when the translation device determines that the second translation condition is reached (for example, the second user ends speaking), the translation device obtains and displays the translated representation data of the second voice data, so that the first user can know what the second user wants to express from the translation device, so that the cross-language communication process between at least two people without any key operation can be better realized, so that the cross-language communication effect can be improved.

[0178] Method embodiment six

[0179] In fact, in order to further improve the user experience, some prompts (for example, text prompts, voice prompts, or vibration prompts, etc.) can be given to the second user (or the first user) when the translation device starts to collect the voice for the second user.

[0180] Based on this, the embodiment of the present application also provides another possible implementation manner of step 24, which can be specifically: when it is determined that the preset recording condition is reached, the translation device sends a second prompt and obtains the second voice data.

[0181] Among them, the second prompt is used to prompt the second user that he / she can start speaking; and the present application embodiment does not limit the second prompt, for example, it can include at least one content of prompt text (for example, the text content of “Say English, please……” shown in the figure), prompt voice (for example, the audio data carrying the voice information of “Say English, please……” shown in the figure), and prompt vibration. Figure 10 Figure 10

[0182] ​​It can be seen that, for the translation device, when the translation device detects that the playing of the translation audio data of the first voice data ends, the translation device can infer that the second user can speak, so the translation device can determine that the preset recording condition is reached, so that the translation device can not only send the second prompt to make the second user know from the second prompt that he / she can start to speak, but also can perform real-time recording for the second user to obtain the second voice data, so that the second voice data can represent the speaking content of the second user, so as to subsequently perform language translation processing on the user semantics carried by the second voice data, thereby effectively avoiding the adverse effects caused by the second user not knowing when to start speaking (for example, wasting of recording resources caused by starting to speak too late), thereby improving the user experience.

[0183] Method embodiment seven

[0184] In fact, for the cross-language communication scene, the cross-language communication content between at least two persons can usually be displayed in a certain default display manner (for example, a bubble picture display manner), so as to further improve the user experience, the present application embodiment also provides another possible implementation manner of the translation method, in which the translation method not only includes all or part of the above steps, but also includes at least one of steps 26-29:

[0185] Step 26: when it is determined that the preset switching condition is reached, the translation device displays the translation representation data of the first voice data in a preset screen display manner.

[0186] The preset switching condition refers to a trigger condition for switching the screen of the translation device from the above-mentioned "screen content display manner corresponding to the third action condition" back to the preset screen display manner, and the preset switching condition is not limited in the present application embodiment, for example, when the "translation representation data" includes translation audio data, the preset switching condition can be specifically that the playing of the translation audio data of the first voice data ends.

[0187] The "preset screen display manner" refers to the default display manner of the screen of the translation device, and the preset screen display manner is not limited in the present application embodiment, for example, it can be implemented by using a bubble picture.

[0188] In addition, the execution time of step 26 is not limited in the present application embodiment, for example, it is later than the execution time of the above-mentioned step "when it is determined that the body movement of the first user meets the third action condition, display the translation representation data of the first voice data in the screen content display manner corresponding to the third action condition".

[0189] Based on the relevant content of step 26 above, it can be seen that for the translation device, when the translation device is displaying the translation representation data of the first voice data according to the screen content display mode corresponding to the third action condition, if the translation device determines that the preset switching condition has been met (e.g., the translation audio data of the first voice data has finished playing), the translation device can display the translation representation data according to the preset screen display mode, so as to achieve the purpose of switching the translation representation data back to the default display mode, which is conducive to improving the display effect of the translation representation data.

[0190] Step 27: When it is determined that the first user's body movements meet the fourth action condition, the translation device displays the translation representation data of the first voice data according to the preset screen display method.

[0191] The fourth action condition describes what physical actions the first user typically performs when the first user manually interrupts the display of the translation representation data of the first voice data. Moreover, the embodiments of this application do not limit the fourth action condition. For example, it can retract the translation device for the first user so that the translation device is no longer facing the second user, and enable the first user to reuse the translation device for voice translation processing.

[0192] It should be noted that the embodiments of this application do not limit the detection method of the above-mentioned "fourth action condition". For example, it can be implemented by means of a distance sensor, a gyroscope, and a gravity sensor.

[0193] Furthermore, the implementation of this application does not limit the execution time of step 27. For example, it may be later than the execution time of the above step "when it is determined that the first user's body movement meets the third action condition, display the translation representation data of the first voice data according to the screen content display method corresponding to the third action condition".

[0194] Based on the relevant content of step 27 above, it can be seen that when the translation device is displaying the translation representation data of the first voice data according to the screen content display mode corresponding to the third action condition, if the translation device detects that the first user's body movement meets the fourth action condition (for example, the first user retracts the translation device), the translation device can determine that the first user wants to manually interrupt the process of displaying the translation representation data to the second user. Therefore, the translation device can display the translation representation data according to the preset screen display mode to achieve the purpose of manually switching the translation representation data back to the default display mode, which is beneficial to improving the display effect of the translation representation data.

[0195] Step 28: When it is determined that the first user's body movements meet the fifth action condition, the translation device hides the translation representation data of the first voice data and displays a preset guide page.

[0196] The fifth action condition is used to describe which body actions the first user usually performs when the first user wants to view the usage guide of the translation device. The fifth action condition is not limited in the embodiments of this application. For example, the fifth action condition can be a gesture swipe-up operation performed on the screen of the translation device.

[0197] The "preset guide page" is used to describe the usage guide of the translation device.

[0198] In addition, the embodiments of this application do not limit the execution time of step 28. For example, step 28 is executed later than step 22 (or step 12) described above.

[0199] Based on the above description of step 28, for the translation device, when the translation device is displaying the translated representation data of the first voice data, if the translation device detects that the body action of the first user meets the fifth action condition (for example, a gesture swipe-up operation), the translation device can determine that the first user wants to view the usage guide of the translation device. Therefore, the translation device can directly hide the translated representation data and display a preset guide page, so that the first user can learn how to use the translation device on the preset guide page.

[0200] Step 29: When it is determined that the body action of the first user meets the sixth action condition, the translation device displays at least one historical voice translation description information.

[0201] The sixth action condition is used to describe which body actions the first user usually performs when the first user wants to view the historical chat record between the first user and the second user. The sixth action condition is not limited in the embodiments of this application. For example, the sixth action condition can be a gesture swipe-down operation performed on the screen of the translation device.

[0202] The "at least one historical voice translation description information" is used to describe the historical chat record between the first user and the second user. The historical voice translation description information is not limited in the embodiments of this application. For example, the historical voice translation description information can include voice data, voice recognition text of the voice data, and translated representation data (for example, translated text data and translated audio data) of the voice data.

[0203] In addition, the embodiments of this application do not limit the execution time of step 29. For example, step 29 is executed later than step 22 (or step 12) described above.

[0204] Based on the related content of step 29, for the translation device, when the translation device is displaying the translation representation data of the first voice data, if the translation device detects that the first user's body movement meets the sixth movement condition (for example, gesture sliding up, etc.), the translation device can determine that the first user wants to view the historical chat record with the second user, so the translation device can display at least one historical voice translation description information, so that the historical voice translation description information can represent the cross-language communication content between the first user and the second user in the historical time period, so as to meet the demand of the first user for viewing the historical chat record, thereby facilitating the improvement of user experience.

[0205] Based on the translation method provided in the above method embodiment, the present embodiment further provides a translation device, which will be explained and described below with reference to the accompanying drawings.

[0206] Device embodiment

[0207] The device embodiment introduces the translation device, and the related content can be referred to the above method embodiment.

[0208] Referring to Figure 11 The figure is a structural schematic diagram of a translation device provided in the present embodiment.

[0209] The translation device 1100 provided in the present embodiment comprises:

[0210] The first acquisition unit 1101 is configured to acquire first voice data when it is determined that the first user's body movement meets the first movement condition; wherein the first voice data carries the speaking content of the first user.

[0211] The first display unit 1102 is configured to acquire and display translation representation data of the first voice data when it is determined that the first translation condition is met; wherein the translation representation data comprises translation text data and / or translation audio data.

[0212] In a possible implementation, the first acquisition unit 1101 is specifically configured to acquire the first voice data when it is determined that the distance between the body pronunciation part of the first user and the translation device meets the first distance condition, or it is determined that the included angle between the translation device and the reference horizontal plane meets the first angle condition.

[0213] In a possible implementation, the translation device comprises a distance sensor and a gyroscope; and the first acquisition unit 1101 is specifically configured to acquire the first voice data when it is determined that the distance description information collected by the distance sensor meets the first information condition, or it is determined that the angle description information collected by the gyroscope meets the second information condition.

[0214] In a possible implementation, the first display unit 1102 is specifically configured to: acquire and display the translated representation data of the first voice data when it is determined that the limb action of the first user meets the second action condition.

[0215] In a possible implementation, the first display unit 1102 is specifically configured to: acquire and display the translated representation data of the first voice data when it is determined that the distance between the articulatory position of the first user and the translation device meets the second distance condition, or it is determined that the included angle between the translation device and the reference horizontal plane meets the second angle condition.

[0216] In a possible implementation, the first display unit 1102 includes:

[0217] The first display sub-unit is configured to display the translated representation data of the first voice data in the screen content display mode corresponding to the third action condition when it is determined that the limb action of the first user meets the third action condition.

[0218] In a possible implementation, the first display sub-unit is specifically configured to: flip the display of the translated representation data of the first voice data when it is determined that the first user displays the screen of the translation device to the second user in a horizontal longitudinal unfolding manner.

[0219] In a possible implementation, the first display sub-unit is specifically configured to: enlarge the display of the translated representation data of the first voice data when it is determined that the first user displays the screen of the translation device to the second user in a vertical horizontal flipping manner.

[0220] In a possible implementation, the first display sub-unit is specifically configured to: enlarge the display of the translated representation data of the first voice data in a horizontal screen mode when it is determined that the first user displays the screen of the translation device to the second user in a horizontal horizontal flipping manner.

[0221] In a possible implementation, the translation apparatus 1100 further includes:

[0222] The vibration reminding unit is configured to control the translation device to vibrate when the translated representation data of the first voice data is displayed in the screen content display mode corresponding to the third action condition.

[0223] In a possible implementation, the translation apparatus 1100 further includes:

[0224] a second obtaining unit, configured to obtain second voice data when it is determined that a preset recording condition is met, wherein the second voice data is used to represent the speaking content of a second user, and the language category to which the second voice data belongs is different from the language category to which the first voice data belongs;

[0225] a second display unit, configured to obtain and display translated representation data of the second voice data when it is determined that a second translation condition is met.

[0226] In a possible implementation, the translation apparatus 1100 further includes:

[0227] a first switching unit, configured to display the translated representation data of the first voice data in a preset display mode when it is determined that a preset switching condition is met.

[0228] In a possible implementation, the translation apparatus 1100 further includes:

[0229] a second switching unit, configured to display the translated representation data of the first voice data in the preset display mode when it is determined that the limb action of the first user meets a fourth action condition.

[0230] In a possible implementation, the translation apparatus 1100 further includes:

[0231] a third display unit, configured to hide the translated representation data of the first voice data and display a preset guide page when it is determined that the limb action of the first user meets a fifth action condition, wherein the preset guide page is used to describe the usage guide of the translation device.

[0232] In a possible implementation, the translation apparatus 1100 further includes:

[0233] a fourth display unit, configured to display at least one historical voice translation description information when it is determined that the limb action of the first user meets a sixth action condition.

[0234] Based on the related content of the translation device 1100, for the translation device 1100 provided by the embodiment of the application, when the translation device 1100 determines that the first user's limb action meets the first action condition (for example, the first user moves the translation device 1100 to the lips, etc.), the translation device 1100 starts to collect sound for the first user to obtain first voice data, so that the first voice data can represent the speaking content of the first user, so that when the translation device 1100 determines that the first translation condition is met (for example, the first user moves the translation device 1100 away from the lips, etc.), the translation device 1100 can determine that the first user ends speaking, and therefore the translation device 1100 can obtain and display the translated representation data (for example, translated text data and / or translated audio data) of the first voice data, so that the first user (or the cross-language communication object of the first user) can understand the translation result of the first voice data from the translation device 1100.

[0235] It can be seen that the translation device 1100 provided by the application can control the language translation process by means of some simple limb actions (for example, moving the translation device 1100 to the user's lips, moving the translation device 1100 away from the user's lips, etc.), so that the user of the translation device 1100 does not need to perform excessive or long key operation in the language translation process, thereby enabling the translation device 1100 to effectively overcome the adverse effects (for example, complicated operation, key finger twitching, etc.) caused by controlling the language translation process by means of key operation, so as to effectively simplify the language translation process, thereby effectively improving the user experience.

[0236] Based on the translation method provided by the above method embodiment, the embodiment of the application further provides a translation device, which will be explained and described below with reference to the accompanying drawings.

[0237] Device embodiment

[0238] The device embodiment introduces the translation device, and the related content is described above.

[0239] Referring to Figure 12 The figure is a structural schematic diagram of a translation device provided by the embodiment of the application.

[0240] The translation device 1200 provided by the embodiment of the application comprises a processor 1201, a memory 1202, and a system bus 1203;

[0241] The processor 1201 and the memory 1202 are connected through the system bus 1203;

[0242] The memory 1202 is configured to store one or more programs, which include instructions that, when executed by the processor 1201, cause the processor 1201 to perform any of the above-described embodiments of the translation method. For ease of understanding, some embodiments will be described below.

[0243] In a possible implementation, the processor 1201 can perform the following steps:

[0244] Upon determining that the limb action of the first user meets the first action condition, the first voice data is acquired; wherein the first voice data carries the speaking content of the first user;

[0245] Upon determining that the first translation condition is met, the translated representation data of the first voice data is acquired and displayed; wherein the translated representation data includes translated text data and / or translated audio data.

[0246] In a possible implementation, the processor 1201 can specifically perform the following steps: upon determining that the distance between the articulatory part of the body of the first user and the translation device meets the first distance condition, or determining that the included angle between the translation device and the reference horizontal plane meets the first angle condition, the first voice data is acquired.

[0247] In a possible implementation, the translation device includes a distance sensor and a gyroscope; and the processor 1201 can perform the following steps: upon determining that the distance description information collected by the distance sensor meets the first information condition, or determining that the angle description information collected by the gyroscope meets the second information condition, the first voice data is acquired.

[0248] In a possible implementation, the processor 1201 can specifically perform the following steps: upon determining that the limb action of the first user meets the second action condition, the translated representation data of the first voice data is acquired and displayed.

[0249] In a possible implementation, the processor 1201 can specifically perform the following steps: upon determining that the distance between the articulatory part of the body of the first user and the translation device meets the second distance condition, or determining that the included angle between the translation device and the reference horizontal plane meets the second angle condition, the translated representation data of the first voice data is acquired and displayed.

[0250] In a possible implementation, the processor 1201 can specifically perform the following steps: upon determining that the limb action of the first user meets the third action condition, the translated representation data of the first voice data is displayed in a screen content display mode corresponding to the third action condition.

[0251] In a possible implementation, the processor 1201 can specifically perform the following steps: when it is determined that the first user displays the screen of the translation device to the second user in a horizontal portrait mode, the processor 1201 can flip the translated representation data of the first voice data.

[0252] In a possible implementation, the processor 1201 can specifically perform the following steps: when it is determined that the first user displays the screen of the translation device to the second user in a portrait horizontal flipping mode, the processor 1201 can enlarge the translated representation data of the first voice data.

[0253] In a possible implementation, the processor 1201 can specifically perform the following steps: when it is determined that the first user displays the screen of the translation device to the second user in a horizontal flipping mode, the processor 1201 can enlarge the translated representation data of the first voice data in a horizontal screen display mode.

[0254] In a possible implementation, the processor 1201 can further perform the following steps:

[0255] When the translated representation data of the first voice data is displayed in the screen content display mode corresponding to the third action condition, the processor 1201 can control the translation device to vibrate.

[0256] In a possible implementation, the processor 1201 can further perform the following steps:

[0257] When it is determined that the preset recording condition is met, the processor 1201 can acquire second voice data; the second voice data can be used to represent the speaking content of the second user; and the language category to which the second voice data belongs can be different from the language category to which the first voice data belongs.

[0258] When it is determined that the second translation condition is met, the processor 1201 can acquire and display the translated representation data of the second voice data.

[0259] In a possible implementation, the processor 1201 can further perform the following steps:

[0260] When it is determined that the preset switching condition is met, the processor 1201 can display the translated representation data of the first voice data in the preset screen display mode.

[0261] In a possible implementation, the processor 1201 can further perform the following steps:

[0262] When it is determined that the limb action of the first user meets the fourth action condition, the processor 1201 can display the translated representation data of the first voice data in the preset screen display mode.

[0263] In a possible implementation, the processor 1201 can further perform the following steps:

[0264] When it is determined that the limb action of the first user meets the fifth action condition, the translated representation data of the first voice data is hidden, and a preset guide page is displayed; wherein the preset guide page is used to describe the usage guide of the translation device.

[0265] In a possible implementation, the processor 1201 can further perform the following steps:

[0266] When it is determined that the limb action of the first user meets the sixth action condition, at least one historical voice translation description information is displayed.

[0267] In addition, in a possible implementation, as shown in Figure 13 The translation device 1200 can further include at least one of a communication module, a distance sensor, a gyroscope, a gravity sensor, and a vibration motor; wherein the processor and the communication module are connected through the system bus; the processor and the distance sensor are connected through the system bus; the processor and the gyroscope are connected through the system bus; the processor and the gravity sensor are connected through the system bus; and the processor and the vibration motor are connected through the system bus.

[0268] The above-mentioned "communication module" is used to realize the communication process between the translation device 1200 and the server, so that the translation device 1200 can use some service items (such as voice recognition service, language translation service, audio synthesis service, etc.) provided by the server; and the communication module is not limited in the embodiments of the present application, for example, it can be implemented in any way that can provide communication service for the translation device 1200 at present or in the future.

[0269] It should be noted that, in Figure 13 The recognition engine is used to provide voice recognition service for the translation device 1200; the translation engine is used to provide language translation service for the translation device 1200; and the synthesis engine is used to provide audio synthesis service for the translation device 1200.

[0270] The above-mentioned "distance sensor" is used to detect the distance between the first user and the translation device 1200 in real time; and the distance sensor is not limited in the embodiments of the present application, for example, it can be implemented in any sensor that can provide distance measurement service for the translation device 1200 at present or in the future.

[0271] The "gyroscope" is used to detect the angle information of the translation device 1200 in real time; and the embodiment of the present application does not limit the gyroscope, for example, it can be implemented by using any device capable of providing angle measurement service for the translation device 1200.

[0272] The "gravity sensor" is used to detect the gravity acceleration description information of the translation device 1200 in real time; and the embodiment of the present application does not limit the gravity sensor, for example, it can be implemented by using any device capable of providing gravity acceleration measurement service for the translation device 1200.

[0273] The "vibration motor" is used to control the translation device 1200 to vibrate; and the embodiment of the present application does not limit the vibration motor, for example, it can be implemented by using any device capable of providing vibration service for the translation device 1200.

[0274] In a possible implementation, the translation device 1200 can implement at least one function shown in (1)-(10) as follows:

[0275] (1) Triggering each process node in the speech translation process for the first user through body movement, for example, it can be judged whether the speech recognition process for the first user can be triggered through the distance sensor and the gravity sensor, it can also be judged whether the language translation process for the content spoken by the first user can be triggered through the distance sensor and the gravity sensor, it can also be judged whether the screen is flipped through the gravity sensor and the gyroscope, so as to automatically pull up the speech recognition recording after the broadcast ends, to start the speech recognition process for the second user, and it can also be judged whether the screen is flipped vertically or horizontally through the gravity sensor and the gyroscope.

[0276] (2) After the screen is flipped, the translation is displayed in full screen and enlarged, and the audio data of the translation is started to be broadcast.

[0277] (3) When broadcasting, the audio and characters are synchronized, and the broadcast content is displayed with word-by-word highlighting.

[0278] (4) After the screen is flipped, the translation is displayed in full screen and enlarged, and the audio data of the translation is started (or continued) to be broadcast, and the broadcast is waited to end, and the speech recording is automatically triggered.

[0279] (5) When the speech recording is automatically triggered, the recognized speech content is displayed on the screen in real time.

[0280] (6) After the speech recognition ends (i.e., the speaker does not speak), the automatic end point is judged (2S), and the automatic translation broadcast is started to be triggered.

[0281] (7) After the broadcast ends, the full-screen display of the translated text pop-up disappears, and the default bubble screen is returned.

[0282] (8) When the broadcast is not over, the distance sensor and the gravity sensor are used to determine whether to close the full-screen pop-up in advance and return to the default bubble screen.

[0283] (9) Only the content of the current dialogue bubble is displayed each time, and gesture sliding down is supported to call up more chat records.

[0284] (10) Gesture sliding up is supported to hide the current dialogue bubble content and display a blank guide page.

[0285] Based on the above description of the translation device 1200, it can be seen that the translation device 1200 provided by the embodiments of the present application can control the language translation process by means of some simple body movements. Moreover, the implementation process can be as follows: when the translation device 1200 determines that the body movement of the first user meets the first action condition (for example, the first user moves the translation device 1200 to the lips, etc.), the translation device 1200 starts to collect the sound for the first user and obtains first voice data, so that the first voice data can represent the speaking content of the first user, so that when the translation device 1200 determines that the first translation condition is met (for example, the first user moves the translation device 1200 away from the lips, etc.), the translation device 1200 can determine that the first user has finished speaking, and therefore the translation device 1200 can obtain and display the translated representation data (for example, translated text data and / or translated audio data) of the first voice data, so that the first user (or the cross-language communication object of the first user) can understand the translation result of the first voice data from the translation device 1200.

[0286] It can be seen that the translation device 1200 provided by the present application can control the language translation process by means of some simple body movements (for example, moving the translation device 1200 to the user's lips, moving the translation device 1200 away from the user's lips, etc.), so that the user of the translation device 1200 does not need to perform too many or too long key operations in the language translation process, thereby enabling the translation device 1200 to effectively overcome the adverse effects (for example, complicated operation, key finger twitching, etc.) caused by controlling the language translation process by means of key operations, so as to effectively simplify the language translation process, thereby effectively improving the user experience.

[0287] Further, the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores instructions, and when the instructions run on a terminal device, the terminal device executes any one of the above-mentioned translation methods.

[0288] Further, the embodiments of the present application further provide a computer program product, which, when running on a terminal device, causes the terminal device to perform any of the above-mentioned implementation methods of the translation method.

[0289] From the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the above-mentioned embodiment methods can be implemented by means of software and the necessary universal hardware platforms. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) execute the methods described in the various embodiments or some parts of the embodiments of the present application.

[0290] It should be noted that the various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts of each embodiment can be mutually referred to. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts are referred to the method part.

[0291] It should also be noted that the terms such as first and second in the present document are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0292] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of translation, characterized by, The method is applied to a translation device comprising a gyroscope and a gravity sensor, and the method comprises: When it is determined that the first user's limb action meets a first action condition, first voice data is acquired; wherein the first voice data carries the speaking content of the first user; When it is determined that a first translation condition is met, translated representation data of the first voice data is acquired and displayed; wherein the translated representation data comprises translated text data and / or translated audio data, and the voice recognition process and the language translation process of the first voice data are both triggered according to information collected by the gravity sensor; The display of the translated representation data of the first voice data comprises: when it is determined that the first user's limb action meets a third action condition, the translated representation data of the first voice data is displayed in a screen content display mode corresponding to the third action condition, which is different from a preset screen display mode; After the translated representation data of the first voice data is displayed in the screen content display mode corresponding to the third action condition, the method further comprises: When a preset switching condition is met, the translated representation data of the first voice data is displayed in the preset screen display mode; and / or when it is determined that the first user's limb action meets a fourth action condition, the translated representation data of the first voice data is displayed in the preset screen display mode; The display of the translated representation data of the first voice data when it is determined that the first user's limb action meets a third action condition comprises: When it is determined according to the information collected by the gyroscope and the information collected by the gravity sensor that the first user displays the screen of the translation device to a second user in a horizontal longitudinal unfolding manner, the translated representation data of the first voice data is flipped and displayed; When it is determined according to the information collected by the gyroscope and the information collected by the gravity sensor that the first user displays the screen of the translation device to a second user in a vertical horizontal flipping manner, the translated representation data of the first voice data is enlarged and displayed; When it is determined according to the information collected by the gyroscope and the information collected by the gravity sensor that the first user displays the screen of the translation device to a second user in a horizontal horizontal flipping manner, the translated representation data of the first voice data is enlarged and displayed in a horizontal screen.

2. The method of claim 1, wherein, The acquisition of the first voice data when it is determined that the first user's limb action meets a first action condition comprises: When it is determined that the distance between the first user's body pronunciation part and the translation device meets a first distance condition, or it is determined that the included angle between the translation device and a reference horizontal plane meets a first angle condition, the first voice data is acquired.

3. The method of claim 1, wherein, The translation device comprises a distance sensor and a gyroscope; The acquisition of the first voice data when it is determined that the first user's limb action meets a first action condition comprises: The first voice data is acquired when it is determined that the distance description information collected by the distance sensor meets a first information condition or that the angle description information collected by the gyroscope meets a second information condition.

4. The method of claim 1, wherein, The method further includes: The translated representation data of the first voice data is acquired and displayed when it is determined that the body action of the first user meets a second action condition.

5. The method of claim 4, wherein, The method further includes: The translated representation data of the first voice data is acquired and displayed when it is determined that the distance between the body pronunciation part of the first user and the translation device meets a second distance condition or that the included angle between the translation device and a reference horizontal plane meets a second angle condition.

6. The method of claim 1, wherein, The method further includes: The translation device is controlled to vibrate when it is determined that the translated representation data of the first voice data is displayed in a screen content display mode corresponding to the third action condition.

7. The method of claim 1, wherein, The method further includes: Second voice data is acquired when it is determined that a preset recording condition is met, wherein the second voice data is used to represent the speaking content of a second user, and the language category to which the second voice data belongs is different from the language category to which the first voice data belongs. Translated representation data of the second voice data is acquired and displayed when it is determined that a second translation condition is met.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: The translated representation data of the first voice data is hidden and a preset guide page is displayed when it is determined that the body action of the first user meets a fifth action condition, wherein the preset guide page is used to describe the usage guide of the translation device. And / or, At least one historical voice translation description information is displayed when it is determined that the body action of the first user meets a sixth action condition.

9. A translation device, characterized by The device is applied to a translation device, and the translation device includes a gyroscope and a gravity sensor. A first acquisition unit is configured to acquire first voice data when it is determined that a body action of a first user meets a first action condition, wherein the first voice data carries the speaking content of the first user. A first display unit is configured to acquire and display translated representation data of the first voice data when it is determined that a first translation condition is met, wherein the translated representation data includes translated text data and / or translated audio data, and the voice recognition process and the language translation process of the first voice data are triggered based on information collected by the gravity sensor. The first display unit is specifically configured to display the translated representation data of the first voice data in a screen content display mode corresponding to the third action condition when it is determined that the body action of the first user meets the third action condition, and the screen content display mode corresponding to the third action condition is different from a preset screen display mode. The translation device further includes: A first switching unit is configured to display the translated representation data of the first voice data in a preset screen display mode when it is determined that a preset switching condition is met. And / or, The second switching unit is configured to display the translated representation data of the first voice data in the preset display mode when it is determined that the limb movement of the first user meets the fourth movement condition. The first display unit is specifically configured to: when it is determined according to the information collected by the gyroscope and the information collected by the gravity sensor that the first user displays the screen of the translation device to the second user in a horizontal portrait mode, flip display the translated representation data of the first voice data; when it is determined according to the information collected by the gyroscope and the information collected by the gravity sensor that the first user displays the screen of the translation device to the second user in a portrait horizontal flip mode, magnify display the translated representation data of the first voice data; and when it is determined according to the information collected by the gyroscope and the information collected by the gravity sensor that the first user displays the screen of the translation device to the second user in a landscape horizontal flip mode, magnify landscape display the translated representation data of the first voice data.

10. A translation device, characterized by The device comprises a processor, a memory, and a system bus; The processor and the memory are connected through the system bus; The memory is configured to store one or more programs, the one or more programs comprising instructions that, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 8.

11. The apparatus of claim 10, wherein, The translation device further comprises at least one of a communication module, a distance sensor, a gyroscope, a gravity sensor, and a vibration motor; The processor and the communication module are connected through the system bus; The processor and the distance sensor are connected through the system bus; The processor and the gyroscope are connected through the system bus; The processor and the gravity sensor are connected through the system bus; The processor and the vibration motor are connected through the system bus.

12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores instructions, when the instructions run on the terminal device, cause the terminal device to perform the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Control method for translation device, translation device, and program

    CN108307659A

  • Real -time pronunciation inter -translation device

    CN206470756U

  • Screen rotating translation machine

    CN209514620U