Presentation method and apparatus, training method and apparatus for semantic unit detection model, electronic device, storage medium, and computer program
By dividing language sequences into semantic units for display, the method stabilizes subtitle display and translation in machine simultaneous interpretation, improving user experience and translation consistency.
Patent Information
- Application Number
- JP2022111826
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-18
- Filing Date
- 2022-07-12
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2042-07-12
AI Technical Summary
Current machine simultaneous interpretation technologies display subtitles based on character or word granularity, leading to unstable and fluctuating translations that disrupt user experience and browsing habits.
The method involves splitting the display target language sequence into semantic units with accurate semantic meaning and displaying them one by one, ensuring stability and consistency in subtitle display.
This approach enhances the stability of subtitle display and translation results, aligning with user browsing habits and reducing real-time fluctuations.
Smart Images

Figure 0007698610000009 
Figure 0007698610000010 
Figure 0007698610000011
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to natural language processing and machine translation technology. More specifically, the present disclosure provides a display method and apparatus, a training method and apparatus for a semantic unit detection model, an electronic device, a storage medium, and a computer program.
Background Art
[0002] With the development of globalization, international exchanges are frequent, the demand for machine simultaneous interpretation is increasing, and there is a broad development space. The display form of machine simultaneous interpretation is to display simultaneous interpretation subtitles on the screen.
Summary of the Invention
Means for Solving the Problems
[0003] The present invention provides a display method and apparatus, a training method and apparatus for a semantic unit detection model, an electronic device, and a storage medium.
[0004] According to a first aspect, a display method is provided, the method includes: obtaining a display target language sequence; splitting the display target language sequence into a plurality of semantic units having semantics; and converting the plurality of semantic units into subtitles and displaying them one by one.
[0005] According to a second aspect, a training method for a semantic unit detection model is provided, the method includes: obtaining a sample language sequence, the sample language sequence including a plurality of elements, each element having an original tag, and the original tag of each element indicating whether an element unit composed of the element and at least one previous element of the element has semantics as a semantic unit; and training a semantic unit detection model using the sample language sequence and the original tags of each element in the sample language sequence.
[0006] According to a third aspect, a display device is provided, which includes a first acquisition module used to acquire a display target language sequence, a splitting module used to split the display target language sequence into a plurality of semantic units having semantics, and a display module used to convert the plurality of semantic units into subtitles and display them one by one.
[0007] According to a fourth aspect, a training device for a semantic unit detection model is provided, which includes a second acquisition module used to acquire a sample language sequence, where the sample language sequence includes a plurality of elements, each element has an original tag, and the original tag of each element indicates whether the element unit composed of the element and at least one previous element of the element is a semantic unit having semantics, and a training module used to train the semantic unit detection model using the sample language sequence and the original tags of each element in the sample language sequence.
[0008] According to a fifth aspect, an electronic device including at least one processor and a memory communicatively connected to the at least one processor is provided. Among them, instructions executable by the at least one processor are stored in the memory, and when the instructions are executed by the at least one processor, the at least one processor can execute the method provided by the present disclosure.
[0009] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to cause a computer to execute the method provided by the present disclosure.
[0010] According to a seventh aspect, a computer program is provided, which realizes the method provided by the present disclosure when executed by a processor.
[0011] It should be understood that the content described in this section is not intended to indicate the key points or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description.
[0012] The drawings are provided to better understand the technical solution of the present application and do not limit the present application.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0014] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. Here, various details of the embodiments of the present disclosure are included for easier understanding, and they should be considered to be exemplary. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clear and concise description, descriptions of well-known functions and configurations are omitted in the following description.
[0015] With the development of globalization, international exchanges are frequent, the demand for machine simultaneous interpretation is increasing, and there is a broad development space. Machine simultaneous interpretation uses speech recognition and machine translation technologies to automatically identify the speech content of the speaker, convert the speech into text, and translate it into the target language. The value of machine simultaneous interpretation products is mainly reflected in solving problems such as cross-language communication and cross-language acquisition. The application scenarios of machine simultaneous interpretation are mainly international conferences, and the mainstream display form in the industry is to display simultaneous interpretation subtitles on the screen.
[0016] Currently, the bilingual subtitles of machine simultaneous interpretation are displayed on the screen in real time in an incremental form according to the granularity of "characters" or "words", the subtitles are updated in real time, and especially the variation probability of some of the latest characters is large. Every time the speaker utters a character, the current speech recognition result is translated in real time, and the translation result remains unchanged until the end of a sentence. The translation effect in the intermediate state is low, and the variation of the translation result is large and unstable.
[0017] Therefore, the current solutions for machine simultaneous interpretation have at least the following problems. The method of displaying the subtitles of simultaneous interpretation on the screen according to the granularity of "characters" and "words" does not conform to human browsing habits. The subtitles fluctuate and blink in real time, with low stability and a poor browsing experience. The translation results are unstable, and the translation results already displayed on the screen are changed in real time with the incremental input of the original text, increasing the browsing burden on the user.
[0018] The embodiments of the present disclosure provide a display method, which divides the display target language sequence into a plurality of semantic units with semantics, converts the plurality of semantic units into subtitles and displays them one by one, ensuring the stability of the subtitle display effect and improving the user's browsing experience.
[0019] In the technical solutions of the present disclosure, the processing of collecting, storing, using, processing, transmitting, providing, and disclosing relevant user personal information all conform to the provisions of relevant laws and regulations and do not violate public order and good customs.
[0020] FIG. 1 is a schematic diagram of an exemplary system architecture to which the display method and / or the training method of the semantic unit detection model according to an embodiment of the present disclosure can be applied. It should be noted that what is shown in FIG. 1 is only an example to which the system architecture of the embodiment of the present disclosure can be applied, which is helpful for those skilled in the art to understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0021] As shown in FIG. 1, the system architecture 100 according to this embodiment can include a plurality of terminal devices 101, a network 102, and a server 103. The network 102 is for providing a medium for a communication link between the terminal device 101 and the server 103. The network 102 can include various connection types such as wired and / or wireless communication links, for example.
[0022] The user can send and receive messages and the like by interacting with the server 103 via the network 102 using the terminal device 101. The terminal device 101 may be various electronic devices, including but not limited to smartphones, tablet computers, laptop portable computers, etc.
[0023] At least one of the display method and the training method of the semantic unit detection model provided by the embodiments of the present disclosure can generally be executed by the server 103. Accordingly, the display device and the training device of the semantic unit detection model provided by the embodiments of the present disclosure can generally be installed in the server 103. The display method and the training method of the semantic unit detection model provided by the embodiments of the present disclosure may be executed by a server or a server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103. Accordingly, the display device and the training device of the semantic unit detection model provided by the embodiments of the present disclosure may be installed in a server or a server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103.
[0024] FIG. 2A is a flowchart of a display method according to an embodiment of the present disclosure. As shown in FIG. 2A, the display method 200 may include operations S210 to S230.
[0025] In operation S210, an object display language sequence is acquired. For example, the object display language sequence may be a real-time audio stream generated by the user's speech, or a text sequence obtained by speech recognition of the audio stream. The acquisition of the object display language sequence may be derived from a publicly available dataset or may be approved by the user.
[0026] In operation S220, the object display language sequence is divided into a plurality of semantic units having semantics.
[0027] For example, the display target language sequence is split according to semantics to obtain a plurality of semantic units having semantics.
[0028] When the display target language sequence is an audio stream, each semantic unit includes at least one audio fragment, and the at least one fragment has an accurate semantic meaning. When the display target language sequence is a text sequence, each semantic unit includes at least one character or word, and the at least one character or word also has an accurate semantic meaning.
[0029] For example, the display target language sequence is
[0030]
Number
[0031] such a text sequence, and the display target language sequence is split into a plurality of semantic units such as [Hello everyone], [Today], [I]... The content in each semantic unit is a collocation or phrase that all has an accurate semantic meaning.
[0032] If the above display target language sequence is split into a plurality of semantic units such as [Everyone], [Good today I]..., the semantic unit [Good today I] does not have an accurate semantic meaning.
[0033] In operation S230, a plurality of semantic units are converted into subtitles and displayed one by one. For example, when the display target language sequence is an audio stream, the audio fragment in each semantic unit of the display target language sequence is converted into characters, and each semantic unit is displayed one by one as a subtitle. When the display target language sequence is a text sequence, the collocation or phrase in each semantic unit of the display target language sequence is used as a subtitle, and each semantic unit is displayed one by one.
[0034] In an embodiment of the present disclosure, subtitles corresponding to each semantic unit are displayed one by one, that is, for display with the semantic unit as the granularity, real-time fluctuations and blinking of the subtitles can be avoided, and the stability of subtitle display can be ensured.
[0035] FIG. 2B is a flowchart of a display method according to an embodiment of the present disclosure. As shown in FIG. 2B, the display method 200’ can include operations S210’ to S230’.
[0036] In operation S210’, a source language sequence and a target language sequence are obtained. For example, in a simultaneous interpretation scenario, it is necessary to identify the speaker's voice stream as the source language sequence and translate it into the target language sequence.
[0037] The type of language adopted by the speaker may be Chinese, English, etc., the source language sequence may be Chinese, English, etc., and correspondingly, the target language sequence may be English, Chinese, etc. For example, the speaker speaks in Chinese, generates a Chinese source language sequence, and can use machine translation to convert the Chinese source language sequence into an English target language sequence.
[0038] Both the source language sequence and the target language sequence may be voice streams, both the source language sequence and the target language sequence may be text sequences, or one of the source language sequence and the target language sequence is a voice stream and the other is a text sequence, and the present disclosure does not limit this.
[0039] In operation S220’, at least one of the source language sequence and the target language sequence is divided into a plurality of semantic units.
[0040] For example, a source language sequence can be divided into a plurality of semantic units having semantics according to semantics. For the target language sequence, each semantic unit of the source language sequence can be translated one by one to obtain the semantic units of the divided target language sequence, or after translating the source language sequence into a complete target language sequence, the target language sequence can be divided into a plurality of semantic units having semantics according to semantics.
[0041] In operation S230’, a plurality of semantic units of the source language sequence and / or the target language sequence are converted into subtitles and displayed one by one.
[0042] For the source language sequence, the semantic units of the source language sequence can be converted into subtitles and displayed one by one with the semantic units as the granularity. For example, the plurality of semantic units of the source language sequence are [Hello everyone.], [Today], [I]..., and the display form on the simultaneous interpretation screen is to display [Hello everyone.], [Today], [I]... in order.
[0043] For the target language sequence, the semantic units of the target language sequence can be converted into subtitles and displayed one by one with the semantic units as the granularity. For example, the plurality of semantic units of the target language sequence are [hello everyone.], [Today], [I]..., and the display form on the simultaneous interpretation screen is to display [hello everyone.], [Today], [I]... in order.
[0044] As can be understood, displaying with stable semantic fragments conforms to the browsing habits of users. However, the embodiments of the present disclosure are not limited thereto. For example, at least one of the source language sequence and the target language sequence can be displayed in real time with "characters" or "words" as the granularity. For example, the source language sequence is displayed in real time with "characters" or "words" as the granularity, and the target language sequence is displayed with semantic units as the granularity.
[0045] In an embodiment of the present disclosure, at least one of the source language sequence and the target language sequence is displayed with semantic units as the granularity, which can improve the stability of subtitle display.
[0046] FIGS. 3A to 3C are schematic diagrams of a display method according to an embodiment of the present disclosure. In an embodiment of the present disclosure, the source language sequence is
[0047]
Number
[0048] Taking as an example that both the source language sequence and the target language sequence are displayed on the screen one by one with semantic units as the granularity. FIGS. 3A to 3B show the screens of three consecutive frames on the screen 300 respectively.
[0049] As shown in FIG. 3A, the source language sequence displayed on the screen 300 in the first frame is the semantic unit "[Good morning, everyone.]". The displayed target language sequence is the semantic unit "[Good Morning.]".
[0050]
Number
[0051] Embodiments of the present disclosure display semantic units on the screen as the granularity, without changing and updating the translated content already displayed on the screen, and guarantee the stability of the translation result already displayed on the screen.
[0052] According to an embodiment of the present disclosure, a semantic unit with semantics can be further understood as a "non-ambiguous semantic unit". When a non-ambiguous semantic unit divides a language sequence into a plurality of non-ambiguous semantic units, in the plurality of non-ambiguous semantic units, the translation result of the previous non-ambiguous semantic unit and the translation result of the subsequent non-ambiguous semantic unit are independent of each other, that is, the translation result of the previous non-ambiguous semantic unit does not change with the translation result of the subsequent non-ambiguous semantic unit. In this way, when displaying with the non-ambiguous semantic unit as the granularity, the translated content already displayed on the screen is not changed and updated, and the stability of the translation result already displayed on the screen is guaranteed.
[0053] Furthermore, the non-ambiguous semantic unit refers to the smallest fragment in which the translation result of the previous non-ambiguous semantic unit does not change with the translation result of the subsequent non-ambiguous semantic unit. In this way, it can be ensured that the delay of the system is small.
[0054] FIG. 4 is a schematic system diagram of a method for dividing a language sequence according to an embodiment of the present disclosure into a plurality of semantic units with semantics.
[0055] As shown in FIG. 4, the system 400 includes a semantic unit detection model 410. The language sequence 401 may be a text sequence. The semantic unit detection model 410 is used to classify each character in the language sequence 401 and obtain the tag of each character. The tag of each character indicates whether the text unit composed of the character and the character before it is a semantic unit with semantics. For example, the tag of a certain character in the language sequence 401 is 1, indicating that the text unit composed of the character and the character before it is a semantic unit with semantics. The tag is 0, indicating that the text unit composed of the character and the character before it is not a semantic unit with semantics.
[0056] For example, the language sequence 401 is "Hello everyone. Today I...", and the language sequence 401 is input into the semantic unit detection model 410. The text sequence 402 with the output tags is "Big[0]Family[0]Good[1]Now[0]Day[1]I[1]...". The tag for "Big" is 0, indicating that "Big" is not a semantic unit with semantics. The tag for "Family" is 0, indicating that "[Big Family]" is not a semantic unit with semantics. This is because the translation result of "[Big Family]" changes due to the following influence. The tag for "Good" is 1, indicating that "[Hello everyone]" is a semantic unit with semantics.
[0057] In the text sequence 402 with tags, splitting is performed at the positions where each tag is 1, and the split result sequence 403 of the language sequence 401 can be obtained. The split result sequence 403 includes a plurality of semantic units with semantics, namely "[Hello everyone]", "[Today]", "[I]",... respectively.
[0058] The embodiments of the present disclosure further provide a training method for the semantic unit detection model. FIG. 5 is a flowchart of a training method for a semantic unit detection model according to an embodiment of the present disclosure.
[0059] As shown in FIG. 5, the training method 500 for the semantic unit detection model includes operations S510 to S520.
[0060] In operation S510, a sample language sequence is obtained. For example, the sample language sequence may be an audio stream or a text sequence. The sample language sequence may be obtained from a publicly available dataset or may be authorized by the user.
[0061] When the sample language sequence is an audio stream, multiple elements in the sample language sequence may be multiple audio fragments, each audio fragment has an original tag, and the original tag of each audio fragment indicates whether the audio unit composed of the audio fragment and the previous audio fragment of the audio fragment has a semantic element with semantics.
[0062] When the sample language sequence is a text sequence, multiple elements in the sample language sequence may be multiple characters or words, each character or word has an original tag, and the original tag of each character or word indicates whether the text unit composed of the character or word and the previous audio fragment of the character or word is a semantic unit with semantics.
[0063]
Number
[0064] For example, input the sample language sequence into the initial semantic unit detection model, output the predicted tag of each character in the sample language sequence, calculate a loss (e.g., cross-entropy loss) based on the difference between the original tag of the sample language sequence in the training data and the predicted tag output by the model, adjust the initial semantic unit detection model based on the loss, obtain an updated semantic unit detection model, and based on the updated semantic unit detection model, return to the step of inputting the sample language sequence into the semantic unit detection model for the next sample language sequence, that is, repeat the above training process until the loss between the original tag of the sample language sequence and the predicted tag output by the model meets a preset condition (e.g., the loss converges), stop the training, and obtain a trained semantic unit detection model.
[0065] The language sequence to be segmented can be input into the above trained semantic unit detection model to obtain the tag of each character in the language sequence.
[0066] An embodiment of the present disclosure can train a semantic unit detection model by using a sample language sequence and tags of each character or word in the sample language sequence, and divide the language sequence into multiple semantic units having meanings based on the trained semantic unit detection model.
[0067] FIG. 6 is a flowchart of a method for identifying the original tag of each character in a sample text sequence according to one embodiment of the present disclosure.
[0068] As shown in FIG. 6, the method includes operations S610 to S640. In operation S610, a target text sequence corresponding to the sample text sequence is obtained.
[0069] For example, the sample language sequence may be a sample text sequence, and the target text sequence corresponding to the sample text sequence may be a text obtained by translating the sample text sequence into a target language. In one embodiment, the sample text sequence may be Chinese and the target text sequence may be English.
[0070]
number
[0071] For the previous i words in the sample text sequence, translate the previous i words into an initial target language fragment. For example, translate the first word "上吾" into "Morning", translate the previous two words "上吾" and "10" into "Morning 10", translate the previous three words "上吾", "10" and "点" into "At 10 am", ... translate the previous seven words into "At 10 am, I went to the park". The translation result of the previous seven words is the translation result of the entire sample text sequence, that is, the sample target text sequence.
[0072] In operation S630, the initial target language fragment of the previous i characters is compared with the target text sequence.
[0073] In operation S640, the original tag of the i-th character in the sample text sequence is identified based on the comparison result.
[0074] For example, compare the initial target language fragment "Morning" of the first word with the target text sequence "At 10 am, I went to the park". In "At 10 am, I went to the park", the target language fragment corresponding to the first word is "At 10 am", and "Morning" is different from "At 10 am", i.e. "Morning" does not match "At 10 am, I went to the park". Therefore, the original tag of the first word can be set to 0, indicating that the text unit [上吾] composed of the first character is not a semantic unit with a meaning.
[0075] For another example, compare the initial target language fragment "Morning 10" of the previous two words with the target text sequence "At 10 am, I went to the park". In "At 10 am, I went to the park", the target language fragment corresponding to the previous two words is "At 10 am", and "Morning 10" is not the same as "At 10 am". Therefore, set the original tag of the second word "10" to 0, indicating that the text unit [上帰り10] composed of the previous two words is not a semantic unit with a meaning.
[0076] For example, compare the initial target language fragment of the previous three words, "At 10 am", with the target text sequence "At 10 am, I went to the park". In "At 10 am, I went to the park", the target language fragment corresponding to the previous three characters is "At 10 am", and the initial target language fragment "At 10 am" is the same as the target language fragment "At 10 am". Therefore, set the original tag of the third word, "o'clock", to 1, indicating that the previous three words, "10 am", are a semantic unit with semantic meaning.
[0077] By analogy, the original tags of the fourth to seventh words can be obtained. As can be understood, when the initial target language fragment of the previous i words is the same as the target language fragment corresponding to the previous i characters in the target text sequence, the original tag of the i-th character is 1. When the initial target language fragment of the previous i words is not the same as the corresponding target language fragment of the previous i words in the target text sequence, the original tag of the i-th word is 0.
[0078]
Number
[0079]
Table 1
[0080]
Number
[0081] Figure 7 is a block diagram of a display device according to an embodiment of the present disclosure. As shown in Figure 7, the display device 700 includes a first acquisition module 701, a splitting module 702, and a display module 703.
[0082] The first acquisition module 701 is used to acquire the display target language sequence. The splitting module 702 is used to split the display target language sequence into a plurality of semantic units having semantics.
[0083] The display module 703 is used to convert a plurality of semantic units into subtitles and display them one by one.
[0084] According to an embodiment of the present disclosure, the display target language sequence includes a source language sequence and a target language sequence, and the target language sequence is obtained by translating the source language sequence.
[0085] The splitting module 702 is used to split at least one of the source language sequence and the target language sequence into a plurality of semantic units having semantics.
[0086] The display module 703 is used to convert a plurality of semantic units of the source language sequence and / or the target language sequence into subtitles and display them one by one.
[0087] According to an embodiment of the present disclosure, in a plurality of semantic units having semantics, the translation result of the semantic unit in the previous context and the translation result of the semantic unit in the subsequent context are independent of each other.
[0088] According to an embodiment of the present disclosure, the splitting module 702 is used to split the display target language sequence into a plurality of semantic units having semantics by using a semantic unit detection model.
[0089] According to an embodiment of the present disclosure, the display target language sequence is a display target text sequence, and the splitting module 702 includes a first input unit, a specific unit, and a splitting unit.
[0090] The first input unit is used to input a text sequence to be displayed into a semantic unit detection model and obtain tags for each character in the text sequence to be displayed. The tag of each character is used to indicate whether a text unit composed of the character and at least one character before the character has semantic meaning as a semantic unit.
[0091] The specific unit is used to identify a target character having a target tag in the text sequence to be displayed. The target tag indicates that a text unit composed of the target character of the target tag and at least one character located before the target character has semantic meaning as a semantic unit.
[0092] The splitting unit is used to perform splitting at the positions of each target character in the text sequence to be displayed, obtain a plurality of text units, and use them as a plurality of semantic units having semantic meaning.
[0093] FIG. 8 is a block diagram of a training device for a semantic unit detection model according to an embodiment of the present disclosure.
[0094] As shown in FIG. 8, the training device 800 for the semantic unit detection model includes a second acquisition module 801 and a training module 802.
[0095] The second acquisition module 801 is used to acquire a sample language sequence. The sample language sequence includes a plurality of elements, each element has an original tag, and the original tag of each element indicates whether an element unit composed of the element and at least one element before the element has semantic meaning as a semantic unit.
[0096] The training module 802 is used to train the semantic unit detection model using the sample language sequence and the original tags of the elements in the sample language sequence.
[0097] According to an embodiment of the present disclosure, the sample language sequence is a sample text sequence, each element in the sample language sequence is each character in the sample text sequence, the length of the sample text sequence is L, and L is an integer greater than or equal to 1.
[0098] The training device 800 of the semantic unit detection model further includes a third acquisition module, a translation module, a comparison module, and a determination module.
[0099] The third acquisition module is used to acquire a target text sequence corresponding to the sample text sequence, and the target text sequence is obtained by translating the sample text sequence.
[0100] The translation module is used to translate the previous i characters in the sample text sequence into an initial target language fragment, where i is an integer greater than or equal to 1 and less than or equal to L.
[0101] The comparison module is used to compare the initial target language fragment of the previous i characters with the target text sequence.
[0102] The determination module is used to determine the original tag of the i-th character in the sample text sequence based on the comparison result.
[0103] According to an embodiment of the present disclosure, the comparison module includes a first comparison unit and a second comparison unit.
[0104] When the initial target language fragment of the previous i characters is the same as the target language fragment corresponding to the previous i characters in the target text sequence, the first comparison unit sets the original tag of the i-th character as a positive sample, and is used to indicate that the text unit composed of the previous i characters in the sample text sequence is a semantic unit having semantics.
[0105] If the second comparison unit determines that the initial target language fragment of the previous i characters is not the same as the target language fragment corresponding to the previous i characters in the target text sequence, the original tag of the i-th character is set as a negative sample and is used to indicate that the text unit composed of the previous i characters in the sample text sequence is not a semantic unit having semantics.
[0106] According to an embodiment of the present disclosure, the sample language sequence is a sample text sequence, each element in the sample language sequence is each character in the sample text sequence, and the training module 802 includes a second input unit and an adjustment unit.
[0107] The second input unit is used to input the sample text sequence into the semantic unit detection model to obtain the predicted tag of each character in the sample text sequence.
[0108] The adjustment unit adjusts the parameters of the semantic unit detection model based on the difference between the original tag and the predicted tag of each character in the sample text sequence, and returns to the step of inputting the sample text sequence into the semantic unit detection model for the next sample text sequence until the difference meets a preset condition.
[0109] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.
[0110] FIG. 9 is a schematic block diagram showing an example of an electronic device 900 capable of implementing an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can further represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The members, their connections and relationships, and their functions shown in this specification are merely illustrative and do not limit the implementation of the present disclosure described and / or claimed in this specification.
[0111] As shown in FIG. 9, the device 900 includes a computing unit 901, which can execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 can further store various programs and data necessary for the operation of the device 900. The computing unit 901, the ROM 902, and the RAM 903 are interconnected via a bus 904. An input / output interface 905 is also connected to the bus 904.
[0112] A plurality of components in the device 900 are connected to the I / O interface 905 and include an input unit 906 such as a keyboard and a mouse, an output unit 907 such as various types of displays and speakers, a storage unit 908 such as a magnetic disk and an optical disk, and a communication unit 909 such as a network card, a modem, and a wireless communication transceiver. The communication unit 909 enables the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0113] The computing unit 901 may be various general-purpose and / or dedicated processing modules having processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a GPU (Graphics Processing Unit), various dedicated artificial intelligence (AI) computing chips, computing units of various operation machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes processing with each of the methods described above, such as the display method and the training method of the semantic unit detection model. For example, in some embodiments, the display method and the training method of the semantic unit detection model may be realized as a computer software program tangibly included in a machine-readable medium such as the storage unit 908. In some embodiments, a part or all of the computer program may be loaded and / or installed in the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the display method and the training method of the semantic unit detection model described above may be executed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the display method and the training method of the semantic unit detection model in any other suitable manner (e.g., via firmware).
[0114] The various embodiments of the systems and techniques described in this specification may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be implemented in one or more computer programs, which may be executed and / or interpreted in a programmable system including at least one programmable processor, which may be a dedicated or general purpose programmable processor, and which receives data and instructions from, and transmits data and instructions to, a memory system, at least one input device, and at least one output device.
[0115] The program code for implementing the methods of the present disclosure may be created in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a dedicated computer, or other programmable data processing device, such that, when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the device, may be executed partly on the device, may be executed partly on the device as an independent software package, and may be executed partly on a remote device or entirely on a remote device or server.
[0116] In the context of this disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use in or in combination with an instruction execution system, apparatus, or electronic device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium include electrical connections made with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0117] To provide for interaction with a user, the computer systems and techniques described herein may be implemented on a computer, which includes a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices may be further provided for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form (including voice input, speech input, or tactile input).
[0118] The systems and techniques described herein can be implemented in a computing system that includes background components (such as a data server), or a computing system that includes middleware components (such as an application server), or a computing system that includes front-end components (such as a user computer having a graphical user interface or a web browser, and the user can interact with embodiments of the systems and techniques described herein through the graphical user interface or the network browser), or a computing system that includes any combination of such background components, middleware components, or front-end components. The components of the system can be connected to each other by digital data communication in any form or medium (such as a communication network). Exemplary communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0119] The computer system may include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by a computer program running on the corresponding computer and having a client-server relationship.
[0120] It should be understood that various forms of the flows shown above may be used, and the steps may be sorted, added, or deleted again. For example, each step described in the present invention may be executed in parallel, sequentially, or in a different order, and the present specification is not limited herein as long as the desired results of the technical solution of the present disclosure can be achieved.
[0121] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.
Claims
1. A display method, which is processed by a processor, obtains a display target language sequence, divides the display target language sequence into a plurality of semantic units having semantics, converts the plurality of semantic units into subtitles and displays them one by one, and displays on the screen with the semantic unit as the granularity, and does not perform change update on the translated content already displayed on the screen, the display target language sequence includes a source language sequence and a target language sequence, and the target language sequence is obtained by translating the source language sequence, converting the plurality of semantic units into subtitles and displaying them one by one includes converting a plurality of semantic units of the source language sequence and / or the target language sequence into subtitles and displaying them one by one, dividing the display target language sequence into a plurality of semantic units having semantics includes dividing the source language sequence into a plurality of source semantic units having semantics, converting the plurality of semantic units into subtitles and displaying them one by one making the source semantic unit and the target semantic unit corresponding to each other into a semantic unit pair, and obtaining a plurality of semantic unit pairs, sequentially displaying each semantic unit pair in the plurality of semantic unit pairs. A display method.
2. In the plurality of semantic units having semantics, the translation results of the semantic units in the previous text and the translation results of the semantic units in the subsequent text are independent of each other The method according to claim 1.
3. The display target language sequence is a display target text sequence, using a semantic unit detection model to divide the display target language sequence into a plurality of semantic units having semantics inputting the display target text sequence into a semantic unit detection model, obtaining tags for each character in the display target text sequence, and the tag of each character indicates whether the text unit composed of the character and at least one character before the character has semantics as a semantic unit, identifying a target character having a target tag in the display target text sequence, and the target tag indicates that the text unit composed of the target character of the target tag and at least one character located before the target character has semantics as a semantic unit. Perform splitting at the positions of each target character in the target text sequence to obtain a plurality of text units, and use them as a plurality of semantic units having the semantics The method according to claim 1
4. A first acquisition module used to acquire a target language sequence to be displayed, A splitting module used to split the target language sequence into a plurality of semantic units having semantics, A display module used to convert the plurality of semantic units into subtitles and display them one by one, and includes The semantic units are used as a granularity to be displayed on the screen, and no change update is performed on the translated content already displayed on the screen, The target language sequence includes a source language sequence and a target language sequence, and the target language sequence is obtained by translating the source language sequence, The display module is used to convert a plurality of semantic units of the source language sequence and / or the target language sequence into subtitles and display them one by one, The splitting module is Split the source language sequence into a plurality of source semantic units having semantics, The display module is Pair the source semantic unit and the target semantic unit that correspond to each other into a semantic unit pair to obtain a plurality of semantic unit pairs, Sequentially display each semantic unit pair in the plurality of semantic unit pairs Display device
5. Among the plurality of semantic units having semantics, the translation results of the semantic units in the previous text and the translation results of the semantic units in the subsequent text are independent of each other The device according to claim 4
6. The target language sequence to be displayed is a target text sequence, The splitting module is Input the target text sequence into a semantic unit detection model and use it to obtain tags for each character in the target text sequence. The tag of each character is a first input unit indicating whether a text unit composed of the character and at least one character before the character has semantics as a semantic unit, It is used to identify a target character having a target tag in the target text sequence. The target tag is a specific unit indicating that a text unit composed of the target character of the target tag and at least one character located before the target character has semantics as a semantic unit A splitting unit that splits at the positions of each target character in the target text sequence to obtain a plurality of text units and is used to form a plurality of semantic units having the semantic meaning, and The apparatus according to claim 4. **Claim 7** At least one processor, and A memory communicatively connected to the at least one processor, wherein The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor can execute the method according to claim 1 An electronic device. **Claim 8** A non-transitory computer-readable storage medium storing computer instructions, wherein The computer instructions are used to cause a computer to execute the method according to claim 1 A readable storage medium. **Claim 9** A computer program that realizes the method according to claim 1 when executed by a processor A computer program.
Citation Information
Patent Citations
Machine translation device, method, and program
JP2016071761A