Character conversion support system and method for supporting character conversion of the character conversion support system

The character conversion support system addresses misinterpretations in character conversion by using a client-server model with inference engines to provide accurate and context-aware conversion candidates, adapting to input changes and allowing real-time correction.

JP7845704B2Active Publication Date: 2026-04-14CONTRACT CO BLINDX LAB
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
CONTRACT CO BLINDX LAB
Filing Date
2024-08-07
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing character conversion systems struggle with accurately converting input hiragana or speech to kanji due to the lack of consideration for context and surrounding conversation, often leading to misinterpretations and errors, especially in languages like Japanese with many homophones and polysemous words, and are prone to input errors.

Method used

A character conversion support system that utilizes a client device and a server device to perform real-time character conversion through an inference engine, which infers conversion strings based on user information, input history, and context, employing large language models to provide highly accurate conversion candidates.

Benefits of technology

The system ensures highly accurate conversion of input characters by leveraging multiple inference engines, adapting to changes in input device, application, and input style, reducing misinterpretations and errors, and allowing for real-time correction and re-evaluation of context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007845704000001
    Figure 0007845704000001
  • Figure 0007845704000002
    Figure 0007845704000002
  • Figure 0007845704000003
    Figure 0007845704000003
Patent Text Reader

Abstract

To acquire and output a conversion character string candidate for input character information from a server device.SOLUTION: The character conversion support system includes client devices CL1 to CLN each having a front end processor that performs character-string processing, and a server device 1 that acquires character strings from the front end processors, wherein the client devices CL1 to CLN are configured to transfer character information input to the front end processors to the server device 1, acquire character-string candidates that are sequentially converted by an inference engine that uses a large-scale language model, and determine the character-string candidates as character strings to be outputted by an outputting unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a character conversion support system in which a client device that can be connected to a specific domain via a predetermined communication medium, a client application that is launched on the client device, and an application server that is launched on a server device cooperate with each other, and a character conversion support method for the character conversion support system.

Background Art

[0002] Conventionally, in a computer system, for character information processing, a FEP (Front End Processor) that converts a key code input while referring to a dictionary library into character information intended is incorporated under the control of an OS. In particular, the FEP (Front End Processor) in character processing plays an important role in a Japanese input system.

[0003] The FEP is responsible for a process of receiving an input from a keyboard and converting it into Japanese characters such as hiragana or kanji. This process includes a step of converting the input roman characters into hiragana and presenting candidate kanji characters and words until the user makes a selection. In addition to the above, the functions of the FEP include a function of presenting and selecting candidate words, and a function of supporting a plurality of input methods (e.g., roman character input, kana input).

[0004] On the other hand, there is a problem that the FEP performs character conversion unintended by the user. This is particularly prominent in the case of technical terms, new words, foreign words, etc.

[0005] Patent Document 1, described below, discloses that, "In order to enable input and display of various characters even when limited computing resources are available, the web browser of client 30 accesses the web server of dictionary center 20 to receive a character input service. The web browser inputs the reading of the character, the web server of dictionary center converts this into a character code, and further converts it into a URL of an image file on the web server of font center 10. As a result, an HTML document containing the URL of the character image file as the source is sent to the browser, the browser retrieves the image file from the web server of font center 10 based on the URL, and displays the character using the retrieved image."

[0006] Furthermore, Patent Document 2 discloses that, in order to convert audio files into text and identify speakers using a learning-type server via the Internet, the speech recognition system 1 consists of a sound collection unit 2 that inputs audio, a processing unit 3 that generates an audio file F from this audio, transmits this audio file F to a character conversion server and a speaker identification server, and receives a text file W and the results of speaker identification H, and a monitor unit 4 that displays the processing results of the processing unit 3. The audio data V separated by the silent periods between speaker H's statements is sent as an audio file F to a character conversion server 5 and a speaker identification server 6, which are cloud services that perform data analysis and interpretation through self-learning based on data collected from a large number of users via the Internet. The received text file W and the results of speaker identification H are then associated with the audio file F and displayed in chronological order on the monitor unit 4. [Prior art documents] [Patent Documents]

[0007] [Patent Document 1] Japanese Patent Publication No. 2001-350747 [Patent Document 2] Japanese Patent Publication No. 2020-060735 [Overview of the project] [Problems that the invention aims to solve]

[0008] It is widely known that in character conversion systems that process Japanese, it is difficult to obtain a perfect sentence from input hiragana, speech, or gestures alone when converting them to kanji. Appropriate expressions cannot be generated without considering a wide range of surrounding context and the situation of the conversation.

[0009] In particular, in languages ​​like Japanese, which have many homophones and polysemous words based on ideographic characters, the input sentence "senshu ga senshu o naoshimashita" (the bow of the ship repaired the bow) can be interpreted as either "the player healed the player" or "the player treated the player," depending on the context. In addition, in actual use, input errors such as typos are often mixed in with the initial input.

[0010] For example, if "senshu ga senshu o naoshimashita" (the captain repaired the captain) is converted to "senshu ga senshu o naoshimashita" (the bow repaired the bow), or if it's a typo and it's written as "senshu ga senshu o naoshimashita" (the ship's bow repaired last week), it might be converted to "sanshu ga senshu o naoshimashita" (last week repaired the ship's owner). In this way, even slight differences in input can greatly change the interpretation (the converted string).

[0011] In particular, when various users connect to a specific domain from client devices and input text information on the provided UI screen, there are issues such as the text not being converted to the string the user intended, or the text conversion process taking a considerable amount of time when long texts are entered.

[0012] The present invention solves the above problems and provides a character conversion support system and a character conversion support method for the character conversion support system, which enable the same user operating the client device to obtain and output highly accurate conversion string candidates for input character information even when the attributes of the input device, the attributes of the application to be launched, the attributes of the service to be launched, and the input method and input style of character information change, by employing conversion candidate strings obtained from an inference engine that infers conversion strings provided by a server device connected to the network for the functional processing of the front-end processor of the client device. [Means for solving the problem]

[0013] The character conversion support system of the present invention, which achieves the above objective, has the following configuration.

[0014] The character conversion support system according to the present invention is a front-end processor that performs string processing via a predetermined communication medium. and an input unit A character conversion support system in which a client device having and a server device that obtains a string from the front-end processor communicate and cooperate in string conversion processing, wherein the client device is The front-end processor acquires the character string and user information input by the input unit, and each time a character is input from the input unit, these are sequentially transferred to the server device. Each time of the transfer, the server device identifies the context based on the input history including the user information and the character string sequentially input from the input unit, rewinds to the beginning part of the sentence input in the past to re-evaluate the context, and based on the re-evaluation result, executes a conversion process in real time by any one of a plurality of inference engines using a large language model, and returns a conversion character string candidate generated by reflecting the identified context and the character string sequentially input from the input unit to the client device as a response. The client device performs a process of displaying the returned conversion character string candidate It is characterized by having a configuration. [Effects of the Invention]

[0015] According to the present invention, by employing conversion candidate strings obtained from an inference engine that infers conversion strings provided by a server device connected to the network, the functional processing of the front-end processor of the client device can be performed by the same user operating the client device, and even if the attributes of the input device, the attributes of the application to be launched, the attributes of the service to be launched, and the input method and style of the character information change, the input character information can be obtained and output with high accuracy as a conversion string candidate. [Brief explanation of the drawing]

[0016] The drawings illustrate specific embodiments of the present invention and include not only essential components of the invention but also optional and preferred embodiments. Block diagram for explaining the configuration of the character conversion support system showing this embodiment.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0017] Next, the best mode for carrying out the present invention will be described with reference to the drawings.

[0018] <Explanation of System Configuration> 〔First Embodiment〕 FIG. 1 is a block diagram for explaining the configuration of the character conversion support system showing this embodiment. In this example, client devices CL1 to CLN that can be connected to a specific domain via a predetermined communication medium, a client application 100 that is launched on the client devices CL1 to CLN, and an application server 11 that is launched on the server device 1 cooperate to support character conversion processing, taking this character conversion support system as an example.

[0019] In FIG. 1, reference numeral 1 denotes a server device that communicates with the client devices CL1 to CLN via a network 21 using a predetermined protocol.

[0020] In server device 1, 11 is an application server that communicates with client applications on client devices CL1 to CLN via network 21. 12 is a Domain Knowledge Server Selection Unit (DKS) that receives user information, service information (uniquely assigned service numbers (Universal Unique IDentifier)), application information, document properties, etc., from a front-end processor (FEP) 102 that performs character conversion processing based on a predetermined language, using information input by the input unit 103 of client devices CL1 to CLN. The DKS 12 then executes a process to select one candidate from the knowledge server group 13 for the second inference processing in the character conversion process.

[0021] 50 is a Domain Name Server (DNS) that provides IP addresses corresponding to domain names entered from the input section 103 of client devices CL1 to CLN.

[0022] Client devices CL1 to CLN connect to specific URLs 1 to URLN via network 22, routed according to the IP addresses (xxx.xxx.xxx) specified by DNS 50. These specific URLs 1 to URLN cover a wide range of genres, and it has been observed that when engaging in conversations or chats appropriate to each URL, specific terminology is likely to be selected, and there are certain trends in the strings used in these conversations.

[0023] In this embodiment, the FEP 102 requests the server device 1 to perform a second inference process in the background, which infers character information in response to the domain name entered by the user of client devices CL1 to CLN from the input unit 103.

[0024] Figure 2 is a block diagram illustrating the software functions based on the hardware resources of the character conversion support system shown in Figure 1, and components identical to those in Figure 1 are denoted by the same reference numerals.

[0025] In Figure 2, the server device 1 comprises multiple second inference units 13-1 to 13-N, each including multiple knowledge servers 13-1-1 to 13-1-N and profiles 13-2-1 to 13-2-N.

[0026] In client devices CL1 to CLN, 101 is a switching unit that performs a process to switch between performing the first inference process locally during character conversion or performing the first inference process on the server device 1 side during character conversion, based on the information (which includes the domain name) input by the input unit 103.

[0027] The FEP102 comprises a first inference unit 102-1 and a storage unit 102-2 that holds user profile information (including a user dictionary). The storage unit 102-2 is configured to hold user information that is transferred to the server device 1. In this embodiment, the user information includes the properties of the document used by the user, the attributes of the service used by the user, and the attributes of the application used by the user, as character processing methods used by the user. Furthermore, the user information held in the storage unit 102-2 is kept in an updatable manner so as to adapt to the environment in which the strings entered by the user are used.

[0028] Figure 3 is a block diagram illustrating the hardware configuration of server device 1 shown in Figure 1. Components identical to those in Figure 1 or Figure 2 are denoted by the same reference numerals.

[0029] In Figure 3, 31 is the communication unit, which performs functional processing to communicate with client devices CL1 to CLN connected to the network 21 using a predetermined protocol.

[0030] 32 is the CPU, which loads the application server 11 (various application programs including the operating system) stored in external memory 33 into RAM 36 and executes it.

[0031] 34 is the inference engine unit, where the second inference units 13-1 to 13-N are deployed. The second inference units 13-1 to 13-N constantly collect and update large amounts of natural language texts and usage examples (corpus) from the social domains in which the user interacts, and are maintained so that the degree of matching of the character conversion candidates evolves.

[0032] 36-1 is the DKS unit, which performs a functional process via the network 21 to pass conversion string candidates inferred by one of the knowledge servers 13-1-1 to 13-1-N selected by the CPU 32 from the inference engine unit 34 to the FEP 102 of one of the client devices CL1 to CLN. This functional process is continuously performed while the user operating the client devices CL1 to CLN is inputting a string via the input unit 310.

[0033] Unit 39, located in the AI ​​Support Department, executes processes that support innovative character conversion in the field of natural language processing (NLP), using large-scale language models (LLMs) built with vast amounts of text data and advanced deep learning techniques. For large-scale language models, T5 (Text-to-Text Transfer Transformer) is used. Numerous derived models based on this architecture have been proposed. T5 is a series of Transformers announced in 2017, and it learns by taking hiragana text as input and outputting text converted to kanji.

[0034] Furthermore, the inference engine of the knowledge server group 13 provided by server device 1 works in cooperation with the AI ​​support unit 39 to perform processing to improve the accuracy of conversion string candidates for kana characters acquired sequentially from client devices CL1 to CLN.

[0035] Alternatively, the functional processing of the AI ​​support unit 39 may be integrated into the inference engine.

[0036] LLM training uses "supervised learning," where a hiragana sentence is taken as input and a corresponding set of kanji sentences (corpus) is provided as training material. In other words, a hiragana sentence is given as input, the output of T5 described above is calculated, and the loss is calculated by comparing it with the training kanji sentences. Based on this loss, the parameters of each parameter of the model are updated according to the gradient of each parameter.

[0037] The inference engine has a corpus for each language community. For example, in formal settings, it uses texts from newspapers, magazines, and books, while in casual settings, it uses conversational texts and social media texts.

[0038] The corpus is constantly updated. This means the inference engine doesn't learn from the user's history, but from the entire society to which the user belongs. To achieve this, the inference engine constantly collects a wide range of information, automatically extracting the corresponding hiragana sentences from mixed kana and kanji sentences through morphological analysis, and then combining them with the original sentence to continuously create "kana sentence and kanji sentence pairs." By collecting, adding to, and updating a large number of these pairs, the learning process is deepened.

[0039] Figure 4 is a block diagram illustrating the hardware resources and software functions of client devices CL1 to CLN shown in Figure 1. Components identical to those in Figure 1 are denoted by the same reference numerals. Client devices CL1 to CLN are assumed to represent various data terminals, including PAD devices, personal computers (PCs), and smartphones.

[0040] Furthermore, in this embodiment, when the kana character conversion mode is selected, the FEP102, via the communication unit 304, transfers the key code input from the input unit 310 to the DKS12 of the server device 1 connected to the network 21, along with the key code information and user information (details will be described later).

[0041] DKS12 selects a second inference unit 13-1 to 13-N to determine a conversion candidate string according to the key code information and user information received from one of the client devices CL1 to CLN. The response unit 36-2 then sends the Japanese string data containing kanji characters converted by one of the second inference units to the FEP 102 of the requesting client device CL1 to CLN via the communication unit 31.

[0042] In this case, FEP102 is controlled to output the string candidates converted by the server device 1 to the text input line on the display screen of display 311, as if they were string candidates converted by itself.

[0043] In this case, server device 1 can freely select one of the second inference units 13-1 to 13-N according to the TPO (T for Time, P for Place, and O for Occasion) of the user operating client devices CL1 to CLN, or it can flexibly switch to string conversion processing by other second inference units even while conversion is in progress.

[0044] Furthermore, DKS12 is configured to recognize, for example, whether the Japanese input entered by the user operating client devices CL1 to CLN is polite Japanese, or special phrases used on social media (strings not used in normal formal documents), such as "buzz" (buzzing), and to allow for the dynamic switching of string conversion processing by a second inference unit even while conversion is in progress.

[0045] In Figure 4, 301 is the CPU, which starts the BIOS stored in ROM 302, controls the I / O of input and output devices, and executes various data processing by running various applications loaded into RAM 303.

[0046] 304 is the communications unit, which communicates with the server device 1 via the network 21 using a predetermined protocol and performs various data processing in the network environment.

[0047] 311 is a display, where kana character information entered by the input unit 310 is passed to the server device 1, and the conversion string candidates generated by the second inference units 13-1 to 13-N of the server device 1 are continuously received and displayed so that the character information is arranged at the position specified by the cursor.

[0048] While the user operates the input unit 310, they continue to receive conversion candidate strings from the server device 1. If the user performs an action to correct the displayed conversion string, the server device 1 further selects one of the knowledge servers 13-1-1 to 13-1-N to handle the most optimal second conversion process.

[0049] [Data processing on the server side] Figure 5 is a flowchart showing an example of a character conversion support method for the character conversion support system illustrating this embodiment. Note that (1) to (9) are data processing on the server device 1 side shown in Figure 1, which is realized by the CPU 32 loading the control program stored in the external memory 33 onto the RAM 36 and processing it.

[0050] First, the CPU 32 obtains user information (including domain information) from the FEP 102 of one of the client devices CL1 to CLN shown in Figure 1, and kana character information input from the input unit 310 (1).

[0051] The CPU32 passes the kana character information and user information obtained from the client devices CL1 to CLN shown in Figure 1 to the DKS12 (2), and the DKS12 selects one of the knowledge servers 13-1-1 to 13-1-N based on the acquired domain information (3).

[0052] Next, the CPU 32 passes the string candidate converted by the selected knowledge server via the network 21 to the FEP 102 of one of the client devices CL1 to CLN (4).

[0053] Next, the CPU 32 determines whether it has received a correction instruction or a re-conversion instruction for the conversion string candidate passed from the FEP 102 of any of the client devices CL1 to CLN (5). If the CPU 32 determines that it has received a correction instruction or a re-conversion instruction, it then determines whether it will continue to select one of the knowledge servers selected in step (3) (6). If the CPU 32 determines that it will not select the same knowledge server, it then controls the DKS 12 to select one of the knowledge servers 13-1-1 to 13-1-N (7). The CPU 32 then returns the process to step (4).

[0054] On the other hand, if in step (6) the CPU 32 decides to continue selecting one of the same knowledge servers 13-1-1 to 13-1-N that was selected in step (3), it returns to step (4).

[0055] On the other hand, in step (5), if a certain amount of time has elapsed since the correction instruction or re-conversion instruction for the conversion string candidate handed over from the FEP102 of any of the client devices CL1 to CLN (8), the CPU 32 determines whether it has continued not to acquire user information (including domain information) from any of the FEP102 of any of the client devices CL1 to CLN shown in Figure 1 and text information input from the input unit 310, that is, whether it has received kana character information from any of the FEP102 of any of the client devices CL1 to CLN (9).

[0056] If the CPU 32 determines that it has received kana character information from the FEP 102 of any of the client devices CL1 to CLN, it returns to step (2) and repeats the same process. On the other hand, if the CPU 32 determines in step (9) that it has not received kana character information from the FEP 102 of any of the client devices CL1 to CLN, it terminates this process.

[0057] [Data processing on client devices CL1 to CLN] Figure 6 is a flowchart showing an example of a character conversion support method for the character conversion support system illustrating this embodiment. Note that (11) to (19) are data processing on the client device CL1 to CLN side shown in Figure 1, which is realized by the CPU 301 loading a control program stored in an external memory (not shown) onto the RAM 303 and processing it.

[0058] First, when each user operating client devices CL1 to CLN instructs access to any of the domains, the CPU 301 launches a browser from external memory (not shown) (11). At that time, each user operates the keyboard of the input unit 310 to input kana characters, for example using the Romanization method (12), and the FEP 102 outputs the kana characters that are being entered sequentially to the server device 1 (13).

[0059] In this process, FEP102 detects user information (document properties used by the user, services used by the user, and applications used by the user) held in the holding unit 102-2 (14), and transfers the detected user information to the server device 1 (15).

[0060] Next, when the CPU 301 obtains a candidate for conversion string from the knowledge server selected by the server device 1 (one of the second inference units 13-1 to 13-N) via the network 21 (16), it displays the converted string on the string input line (indicated by the cursor) specified by the browser on the display 311 (17).

[0061] Next, the CPU 301 determines whether each user has given an editing instruction such as correcting, modifying, or deleting the converted string displayed on the browser of the display 311 (18). If it determines that an editing instruction has been given, it returns to step (13).

[0062] On the other hand, in step (18), if the CPU 301 determines that no editing instruction has been given, it further determines whether kana characters have been entered from the input unit 310 (19). If the CPU 301 determines that no character code has been entered, it terminates the process of obtaining string candidates by the server device 1. If it determines that a character code has been entered, it returns to step (13) and repeats the same process.

[0063] [Effects of the First Embodiment] According to this embodiment, the functional processing of the front-end processor 102 of client devices CL1 to CLN is performed using conversion candidate strings obtained from multiple inference engines (inference engines utilizing LLM) that infer conversion strings provided by the server device 1 connected to the network 21. This allows the same user operating client devices CL1 to CLN to obtain and output highly accurate conversion string candidates for the input character information, even when the attributes of the input device, the attributes of the application being launched, the attributes of the service being launched, and the input method and style of the character information change in various character processing environments.

[0064] Furthermore, the system operator may freely choose whether to offer the free service of providing string conversion candidates provided by the server device 1 to general users, or to offer it as a paid service with customized functions for specific users.

[0065] Furthermore, while the system described above uses Japanese as an example of a language model, it is also possible to target major foreign language models for the language model. Furthermore, in environments where users of different languages ​​exchange text information, such as through chat, it is possible to freely construct an environment where similar string conversion processing can be performed by using a system that supports translation models.

[0066] [Second Embodiment] In the above embodiment, an example was described assuming that short sentences are input from client devices CL1 to CLN, but of course, the present invention can also be applied to longer contexts.

[0067] In the above system, if server device 1 interprets a long context, it is expected that the response speed until the conversion is visible will be delayed, in exchange for higher conversion accuracy.

[0068] Therefore, each time a user operating client devices CL1 to CLN reads a single character (1 token) input from the input unit 310, the FEP 102 immediately sends the input character code and string conversion request to the second inference units 13-1 to 13-N of server device 1 via the network 21.

[0069] In this way, when server device 1 receives a character code and string conversion request, the second inference units 13-1 to 13-N respond to client devices CL1 to CLN with the conversion process to a string and the character code indicating the generated string.

[0070] In this process, server device 1 controls the string conversion process so that it rewinds to the beginning of each paragraph (translation unit) of the document before starting.

[0071] Therefore, the beginning of each paragraph will be repeatedly converted to a string (rewinded), but since the input characters become longer with each conversion, it is expected that the conversion accuracy will improve.

[0072] As a result, client devices CL1 to CLN will display converted strings in a way that, as input progresses, words that were initially mistranslated will automatically be corrected and rewritten appropriately.

[0073] With this type of conversion process, if the character input processing on the client devices CL1 to CLN is fast, the burden on the string inference and conversion processing on the server device 1 increases. Therefore, the server device 1 controls the system to temporarily hold requests from each user in storage until the ongoing string conversion process is completed.

[0074] Then, when starting the next process, instead of executing the pending requests sequentially, the system executes only the last request and deletes any previously pending requests.

[0075] [Effects of the second embodiment] As a result, although it appears to client devices CL1 to CLN that server device 1 is converting one character at a time in parallel, server device 1 only needs one thread per user.

[0076] Furthermore, since client devices CL1 to CLN intermittently obtain the results of the string conversion process for long texts from server device 1, the display of client devices CL1 to CLN is controlled so that any newly added hiragana strings from the last string sent to server device 1 are concatenated to the end of the last converted text.

[0077] In this way, even if there is conversion latency in the server device 1 during the long chunk inference process, the CPU 301 controls the display on the client device CL1~CLN's display 311 so that the hiragana input is automatically converted into appropriate sentences without delay.

[0078] The rewind process is generally performed on a paragraph-by-paragraph basis, but it can be adjusted at a minimum of one sentence and at a maximum of the entire document being processed, based on the additional status on the server device 1.

[0079] [Third Embodiment] In the above embodiment, an example was described assuming that short text is input from client devices CL1 to CLN. However, the display 311 of client devices CL1 to CLN may be configured to control the re-conversion process even if the conversion process has already been completed for converted strings, etc.

[0080] As mentioned above, even after performing various string conversion processes, it is conceivable that contextual misinterpretations may still remain.

[0081] In this case, the system may be configured to allow each user operating client devices CL1 to CLN to utilize the intermediate results of the inference process by controlling the manner of the string conversion process, thereby maximizing the use of the intermediate results and reducing the burden of string input processing on the user.

[0082] Specifically, in server device 1, the second inference units 13-1 to 13-N reserve a conversion buffer on memory resources (not shown), and buffer all the plaintext before the current paragraph conversion (raw hiragana strings entered by each user and the trajectory of flicks) on this buffer, so that it is possible to roll back (so-called undo) even after the conversion has been confirmed.

[0083] Therefore, in server device 1, the second inference units 13-1 to 13-N maintain a confidence level (0.0 to 1.0 (probability notation)) of the inference result for each phrase, and instruct the FEP102 of client devices CL1 to CLN to highlight a part of the generated document according to this confidence level.

[0084] [Effects of the third embodiment] This allows each user of client devices CL1 to CLN to quickly identify the areas to be corrected on the screen of display 311, return to that point, selectively revert to the pre-conversion state on a phrase-by-phrase basis, and then proceed to the appropriate correction work.

[0085] Therefore, unlike simple kana-kanji conversion, this embodiment allows users to rewind and edit even after the character conversion process has been finalized, thus reducing the psychological burden on each user when rewinding.

[0086] Note that the highlighting function based on the confidence level mentioned above is not mandatory. Alternatively, when a user manipulates the cursor on the screen, control may be implemented to rewind only the relevant phrase and selectively display only the string before conversion.

[0087] Furthermore, the string conversion inference process performed by the second inference units 13-1 to 13-N in server device 1 may involve simple input errors or insufficient expression in the original text or conversation.

[0088] Therefore, even if only the grammatical errors are corrected, by controlling the system to allow re-evaluation on a paragraph-by-paragraph basis, the improved accuracy of the correction from one part can ripple through to other parts, automatically proofreading the surrounding text.

[0089] [Fourth Embodiment] In the above embodiment, an example was described assuming that short sentences are input from client devices CL1 to CLN. However, when each user inputs text from the input unit 310, a lightweight first inference unit 102-1 may be provided on the client devices CL1 to CLN to flexibly handle the situation, such as removing typos and noise before the user input.

[0090] In this way, instead of relying solely on the string conversion inference processing by the server device 1, the first inference unit 102-1 on the client devices CL1 to CLN cooperates with the inference processing on the server device 1. This eliminates user-specific key input errors, restricts the transfer of user-specific strings that interfere with inference to the second inference units 13-1 to 13-N to the server device 1, and prevents contamination of the inference processing in the second inference units 13-1 to 13-N, thereby enabling the acquisition of highly accurate string conversion processing results from the server device 1.

[0091] [Effects of the fourth embodiment] In this case, the local first inference unit 102-1 on the client device CL1 to CLN side, assuming the input unit 310 is a keyboard, does not target the Romanized string but directly acquires information about the keystrokes themselves and applies a predetermined inference process specific to each user, thereby enabling it to extract a more accurate string and reflect it in the inference process on the server device 1 side.

[0092] Thus, when the character input mode can selectively switch between direct input using a keyboard and flick input, by selecting the second inference units 13-1 to 13-N according to each input mode, an improvement in character string conversion accuracy can be expected for each input mode.

[0093] 〔Effect of the Fourth Embodiment〕 In the above embodiment, an example assuming the case where short texts are input from the client devices CL1 to CLN has been described. However, when each user performs key input from the input unit 310, in order to perform cleansing such as correcting typos and removing noise in the front stage of user input, a lightweight first inference unit 102-1 may be provided on the client devices CL1 to CLN side so as to be configured to be able to handle flexibly.

[0094] Note that the first inference unit 102-1 is secured for each user on the client devices CL1 to CLN side and is provided with a dedicated personal dictionary memory.

[0095] 〔Fifth Embodiment〕 In the above embodiment, an example assuming the case where short texts are input from the client devices CL1 to CLN to the keyboard of the input unit 310 has been described. In particular, key touches have individual differences for each person in terms of the movement of the fingertips and the speed of the fingertips hitting. For example, by measuring and learning the fluctuations in input and the frequency of typos specific to Roman character input and executing control to reflect them in the character string conversion candidates, an improvement in conversion accuracy can be expected.

[0096] In particular, for the input of "ぃ", since the characteristics of key input types are different between users who input "xi" and users who input "li", control is executed so that both "船員" and "船員n" output conversion character string candidates of "船員".

[0097] 〔Effect of the Fifth Embodiment〕 According to the present embodiment, in the second inference units 13-1 to 13-N that use the LLM on the server device 1 side, input character strings that may cause large error generation can be corrected in advance.

[0098] Furthermore, in client devices CL1 to CLN that employ an input unit 310 for flick input of strings, it becomes possible to correct errors specific to touch mistakes in the placement of kana characters.

[0099] [Sixth Embodiment] In this embodiment, when a keyboard is used in client devices CL1 to CLN, the correlation between the time interval of keystrokes and the delimiter of a phrase (token) may be used to control the execution of the following process when the length of the keystroke interval exceeds a certain threshold.

[0100] Specifically, in server device 1, when each user of client devices CL1 to CLN uses the keyboard, if the length of the interval between keystrokes exceeds a certain threshold, a control is executed to insert a hidden delimiter.

[0101] [Effects of the 6th Embodiment] According to this embodiment, although user input is not echoed, by inserting the above-mentioned delimiter when the string is internally broken down into phrases, it is possible to suppress a chain reaction of misinterpretations that arises from errors in phrase division that span conjunctions such as "wa" and "e," which are likely to occur when converting long sentences all at once. Furthermore, the same issue is expected to occur with voice input as well, so the same countermeasures can be expected to have the same effect.

[0102] [Seventh Embodiment] In this embodiment, if there are multiple application services used by the same user in client devices CL1 to CLN, it is expected that the input tendencies will change even for the same user. Therefore, the server device 1 may be configured to perform control to switch the selection order among the second inference units 13-1 to 13-N that utilize LLM, so as to use the higher-ranking second inference unit.

[0103] Specifically, for example, when the same user performs text input for business purposes, the system will control the selection order so that it uses the highest-ranking second inference unit among the second inference units 13-1 to 13-N to select the target document name and the strings obtained during the process of converting from the beginning to the paragraphs.

[0104] [Effects of the 7th Embodiment] According to this embodiment, from the perspective of each user operating client devices CL1 to CLN, it is expected that natural-sounding text will be generated according to the type of application being run, even though they are creating documents using the same interface from the same document creation software or spreadsheet software. In addition, when RAG, which combines external information retrieval, is adopted, it becomes unnecessary to go through AI chat during the document creation process, improving convenience. Here, RAG is an abbreviation for Retrieval-Augmented Generation, which means search augmentation generation.

[0105] On the other hand, for conversation and social media use, it becomes possible to perform long string inference processing across conversations, and simultaneously improve the accuracy of proper noun conversion by taking life logs.

[0106] Furthermore, by having the client application 100 provide the conversation partner's text, it becomes possible to include the previous conversation text from the previous partner or group in the string conversion inference. This is expected to effectively suppress misinterpretations and accidental messages that frequently occur in conversation and social networking applications, especially because each message is short.

[0107] Furthermore, in the character conversion support system, the functions of the second inference units 13-1 to 13-N in server device 1 and the first inference unit 102-1 in client devices CL1 to CLN are clearly separated. This configuration eliminates the need for server device 1 to acquire user-specific user information (information derived from typing habits) and manage its history, thus preventing information linked to individual attributes (individual profiles) from leaking from server device 1.

[0108] Furthermore, if it is necessary to refer to the activity history of each user of client devices CL1 to CLN for business purposes, it is also possible to incorporate a strict account management function into the system independently of this. Furthermore, information linked to individual attributes is limited to low-level statistical information such as typing habits, and this information is stored only on the individual client devices CL1 to CLN, and can be controlled so that server device 1 does not need to know about it.

[0109] Furthermore, in this embodiment, the output unit that outputs the string candidates converted by the server device 1 is not limited to a display unit, but may be a device that outputs the converted string candidates as audio information in each language (speaker, earphone with microphone, goggle device, etc.).

[0110] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program (a program conforming to the flowcharts shown in Figures 5 and 6 above). It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0111] The disclosure relating to the present invention described above can be summarized to at least the following:

[0112] (1) A front-end processor that performs string processing via a predetermined communication medium.

Figure 6

[0116] ( 2 ) The front-end processor acquires the character string and user information input by the input unit, and each time a character is input from the input unit, these are sequentially transferred to the server device. Each time of the transfer, the server device identifies the context based on the input history including the user information and the character string sequentially input from the input unit, rewinds to the beginning part of the sentence input in the past to re-evaluate the context, and based on the re-evaluation result, executes a conversion process in real time by any one of a plurality of inference engines using a large language model, and returns a conversion character string candidate generated by reflecting the identified context and the character string sequentially input from the input unit to the client device as a response. The client device performs a process of displaying the returned conversion character string candidate The text is entered via keyboard input or flick input.

[0117] ( 3 The server device is characterized by having a configuration that selects one inference engine from among a plurality of inference engines based on the character processing pattern included in the user information.

[0118] ( 4 The character processing methods included in the user information are characterized by including the properties of the document used by the user, the attributes of the service used by the user, and the attributes of the application used by the user.

[0119] ( 5 ) A front-end processor that performs string processing via a predetermined communication medium. input by the input unit A client device having [a certain feature] and a server device that obtains a string from the front-end processor communicate to coordinate string conversion processing. and an input unit A character conversion support method, sentence Note: Client device before in , The front-end processor acquires the string and user information input by the input unit, and sequentially transfers them to the server device each time a character is input from the input unit. The server device identifies the context based on the input history, including the user information and the strings sequentially input from the input unit, and re-evaluates the context by rewinding to the beginning of previously input sentences. Based on the re-evaluation result, it performs a real-time conversion process using one of multiple inference engines employing large-scale language models. It returns a conversion string candidate, generated reflecting the identified context and the strings sequentially input from the input unit, as a response to the client device. The client device then performs a process to display the returned conversion string candidate. It is characterized by the following: [Industrial applicability]

[0120] In the system described above, the language conversion process is initiated by passing a key code to FEP102 installed on client devices CL1 to CLN. However, this is a method mainly used as preprocessing in the field of natural language processing, and it is also possible to construct a system that incorporates an analysis means, such as MeCab, which decomposes a text into morphemes (the smallest units in which a word has meaning) based on the grammar and part-of-speech information of the target language.

[0121] Furthermore, the multiple second inference units 13-1 to 13-N on the server device 1 learn the converted strings and the degree of modification, either independently or in cooperation with each other. This allows them to convert to specific strings that are tailored to each user's requirements, enabling them to evolve to present a timely array of conversion string candidates that are adapted to user-specific expressions and offer a wide range of versatile conversion candidates.

[0122] Furthermore, since the language types that are valid on client devices CL1 to CLN are already set up, by arranging them on the server device 1's multiple second inference units 13-1 to 13-N according to the language used, it is possible to build a system that incorporates worldwide translation processing even when different languages ​​are used between client devices.

[0123] Furthermore, by adding features such as automatic typographical error correction, OS-specific language processing functions, automatic conversion trigger functions, customization options for individual user environments, context recognition functions, modeless functionality, and the ability to independently evolve the DNN at the appropriate time, it is possible to further improve user convenience. [Explanation of symbols]

[0124] 1 Server device CL1~CLN Client Devices 21 Network 50 DNS

Claims

1. A character conversion support system in which a client device having a front-end processor and input unit for character processing and a server device that acquires character strings from the front-end processor communicate via a predetermined communication medium to coordinate character conversion processing, The aforementioned client device The front-end processor acquires the string and user information input by the input unit. Each time a character is input from the input unit, these characters are sequentially transferred to the server device. The server device is Each time the aforementioned transfer is performed, the context is identified based on the user information and the input history, which includes strings sequentially entered from the input unit. Rewind to the beginning of the previously entered sentence and re-evaluate the context. Based on the re-evaluation results, the conversion process is performed in real time using one of several inference engines that employ large-scale language models. The system returns a candidate for a converted string, generated in accordance with the identified context and the strings sequentially input from the input unit, to the client device as a response. The character conversion support system is characterized by having a client device that performs processing to display the returned conversion string candidates.

2. The character conversion support system according to Claim 1, characterized in that the character entered by the input unit is entered by keyboard input or flick input.

3. The server device is The character conversion support system according to claim 1, characterized in that it includes a configuration to select one inference engine from among a plurality of inference engines based on the character processing method included in the user information.

4. The character processing method included in the user information is characterized in that it includes the properties of the document used by the user, the attributes of the service used by the user, and the attributes of the application used by the user, as described in claim 1.

5. A character conversion support method in which a client device having a front-end processor and input unit for character processing and a server device that acquires a character from the front-end processor communicate via a predetermined communication medium to coordinate character conversion processing, In the aforementioned client device, The front-end processor acquires the string and user information input by the input unit. Each time a character is input from the input unit, these characters are sequentially transferred to the server device. In the aforementioned server device, Each time the aforementioned transfer is performed, the context is identified based on the user information and the input history, which includes strings sequentially entered from the input unit. Rewind to the beginning of the previously entered sentence and re-evaluate the context. Based on the re-evaluation results, the conversion process is performed in real time using one of several inference engines that employ large-scale language models. The system returns a candidate for a converted string, generated in accordance with the identified context and the strings sequentially input from the input unit, to the client device as a response. In the aforementioned client device, A character conversion support method characterized by performing a process to display the returned conversion string candidates.

Citation Information

Patent Citations

  • Device and method for processing character

    JP2001350747A

  • Communication device, communication system, and character string conversion method

    JP2011215911A

  • Voice recognition system

    JP2020060735A