Chinese character input system automatic prediction method and device based on large model
Through large-scale model pre-training and direct preference optimization algorithm fine-tuning, combined with the pinyin-Chinese character pair dataset, the accuracy and word order prediction problems of the Chinese pinyin input method were solved, achieving higher Chinese character input accuracy and personalized prediction.
Patent Information
- Application Number
- CN202510528900.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-09-19
AI Technical Summary
Existing Chinese character pinyin input methods have problems such as high error rate, poor fuzzy pinyin processing and weak word order prediction ability, making it difficult to accurately convert Chinese characters in long sentences or complex phrases.
A large model is used for pre-training and direct preference optimization algorithm fine-tuning. Combined with the training dataset of pinyin-Chinese character pairs, the candidate Chinese character sequences are reordered and screened through the large model to improve the accuracy of Chinese character prediction.
It significantly improves the accuracy of the Chinese character input system, reduces the misselection of homophones, enhances the understanding of contextual semantics, provides personalized prediction services, and improves user experience.
Smart Images

Figure CN120669867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural speech processing, and in particular to a large-model-based automatic prediction method and device for a Chinese character input system. Background Art
[0002] The existing Chinese character pinyin input method is one of the most commonly used input methods for Chinese users. It realizes text input by converting the pinyin string input by the user into the corresponding Chinese character sequence. However, the traditional Chinese character input system has many shortcomings.
[0003] First, the error rate is high: when the user's pinyin spelling is incorrect or there are uncommon words, the input method often cannot provide the correct candidate, resulting in incorrect output. Second, the handling of ambiguous pinyin is poor: many users may not distinguish certain subtle differences in pronunciation (such as "z / zh", "c / ch" or tones) when inputting. Traditional algorithms have difficulty effectively identifying such ambiguous pinyin, and the accuracy of candidate characters is reduced. Third, the word order prediction ability is weak: traditional pinyin input methods are mainly based on fixed N-gram models (such as bigrams or trigrams) to predict the next character or word. They have difficulty taking advantage of long-distance context. As a result, when inputting long sentences or complex phrases, the order and collocation of candidate words may not be reasonable, and the user's intention cannot be fully understood. Summary of the Invention
[0004] The present invention provides a large-model-based automatic prediction method and device for a Chinese character input system, which are used to solve the defects of the Chinese character input system in the prior art in terms of accuracy and improve the accuracy of the Chinese character input system.
[0005] The present invention provides a large-model-based automatic prediction method for a Chinese character input system, comprising: Obtain multiple candidate Chinese character sequences corresponding to the pinyin sequence; Inputting the multiple candidate Chinese character sequences and the pinyin sequence into a large model, reordering and / or screening the multiple candidate Chinese character sequences to obtain a final multiple candidate Chinese character sequences; The large model is pre-trained based on a training data set including pinyin-Chinese character pairs and fine-tuned based on a direct preference optimization algorithm.
[0006] In some embodiments, the inputting of the multiple candidate Chinese character sequences and the pinyin sequence into the large model, reordering and / or screening the multiple candidate Chinese character sequences to obtain the final multiple candidate Chinese character sequences includes: Combining the pinyin sequence with each of the candidate Chinese character sequences to obtain multiple combined sequences; Calculating the semantic score of each candidate Chinese character sequence according to each combined sequence; According to the semantic score of each candidate Chinese character sequence, the multiple candidate Chinese character sequences are reordered and / or screened to obtain a final multiple candidate Chinese character sequences.
[0007] In some embodiments, the inputting of the multiple candidate Chinese character sequences and the pinyin sequence into the large model, reordering and / or screening the multiple candidate Chinese character sequences to obtain the final multiple candidate Chinese character sequences includes: Performing Chinese character prediction on the pinyin sequence to obtain multiple predicted Chinese character sequences; According to the multiple predicted Chinese character sequences, the multiple candidate Chinese character sequences are reordered and / or screened to obtain a final multiple candidate Chinese character sequences.
[0008] In some embodiments, the method further comprises: The plurality of candidate Chinese character sequences are reordered and / or screened based on a word frequency library; the word frequency library records the occurrence probabilities of Chinese characters and / or words and the co-occurrence relationships between Chinese characters and / or words.
[0009] In some embodiments, the method comprises: Segmenting the Chinese text corpus and converting the segmented Chinese text corpus into corresponding pinyin sequences to construct a pinyin-Chinese character pair dataset; Pinyin-Chinese character pairs of incorrect pinyin and correct Chinese characters and Pinyin-Chinese character pairs of ambiguous pinyin and correct Chinese characters are introduced into the pinyin-Chinese character pair data set to obtain a training data set.
[0010] In some embodiments, the method further comprises: The large model is optimized based on the candidate Chinese character sequence selected by the user.
[0011] The present invention also provides a large-model-based automatic prediction device for a Chinese character input system, comprising: An acquisition module, used to obtain multiple candidate Chinese character sequences corresponding to the pinyin sequence; a first processing module, configured to input the plurality of candidate Chinese character sequences and the pinyin sequence into a large model, and reorder and / or filter the plurality of candidate Chinese character sequences to obtain a final plurality of candidate Chinese character sequences; The large model is pre-trained based on a training data set including pinyin-Chinese character pairs and fine-tuned based on a direct preference optimization algorithm.
[0012] The present invention also provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the automatic prediction method for a Chinese character input system based on a large model as described above is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements any of the above-mentioned large model-based automatic prediction methods for Chinese character input systems.
[0014] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described large-model-based automatic prediction methods for a Chinese character input system.
[0015] The present invention provides a large-model-based automatic prediction method and device for a Chinese character input system. The large model is pre-trained using a training data set including pinyin-Chinese character pairs, so that the large model has the basic ability to generate Chinese characters from pinyin. The large model is then fine-tuned using a direct preference optimization algorithm to make the Chinese character prediction results of the large model more in line with user expectations. The fine-tuned large model is used to reorder and / or screen multiple candidate Chinese character sequences corresponding to the pinyin sequence, thereby improving the overall quality of the candidate Chinese character sequences and thus improving the Chinese character prediction accuracy of the Chinese character input system. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 It is a flow chart of the automatic prediction method of the Chinese character input system based on the large model provided by the present invention.
[0018] Figure 2 It is a structural diagram of the automatic prediction device of the Chinese character input system based on the large model provided by the present invention.
[0019] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0020] Existing Chinese Pinyin input methods need to improve their accuracy. With the development of artificial intelligence and deep learning, large models have demonstrated powerful language understanding and generation capabilities in natural language processing tasks. Applying large models to Chinese Pinyin input methods is expected to significantly improve the accuracy of Pinyin-to-Chinese character conversion.
[0021] Therefore, there is an urgent need for a new technical solution to overcome the deficiencies of traditional Chinese character input systems, make full use of the semantic understanding ability of large models, and improve the error tolerance and prediction ability of Chinese pinyin input methods. The present invention is proposed based on this consideration, aiming to provide a Chinese character input system combined with a large model to significantly improve the accuracy of pinyin-to-Chinese character conversion.
[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0023] Figure 1 It is a schematic flowchart of an automatic prediction method for a Chinese character input system based on a large model provided by the present invention. As Figure 1 shown, the present invention provides an automatic prediction method for a Chinese character input system based on a large model, including: Step 110, obtaining a plurality of candidate Chinese character sequences corresponding to a pinyin sequence.
[0024] Specifically, a Chinese character sequence refers to a sequence formed by arranging Chinese characters in a certain order, and it can be simply understood that a Chinese character sequence is a word. For example, the Chinese character sequence of "artificial intelligence" is "人", "工", "智", "能".
[0025] The manner of obtaining a plurality of candidate Chinese character sequences corresponding to a pinyin sequence can be: using a traditional Chinese pinyin input method, according to the pinyin sequence input by the user, outputting a plurality of candidate Chinese character sequences corresponding to the pinyin.
[0026] For example, when the user inputs the pinyin sequence "jintian", the traditional Chinese pinyin input method may successively output several candidate Chinese character sequences such as "今天", "金田", "紧天", etc., and use these several candidate Chinese character sequences as the preliminary prediction results.
[0027] Step 120, inputting the plurality of candidate Chinese character sequences and the pinyin sequence into a large model, reordering and / or screening the plurality of candidate Chinese character sequences to obtain a final candidate Chinese character sequence; wherein, the large model is pre-trained based on a training data set including pinyin-Chinese character pairs and fine-tuned based on the direct preference optimization algorithm.
[0028] Specifically, select a large model suitable for Chinese (such as the Transformer model), and continue to train it on this basis to make it adapt to the pinyin-to-Chinese character conversion task.
[0029] The specific process of large model pre-training can be as follows: taking the pinyin sequence as the input and the corresponding Chinese character sequence as the target output, pre-training is carried out through a large number of pinyin-Chinese character pairs in the training dataset. The model gradually learns the mapping relationship between pinyin and Chinese characters internally and enhances its "understanding" ability of pinyin input. For example, when providing the pinyin sequence "zhong guo" to the large model, it can predict and output the correct Chinese character sequence "中国".
[0030] After large-scale pre-training, the large model already has the basic ability to generate Chinese characters from pinyin and retains the understanding of Chinese semantics and context, which lays a foundation for further optimization.
[0031] On the basis of pre-training, the Direct Preference Optimization (DPO) algorithm is used to fine-tune the pre-trained large model to enhance the large model's preference for correct candidate Chinese characters.
[0032] The specific process of using the DPO algorithm to fine-tune the pre-trained large model can be as follows: using the training data containing candidate preferences, that is, for each pinyin sequence, preparing multiple candidate Chinese character sequences, taking the candidate Chinese character sequence that best conforms to the semantics or user intention among the multiple candidate Chinese character sequences as the preferred Chinese character sequence, and marking the preferred Chinese character sequence. When fine-tuning, the large model will convert these preference information into loss function signals and directly optimize the model parameters, so that when facing multiple candidate Chinese character sequences corresponding to the same pinyin, it gives a higher probability to the preferred candidate Chinese character sequence.
[0033] Compared with traditional reinforcement learning methods, the DPO algorithm does not require explicit training of the reward model, but directly adjusts the main model by comparing the quality of candidate outputs, making the fine-tuning process more stable and efficient. Through DPO fine-tuning, the large model can more accurately distinguish homophonic candidate words, reduce the situation of选错汉字序列 (choosing the wrong Chinese character sequence due to the same pronunciation), and learn to select the correct Chinese character sequence that conforms to the context semantics and common sense among multiple candidates.
[0034] The final large model is obtained through pre-training and DOP fine-tuning. This model integrates the language knowledge of massive data and the mapping relationship between pinyin and Chinese characters, and is closer to the user's expectation of correct output in the actual input method scenario after preference fine-tuning.
[0035] After obtaining the final large model, multiple candidate Chinese character sequences and pinyin sequences are input into the large model. Based on the context and its own language knowledge, the large model evaluates the rationality and correctness of multiple candidate Chinese character sequences, so as to re-rank and / or screen multiple candidate Chinese character sequences to obtain the final multiple candidate Chinese character sequences.
[0036] The final candidate Chinese character sequences are presented to the user in a list format. Typically, the final multiple candidate Chinese character sequences are presented according to the re-sorted ranking results, with the most likely candidate sequence being presented as the first candidate, while other candidate sequences are presented as alternatives for the user to choose from. This multi-candidate output also improves the user's interactive experience: if the first candidate isn't the desired word, the alternative list is more likely to contain the correct word that meets the user's expectations.
[0037] The automatic prediction method for a Chinese character input system based on a large model provided by the present invention pre-trains the large model through a training data set including pinyin-Chinese character pairs, so that the large model has the basic ability to generate Chinese characters from pinyin. The large model is then fine-tuned through a direct preference optimization algorithm to make the Chinese character prediction results of the large model more in line with user expectations. The fine-tuned large model is used to reorder and / or screen multiple candidate Chinese character sequences corresponding to the pinyin sequence, thereby improving the overall quality of the candidate Chinese character sequences, thereby improving the Chinese character prediction accuracy of the Chinese character input system.
[0038] In some embodiments, the large model-based automatic prediction method for Chinese character input system provided by the present invention further includes: Segment the Chinese text corpus and convert the segmented Chinese text corpus into corresponding pinyin sequences to construct a pinyin-Chinese character pair dataset; Pinyin-Chinese character pairs consisting of incorrect pinyin and correct Chinese characters and Pinyin-Chinese character pairs consisting of ambiguous pinyin and correct Chinese characters are introduced into the pinyin-Chinese character pair dataset to obtain a training dataset.
[0039] Specifically, a massive amount of Chinese text corpus is collected, including texts from general fields and various vertical fields, such as news, social media posts, literary works, etc., to ensure that the corpus covers different styles and topics.
[0040] The collected Chinese text corpus is segmented into sentences or phrases. A standard phonetic-to-text conversion tool is used to convert the segmented Chinese text corpus into corresponding phonetic sequences. A mapping relationship is established between the segmented Chinese text corpus and the corresponding phonetic sequences, resulting in a phonetic-to-Chinese character pair dataset. The necessary word segmentation information is retained during the conversion process to ensure accurate alignment between phonetic and Chinese characters.
[0041] To improve the model's robustness to actual user input, we simulated incorrect and ambiguous pinyin. For example, for easily confused initials and finals (such as "n" and "l," "ang" and "an"), we randomly replaced the pinyin in some samples to simulate possible user input errors. We also simulated the omission or mislabeling of tones. Furthermore, we included data on initial abbreviations (for example, entering "Beijing" as "bj") and continuous pinyin.
[0042] The training dataset is generated by introducing pinyin-Chinese character pairs containing incorrect pinyin and correct Chinese characters, as well as pinyin-Chinese character pairs containing ambiguous pinyin and correct Chinese characters. This training dataset is used to pre-train the model, enabling it to recover correct Chinese characters even under noisy conditions.
[0043] The large-model-based automatic prediction method for a Chinese character input system provided by the present invention provides a rich and diverse training sample for the model by introducing pinyin-Chinese character pairs of incorrect pinyin and correct Chinese characters and pinyin-Chinese character pairs of ambiguous pinyin and correct Chinese characters into a pinyin-Chinese character pair dataset, thereby improving the robustness of the model and enabling the model to recover correct Chinese characters under noisy conditions.
[0044] In some embodiments, multiple candidate Chinese character sequences and the pinyin sequence are input into a large model, and the multiple candidate Chinese character sequences are reordered and / or screened to obtain a final multiple candidate Chinese character sequences, including: Combine the pinyin sequence with each candidate Chinese character sequence to obtain multiple combination sequences; According to each combination sequence, the semantic score of each candidate Chinese character sequence is calculated; According to the semantic score of each candidate Chinese character sequence, multiple candidate Chinese character sequences are reordered and / or screened to obtain a final candidate Chinese character sequence.
[0045] Specifically, the pinyin sequence is combined with each candidate Chinese character sequence to obtain multiple combined sequences. In each combined sequence, a special marker is used to separate the pinyin sequence and the candidate Chinese character sequence, allowing the model to clearly distinguish between the pinyin sequence and the candidate Chinese character sequence while preserving semantic relevance. For example, the separator "<|pinyin|>" is used, such as: candidate Chinese character sequence <|pinyin|> pinyin sequence.
[0046] For each combination sequence, the large model calculates the semantic plausibility of the candidate Chinese character sequence in the combination sequence based on the pinyin sequence portion of the combination sequence, and obtains the semantic score of the candidate Chinese character sequence in the combination sequence. For example, based on the pinyin sequence portion of the combination sequence, the generation probability of the candidate Chinese character sequence in the combination sequence is calculated and the generation probability is used as the semantic score.
[0047] After obtaining the semantic score of each candidate Chinese character sequence, candidate Chinese character sequences with semantic scores greater than or equal to a threshold are screened out, and the screened candidate Chinese character sequences are sorted from large to small according to the semantic scores to obtain the final candidate Chinese character sequence.
[0048] When each candidate Chinese character sequence is greater than the threshold, no screening is required; when the original order of multiple candidate Chinese character sequences conforms to the order from largest to smallest according to the semantic score, no re-ordering is required.
[0049] The large-model-based automatic prediction method for a Chinese character input system provided by the present invention combines a pinyin sequence with each candidate Chinese character sequence to obtain multiple combined sequences; calculates the semantic score of the candidate Chinese character sequence in each combined sequence; and reorders and / or screens the multiple candidate Chinese character sequences based on the semantic score of each candidate Chinese character sequence, placing candidate Chinese character sequences that are more in line with the context or common usage at the front, significantly reducing the error rate of pinyin-to-Chinese character conversion and further improving the accuracy of pinyin-to-Chinese character conversion.
[0050] In some embodiments, multiple candidate Chinese character sequences and the pinyin sequence are input into a large model, and the multiple candidate Chinese character sequences are reordered and / or screened to obtain a final multiple candidate Chinese character sequences, including: Perform Chinese character prediction on the pinyin sequence to obtain multiple predicted Chinese character sequences; According to the multiple predicted Chinese character sequences, the multiple candidate Chinese character sequences are reordered and / or screened to obtain the final multiple candidate Chinese character sequences.
[0051] Specifically, the large model directly predicts Chinese characters for the pinyin sequence, obtaining multiple predicted Chinese character sequences. Based on the multiple predicted Chinese character sequences, multiple candidate Chinese character sequences are screened so that each candidate Chinese character sequence after screening falls within the predicted Chinese character sequence. The screened candidate Chinese character sequences are then reordered according to the sort order of the multiple predicted Chinese character sequences to obtain the final multiple candidate Chinese character sequences.
[0052] The large-model-based automatic prediction method for a Chinese character input system provided by the present invention directly predicts Chinese characters for pinyin sequences through a large model to cover possible Chinese character sequences. Based on multiple predicted Chinese character sequences, multiple candidate Chinese character sequences are reordered and / or screened to eliminate low-quality or unreasonable candidate Chinese character sequences, thereby improving the overall quality of the candidate Chinese character sequences and further improving the accuracy of converting pinyin to Chinese characters.
[0053] Because the large model learns more global grammatical and semantic information during training, it can detect subtle differences that statistical models struggle to detect, such as distinguishing common collocations from rare ones and understanding long-range dependencies in long sentences. Through the large model's predictions, some candidates that originally scored low in the statistical model but are more semantically correct may be promoted to preferred results.
[0054] In some embodiments, the large model-based automatic prediction method for Chinese character input system provided by the present invention further includes: Based on a word frequency library, multiple candidate Chinese character sequences are reordered and / or screened; the word frequency library records the occurrence probabilities of Chinese characters and / or words and the co-occurrence relationships between Chinese characters and / or words.
[0055] Specifically, a word frequency database is constructed, which records the occurrence probability of Chinese characters and / or words (i.e., the occurrence probability of a single Chinese character, the occurrence probability of a multi-character word, and the occurrence probability of a single Chinese character and a multi-character word together), as well as the co-occurrence relationship between Chinese characters and / or words (i.e., the co-occurrence relationship between single Chinese characters, the co-occurrence relationship between a single Chinese character and a multi-character word, and the co-occurrence relationship between multi-character words).
[0056] Before inputting multiple candidate Chinese character sequences and pinyin sequences into the large model, the multiple candidate Chinese character sequences can be reordered and / or screened based on the word frequency library, that is, very rare candidate Chinese character sequences or candidate Chinese character sequences that do not conform to grammatical combinations are eliminated, and candidate Chinese character sequences that are high-frequency and consistent with the context are retained first; the remaining candidate Chinese character sequences are reordered from high to low according to the probability of occurrence.
[0057] The large-model-based automatic prediction method for a Chinese character input system provided by the present invention ensures that the generated candidate Chinese character sequences are closer to commonly used expressions by weighting the word frequency information in the word frequency library, thereby improving the accuracy and practicality of subsequent predictions.
[0058] In some embodiments, the large model-based automatic prediction method for Chinese character input system provided by the present invention further includes: The large model is optimized based on the candidate Chinese character sequences selected by the user.
[0059] Specifically, after the user selects a candidate Chinese character sequence, the candidate sequence selected by the user is fed back to the large model, and the large model is optimized using the candidate sequence selected by the user, that is, the model parameters are updated. The word frequency information in the word frequency library can also be updated based on the candidate sequence selected by the user.
[0060] The automatic prediction method for a Chinese character input system based on a large model provided by the present invention optimizes the large model through the candidate Chinese character sequence selected by the user, so that the candidate Chinese character sequence subsequently predicted by the large model is more in line with the user's personal habits, providing personalized prediction services and improving user experience.
[0061] The following describes the automatic prediction device for a Chinese character input system based on a large model provided by the present invention. The automatic prediction device for a Chinese character input system based on a large model described below and the automatic prediction method for a Chinese character input system based on a large model described above can refer to each other.
[0062] Figure 2 This is a schematic diagram of the structure of the automatic prediction device for Chinese character input system based on the large model provided by the present invention. Figure 2 As shown, the present invention provides an automatic prediction device for a Chinese character input system based on a large model, comprising: An acquisition module 210 is used to acquire multiple candidate Chinese character sequences corresponding to the pinyin sequence; A first processing module 220 is configured to input the plurality of candidate Chinese character sequences and the pinyin sequence into a large model, and reorder and / or filter the plurality of candidate Chinese character sequences to obtain a final plurality of candidate Chinese character sequences; The large model is pre-trained based on a training data set including pinyin-Chinese character pairs and fine-tuned based on a direct preference optimization algorithm.
[0063] In some embodiments, the first processing module 220 is specifically configured to: Combining the pinyin sequence with each of the candidate Chinese character sequences to obtain multiple combined sequences; Calculating the semantic score of each candidate Chinese character sequence according to each combined sequence; According to the semantic score of each candidate Chinese character sequence, the multiple candidate Chinese character sequences are reordered and / or screened to obtain a final multiple candidate Chinese character sequences.
[0064] In some embodiments, the first processing module 220 is specifically configured to: Performing Chinese character prediction on the pinyin sequence to obtain multiple predicted Chinese character sequences; According to the multiple predicted Chinese character sequences, the multiple candidate Chinese character sequences are reordered and / or screened to obtain a final multiple candidate Chinese character sequences.
[0065] In some embodiments, the apparatus further comprises: The second processing module is used to reorder and / or screen the multiple candidate Chinese character sequences based on a word frequency library; the word frequency library records the occurrence probability of Chinese characters and / or words and the co-occurrence relationship between Chinese characters and / or words.
[0066] In some embodiments, the apparatus further comprises: A construction module is used to segment the Chinese text corpus and convert the segmented Chinese text corpus into corresponding pinyin sequences to construct a pinyin-Chinese character pair dataset; The introduction module is used to introduce the pinyin-Chinese character pairs of incorrect pinyin and correct Chinese characters and the pinyin-Chinese character pairs of ambiguous pinyin and correct Chinese characters into the pinyin-Chinese character pair data set to obtain a training data set.
[0067] In some embodiments, the apparatus further comprises: The optimization module is used to optimize the large model based on the candidate Chinese character sequence selected by the user.
[0068] It should be noted here that the above-mentioned large-model-based automatic prediction device for Chinese character input system provided by the present invention can implement all the method steps implemented in the above-mentioned method embodiment, and can achieve the same technical effect. The parts and beneficial effects that are the same as the method embodiment in this embodiment will not be described in detail here.
[0069] Figure 3 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 3 As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute a large-model-based automatic prediction method for a Chinese character input system, the method comprising: obtaining multiple candidate Chinese character sequences corresponding to a pinyin sequence; inputting the multiple candidate Chinese character sequences and the pinyin sequence into the large model, reordering and / or screening the multiple candidate Chinese character sequences to obtain a final multiple candidate Chinese character sequences; wherein the large model is pre-trained based on a training data set including pinyin-Chinese character pairs and fine-tuned based on a direct preference optimization algorithm.
[0070] Furthermore, the logic instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0071] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the automatic prediction method of the Chinese character input system based on the large model provided by the above methods, the method including: obtaining multiple candidate Chinese character sequences corresponding to the pinyin sequence; inputting the multiple candidate Chinese character sequences and the pinyin sequence into the large model, reordering and / or screening the multiple candidate Chinese character sequences to obtain the final multiple candidate Chinese character sequences; wherein the large model is pre-trained based on a training data set including pinyin-Chinese character pairs, and fine-tuned based on a direct preference optimization algorithm.
[0072] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the large-model-based automatic prediction method for a Chinese character input system provided by the above-mentioned methods, the method comprising: obtaining multiple candidate Chinese character sequences corresponding to a pinyin sequence; inputting the multiple candidate Chinese character sequences and the pinyin sequence into a large model, reordering and / or screening the multiple candidate Chinese character sequences to obtain a final multiple candidate Chinese character sequences; wherein the large model is pre-trained based on a training data set including pinyin-Chinese character pairs, and fine-tuned based on a direct preference optimization algorithm.
[0073] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0074] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for automatic prediction of Chinese character input system based on a large model, characterized in that: include: Obtain multiple candidate Chinese character sequences corresponding to the pinyin sequence; Inputting the multiple candidate Chinese character sequences and the pinyin sequence into a large model, reordering and / or screening the multiple candidate Chinese character sequences to obtain a final multiple candidate Chinese character sequences; The large model is pre-trained based on a training data set including pinyin-Chinese character pairs and fine-tuned based on a direct preference optimization algorithm.
2. The automatic prediction method for Chinese character input system based on large model according to claim 1 is characterized in that: The step of inputting the plurality of candidate Chinese character sequences and the pinyin sequence into a large model, reordering and / or screening the plurality of candidate Chinese character sequences, and obtaining a final plurality of candidate Chinese character sequences comprises: Combining the pinyin sequence with each of the candidate Chinese character sequences to obtain multiple combined sequences; Calculating the semantic score of each candidate Chinese character sequence according to each combined sequence; According to the semantic score of each candidate Chinese character sequence, the multiple candidate Chinese character sequences are reordered and / or screened to obtain a final multiple candidate Chinese character sequences.
3. The automatic prediction method for Chinese character input system based on large model according to claim 1 is characterized in that: The step of inputting the plurality of candidate Chinese character sequences and the pinyin sequence into a large model, reordering and / or screening the plurality of candidate Chinese character sequences, and obtaining a final plurality of candidate Chinese character sequences comprises: Performing Chinese character prediction on the pinyin sequence to obtain multiple predicted Chinese character sequences; According to the multiple predicted Chinese character sequences, the multiple candidate Chinese character sequences are reordered and / or screened to obtain a final multiple candidate Chinese character sequences.
4. The automatic prediction method for Chinese character input system based on large model according to claim 1 is characterized in that: The method further comprises: The plurality of candidate Chinese character sequences are reordered and / or screened based on a word frequency library; the word frequency library records the occurrence probabilities of Chinese characters and / or words and the co-occurrence relationships between Chinese characters and / or words.
5. The automatic prediction method for Chinese character input system based on large model according to claim 1 is characterized in that: The method comprises: Segmenting the Chinese text corpus and converting the segmented Chinese text corpus into corresponding pinyin sequences to construct a pinyin-Chinese character pair dataset; Pinyin-Chinese character pairs of incorrect pinyin and correct Chinese characters and Pinyin-Chinese character pairs of ambiguous pinyin and correct Chinese characters are introduced into the pinyin-Chinese character pair data set to obtain a training data set.
6. The automatic prediction method for Chinese character input system based on a large model according to any one of claims 1 to 5, characterized in that: The method further comprises: The large model is optimized based on the candidate Chinese character sequence selected by the user.
7. An automatic prediction device for Chinese character input system based on a large model, characterized in that: include: An acquisition module, used to acquire multiple candidate Chinese character sequences corresponding to the pinyin sequence; a first processing module, configured to input the plurality of candidate Chinese character sequences and the pinyin sequence into a large model, and reorder and / or filter the plurality of candidate Chinese character sequences to obtain a final plurality of candidate Chinese character sequences; The large model is pre-trained based on a training data set including pinyin-Chinese character pairs and fine-tuned based on a direct preference optimization algorithm.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the automatic prediction method of the Chinese character input system based on the large model as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the automatic prediction method of the Chinese character input system based on the large model as described in any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the automatic prediction method of the Chinese character input system based on the large model as described in any one of claims 1 to 6 is implemented.