Text input method, apparatus, device, medium, and program product

By collecting and analyzing multimodal contextual information of user input text, and combining a global base model and a personalized model, the scores of candidate content are calculated, which solves the problem of inaccurate input prediction content in existing technologies and achieves higher prediction accuracy.

CN122491264APending Publication Date: 2026-07-31CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNITED NETWORK COMM GRP CO LTD
Filing Date
2026-04-16
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of inferring user input needs from the context information corresponding to text information is low, resulting in the input prediction content failing to accurately match user needs.

Method used

Collect textual context information and multimodal context information of user input text, obtain unified multimodal features, use a global basic model and a personalized lightweight model to calculate the global basic score and personalized score of candidate content, and combine the input scene information to determine the final score to improve prediction accuracy.

Benefits of technology

It improves the accuracy of input prediction, making it consistent with common language logic, popular usage habits, and individual user habits, while also matching the current input scenario.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491264A_ABST
    Figure CN122491264A_ABST
Patent Text Reader

Abstract

This application provides a text input method, apparatus, device, medium, and program product, relating to the field of artificial intelligence technology, for improving the accuracy of input prediction content. The specific technical solution includes: collecting input context information and multimodal context information of the user-input text; obtaining unified multimodal features based on the input context information and multimodal context information; inputting the unified multimodal features into a global base model to obtain a global base score corresponding to each candidate content in the candidate pool; inputting the unified multimodal features into a personalized lightweight model to obtain a personalized score corresponding to each candidate content in the candidate pool; determining the input scenario information of the user-input text based on the unified multimodal features; and determining the predicted text of the user-input text based on the global base score, personalized score, and input scenario information. This application is applied to scenarios involving text prediction of user-input text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a text input method, apparatus, device, medium, and program product. Background Technology

[0002] With the rapid development of intelligent human-computer interaction technology, it has become increasingly important to accurately predict user input through intelligent input prediction technology in order to improve the convenience of user interaction with devices.

[0003] In related technologies, contextual information corresponding to the text information currently input by the user is obtained, and then the user's input needs are inferred based on the contextual information, and corresponding input prediction content is generated.

[0004] However, in the process of implementing the above scheme, the inference accuracy is low when inferring the user's input needs from the context information corresponding to the text information, resulting in the input prediction content failing to accurately match the user's needs. Therefore, the generated input prediction content in the related technologies is not accurate enough. Summary of the Invention

[0005] This application provides a text input method, apparatus, device, medium, and program product for improving the accuracy of input prediction content.

[0006] In a first aspect, embodiments of this application provide a text input method, the method comprising: collecting input context information and multimodal context information of user-input text, the multimodal context information including at least one of the following: client's geographical location information, client's time scenario information, client's user preference profile information, and client's application context information; obtaining unified multimodal features based on the above input context information and the above multimodal context information; inputting the unified multimodal features into a global base model to obtain a global base score corresponding to each candidate content in a candidate pool, the global base score representing the degree to which the candidate content conforms to common language logic and popular usage habits; inputting the unified multimodal features into a personalized lightweight model to obtain a personalized score corresponding to each candidate content in the above candidate pool, the personalized score representing the degree to which the candidate content conforms to the user's personal usage habits; determining the input scenario information of user-input text based on the above unified multimodal features; determining the final score of each candidate content in the above candidate pool based on the above global base score, the above personalized score, and the above input scenario information; and determining the predicted text of user-input text based on the final score of each candidate content; wherein the above candidate pool includes at least two candidate contents.

[0007] The technical solution provided in this application brings at least the following beneficial effects: In this solution, while collecting the text context information of the input text, multimodal context information representing the current input scenario information is also collected. Based on the text context information and multimodal context information of the input text, multimodal features that can represent both the language logic of the input text and the current input scenario are obtained. Based on these multimodal features, a global basic score reflecting whether each candidate content conforms to general language logic and popular usage habits, a personalized score reflecting whether each candidate content conforms to the user's personal usage habits, and input scenario information representing the scenario when the user inputs the text are obtained. Based on the global basic score, personalized score, and input scenario information, the predicted input text that conforms to both general language logic and popular usage habits, as well as the user's personal usage habits, and also matches the current input scenario is determined. Thus, the accuracy of the predicted input content is improved.

[0008] One possible implementation involves determining the final score for each candidate content in the candidate pool based on the global base score, the personalized score, and the input scenario information. This includes: determining a scenario enhancement score for each candidate content in the candidate pool based on the matching degree between each candidate content and the input scenario information; and determining the final score for each candidate content in the candidate pool based on the global base score, the personalized score, and the scenario enhancement score. The scenario enhancement score represents the degree of matching between the candidate content and the input scenario of the user's input text.

[0009] Another possible implementation involves determining the final score for each candidate content in the candidate pool based on the global base score, the personalized score, and the scene enhancement score corresponding to each candidate content in the candidate pool. This includes: determining the first weight, second weight, and third weight corresponding to the first candidate content based on weight information, wherein the weight information includes at least one of the following: scene matching degree, historical selection probability, language confidence, and personal usage pattern; calculating the final score corresponding to the first candidate content based on the first weight, the second weight, the third weight, and the global base score, the personalized score, and the scene enhancement score corresponding to the first candidate content, wherein the first candidate content is any candidate content in the candidate pool.

[0010] Another possible implementation is to determine the predicted text of the user input text based on the final score of each candidate content, including: determining the candidate content in the candidate pool whose final score is greater than the score threshold as the predicted text; the predicted text includes at least one of the following: next word prediction, input information completion prediction, and scene-related content prediction.

[0011] Another possible implementation method includes updating the user preference profile information and the personalized lightweight model based on the target text selected by the user from the predicted text.

[0012] Secondly, embodiments of this application provide a text input device, including: an acquisition module and a processing module; the acquisition module is configured to collect input context information and multimodal context information of user-input text, the multimodal context information including at least one of the following: the client's geographical location information, the client's time scene information, the client's user preference profile information, and the client's application context information; and, based on the above input context information and the above multimodal context information, acquire unified multimodal features; and, input the above unified multimodal features into a global base model to acquire a global base score corresponding to each candidate content in the candidate pool, the global base score representing the candidate content's conformance to the global base model. The processing module is used to input the unified multimodal features into a personalized lightweight model to obtain a personalized score for each candidate content in the candidate pool, wherein the personalized score represents the degree to which the candidate content conforms to the user's personal usage habits; and, based on the unified multimodal features, to determine the input scenario information of the user input text; and, based on the global base score, the personalized score, and the input scenario information, to determine the final score of each candidate content in the candidate pool; and, based on the final score of each candidate content, to determine the predicted text of the user input text; wherein the candidate pool includes at least two candidate contents.

[0013] One possible implementation is that the aforementioned processing module is specifically used to: determine the scene enhancement score corresponding to each candidate content in the aforementioned candidate pool based on the matching degree between each candidate content in the aforementioned candidate pool and the aforementioned input scene information; and determine the final score corresponding to each candidate content in the aforementioned candidate pool based on the aforementioned global base score, the aforementioned personalized score, and the aforementioned scene enhancement score; wherein, the aforementioned scene enhancement score represents the degree of matching between the aforementioned candidate content and the input scene of the aforementioned user input text.

[0014] Another possible implementation, the aforementioned processing module is specifically used to: determine the first weight, second weight, and third weight corresponding to the first candidate content based on the weighting information, wherein the weighting information includes at least one of the following: scene matching degree, historical selection probability, language confidence, and personal usage pattern; calculate the final score corresponding to the first candidate content based on the first weight, the second weight, the third weight, and the global base score, the personalized score, and the scene enhancement score corresponding to the first candidate content, wherein the first candidate content is any candidate content in the candidate pool.

[0015] Another possible implementation is that the above processing module is specifically used to: determine the candidate content in the above candidate pool whose final score is greater than the score threshold as the predicted text; the above predicted text includes at least one of the following: next word prediction, input information completion prediction, and scene-related content prediction.

[0016] In another possible implementation, the aforementioned processing module is also used to update the aforementioned user preference profile information and the aforementioned personalized lightweight model based on the target text selected by the user from the aforementioned predicted text.

[0017] Thirdly, this application provides an electronic device comprising: a processor and a memory; the memory stores a program or instructions executable on the processor, wherein the program or instructions, when executed by the processor, implement the method of the first aspect described above.

[0018] Fourthly, this application provides a readable storage medium on which a program or instructions are stored, which, when executed by a computer, implement the method of the first aspect described above.

[0019] Fifthly, this application provides a computer program product stored in a storage medium, which, when executed by a computer, implements the method described in the first aspect.

[0020] In a sixth aspect, embodiments of this application provide a chip including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the method described in the first aspect.

[0021] The beneficial effects of the second to sixth aspects mentioned above are described in the corresponding description of the first aspect and will not be repeated here. Attached Figure Description

[0022] Figure 1 A schematic diagram of the network architecture for a text input method application provided in this application embodiment;

[0023] Figure 2A flowchart illustrating a text input method provided in an embodiment of this application;

[0024] Figure 3 A flowchart illustrating another text input method provided in an embodiment of this application;

[0025] Figure 4 A flowchart illustrating yet another text input method provided in an embodiment of this application;

[0026] Figure 5 A flowchart illustrating yet another text input method provided in an embodiment of this application;

[0027] Figure 6 A flowchart illustrating yet another text input method provided in an embodiment of this application;

[0028] Figure 7 A flowchart illustrating the implementation process of a text input method provided in this application embodiment;

[0029] Figure 8 This is a schematic diagram of the structure of a text input system provided in an embodiment of this application;

[0030] Figure 9 This is a schematic diagram of the structure of a text input device provided in an embodiment of this application;

[0031] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0032] The text input method, apparatus, device, medium, and program products provided in this application will now be described in detail with reference to the accompanying drawings.

[0033] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0034] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0035] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0036] In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0037] The embodiments of this application provide a text input method, apparatus, device, medium, and program product that can be applied to scenarios involving text prediction of user-inputted text.

[0038] With the improvement of mobile terminal computing power, more and more intelligent input prediction capabilities are migrating from the cloud to the client side for inference. Existing technologies mainly include the following input prediction methods: input prediction technology based on statistical language models, general language model prediction technology based on deep learning, and personalized prediction models based on user historical behavior.

[0039] However, in existing input prediction methods, the inference accuracy is low when inferring user input needs from the context information corresponding to text information, resulting in the obtained input prediction content failing to accurately match the user's needs. Therefore, the generated input prediction content in related technologies is not accurate enough.

[0040] To address the aforementioned technical problems, embodiments of this application provide a text input method, apparatus, device, medium, and program product. While collecting the text context information of the input text, it also collects multimodal context information representing the current input scenario. Based on the text context information and multimodal context information of the input text, it obtains multimodal features that represent both the language logic of the input text and the current input scenario. Then, based on these multimodal features, it obtains a global base score reflecting whether each candidate content conforms to general language logic and common usage habits, a personalized score reflecting whether each candidate content conforms to the user's personal usage habits, and input scenario information representing the scenario in which the user inputs the text. Based on the global base score, personalized score, and input scenario information, it determines the predicted input text that conforms to both general language logic and common usage habits, as well as the user's personal usage habits, and also matches the current input scenario. This improves the accuracy of the predicted input content.

[0041] The text input method, apparatus, device, medium, and program product provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0042] Figure 1 This illustration shows a network architecture for a text input method application provided in an embodiment of this application. For example... Figure 1 As shown, the network architecture includes a text input device 101 and a terminal device 102. The text input device 101 and the terminal device 102 are interconnected.

[0043] In some embodiments, the text input device 101 may be a server, a computer, or a processor or processing unit within a server or computer. The server may be a single server or a server cluster comprising multiple servers. It should be noted that the embodiments of this application do not limit the specific device form of the text input device 101. Figure 1 The text input device 101 is used as an example of a single server.

[0044] In some embodiments, the terminal device may be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, personal computer (PC), ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc., and the embodiments of this application do not specifically limit it. Figure 1 The example shown is a mobile phone, with terminal device 102 as an example.

[0045] In some embodiments, the text input device 101 collects input context information and multimodal context information of the user's input text. The multimodal context information includes at least one of the following: the client's geographic location information, the client's time context information, the client's user preference profile information, and the client's application context information. Based on the above input context information and the above multimodal context information, a unified multimodal feature is obtained. The unified multimodal feature is input into a global base model to obtain a global base score corresponding to each candidate content in the candidate pool. This global base score represents the degree to which the candidate content conforms to common language logic and popular usage habits. A multimodal feature input personalized lightweight model is used to obtain a personalized score for each candidate content in the candidate pool, which represents the degree to which the candidate content conforms to the user's personal usage habits; based on the unified multimodal features, the input scenario information of the user's input text is determined; based on the global base score, the personalized score, and the input scenario information, the final score of each candidate content in the candidate pool is determined; based on the final score of each candidate content, the predicted text of the user's input text is determined and sent to the terminal device; the terminal device 102 receives the predicted text sent by the text input device 101 and displays the predicted text.

[0046] It should be noted that the network architecture described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As network architectures evolve, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0047] See Figure 2 This is a flowchart illustrating a text input method provided in an embodiment of this application. Figure 2As shown, the text input method provided in this application embodiment can be implemented by the above-mentioned text input device, specifically including the following steps 201 to 207.

[0048] Step 201: The text input device collects the input context information and multimodal context information of the user's input text.

[0049] In some embodiments, the user input text described above represents text that the user has already entered.

[0050] In some embodiments, the above-mentioned input context information may also be referred to as text input sequence information.

[0051] In some embodiments, the input context information includes: text information in the current input box, contextual word order or sentence order information, user's recent input history information, and input box type information. Of course, the input context information may also include input context information of other user-inputted text, which can be determined according to actual needs, and this application does not limit this.

[0052] In some embodiments, the above-mentioned input box type information is used to indicate the type of the input box.

[0053] In some embodiments, the text input device may determine the type of the input box based on the type of text that the user needs to enter into the input box or the function type of the input box.

[0054] In some embodiments, the type of the input box includes at least one of the following: address input box, name input box, search input box, and chat message input box. Of course, the type of input box may also include other types, which can be determined according to actual needs, and this application does not limit this.

[0055] In some embodiments, the multimodal context information described above is used to characterize the environment and personalization features of the user when inputting the user input text.

[0056] In some embodiments, the multimodal context information includes at least one of the following: the client's geographic location information, the client's time context information, the client's user preference profile information, and the client's application context information. Of course, the multimodal context information may also include other information, which can be determined according to actual needs, and this application does not limit this.

[0057] In some embodiments, the geolocation information of the client refers to the location information of the client when the user enters the user input text.

[0058] In some embodiments, the aforementioned client-side time context information refers to the time information when the user inputs the aforementioned user input text.

[0059] In some embodiments, the time-scene information of the client mentioned above includes at least one of the following: time period information, holiday information, and user's daily routine information. Of course, the time-scene information of the client mentioned above may also include other information, which can be determined according to actual needs, and this application does not limit it.

[0060] In some embodiments, the aforementioned user preference profile information for the client refers to a comprehensive set of information constructed by analyzing long-term user behavior data to accurately depict the user's personalized preferences.

[0061] In some embodiments, the user preference profile information of the aforementioned client includes at least one of the following: frequently used address information, frequently used contact information, keyword preference statistics, and natural language usage habits. Of course, the user preference profile information of the aforementioned client may also include other information, which can be determined according to actual needs, and this application does not limit this.

[0062] In some embodiments, the application context information of the client described above may also be referred to as the current application context information.

[0063] In some embodiments, the application context information of the client mentioned above refers to the application information of the application that is running in the foreground or background when the user enters the user input text.

[0064] In some embodiments, the application information mentioned above includes at least one of the following: application type information, application name information, and application runtime information. Of course, the application information may also include other application information, which can be determined according to actual needs, and this application does not limit this.

[0065] Step 202: The text input device obtains unified multimodal features based on input context information and multimodal context information.

[0066] In some embodiments, the above-described unified multimodal feature representation integrates comprehensive features that combine text input context and multimodal context information.

[0067] In some embodiments, the unified multimodal features mentioned above include at least one of the following: text vector features, time vector features, geographic vector features, application vector features, and user profile vector features. Of course, the unified multimodal features mentioned above may also include other features, which can be determined according to actual needs, and this application does not limit them.

[0068] In some embodiments, the text input device can map the input context information and the multimodal context information to a unified semantic vector space to obtain the unified multimodal features.

[0069] Specifically, the text input device can convert the aforementioned text input sequence information into text vectors through an embedding layer to obtain the aforementioned text vector features, periodically encode the aforementioned client's time scene information to obtain the aforementioned time vector features, regionally discretize and encode the aforementioned client's geographical location information to obtain the aforementioned geographical vector features, classify and encode the aforementioned client's application context information to obtain the aforementioned application vector features, and encode the aforementioned client's user preference profile information to obtain the aforementioned user profile vector features. Then, feature fusion processing is performed on the aforementioned text vector features, the aforementioned time vector features, the aforementioned geographical vector features, the aforementioned application vector features, and the aforementioned user profile vector features to obtain the aforementioned unified multimodal features.

[0070] In some embodiments, for different types of data in the above-mentioned input context information and multimodal context information, the text input device uses a differential encoding method to encode them to obtain their corresponding feature vectors.

[0071] In some embodiments, the differentiated encoding methods described above include at least one of the following: token embedding, geospatial embedding, time slot embedding, user preference embedding, and app context embedding. Of course, other encoding methods may also be included, which can be determined according to actual needs, and this application does not limit them.

[0072] In some embodiments, for text context data, the text input device uses Token Embedding encoding; for geolocation data, the text input device uses Geo Embedding encoding; for time-scene data, the text input device uses Time Slot Embedding encoding; for user profile data, the text input device uses User Preference Embedding encoding; and for application context data, the text input device uses App Context Embedding encoding. Finally, the text input device concatenates the feature vectors encoded by each type of data to form a unified vector, i.e., the aforementioned unified multimodal feature.

[0073] In some embodiments, the text input device forms the aforementioned unified multimodal feature by concatenating the feature vectors encoded from each type of data. Specifically, the text input device can obtain the aforementioned unified multimodal feature by concatenating the vectors of FeatureVector = concat(TextEmb, GeoEmb, TimeEmb, UserEmb, AppEmb). Here, FeatureVector is the aforementioned unified multimodal feature; TextEmb represents the feature vector encoded using Token Embedding for the aforementioned text context data; GeoEmb represents the feature vector encoded using Geo Embedding for the aforementioned geographic location data; TimeEmb represents the feature vector encoded using Time Slot Embedding for the aforementioned time scene data; UserEmb represents the feature vector encoded using User Preference Embedding for the aforementioned user profile data; and AppEmb represents the feature vector encoded using App Context Embedding for the aforementioned application context data.

[0074] Step 203: The text input device inputs the above-mentioned unified multimodal features into the global basic model to obtain the global basic score corresponding to each candidate content in the candidate pool.

[0075] In some embodiments, the aforementioned global base model is used to provide general prediction capabilities, namely, based on unified multimodal features, quantifying the degree to which each candidate content in the candidate pool conforms to common language logic and popular usage habits.

[0076] In some embodiments, the global base model described above can be a pre-trained lightweight language prediction model.

[0077] In some embodiments, the global basic model described above can be any of the following: a lightweight transformer (LiteTransformer), a long short-term memory network (LSTM / Bi-LSTM), a mobile transformer (Mobile BERT), or a tiny transformer (TinyBERT). Of course, the global basic model can also be other models, which can be determined according to actual needs, and this application does not limit this.

[0078] In some embodiments, the aforementioned global baseline score represents the degree to which candidate content conforms to common language logic and popular usage habits. Specifically, the higher the global baseline score of a candidate content, the more it conforms to common language logic and popular usage habits.

[0079] Step 204: The text input device inputs the unified multimodal features into the personalized lightweight model to obtain the personalized score corresponding to each candidate content in the above candidate pool.

[0080] In some embodiments, the above-mentioned personalized lightweight model employs lightweight parameters such as Low-Rank Adaptation (LoRA) and adapter modules.

[0081] In some embodiments, the personalized lightweight model described above stores only the preferences of individual users.

[0082] In some embodiments, the aforementioned personalized lightweight model resides locally.

[0083] In some embodiments, the above-mentioned personalized lightweight model only learns user-unique patterns (such as frequently used names, addresses, professional terms, etc.), making the prediction results obtained by the above-mentioned personalized lightweight model more closely reflect the user's personal habits.

[0084] In some embodiments, the personalized score represents the degree to which candidate content conforms to the user's personal usage habits.

[0085] Step 205: The text input device determines the input scenario information of the user's input text based on unified multimodal features.

[0086] In some embodiments, the above-mentioned input scenario information is used to characterize the input scenario when the user inputs the above-mentioned user input text.

[0087] Step 206: The text input device determines the final score of each candidate content in the above candidate pool based on the global basic score, the personalized score, and the input scenario information.

[0088] In some embodiments, the final score of a candidate content represents the degree to which the candidate content matches the predicted content of the user input text. Specifically, the higher the final score of a candidate content, the higher the degree to which the candidate content matches the predicted content of the user input text.

[0089] In some embodiments, the text input device may determine the scene enhancement score of each candidate content based on the input scene information, and then determine the final score of each candidate content based on the global base score, personalized score and scene enhancement score of each candidate content.

[0090] It should be noted that the specific implementation process of determining the scene enhancement score of each candidate content based on the input scene information by the text input device, and then determining the final score of each candidate content based on the global basic score, personalized score and scene enhancement score of each candidate content, can be found in the relevant description in the following embodiments. To avoid repetition, this application will not elaborate on it here.

[0091] In some embodiments, combined with Figure 2 ,like Figure 3 As shown, step 206 above can be implemented through steps 206a and 206b.

[0092] Step 206a: The text input device determines the scene enhancement score corresponding to each candidate content in the candidate pool based on the matching degree between each candidate content in the candidate pool and the input scene information.

[0093] In some embodiments, the text input device may determine the scene enhancement score for each candidate content based on the degree of matching between each candidate content and the input scene information.

[0094] For example, assuming that the input scenario information represents the user's input scenario as working at the company, then the scenario enhancement scores of the work-related candidate content in the candidate pool are higher.

[0095] For example, assuming that the input scenario information represents the user's input scenario as resting at home, then the scenario enhancement scores of the candidate content in the candidate pool that are related to life are higher.

[0096] Step 206b: The text input device determines the final score corresponding to each candidate content in the candidate pool based on the global base score, personalized score, and scene enhancement score corresponding to each candidate content in the candidate pool.

[0097] The aforementioned scenario enhancement score represents the degree of matching between the candidate content and the input scenario of the user's input text.

[0098] In some embodiments, the text input device can calculate the final score corresponding to the first candidate content based on the first weight, the second weight, the third weight, and the global base score, personalized score, and scene enhancement score corresponding to the first candidate content. The first weight, the second weight, and the third weight respectively represent the importance of the global base score, the personalized score, and the scene enhancement score.

[0099] It should be noted that the specific implementation process for calculating the final score corresponding to the first candidate content based on the first weight, the second weight, the third weight, and the global basic score, personalized score, and scene enhancement score corresponding to the first candidate content by the text input device can be found in the relevant description in the following embodiments. To avoid repetition, this application will not elaborate on it here.

[0100] In some embodiments, combined with Figure 3 ,like Figure 4 As shown, step 206b can be implemented through steps 206b1 and 206b2.

[0101] Step 206b1: The text input device determines the first weight, second weight, and third weight corresponding to the first candidate content based on the weight information.

[0102] In some embodiments, the weighting information includes at least one of the following: scenario matching degree, historical selection probability, language confidence degree, and personal usage pattern.

[0103] In some embodiments, the scene matching degree described above represents the degree to which the candidate content matches the current input scene.

[0104] In some embodiments, the aforementioned historical selection probability represents the frequency percentage of times a user selects similar candidate content in past input scenarios, reflecting the user's preference for a specific type of candidate content.

[0105] In some embodiments, the above-mentioned language confidence level represents the reliability of candidate content in following general language expression logic, grammatical rules, and semantic rationality.

[0106] In some embodiments, the aforementioned personal usage patterns represent stable input behavior patterns and preference characteristics that users have formed over a long period of time.

[0107] In some embodiments, the aforementioned first weight is used to weight the global base score corresponding to the candidate content, so as to reflect the proportion of the global base score in the final score calculation.

[0108] In some embodiments, the second weight is used to weight the personalized scores corresponding to the candidate content in order to reflect the proportion of personalized scores in the final score calculation.

[0109] In some embodiments, the aforementioned third weight is used to weight the scene enhancement score corresponding to the candidate content, so as to reflect the proportion of the scene enhancement score in the final score calculation.

[0110] Step 206b2: The text input device calculates the final score corresponding to the first candidate content based on the first weight, the second weight, the third weight, and the global basic score, personalized score, and scene enhancement score corresponding to the first candidate content.

[0111] The first candidate is any one of the candidate contents in the aforementioned candidate pool.

[0112] In some embodiments, the text input device can calculate the final score corresponding to the first candidate content using the following fusion formula, namely Formula 1.

[0113] FinalScore=w1*GlobalModelScore+w2*PersonalModelScore+w3*ContextBoost (1)

[0114] Among them, FinalScore represents the final score of the candidate content; w1 represents the first weight; GlobalModelScore represents the global base score of the candidate content; w2 represents the second weight; PersonalModelScore represents the personalized score of the candidate content; w3 represents the third weight; and ContextBoost represents the scene enhancement score of the candidate content.

[0115] In some embodiments, the text input device may determine the first weight and the second weight based on the historical selection probability.

[0116] In some embodiments, the text input device may determine the aforementioned third weight based on the current geographical location, time context, and application type triggering context.

[0117] In this way, the text input device can determine the weights of the global base score, the personalized score, and the scene enhancement score, and then weight the scores according to the weights of each score. Based on the global base score, the personalized score, and the input scene information, it can determine the input prediction text that conforms to the logic of common language and the usage habits of the general public, as well as the user's personal usage habits, and matches the current input scene. This improves the accuracy of the input prediction content.

[0118] In this way, the text input device can determine the scene enhancement score, which represents the degree of matching between each candidate content and the scene of the user's input text, based on the matching degree between each candidate content in the candidate pool and the input scene information. Thus, based on the global base score, the personalized score, and the input scene information, the device can determine the input prediction text that conforms to the logic of general language and the usage habits of the general public, as well as the user's personal usage habits, and also matches the current input scene. This improves the accuracy of the input prediction content.

[0119] Step 207: The text input device determines the predicted text of the user input text based on the final score of each candidate content.

[0120] The aforementioned candidate pool includes at least two candidate entries.

[0121] In some embodiments, the text input device may determine the predicted text of the user input text from the candidate pool based on the final score and score threshold of each candidate content.

[0122] It should be noted that the specific implementation process of the text input device determining the predicted text of the user input text from the above candidate pool based on the final score and score threshold of each candidate content can be found in the relevant description in the following embodiments. To avoid repetition, this application will not elaborate on it here.

[0123] In some embodiments, combined with Figure 3 ,like Figure 5 As shown, step 207 above can be implemented through step 207a as follows.

[0124] Step 207a: The text input device determines the candidate content in the above candidate pool whose final score is greater than the score threshold as the predicted text.

[0125] In some embodiments, the above-mentioned score threshold can be a fixed value, such as 60. Of course, the above-mentioned score threshold can also be other values ​​preset by the system. The specific value can be determined according to actual needs, and this application does not limit it.

[0126] In some embodiments, the text input device may calculate the final score corresponding to each candidate content in the memory pool, and compare the final score corresponding to each candidate content with the score threshold, thereby selecting candidate content in the memory pool whose final score is greater than the score threshold as predicted text.

[0127] In some embodiments, the predicted text includes at least one of the following: next word prediction, input information completion prediction, and scene-related content prediction.

[0128] In some embodiments, the predicted text may also include suggestions for filling in information.

[0129] In some embodiments, the text input device may display the predicted text described above.

[0130] In some embodiments, the text input device may determine the target predicted text as the target text that meets the user's needs based on the user's input of the target predicted text in the predicted text.

[0131] In this way, the text input device can determine the input prediction text based on the final score and score threshold of the candidate content. This text prediction text not only conforms to the logic of general language and the usage habits of the general public, but also conforms to the user's personal usage habits, and matches the current input scenario. This improves the accuracy of the input prediction content.

[0132] The text input method provided in this application collects not only the text context information of the input text but also multimodal context information representing the current input scenario. Based on the text context information and multimodal context information of the input text, it obtains multimodal features that can represent both the language logic of the input text and the current input scenario. Based on these multimodal features, it obtains a global base score reflecting whether each candidate content conforms to general language logic and common usage habits, a personalized score reflecting whether each candidate content conforms to the user's personal usage habits, and input scenario information representing the scenario in which the user inputs the text. Based on the global base score, personalized score, and input scenario information, it determines the input prediction text that conforms to general language logic and common usage habits, conforms to the user's personal usage habits, and matches the current input scenario. This improves the accuracy of the input prediction content.

[0133] In some embodiments, combined with Figure 2 ,like Figure 6 As shown, the text input method provided in this application embodiment may further include the following step 208.

[0134] Step 208: The text input device updates the user preference profile information and the personalized lightweight model based on the target text selected by the user from the predicted text.

[0135] In some embodiments, the text input device may record the target text and save it in a local feature cache for updating the user profile information and updating the model parameters of the personalized lightweight model.

[0136] In some embodiments, the text input device can update user preference profile information and personalized lightweight model through the following process: first, record the predicted hit event, then update the user profile information based on the predicted hit event, and incrementally update the personalized lightweight model. At the same time, a time decay mechanism is introduced to eliminate user personalized features that have not been used for a long time.

[0137] In some embodiments, the text input device can perform the above update process on the terminal side.

[0138] In this way, the text input device can update user profile information and the parameters of the personalized lightweight model that represents the user's personalized input habits in real time, ensuring that the generated predicted text always conforms to the user's personalized usage habits and improving the accuracy of the generated text prediction content.

[0139] It should be noted that step 208 can be performed after step 207.

[0140] The text input method of this application will be described below through specific embodiments.

[0141] For example, such as Figure 7 As shown, the implementation process of the text input method provided in this application embodiment includes the following S1 to S8:

[0142] S1. The text input device collects real-time input context.

[0143] S2. The text input device performs multimodal context acquisition, collecting geographic location information, time scene information, user profile information, and application context information.

[0144] S3. The text input device performs multimodal feature encoding.

[0145] S4. The text input device performs global basic model reasoning.

[0146] S5. Text input device performs personalized model reasoning.

[0147] S6. The text input device performs dynamic weighted decision fusion, that is, the text input device performs scene adaptation, confidence assessment and resource consumption optimization.

[0148] S7. The text input device generates the prediction results.

[0149] S8. The text input device performs incremental learning processing based on user feedback to update the personalized model.

[0150] It should be noted that after the text input device completes step S8, it can jump to step S5.

[0151] Thus, while collecting the text context information of the input text, multimodal context information representing the current input scenario is also collected. Based on the text context information and multimodal context information of the input text, multimodal features that can represent both the language logic of the input text and the current input scenario are obtained. Based on these multimodal features, a global basic score reflecting whether each candidate content conforms to general language logic and common usage habits, a personalized score reflecting whether each candidate content conforms to the user's personal usage habits, and input scenario information representing the scenario in which the user inputs the text are obtained. Based on the global basic score, personalized score, and input scenario information, the predicted input text that conforms to general language logic and common usage habits, conforms to the user's personal usage habits, and matches the current input scenario is determined. This improves the accuracy of the predicted input content. In addition, in the text input method proposed in this application, by updating user profile information and personalized lightweight models in real time, the predicted text generated by the text input device always conforms to the user's personalized usage habits, thus improving the accuracy of the predicted text content generated by the text input device.

[0152] It should be noted that the descriptions of each step S1 to S8 in this embodiment can be found in the descriptions in the above embodiments, and will not be repeated here.

[0153] The steps included in the text input method proposed in this application are described in detail below:

[0154] Step 1: The text input device collects real-time text input context, which includes: the text in the current input box, the context word order or sentence order, the user's recent input history, and the input box type (e.g., address type, name type, search box type, chat input box type, etc.). The input context collected by the text input device is used to construct real-time language sequence features.

[0155] Step 2: The text input device collects multimodal contextual information, which includes multi-dimensional information. Specifically, this multimodal information includes: geographic location information, time context information, user preference profile information, and application context information. Specifically, the geographic location information can include Global Positioning System (GPS) information and latitude and longitude information. The geographic location information collected by the text input device is automatically encoded into a feature vector (embedding) to identify the geographic context information when the user inputs text; the time information collected by the text input device includes: time period information, holiday information, and user daily behavior patterns (e.g., the user inputs "daily report" every morning); the user preference profile collected by the text input device includes: frequently used contact information (e.g., names of frequently used contacts), frequently used address information, keyword preference statistics, and natural language usage standardization information; the application context information collected by the text input device includes: application category information of the currently opened application, and the purpose of the input box (e.g., search, form, chat, notes, etc.); the text input device will subsequently merge the collected multimodal context information and text input context information into a unified multimodal vector.

[0156] Step 3: The text input device performs multimodal feature encoding. The text input device will uniformly encode the aforementioned text input context information and the aforementioned multimodal context information. Specifically, for the text data in the aforementioned text input context information and the aforementioned multimodal context information, the text input device encodes it using token embedding to obtain a text feature vector; for the geographic location data in the aforementioned text input context information and the aforementioned multimodal context information, the text input device encodes it using geo embedding to obtain a geographic location feature vector; for the time scene data in the aforementioned text input context information and the aforementioned multimodal context information, the text input device encodes it using time slot embedding to obtain a time scene feature vector; for the application context data in the aforementioned text input context information and the aforementioned multimodal context information, the text input device encodes it using embedding to obtain an application context feature vector; for the user preference profile data in the aforementioned text input context information and the aforementioned multimodal context information, the text input device encodes it using embedding to obtain a user preference profile feature vector. The text input device performs feature fusion processing on the aforementioned text feature vector, the aforementioned geographic location feature vector, the aforementioned time scene feature vector, the aforementioned application context feature vector, and the aforementioned user preference profile feature vector to obtain a multimodal feature vector.

[0157] Step 4: The text input device inputs the aforementioned multimodal feature vectors into the global base model for inference to obtain the global base score corresponding to each candidate content in the candidate pool. The aforementioned global base model is a pre-trained lightweight language prediction model, which includes at least one of the following: Lite Transformer, LSTM / Bi-LSTM, MobileBERT, and TinyBERT.

[0158] Step 5: The text input device inputs the aforementioned multimodal feature vectors into the personalized lightweight model for inference. This personalized lightweight model uses lightweight parameters such as LoRA or Adapter, and it only stores the preferences of individual users. Furthermore, the personalized lightweight model resides locally and does not need to be uploaded to the cloud. The aforementioned personalized lightweight model has the following characteristics: fast convergence, learning only user-unique patterns (e.g., frequently used names, addresses, professional vocabulary), and prediction results that are closer to individual habits.

[0159] Step 6: The text input device performs scene-aware fusion and weight adjustment. Specifically, the text perception device can dynamically adjust the model weights according to the current scene. For example, in the office, it prioritizes predicting work-related words (weights biased towards the personalized model); at home, it prioritizes predicting everyday words; and at the airport, it increases the probability of words such as "flight number" and "gate". The text input device determines the first, second, and third weights based on the weight information and calculates the final score for each candidate in the candidate pool using the aforementioned fusion formula.

[0160] Step 7: The text input device generates predicted content. This predicted content includes, but is not limited to, the following: next word prediction, next character or sentence completion, fill-in suggestions (e.g., address, name), and scene-related content (e.g., airport, conference, restaurant) provided to the user as input suggestions.

[0161] Step 8: Local incremental learning based on user feedback. The text input device records the prediction results selected by the user and saves them in a local feature cache. This cache is used to update the aforementioned user profile information and the personalized model parameters. Specifically, when the user selects a prediction, the text input device records the prediction hit event and updates the user profile vector based on the prediction hit event, as well as incrementally updating the personalized model parameters. Furthermore, this application introduces a time decay mechanism, meaning the text input device discards features that have not been used for a long time.

[0162] For example, the text input method proposed in this application will be described below using the execution of the text input method on an iOS client as an example. Specifically, the execution of the text input method proposed in this application on an iOS client includes the following process:

[0163] Step 1: Predictive triggering and scenario determination based on input events.

[0164] Specifically, in the iOS client, the prediction process is triggered by listening to text input events and detecting character changes or cursor movement. The text input device first determines the current input scenario type based on the business identifier of the input control, the page context, and placeholder information, which is used for subsequent model feature selection and parameter configuration.

[0165] Step 2: Multimodal context acquisition and structured modeling.

[0166] Specifically, after prediction is triggered, the client synchronously collects multimodal context information, including the following: the currently input text sequence, the current system time, the acquired latitude and longitude information, the current application module and page identifier, and locally stored user profile features. This information is preprocessed and organized into a structured feature dictionary, serving as the input feature set for both the general basic prediction model and the personalized lightweight large model.

[0167] Step 3: Feature Vectorization and Large Model Input Mapping

[0168] Specifically, the client has a built-in feature encoding module that vectorizes features of different modalities: text input sequences are mapped to word or subword embedding vectors; time, location, and application context are mapped to discrete or continuous embedding vectors; and user profile features are mapped to personalized bias vectors. These vectors are then concatenated in a predefined order and mapped to the multi-input feature interface defined by the general basic prediction model.

[0169] Step 4: Load and maintain at least two model instances in the multi-model collaborative inference client based on the general basic large model and the personalized lightweight large model.

[0170] Specifically, the aforementioned general-purpose basic prediction model is used to provide stable general language prediction capabilities; the aforementioned personalized lightweight prediction model is used to characterize users' long-term and short-term input preferences. During the prediction phase, the client simultaneously inputs the fused multimodal features into the aforementioned general-purpose basic model and personalized lightweight model, and obtains candidate prediction results and their probability distributions through the inference interfaces of the general-purpose basic model and personalized lightweight model.

[0171] Step 5: Prediction results fusion and candidate output.

[0172] Specifically, the text input device, through the client, performs a weighted fusion of the prediction results from the general model and the personalized model according to the fusion strategy configured for the current scene, and outputs candidate results after sorting them according to the prediction probability. The prediction results are presented to the user in real time through the input method candidate bar or the input control auxiliary view.

[0173] Step Six: Incremental training and model update on the edge based on the general basic large model and the personalized lightweight large model.

[0174] Specifically, when a user selects a prediction result, the client treats this selection as a positive sample: it encapsulates the input context and the user's selection result into a training sample; it writes the sample to a local training sample cache; it calls Core ML's model update interface to perform incremental training on the personalized lightweight model; after updating the model parameters, it replaces the model instance in memory for subsequent inference. This training process is executed asynchronously in a background thread and does not affect the foreground input experience.

[0175] Step 7: Training sample and model lifecycle management.

[0176] Specifically, to control resource consumption, the text input device implements constraint management during the training process: it employs time decay and capacity capping mechanisms to eliminate old training samples; when device performance or battery power is insufficient, it pauses model updates and only performs inference; and it periodically performs lightweight refactoring of the personalized model to prevent overfitting. By combining the multimodal context fusion algorithm with the edge-side inference and incremental training capabilities of the general-purpose basic large model and the personalized lightweight large model, this application can complete input prediction and adaptive model updates locally on the iOS client, achieving a privacy-preserving, real-time responsive, and continuously evolving personalized intelligent input effect, demonstrating clear engineering feasibility and verifiable technical effects.

[0177] In some embodiments, the text input method proposed in this application involves the following algorithms: a multimodal context feature construction algorithm, a multimodal attention fusion algorithm, a dual-model input prediction algorithm, a scene-aware prediction result fusion algorithm, a user selection feedback and incremental learning algorithm, and a time decay and model stability control algorithm. Each algorithm involved in this application is described in detail below:

[0178] Algorithm 1: Multimodal Context Feature Construction Algorithm

[0179] Specifically, the input information for this multimodal context feature construction algorithm includes: the current input character sequence, time information, location information, application type, and user profile. The text input device can generate multimodal features based on this algorithm through the following steps: segmenting or encoding the input character sequence to generate a text vector; mapping the time information to discrete time periods (e.g., hourly segments, weekdays, or weekends) and generating a time vector; mapping latitude and longitude information to semantic location labels and generating a location vector; encoding the application type and input scenario into an application context vector; extracting user preference vectors from the local user profile; and combining these vectors into a feature set, i.e., multimodal features.

[0180] Algorithm 2: Multimodal Attention Fusion Algorithm

[0181] Specifically, the input information of this multimodal attention fusion algorithm includes the aforementioned feature set, i.e., multimodal features. The text input device can generate a multimodal fusion vector based on this multimodal attention fusion algorithm through the following steps: using a text vector as the query vector; calculating the similarity score between the text vector and the other modal vectors respectively; normalizing the similarity scores to obtain the weights of each modality; weighting and summing the modal vectors according to their weights; and concatenating the weighted result with the text vector to generate the final multimodal fusion vector.

[0182] Algorithm 3: Dual-model input prediction algorithm

[0183] Specifically, the input information of the dual-model input prediction algorithm includes the aforementioned final multimodal fusion vector. The text input device can obtain the predicted probability distribution of each candidate content based on the dual-model input prediction algorithm through the following process: inputting the final multimodal fusion vector into the global base model to obtain the predicted probability distribution; inputting the final multimodal fusion vector into the personalized lightweight model to obtain the predicted probability distribution; and retaining the candidate content with the highest predicted probability distribution and its probability.

[0184] Algorithm 4: Scene-Aware Prediction Result Fusion Algorithm

[0185] Specifically, the input information for the scene-aware prediction result fusion algorithm includes: the prediction probability distribution of candidate content and the current context information. The text input device can generate prediction results through the following steps: determine the current scene type based on the current context information (e.g., address input, search input); read the corresponding initial fusion weights from the local configuration table; fine-tune the weights based on historical prediction hit rates; calculate the final prediction score according to the formula; and output the prediction results in sorted order based on the final scores.

[0186] Algorithm 5: User Selection Feedback and Incremental Learning Algorithm

[0187] Specifically, the input information for the user selection feedback and incremental learning algorithm includes the prediction result selected by the user. The text input device can update the user profile and the personalized lightweight model through the following steps: marking the prediction result selected by the user as a positive sample; updating the statistical frequency of the corresponding entity or phrase in the user profile; performing a small-step gradient update on the trainable parameters of the personalized lightweight model; and performing time decay processing on parameters that have been missing for a long time.

[0188] Algorithm 6: Time Decay and Model Stability Control Algorithm

[0189] Specifically, the text input device can periodically calculate the last hit time of each user feature using the time decay and model stability control algorithm; and decay its weight when it exceeds a preset time threshold; or remove the corresponding feature from the user profile when the weight is below the threshold.

[0190] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there is no conflict, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0191] Figure 8 This is a schematic diagram of the structure of a text input system provided in an embodiment of this application. Figure 8 As shown, the text input system 800 may include: a text context acquisition module 801, a geographic location acquisition module 802, a time scene acquisition module 803, a multimodal feature encoding module 804, a global basic prediction model 805, a scene perception fusion and weight decision module 806, a prediction output module 807, a user selection feedback and incremental learning module 808, a personalized lightweight model 809, and a user preference profile module 810.

[0192] The text context acquisition module 801 is used to acquire the input context information of the user's input text and is applied to step 201 and its related schemes. The geographic location acquisition module 802 is used to acquire the client's geographic location information from the multimodal context information and is applied to step 201 and its related schemes. The time scene acquisition module 803 is used to acquire the client's time scene information from the multimodal context information and is applied to step 201 and its related schemes. The multimodal feature encoding module 804 is used to perform data fusion on the acquired client's geographic location information and client's time scene information to obtain the unified multimodal feature. The global basic prediction model 805 is used to determine the global basic score of each candidate content in the candidate pool based on the unified multimodal feature and through the scene-aware fusion and weight decision module 806. The aforementioned prediction output module 807 is used to generate predicted text of the user input text, and is applied to step 207 and the related schemes in step 207; the aforementioned user selection feedback and incremental learning module 808 is used to update the personalized lightweight model based on the predicted text generated by the aforementioned prediction output module 807 and according to the user's selection, and is applied to step 208 and the related schemes in step 208; the aforementioned personalized lightweight model 809 is used to calculate the personalized score of each candidate content in the memory pool, and is applied to step 204 and the related schemes in step 204.

[0193] It should be noted that for a detailed explanation of the steps performed by each module and their beneficial effects, please refer to the description in the above embodiments, which will not be repeated here.

[0194] As can be seen, the above mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the embodiments of this application provide corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the modules and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0195] This application embodiment can divide the text input device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. Optionally, the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0196] In some embodiments, this application also provides a text input device. The text input device may include one or more functional modules for implementing the text input method of the above method embodiments.

[0197] For example, Figure 9 This is a schematic diagram of the structure of a text input device provided in an embodiment of this application. Figure 9 As shown, the text input device 900 includes an acquisition module 901 and a processing module 902.

[0198] The aforementioned acquisition module 901 is used to collect input context information and multimodal context information of user input text. This multimodal context information includes at least one of the following: the client's geographic location information, the client's time context information, the client's user preference profile information, and the client's application context information; and, based on the aforementioned input context information and multimodal context information, to acquire unified multimodal features. The aforementioned processing module 902 is used to input the aforementioned unified multimodal features into a global base model to obtain a global base score corresponding to each candidate content in the candidate pool. This global base score represents that the candidate content conforms to common language logic and is widely used. The process involves: determining the degree to which the user's input text conforms to their personal usage habits; inputting the aforementioned unified multimodal features into a personalized lightweight model to obtain a personalized score for each candidate content in the aforementioned candidate pool, whereby the personalized score characterizes the degree to which the candidate content conforms to the user's personal usage habits; determining the input scenario information of the aforementioned user input text based on the aforementioned unified multimodal features; determining the final score of each candidate content in the aforementioned candidate pool based on the aforementioned global base score, the aforementioned personalized score, and the aforementioned input scenario information; and determining the predicted text of the aforementioned user input text based on the final score of each candidate content; wherein the aforementioned candidate pool includes at least two candidate contents.

[0199] In the text input device provided in this application, this solution collects not only the text context information of the input text but also multimodal context information representing the current input scenario. Based on the text context information and multimodal context information of the input text, it obtains multimodal features that can represent both the language logic of the input text and the current input scenario. Based on these multimodal features, it obtains a global basic score reflecting whether each candidate content conforms to general language logic and common usage habits, a personalized score reflecting whether each candidate content conforms to the user's personal usage habits, and input scenario information representing the scenario in which the user inputs the text. Based on the global basic score, personalized score, and input scenario information, it determines the input prediction text that conforms to both general language logic and common usage habits, as well as the user's personal usage habits, and also matches the current input scenario. This improves the accuracy of the input prediction content.

[0200] In some embodiments, the processing module 902 is specifically configured to: determine the scene enhancement score corresponding to each candidate content in the candidate pool based on the matching degree between each candidate content in the candidate pool and the input scene information; and determine the final score corresponding to each candidate content in the candidate pool based on the global basic score, the personalized score, and the scene enhancement score corresponding to each candidate content in the candidate pool; wherein the scene enhancement score represents the degree of matching between the candidate content and the input scene of the user input text.

[0201] In other embodiments, the processing module 902 is specifically used to: determine the first weight, second weight, and third weight corresponding to the first candidate content based on the weight basis information, wherein the weight basis information includes at least one of the following: scene matching degree, historical selection probability, language confidence degree, and personal usage pattern; calculate the final score corresponding to the first candidate content based on the first weight, the second weight, the third weight, and the global basic score, the personalized score, and the scene enhancement score corresponding to the first candidate content, wherein the first candidate content is any one of the candidate contents in the candidate pool.

[0202] In some other embodiments, the processing module 902 is specifically used to: determine the candidate content in the candidate pool whose final score is greater than the score threshold as the predicted text; the predicted text includes at least one of the following: next word prediction, input information completion prediction, and scene-related content prediction.

[0203] In some other embodiments, the processing module 902 is further configured to update the user preference profile information and the personalized lightweight model based on the target text selected by the user from the predicted text.

[0204] It should be noted that the text input device can implement all the processes implemented in the above method embodiments and achieve the same beneficial effects. To avoid repetition, it will not be described again here.

[0205] In the case where the functions of the integrated modules described above are implemented in hardware, this application provides a possible structural schematic diagram of the electronic device involved in the above embodiments. For example... Figure 10 As shown, the electronic device 90 includes: a processor 92, a communication interface 93, and a bus 94. Optionally, the electronic device 90 may also include a memory 91.

[0206] Processor 92 may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 92 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 92 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0207] Communication interface 93 is used to connect with other devices via a communication network. This communication network can be Ethernet, wireless access network, wireless local area network (WLAN), etc.

[0208] The memory 91 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0209] As one possible implementation, the memory 91 can exist independently of the processor 92. The memory 91 can be connected to the processor 92 via a bus 94 and is used to store instructions or program code. When the processor 92 calls and executes the instructions or program code stored in the memory 91, it can implement the text input method provided in the embodiments of this application.

[0210] In another possible implementation, memory 91 can also be integrated with processor 92.

[0211] Bus 94 can be an Extended Industry Standard Architecture (EISA) bus, etc. Bus 94 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0212] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the service calling device can be divided into different functional modules to complete all or part of the functions described above.

[0213] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described text input method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0214] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0215] This application also provides a readable storage medium storing a program or instructions that, when executed by a computer, implement the text input method provided in the above embodiments. It is understood that all or part of the processes in the above method embodiments can be executed by computer instructions instructing related hardware; the readable storage medium can be any of the foregoing embodiments or memory; the readable storage medium can also be an external storage device of the service invocation device, such as a pluggable hard drive, Smart MediaCard (SMC), Secure Digital (SD) card, flash card, etc., equipped on the service invocation device. Further, the readable storage medium can include both internal storage units of the service invocation device and external storage devices. The readable storage medium is used to store the computer program and other programs and data required by the service invocation device. The readable storage medium can also be used to temporarily store data that has been output or will be output.

[0216] This application also provides a computer program product, which is stored in a storage medium and implements the text input method provided in the above embodiments when executed by a computer.

[0217] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0218] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0219] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A text input method, characterized by, include: The system collects input context information and multimodal context information of user input text. The multimodal context information includes at least one of the following: the client's geographical location information, the client's time context information, the client's user preference profile information, and the client's application context information. Based on the input context information and the multimodal context information, unified multimodal features are obtained; The unified multimodal features are input into the global base model to obtain the global base score corresponding to each candidate content in the candidate pool. The global base score represents the degree to which the candidate content conforms to the logic of common language and the usage habits of the general public. The unified multimodal features are input into the personalized lightweight model to obtain the personalized score corresponding to each candidate content in the candidate pool. The personalized score represents the degree to which the candidate content conforms to the user's personal usage habits. Based on the unified multimodal features, the input scenario information of the user input text is determined; Based on the global base score, the personalized score, and the input scenario information, the final score of each candidate content in the candidate pool is determined; Based on the final score of each candidate content, the predicted text of the user input text is determined; The candidate pool includes at least two candidate contents.

2. The text input method of claim 1, wherein, The step of determining the final score for each candidate content in the candidate pool based on the global base score, the personalized score, and the input scenario information includes: Based on the matching degree between each candidate content in the candidate pool and the input scene information, the scene enhancement score corresponding to each candidate content in the candidate pool is determined; Based on the global base score, the personalized score, and the scene enhancement score corresponding to each candidate content in the candidate pool, the final score corresponding to each candidate content in the candidate pool is determined; The scene enhancement score represents the degree of matching between the candidate content and the input scene of the user input text.

3. The text input method of claim 2, wherein, The step of determining the final score for each candidate content in the candidate pool based on the global base score, the personalized score, and the scene enhancement score for each candidate content in the candidate pool includes: Based on the weighting information, the first weight, second weight and third weight corresponding to the first candidate content are determined. The weighting information includes at least one of the following: scene matching degree, historical selection probability, language confidence degree and personal usage pattern. Based on the first weight, the second weight, the third weight, and the global base score, the personalized score, and the scene enhancement score corresponding to the first candidate content, the final score corresponding to the first candidate content is calculated, wherein the first candidate content is any candidate content in the candidate pool.

4. The text input method of claim 2, wherein, Determining the predicted text of the user input text based on the final score of each candidate content includes: Candidate content in the candidate pool whose final score is greater than the score threshold is determined as the predicted text; The predicted text includes at least one of the following: next word prediction, input information completion prediction, and scene-related content prediction.

5. The text input method according to any one of claims 1 to 4, wherein, The method further includes: Based on the target text selected by the user from the predicted text, the user preference profile information and the personalized lightweight model are updated.

6. A text input device, characterized by include: Acquisition module and processing module; The acquisition module is used to collect input context information and multimodal context information of user input text. The multimodal context information includes at least one of the following: the client's geographical location information, the client's time scene information, the client's user preference profile information, and the client's application context information. as well as, Based on the input context information and the multimodal context information, unified multimodal features are obtained; and, The processing module is used to input the unified multimodal features into the global base model and obtain the global base score corresponding to each candidate content in the candidate pool. The global base score represents the degree to which the candidate content conforms to the logic of common language and the usage habits of the general public. as well as, The unified multimodal features are input into the personalized lightweight model to obtain the personalized score corresponding to each candidate content in the candidate pool. The personalized score represents the degree to which the candidate content conforms to the user's personal usage habits. as well as, Based on the unified multimodal features, the input scenario information of the user input text is determined; as well as, Based on the global base score, the personalized score, and the input scenario information, the final score of each candidate content in the candidate pool is determined; as well as, Based on the final score of each candidate content, the predicted text of the user input text is determined; The candidate pool includes at least two candidate contents.

7. The text input device of claim 6, wherein, The processing module is specifically used for: Based on the matching degree between each candidate content in the candidate pool and the input scene information, the scene enhancement score corresponding to each candidate content in the candidate pool is determined; Based on the global base score, the personalized score, and the scene enhancement score corresponding to each candidate content in the candidate pool, the final score corresponding to each candidate content in the candidate pool is determined; The scene enhancement score represents the degree of matching between the candidate content and the input scene of the user input text.

8. The text input apparatus according to claim 7, wherein The processing module is specifically used for: Based on the weighting information, the first weight, second weight and third weight corresponding to the first candidate content are determined. The weighting information includes at least one of the following: scene matching degree, historical selection probability, language confidence degree and personal usage pattern. Based on the first weight, the second weight, the third weight, and the global base score, the personalized score, and the scene enhancement score corresponding to the first candidate content, the final score corresponding to the first candidate content is calculated, wherein the first candidate content is any candidate content in the candidate pool.

9. The text input apparatus according to claim 7, wherein The processing module is specifically used for: Candidate content in the candidate pool whose final score is greater than the score threshold is determined as the predicted text; The predicted text includes at least one of the following: next word prediction, input information completion prediction, and scene-related content prediction.

10. The text input device according to any one of claims 6 to 9, characterized in that, The processing module is also used to update the user preference profile information and the personalized lightweight model based on the target text selected by the user from the predicted text.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the text input method as described in any one of claims 1 to 5.

12. A readable storage medium, characterized by, The readable storage medium stores a program or instructions that, when executed by a computer, implement the text input method as described in any one of claims 1 to 5.

13. A computer program product, characterised in that, The computer program product is stored in a storage medium, and when executed by a computer, the computer program product implements the text input method as described in any one of claims 1 to 5.