A bidirectional context-aware text completion method and computer equipment
By acquiring bidirectional contextual information and multimodal feature representations of the input text, the system adaptively generates completion content, solving the operational complexity problem caused by LLM's reliance on user input and improving the efficiency and accuracy of text completion.
Patent Information
- Application Number
- CN202511395883.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Existing Large Language Models (LLMs) rely on user input for text completion, which leads to complex operations, low interaction efficiency, and negatively impacts user experience.
By acquiring bidirectional contextual information from the input text, and combining multimodal feature representations and attention weight distribution, the system adaptively generates complete content, reducing user operation steps and improving interaction efficiency.
It enables the rapid and accurate generation of context-appropriate completion content without requiring additional user input, improving the accuracy and comprehensiveness of natural language completion and adapting to various scenarios.
Smart Images

Figure CN120873767B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing, and in particular to a text completion method and computer device based on bidirectional context awareness. Background Technology
[0002] With the development and application of Internet and artificial intelligence technologies, Large Language Models (LLMs) have been widely used in scenarios such as education / training, online social networking, and content creation. Through corresponding prompts (such as user-inputted instructions or questions), LLMs are guided to identify user intent and interaction needs, and generate and provide responses that meet those needs.
[0003] It is evident that in the current question-and-answer interaction mode based on LLM, the response of LLM depends entirely on user input. Thus, when it is necessary to enrich and enhance the input text, the user needs to input additional completion requirements, forming an input LLM model for completion prompts to generate and return the corresponding completion content. This completion content is then manually inserted into the input text, which is a complex process that results in low interaction efficiency and affects the user experience. Summary of the Invention
[0004] In view of the above problems, this application provides the following solution:
[0005] The first aspect of this application provides a text completion method based on bidirectional context awareness, the method comprising:
[0006] In response to a completion trigger signal for the input text presented on the interactive interface, bidirectional contextual information in the input text is obtained; the bidirectional contextual information includes the preceding text segment before the input marker position and the following text segment after the input marker position.
[0007] Based on the preceding and following fragments, a fused contextual feature representation is generated;
[0008] Obtain a multimodal feature representation associated with the context feature representation; the multimodal feature representation includes feature representations of at least one multimodal information from interactive behavior data, text data, image data, audio data, and video data;
[0009] A correlation analysis is performed on the context feature representation and the multimodal feature representation to obtain the corresponding attention weight distribution; the attention weight distribution can indicate the degree of contribution of different multimodal information to the generation of the completion content for the input text;
[0010] Based on the attention weight distribution, the multimodal feature representation, and the context feature representation, complete content is generated for the input text;
[0011] The completed content is displayed on the interactive interface.
[0012] A second aspect of this application also provides a computer device, the computer device comprising: at least one memory and at least one processor, wherein:
[0013] The memory is used to store multiple computer instructions;
[0014] The processor is configured to execute the computer instructions to implement the various steps of the bidirectional context-aware text completion method provided in the first aspect of this application.
[0015] Therefore, in the bidirectional context-aware text completion method proposed in this application, to prevent frequent or accidental triggering of the completion task, the completion task for the input text is only executed upon response to a completion trigger signal. Combined with the FIM model mechanism, bidirectional contextual information in the input text is obtained. After generating a fused contextual feature representation, it is combined with multimodal feature representations associated with this contextual feature representation to comprehensively, accurately, and flexibly understand the user's completion intent in the current scenario. Through correlation analysis between the two, the attention weight distribution of the multimodal feature representation is obtained to understand the contribution of different multimodal information to the generated completion content. Then, based on this attention weight distribution, multimodal feature representation, and contextual feature representation, the generated completion content is ensured to be both context-appropriate and logically consistent, improving the accuracy and comprehensiveness of natural language completion and adapting to various scenarios. Furthermore, it eliminates the need for users to switch applications or require additional input for completion, reducing operational steps and improving user experience. Attached Figure Description
[0016] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0017] Figure 1 This is a flowchart illustrating the bidirectional context-aware text completion method proposed in Embodiment 1 of this application.
[0018] Figure 2 This is a flowchart illustrating the bidirectional context-aware text completion method proposed in Embodiment 2 of this application;
[0019] Figure 3This is a flowchart illustrating the bidirectional context-aware text completion method proposed in Embodiment 3 of this application;
[0020] Figure 4 A schematic diagram of a configuration interface for implementing personalized configuration in this application;
[0021] Figure 5 This is a flowchart illustrating the bidirectional context-aware text completion method proposed in Embodiment 4 of this application;
[0022] Figure 6 This is a schematic diagram of the hardware structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0023] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only for explaining specific embodiments and is not intended to limit the application. The terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence; they can be interchanged where appropriate. This is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0024] To address the problems raised in the background section, this application proposes a bidirectional context-aware text completion method. The following detailed description, in conjunction with the accompanying drawings, provides an embodiment of the bidirectional context-aware text completion method of this application.
[0025] Reference Figure 1 This is a flowchart illustrating the bidirectional context-aware text completion method proposed in Embodiment 1 of this application. This method can be applied to computer devices, such as cloud servers or local terminal devices with sufficient computing resources. Figure 1 As shown, the text completion method based on bidirectional context awareness proposed in this embodiment may include, but is not limited to:
[0026] Step S11: In response to the completion trigger signal of the input text presented on the interactive interface, obtain bidirectional context information in the input text; the bidirectional context information includes the preceding text segment before the input marker position and the following text segment after the input marker position.
[0027] In natural language processing scenarios, after users access the corresponding interactive interface (such as a created conversation interface, text editing interface, or content creation interface) through professional applications or web pages, they can directly input (such as using input components like a mouse, keyboard, or stylus) questions, task descriptions, social content, or text data converted from audio data collected in real time by an audio collector, or text data converted from sign language video data collected in real time by a video collector, etc., which can be presented as input text on the interactive interface. This application does not restrict the content of the input text or its acquisition and presentation methods, and can be determined as appropriate.
[0028] For the input text presented in the current interactive interface, if the user wants to add content to the input text (i.e., complete the content), such as when the comment on a hotel / friend's post is short, lacks originality / sincerity, etc., and wants to enrich the already entered comment content, the method proposed in this application will be automatically triggered. This method adaptively and comprehensively understands the user's intent to efficiently and accurately generate complete content that meets the completion requirements of the input text in the current scenario. Therefore, this application proposes an adaptive triggering mechanism that can automatically identify the need for completion operation on the input text presented in the interactive interface and its related information. If necessary, it will automatically generate or respond to the corresponding completion trigger signal to trigger the execution of subsequent method steps.
[0029] As can be seen, this application automatically identifies and generates a completion trigger signal directly when the input text is presented on the output interactive interface, in response to perform the completion operation. This eliminates the need for the user to switch to dedicated software such as an LLM-based AI assistant or a smart question-and-answer window, and to additionally input the actual completion requirements for the input text (such as comment content), significantly reducing user operation steps. It should be noted that this application does not limit the method of generating the completion trigger signal or its signal form; it can be implemented based on a pre-configured / adjusted trigger mechanism / condition. The embodiments of this application will not be detailed here.
[0030] Furthermore, this application proposes to achieve adaptive natural language semantic completion by integrating a multimodal FIM (Fill-in-the-Middle, a method for training and operating Large Language Models (LLM)) model. This breaks through the strict left-to-right generation order of traditional models (such as general generative models), enabling fast, accurate, and flexible completion of user intent in input text during natural language input, improving interaction efficiency and quality, and adapting to various scenarios.
[0031] Therefore, when a computer device determines that a user wants to perform the FIM function (i.e., generate missing content based on context information) on the input text, it can use a matching extraction method based on the conditional content that generates the completion trigger signal to obtain the local text before and after the content to be completed from the input text, denoted as bidirectional context information. This bidirectional context information is used to infer the missing content, i.e., the completion content of the input text, according to the method described below. This application does not limit the method for obtaining bidirectional context information. Wherein, if the input identifier position represents the display position of the completion content, the bidirectional context information includes the preceding text segment (also called the prefix) before the input identifier position and the following text segment (also called the suffix) after the input identifier position, etc., to form bidirectional context information in a preset data format according to the input requirements of the subsequent model.
[0032] Step S12: Generate a fused contextual feature representation based on the preceding and following fragments;
[0033] Step S13: Obtain the multimodal feature representation associated with the context feature representation; the multimodal feature representation is a feature representation of at least one of the following multimodal information: interactive behavior data, text data, image data, audio data, and video data.
[0034] To achieve a more comprehensive, three-dimensional, and in-depth understanding of the user's completion intent, this application combines multimodal information associated with bidirectional contextual information to achieve accurate and effective completion of input text. Compared to the LLM model, which relies solely on user input to generate completion content, this application uses multimodal information to complete user intent. For example, by combining user behavior data related to browsing hotel services, it is easier to accurately supplement service-related descriptions when completing hotel reviews. This application does not restrict the implementation method of obtaining multimodal information. Furthermore, since the bidirectional contextual information obtained from different input texts in real-world scenarios often differs, and the associated multimodal information also differs, even for the same input text from different times / users within the same scenario, the obtained multimodal information often varies, such as differences caused by different user interactions. Therefore, the number of modalities, categories (e.g., images, text, audio, video, or behavioral sequences), and content of the associated multimodal information obtained for different users' input texts may differ depending on the specific bidirectional contextual information. This application can determine these differences based on the actual situation.
[0035] In some embodiments, during a user's operation on an interactive platform in various scenarios, various interactive behavior data generated can be recorded, such as browsing trajectories, search history / results (e.g., structured behavioral sequences such as crawled scene images, videos, or audio files to represent the user's historical intent, operating habits, and preferences), likes, comments, etc., and stored as historical data in a database. This allows for the retrieval of multimodal information associated with bidirectional contextual information or input text from the database or other data sources. This application does not limit the retrieval implementation method. For example, taking a hotel review scenario as an example, the input text is a natural language description of the user's experience at the hotel. Based on the user's identity, hotel services, reviews, and other information, relevant searches can be performed on the historical data of the hotel review platform. If the retrieved data shows that the user has recently frequently browsed hotel content related to high-quality services (i.e., user interaction behavior data associated with hotel services), this can be used as multimodal information to supplement the input text with content related to the user's service attitude, thereby enriching the hotel review content.
[0036] Optionally, this application can perform semantic analysis on the input text or bidirectional context information to determine its explicit reference content (such as one or more of the referenced file name / path, timestamp, spatial ID, and data type) or implicit semantic reasoning results (such as visual / auditory / behavioral keywords or domain context), thereby obtaining corresponding multimodal information. In some embodiments, this application can also obtain multimodal information associated with the bidirectional context information based on the context-aware approach of the interactive application environment. In one possible implementation, the currently active application can be determined. For example, if another application is already running (an application not belonging to the interactive interface, which can be denoted as the second application), the application belonging to the interactive interface (which can be denoted as the first application) can be run simultaneously, and information interaction can be performed on the interactive interface. Both the first application and the second application can be referred to as active applications. Then, multimodal information can be obtained from the current processing data of the second application. In another possible implementation, this application can also obtain multimodal information associated with the bidirectional context information from historical data generated by the user's recent operations. If the multimodal information involves privacy data that needs to be protected (such as camera footage, user's private documents, user identity information, etc., which can be flexibly configured by the user), permission verification is required first to improve data security.
[0037] To enable the model to recognize bidirectional contextual information and multimodal information, a pre-trained word embedding model is typically used before inputting this information into the model. This converts the corresponding modal information into vector representations, resulting in contextual feature representations and multimodal feature representations. The word embedding model can be a known general-purpose word embedding model, such as a Transformer-based cross-modal model, a multimodal encoder-decoder model, or an audio-text model. Depending on the specific task requirements, one or more combinations of these general-purpose word embedding models can be flexibly selected to vectorize bidirectional contextual information and various multimodal information. The implementation process is not detailed in this application.
[0038] Preferably, to improve processing efficiency, for various multimodal information, including but not limited to those described above, the information is pre-converted into corresponding modal feature representations using a word embedding model and then stored, or each modal feature representation is stored independently. In this way, after triggering the text completion task, there is no need to perform online vectorization processing; vector similarity calculation can be used to directly select modal feature representations associated with the context feature representations from the stored modal feature representations to construct the aforementioned multimodal feature representations. Optionally, this application can construct a multimodal feature cache pool configured with a TTL (Time To Live, i.e., the maximum validity period of cached data) to cache modal feature representations generated within the most recent TTL (e.g., TTL = 30 minutes or 1 hour, depending on the situation). Modal feature representations whose actual cache duration exceeds the TTL are cleared / deleted to ensure the freshness of the cached modal feature representations. This automated cache management releases cache space promptly, prevents infinite accumulation, and avoids wasting cache resources. This application does not limit the caching method for each modal feature representation.
[0039] Step S14: Perform correlation analysis on the context feature representation and multimodal feature representation to obtain the corresponding attention weight distribution; this attention weight distribution can indicate the degree of contribution of different multimodal information to the generation of the complete content for the input text;
[0040] Following the above analysis, since the modal information indicated by the multimodal feature representations contributes differently to the actual completion intention, it is desirable to reduce the modal information with low contribution and increase the modal information with high contribution during the completion content generation process. Therefore, this application can perform correlation analysis on the context feature representation and the multimodal feature representation to determine which part of the obtained multimodal feature representation is highly relevant to the completion task and which part is not. This allows for the allocation of higher attention weights to highly relevant feature representations and lower attention weights to less relevant feature representations, so that the subsequent completion content generation process adaptively focuses on the feature representation with high attention weights and ensures that the feature representation conforms to the semantic logic of the input text while meeting the actual completion requirements of the input text. Therefore, in the implementation of step S14, a cross-modal attention mechanism can be used. The obtained attention weight distribution can be a probability distribution of each attention weight value within (0,1), with the sum of all attention weight values being 1. This application does not limit the implementation method.
[0041] Step S15: Based on attention weight distribution, multimodal feature representation and contextual feature representation, generate complete content for the input text;
[0042] Step S16: Display the completed content on the interactive interface.
[0043] Following the above analysis, for the task of text completion in natural language processing, this application not only considers the contextual feature representation (bidirectional contextual information) obtained from the input text, but also the associated multimodal feature representation (multimodal information) and its contribution to the completed content. This ensures that the generated completed content is both context-appropriate and logically sound, improving the relevance and rationality of text completion and achieving more accurate natural language completion. For example, in the hotel review scenario, by combining user behavior data on browsing hotel service-related content, it is easier to accurately supplement service-related descriptions when completing hotel reviews, enriching the hotel review content. This solves the problem in text completion methods based on local word frequency statistics of N-gram models, which lack the ability to analyze the global semantics of long texts, making it difficult to understand the overall semantics and contextual logic of the text, resulting in the inability to accurately complete content such as missing service attitude descriptions in hotel reviews, thus affecting the accuracy of text completion.
[0044] The generated completion content can be directly inserted into the input marker position in the input text presented on the interactive interface to update the input text, or the generated completion content can be presented on the interactive interface in the form of a pop-up window, etc., so that users can flexibly select the completion content that meets the actual completion needs to insert into the input text. Alternatively, the input text can be updated based on the completion content, and the updated input text can be presented on the interactive interface. This application does not restrict the method of presenting the completion content on the interactive interface, and it can be determined as appropriate.
[0045] As can be seen, in any natural language input scenario in the field of natural language processing, especially in online social scenarios, this application can automatically trigger the input text completion task without switching to other applications or requiring additional input completion, even when the input text is presented in the output interactive interface. It quickly, flexibly, and intelligently acquires the bidirectional contextual information and the corresponding feature representations of their associated multimodal information for the content to be completed, thereby assisting in a comprehensive, flexible, and accurate understanding of the user's completion needs in the current scenario. This improves the accuracy of the generated completion content and solves the problems of cumbersome operation, low interaction efficiency, poor flexibility, and inaccurate understanding of user intent, resulting in low accuracy of the generated completion content, inherent in the question-and-answer interaction mode of large language models in input text completion applications. (Refer to...) Figure 2 This is a flowchart illustrating the bidirectional context-aware text completion method proposed in Embodiment 2 of this application. This embodiment describes an optional implementation method of how an adaptive completion task triggering mechanism is used to generate a completion trigger signal for the input text in the method proposed above. Figure 2 As shown, this optional implementation method may include, but is not limited to:
[0046] Step S21: In response to an input operation on the output interactive interface, the input text is presented on the interactive interface, and the input speed of the input text is monitored.
[0047] Step S22: In response to the completion position selection operation of the input text, determine and monitor the duration of the current display position of the input identifier in the input text;
[0048] Step S23: Based on the input speed, adjust the completion trigger threshold for the input identifier position; the input speed and the completion trigger threshold are negatively correlated.
[0049] Step S24: In response to the fact that the dwell time of the input identifier is greater than the completion trigger threshold, a completion trigger signal is generated for the input text.
[0050] In this embodiment, to avoid accidentally triggering completion tasks on the input text and affecting the user's interactive experience, this application pre-configures the conditions for generating completion trigger signals. In actual scenarios, completion trigger signals are only generated when the corresponding conditions are met; otherwise, no completion trigger signal is generated, and the completion task is not performed on the input text, allowing the user to continue performing natural language text input operations. Based on this, this embodiment proposes automatically controlling whether to execute the input text completion task based on the dwell time of the input identifier (such as the cursor corresponding to the mouse and keyboard or other identifiers configured for other input components). Thus, if the input identifier position (i.e., the current display position of the input identifier presented on the interactive interface) remains unchanged for a preset time (which can be configured or adjusted according to the actual scenario; this application does not limit its numerical value), it is assumed that the user wants to add content at that input identifier position to enrich the already entered text, and a completion trigger signal can be generated. Conversely, if the input marker position changes within the preset time period, meaning the input marker stays at the current display position for less than the preset time period, it can be assumed that the user does not currently want to perform the text completion task, and no processing is required.
[0051] In one possible implementation, the aforementioned preset duration can be a pre-configured fixed value for the completion trigger threshold. However, considering that different users or different scenarios have different tolerances for the dwell time at the input marker position when triggering the completion task, this application can dynamically adjust the preset duration, i.e., the completion trigger threshold, based on the input speed of the input text, so that the adjusted completion trigger threshold is more suitable for the user's input state in the current interaction scenario. Therefore, this application can monitor the user's input speed (characters / second, which can be denoted as...) Based on the input speed, the completion trigger threshold for the input identifier position can be adjusted. For example, when the input speed is fast, a smaller completion trigger threshold can be configured; when the input speed is slow, a larger completion trigger threshold can be configured. This application does not limit the size of the completion trigger threshold in these two cases, and users can personalize the configuration in the visual interface.
[0052] Preferably, this application can monitor the user's input speed of the input text (or the historical input speed obtained by the computer device from the statistics of historical input text, etc.) to determine the current completion trigger threshold based on the input speed. Then, the duration of the monitored input marker's current display position in the input text can be compared with the completion trigger threshold. If the duration is less than or equal to the completion trigger threshold, no completion trigger signal will be generated, and no completion task will be performed on the input text. In this case, no processing is required. Otherwise, a completion trigger signal will be generated to respond to the completion trigger signal and the subsequent completion task will be performed according to the method proposed in the embodiments of this application.
[0053] Among them, the completion trigger threshold (which can be denoted as) Adjustments may be made in accordance with, but are not limited to: This formula represents the adjustment method. In this formula... The sigmoid function can be represented as an "S"-shaped curve that smoothly transitions from 0 to 1, thus controlling the input speed. Normalization is applied to values between (0,1). Based on this, the input speed... When the threshold is greater than 5, the completion trigger threshold can be adjusted. The input speed decreased to 300ms. When the threshold is ≤5, the completion trigger threshold can be adjusted. The threshold is increased to 450ms to avoid frequently triggering the completion task on the input text. However, it is not limited to the two completion trigger thresholds described in this embodiment. The value can be determined depending on the situation.
[0054] In some embodiments, this application can also configure another adaptive completion triggering mechanism by detecting whether there are semantic breakpoints in the input text. Specifically, this application can generate a completion trigger signal for the input text upon detecting the presence of a semantic integrity identifier. This semantic integrity identifier is not only a formal breakpoint (a pause / separator to visually or syntactically divide the input text into two parts), but can also mark the boundary of a relatively complete semantic unit. This allows the application to utilize existing, conventional language format markers (punctuation) and user behavior (pauses) to infer when the user's intent and the text's semantic structure might transition or pause, predicting whether the user currently wants to add completion content to enrich the input text. Based on this, this application can pre-configure the various punctuation marks included in the semantic integrity identifier, such as... Based on this, if any of the punctuation marks (period, comma, and semicolon) appears in the input text and the input after that punctuation mark is blank, it indicates that a semantic breakpoint has been created at that punctuation mark, and the user may need to insert completion content here.
[0055] In some embodiments, this application can also predict whether the user wants to autocomplete the input text by combining the matching of the input text with domain keywords. The domain keywords can be determined through methods such as word frequency statistics based on historical data from different domains, for example, by using words with TF-IDF (Term Frequency-Inverse Document Frequency) values greater than a threshold as keywords, or by using a pre-trained keyword recognition model. Then, the keywords corresponding to different domains can be stored in a dictionary or other way for later retrieval. Thus, in response to the input text containing at least one keyword related to a preset domain, such as "hotel" and "recommendation" in a hotel review scenario, a completion trigger signal is generated to perform the input text completion task. Optionally, the domain weight of each word appearing in the input text can be determined by TF-IDF calculation (such as TF-IDF value or its normalized weight value). If the domain weight is greater than a threshold (such as 0.6 or 0.65, etc., this application does not limit its value), it indicates that the corresponding word belongs to the keywords of the corresponding preset domain; otherwise, it is determined that the corresponding word does not belong to the keywords.
[0056] In some embodiments, this application can also predict whether a user wants autocomplete on the input text by analyzing the user's historical interaction behavior. To this end, a response trigger signal for the input text can be generated in response to the detection of interaction behavior data in the user's historical interaction behavior data that satisfies association conditions (such as strong correlation) with the input text. Optionally, this application can determine the degree of association between each historical interaction behavior data and the input text, or the confidence level of a strong correlation between historical interaction behavior data and the input text, through similarity calculation methods or other association rules. If it is greater than an association threshold (e.g., confidence level conf greater than 0.7), it indicates that the historical interaction behavior data satisfies the association conditions with the input text; otherwise, it indicates that the historical interaction behavior data does not satisfy the association conditions with the input text.
[0057] Preferably, to avoid frequent triggering of the completion task, a corresponding completion trigger probability can be configured for at least one of the above conditions. When the condition is met, based on the completion trigger probability, it can be determined whether to generate a completion trigger signal for the input text. For example, based on the completion trigger probability, a random number is generated. If the random number is greater than the completion trigger probability, the completion trigger signal is generated; otherwise, no completion trigger signal is generated. This application does not limit the value of the completion trigger probability or its configuration / adjustment implementation method.
[0058] Optionally, this application can configure corresponding response priorities for various conditions of the adaptive completion triggering mechanism for input text. Based on the response priorities, corresponding completion trigger probabilities are assigned to different conditions. Based on the completion trigger probabilities, it is determined whether a completion trigger signal is generated when responding to the corresponding condition. The control process is not detailed here. For example, following the condition order described above, response priorities are configured from high to low. Conditions with higher response priorities can be configured with larger completion trigger probabilities; conversely, conditions with lower response priorities can be configured with smaller completion trigger probabilities. For instance, the completion trigger probabilities configured for the four conditions described above can be 1.0, 0.8, 0.6, 0.4, etc., but are not limited to these.
[0059] Optionally, to improve the reliability and personalization of the adaptive completion triggering mechanism, the completion triggering probability corresponding to each of the above conditions can be dynamically adjusted based on the user's feedback on the historical completion content. This acceptance feedback can indicate whether the user approves of the completion content generated by the method provided in this application, such as accepting or rejecting (which is an explicit feedback that can be achieved by clicking the corresponding function button) the completion content presented on the interactive interface, or modifying the presented completion content or the completion content after inserting input text (which is an implicit feedback), etc. This application does not limit the content and representation of the acceptance feedback.
[0060] In one possible implementation, this application can determine the user's historical acceptance rate of historically completed content based on user feedback. For example, it can determine the ratio of the number of times the user accepted historically completed content (A, which includes the number of triggers of explicit and implicit feedback described above) within a statistical time window to the total number of triggers (N, which includes the total number of acceptances, rejections, and modifications within the statistical time window, i.e., the total number of triggers of all explicit and implicit feedback), i.e., A / N. Regarding the statistics of trigger counts, a corresponding counter can be configured, i.e., N+=1. In this case, the feedback content triggered each time can be recorded to identify the number of acceptances (A). Alternatively, an acceptance counter and a total counter can be configured to record the number of acceptances (A) and the total number of triggers (N) respectively to determine the user's historical acceptance rate. Subsequently, the completion trigger probability for each condition can be dynamically adjusted based on the user's historical acceptance rate. If the historical acceptance rate is low (less than the acceptance rate threshold; this application does not limit the size of this threshold and it can be determined as needed), the completion trigger probability can be reduced to allow the user to focus more intently on continuous writing. Conversely, if the historical acceptance rate is high (greater than or equal to the acceptance rate threshold), the completion trigger probability can be increased, prioritizing the completion of the input text and improving the user experience. Of course, this application can also dynamically adjust the completion trigger probability of corresponding conditions based on the historical acceptance rate corresponding to each historical completed content. Alternatively, based on the comparison results of the historical acceptance rates corresponding to each condition, the response priority and completion trigger probability of the corresponding conditions can be adjusted so that the adjusted response priority is positively correlated with the completion trigger probability.
[0061] In one possible implementation, this application can also dynamically adjust the completion trigger probability based on different adjustment rules within different ranges of historical acceptance rates. Thus, if the historical acceptance rate is determined to be less than a first acceptance rate threshold, the first adjustment rule is used to adjust the completion trigger probability; if the historical acceptance rate is determined to be greater than a second acceptance rate threshold (which is greater than the first acceptance rate threshold, and this application does not limit the magnitude of these two acceptance rate thresholds), the second adjustment rule is used to adjust the completion trigger probability. The first adjustment rule and the second adjustment rule are different. For example, assume that... Indicates historical acceptance rate, by Indicates the probability of triggering completion, if , can be according to This formula represents the first adjustment rule, enabling dynamic adjustment of the completion trigger probability. If , can be according to The second adjustment rule represented by this formula enables dynamic adjustment of the completion trigger probability, but it is not limited to the adjustment rule described in this embodiment. The constants in the corresponding adjustment rule can be adjusted according to the actual situation.
[0062] The adaptive completion triggering mechanism proposed in this application can dynamically adjust parameters such as the completion triggering threshold, the window size of the context window, etc., based on input speed, input length, historical acceptance rate, etc., through the visualization method described above, and combine one or more conditions to determine whether to perform the completion task of the input text presented on the interactive interface, thereby achieving flexible and intelligent completion triggering and adapting to the usage habits of different users.
[0063] Reference Figure 3 This is a flowchart illustrating the bidirectional context-aware text completion method proposed in Embodiment 3 of this application. Figure 3 As shown, this method may include, but is not limited to:
[0064] Step S31: In response to the completion trigger signal of the input text presented on the interactive interface, obtain bidirectional context information in the input text; the bidirectional context information includes the preceding paragraph before the input marker position and the following paragraph after the input marker position, and the length of each of the preceding paragraph and the following paragraph is associated with the input length of the input text.
[0065] In this embodiment, the size of the context window (a sliding window, but not limited to this method of extracting context fragments via a sliding window) can be determined based on the input length of the input text to be completed. Then, based on the context window of that size and the current input identifier position, the preceding and following fragments can be extracted from the input text. Thus, as the input length of the text changes, this application can dynamically adjust the size of the context window to dynamically adjust the length of the context fragments, thereby focusing on key fragments using an attention mechanism and improving the relevance of the completed content to the input text.
[0066] Based on this, in one possible implementation method, it can be carried out according to, but is not limited to: This formula represents a calculation method that uses the input length L (number of characters) of the input text to determine the window size of the context window. In this formula, log represents the logarithmic function, used to model the user's need to understand the input text. The 800 in the formula is the window size. The maximum value of the context window can be determined based on the resource limitations of the computer device (such as memory and computing power) or the maximum context length of the text completion model, but it is not limited to 800. The 300 in the formula can be the base window size or offset to ensure a minimum window even for very short input text; the 50 in the formula is a scaling factor used to control the influence of the input text length L on the window size W. The larger the scaling factor, the greater the impact of changes in L on W, and it is not limited to the value of the scaling factor 50. It can be seen that, considering the characteristics of the logarithmic function, for shorter input text, increasing the input length slightly requires a significant increase in the context window to better understand user needs; for longer input text, which already contains sufficient information, further increases in length only require a small increase, or even no increase in the context window at all. Furthermore, the rate of change of the context window size varies for input text of different levels / ranges of length. Therefore, according to the window size calculation method expressed in the formula above, a context window of 300 characters can be determined for shorter input text, and a 300-character context fragment can be extracted from the input text accordingly; for longer input text, the context window size can be extended to a maximum of 800 characters, and a maximum of 800-character context fragment can be extracted from the input text accordingly. For example, if L=1000, W=645; L=5000, W=726; L=50000 or greater, W=800, in order to save computing resources and prevent the processing from becoming slow or memory overflow by limiting the window size.
[0067] In some embodiments, the variable parameters included in each condition of the adaptive completion triggering mechanism described above (such as preset duration, semantic integrity identifier, keyword matching degree / TF-IDF value, relevance / confidence threshold, and other completion triggering thresholds, and turning the corresponding conditions on / off) can be adjusted or configured through a visual interface to meet the user's personalized interaction settings requirements. Figure 4 The configuration interface shown in the diagram allows users to enable or disable various conditions within the corresponding windows (e.g., Figure 4 Each condition corresponds to a gray slider on the right, which can be slid left or right to control whether the condition is turned on or off. The variable parameters in the enabled conditions can be configured or adjusted to personalize the adaptive completion trigger mechanism. Optionally, after adjusting or configuring the variable parameters for each condition, a corresponding condition configuration file can be generated. During the execution of this application's method, the corresponding condition configuration file is read sequentially according to the response priority of each condition, and the method described above is used to determine whether the corresponding condition is met to generate a completion trigger signal for the input text. However, this storage method is not limited to this.
[0068] Step S32: Using the context encoding module in the trained text completion model, a context feature representation is generated based on the semantic dependency relationship between the preceding and following fragments.
[0069] In this embodiment, the text completion model can be trained based on the FIM architecture and may include at least a context encoding module, a multimodal encoding module, a multimodal feature fusion module, and a completion decoding module. To improve the inference speed of the text completion model and reduce resource consumption, the trained text completion model, which is a large model (i.e., has a large number of parameters and requires a large amount of resources to run), can be compressed. For example, knowledge distillation or other quantization methods can be used to compress the model. The compressed model (a small model with a significantly reduced number of parameters, which basically retains the performance of the model before compression) can be deployed on a computer device to perform the text completion task on the input text (i.e., a lightweight generator), thereby improving efficiency while ensuring the accuracy of the generated completed content.
[0070] Based on this, this application can convert the extracted context / text fragments into feature representations that the model can recognize, using a context encoding module. Since the context / text fragments are text segments, a pre-trained text encoder can be used as the context encoding module to implement step S32. This text encoder can be a general pre-trained language model that analyzes the semantic dependencies between the context and text fragments to generate and output a context feature representation that integrates context semantic information. For example, this application can segment the input text to obtain a series of token (smallest semantic unit) IDs, denoted as the input text sequence, which represents... The elements can represent n tokens obtained from word segmentation. The input text sequence is divided into two parts according to the input identifier position (which can be denoted as the position of the k-th token), namely the preceding segment (prefix sequence). ) and the following fragment (suffix sequence) ), respectively represented as: and The two tokens are then input into the context encoding module (such as a text encoder based on the Transformer architecture) according to, but not limited to, the format described above. The first layer (word embedding layer) of this context encoding module converts each input token into an initial word embedding vector sequence. This is then combined with the positional encoding of the input identifier position and input into the Transformer network of the context encoding module (which typically includes a multi-head attention mechanism and a feedforward neural network) for processing. For example, token-level dependency analysis (such as deep logical relationships like causality, contrast, or parallelism) is performed between the initial word embedding vector sequences corresponding to the preceding and following paragraphs to achieve joint encoding of these two initial word embedding vector sequences, generating a new word embedding vector sequence that integrates contextual semantic dependencies, i.e., the context feature representation (which can be denoted as...). This application does not provide a detailed description of the joint coding implementation process.
[0071] Optionally, the word embedding layer described above can also be used as an independent word embedding module, converting the preceding and following text segments into corresponding initial word embedding vector sequences. The input context encoding module (such as the text encoder described above) then jointly encodes the two initial word embedding vector sequences to obtain a context feature representation. That is, the hidden state sequence , It represents the i-th input hidden state vector (i.e., the encoding vector / feature representation of the i-th token), which is a matrix with feature dimension d (such as semantics, syntax, and other dimensions).
[0072] Step S33: Through the multimodal coding module in the text completion model, feature extraction is performed on the multimodal information associated with bidirectional contextual information to obtain multimodal feature representation;
[0073] In this embodiment, the multimodal encoding module may include encoders corresponding to different modalities, such as image encoders, audio encoders, and image-text encoders. Each module converts the corresponding type of multimodal information (or preprocessed multimodal information after word segmentation / segmentation) into initial word embedding vectors through its respective word embedding layer or independent word embedding module. Then, the encoders of different modalities encode and concatenate the initial word embedding vectors of the corresponding modalities to initially obtain a multimodal feature representation. This means that the encoded feature vectors of each modality are arranged according to the original order of the multimodal information to form a multimodal feature sequence. Each element represents the encoded feature vector of the corresponding basic unit (such as a token, an image patch / region, a video frame / segment, or an audio spectrogram window / time slice). It should be understood that if the modal feature representations are pre-cached in a feature cache pool using, but not limited to, the encoding method described above, the multimodal feature representations associated with the context feature representations can be directly obtained without using a multimodal encoding module to encode the multimodal information online. In this case, step S33 can be skipped, and step S34 can be executed directly.
[0074] Step S34: Through the multimodal feature fusion module in the text completion model, the multimodal feature representation is mapped to a feature space that matches the context feature representation to obtain a new multimodal feature representation;
[0075] To achieve joint analysis of multimodal features, the feature dimensions of multimodal feature representations directly encoded by a multimodal coding module often differ from those of the context feature representations, making direct fusion impossible. Therefore, a multimodal feature fusion module can be used to map the multimodal feature representations to a feature space that matches the context feature representations, achieving feature dimension alignment (which may include semantic space alignment) between the multimodal and context feature representations. The resulting new multimodal feature representation can be denoted as... .in, It can represent a feature representation (encoded feature vector) that is aligned with the feature dimensions of the context feature representation to represent j basic units. and They are all vector matrices with a feature dimension of d. It is an m x d matrix for subsequent fusion processing. Optionally, this application can use a pre-constructed modality mapping matrix. Convert B to This involves mapping the feature representations of each modality to the same dimensional space (i.e., a feature space that matches the context feature representations). The implementation process will not be detailed in this application.
[0076] Step S35: The correlation analysis between the context feature representation and the new multimodal feature representation is performed through the multimodal feature fusion module to obtain the corresponding attention weight distribution;
[0077] Step S36: The multimodal feature fusion module generates a fused feature representation based on the attention weight distribution, the new multimodal feature representation, and the context feature representation.
[0078] This application can implement step S35 through a cross-modal attention mechanism. Here, the contextual feature representation can be used as the query matrix Q, and the new multimodal feature representation as the key-value matrix K. The similarity between Q and K is obtained through methods such as, but not limited to, dot product operations or cosine similarity calculations. This similarity represents the correlation between the corresponding contextual feature representation and the new multimodal feature representation. Based on each similarity, an attention weight distribution for each K is generated, indicating which multimodal information is more important for generating the completed content. The implementation process is not detailed in this application. Afterwards, based on the attention weight distribution, the new multimodal feature representation and the contextual feature representation can be fused (the new multimodal feature representation is then used as the value matrix V, and a weighted sum of V is performed) to obtain a fused feature representation. Let the attention weight distribution be denoted as... ,Right now arrive The process of obtaining the attention weights and fused feature representation F can be expressed as: This enhances the three-dimensional understanding of the contextual semantics.
[0079] Step S37: Based on the fused feature representation and context feature representation, the completion decoding module in the text completion model generates multiple candidate completion contents and their corresponding completion confidence for the input text.
[0080] This application allows the feature representations F and H to be used together as input to the completion decoding module (i.e., the completion decoder) to comprehensively and deeply understand the user's completion needs and accurately generate the corresponding completion content, without requiring the user to input additional completion requirements. To further improve the accuracy and personalization of the input text completion, the text completion model can generate multiple candidate completion contents according to the method described above, and then use a pre-trained ranking model (i.e., a scoring model for accuracy evaluation, which can be represented as...) Each candidate completion is scored to obtain a corresponding score c, which is the completion confidence c, representing the probability that the corresponding candidate completion meets the user's completion requirements for the input text. This application does not describe in detail the generation process of each candidate completion and the method for obtaining the completion confidence.
[0081] In some embodiments, regarding the implementation process of the completion decoding module generating the aforementioned completed content / candidate completed content, in order to reduce the probability of illusion in the text completion model, ensure that the generated completed content is logically consistent, and improve the completion quality, this application introduces knowledge graph embedding features into the completion decoding module. This knowledge graph can be constructed based on the entities and entity relationships within the context of the input text, ensuring that the knowledge graph contains entity relationships. For example, in the hotel review completion scenario, the pre-constructed knowledge graph contains entity relationships such as "service attitude → enthusiasm" and "service attitude → professionalism." Therefore, this application can utilize entity relationship features... The embedded features of the knowledge graph are represented by the output features obtained from the completion decoding module. The system integrates features to ensure that the generated completion content conforms to the true semantic association. Based on this, this application decodes the fused feature representation and the context feature representation through a completion decoding module. After obtaining the output feature representation, a weighted feature fusion is performed on the output feature representation and the knowledge graph embedding features of the scene to which the input text belongs, based on a gating mechanism, to obtain the fused output feature. Based on the fused output feature, the completion content of the input text is predicted, and the corresponding prediction probability distribution is obtained.
[0082] For the aforementioned gating mechanism, a gating vector can be defined. ,Right now ,in Indicates a fully connected layer. The activation function (such as the sigmoid function) is represented as shown in the formula, and the output features are represented by a fully connected layer and the activation function. and knowledge graph embedding features Then, the parts are assembled. Each element in the array takes a value between 0 and 1, and is used to control the degree of integration of knowledge graph embedding features. Based on this, it is possible to achieve... and The weighted feature fusion. Based on this, this application can be carried out in accordance with, but is not limited to, the following: This formula represents the method of implementation, where ⊙ denotes element-wise multiplication. This indicates the fusion of output features, thereby enabling selective transfer of knowledge graph embedding features. Specifically, when G approaches 1, the text completion model relies more on knowledge graph embedding features when generating completed content; when G approaches 0, the text completion model tends to complete the original output of the decoding module. Thus, based on the fused output features When generating the completion content, higher prediction probabilities can be assigned to words that are more in line with real-world logic, representing the probability that the word is used to form the completion content. That is, the prediction probability distribution of different words used to form the completion content. Based on this prediction probability distribution, words with high prediction frequency (words that are in line with real-world logic) are sampled to form the completion content.
[0083] In some embodiments, to control the diversity of completions, this application can use a dynamic temperature coefficient (i.e., a parameter that can affect the predicted probability distribution output by the text completion model, which can be denoted as the probability adjustment coefficient). The lower the probability adjustment coefficient, the more consistent and rigorous the content output by the text completion model; the higher the probability adjustment coefficient, the more open and imaginative the content output by the text completion model. Even for the same text completion model, different probability adjustment coefficients have a significant impact on the effect of the completed content generated by the text completion model. Therefore, this application can generate a probability adjustment coefficient based on the entropy value of the predicted probability distribution, adjust the predicted probability distribution according to the probability adjustment coefficient, and thus determine the completed content 1 for the input text based on the adjusted predicted probability distribution. The appropriate probability adjustment coefficient (denoted as...) can be determined by referring to, but is not limited to, the method expressed in formula (1). To achieve the prediction of probability distribution Adjustments:
[0084] (1)
[0085] In formula (1), the adjustment coefficient It can be 0.3 or other values, which can be determined based on experience or experiments. This can represent the entropy value of the predicted probability distribution. As shown in the formula, when this entropy value is low (high contextual certainty), the obtained... The entropy is approached to 0.5 to enhance the stability of the completed content. At a higher entropy value (where the contextual semantics are ambiguous), the obtained... Approaching version 1.0 to improve the diversity of completions. Therefore, it can be seen that during the generation of completion content by the completion decoding module of this application, the diversity of completion content is controlled through the optimization method of dynamic temperature coefficient and knowledge graph embedding features, ensuring that the completion content is both context-appropriate and logically sound, thereby improving the relevance and rationality of the completion.
[0086] Step S38: Determine whether there is a target completion content among multiple candidate completion contents with a completion confidence level greater than or equal to the confidence level threshold. If yes, proceed to step S39; otherwise, proceed to step S310.
[0087] Step S39: Insert the target completion content into the input marker position of the interactive interface to update the input text;
[0088] Step S310: Output a completion prompt window on the interactive interface to present multiple candidate completion contents;
[0089] In step S311, in response to the selection operation of candidate completion content, the selected candidate completion content is inserted as the target completion content into the input identifier position to update the input text.
[0090] In practical applications of this application, as analyzed above, the highest completion confidence level (denoted as c) can be selected. * The corresponding candidate completion content is output as the target completion content, which is then presented on the interactive interface. Preferably, this application can further determine whether the highest completion confidence level is greater than or equal to a confidence threshold (such as 0.85 or 0.9, which can be determined based on the actual scenario's requirements for completion accuracy; the stricter the requirement for completion accuracy, the higher the confidence threshold can be configured). If the highest completion confidence level is greater than or equal to the confidence threshold, it can be directly used as the target completion content and inserted into the input marker position of the interactive interface. At this time, a prompt message for inserting the target completion content can also be output. The user can choose whether to insert the target completion content into the input marker position of the input text according to the actual situation, so that if it does not meet the user's needs, the insertion operation of the target completion content can be canceled immediately.
[0091] Based on the above analysis, if the highest completion confidence is less than the confidence threshold, it indicates that the completion confidence of each of the multiple candidate completion contents is low. These multiple candidate completion contents can be output for the user to choose from. In this case, a completion prompt window is output in the form of a pop-up (such as a semi-transparent pop-up on the interactive interface), presenting multiple candidate completion contents. The user selects one candidate completion content from these (e.g., by clicking on its display area, or by controlling the selection arrow or other indicators to switch to its display area; this application does not restrict the selection method) as the target completion content, which is then inserted at the input indicator position of the input text.
[0092] Optionally, multiple candidate completion content can be presented in a hierarchical manner in the completion suggestion window. For example, the candidate completion content corresponding to the highest completion confidence can be presented in a first display state, while other candidate completion content can be presented in a second display state. The first and second display states have higher visual salience. For example, the first display state can be a highlighted state (or a highlighted magnified display state), and the second display state can be a grayscale state (or a grayscale reduced display state). Alternatively, different background colors can be used to distinguish between the two types of candidate completion content, so that users will first pay attention to the candidate completion content corresponding to the highest completion confidence, increasing the probability that it will be selected as the target completion content. However, this is not limited to the hierarchical presentation method proposed in this application. Optionally, other candidate completion content can be further divided into different levels to be presented in different display states. The implementation process is similar and will not be described in detail here.
[0093] In some embodiments, to address the diverse language styles present in scenarios such as social media (e.g., formal / casual / colloquial, humorous / positive / negative emotional tones, professional or specific personas), this application can further acquire the language style of the input text or a pre-configured language style to improve the model's generalization ability. This allows for the generation of candidate completion content based on that language style during the candidate completion generation process, resulting in candidate completion content with that specific language style. Therefore, in one possible implementation, the language style of the completion content can be configured through a visual interface, such as... Figure 4 As shown, users can select or input a language style to generate a corresponding style configuration file for storage. During the completion content generation process, this configuration file can be read to form language style prompts, thus converting the directly generated candidate completion content into candidate completion content for the corresponding language style. Figure 4 Adaptive language can represent the language style of the input text or the language style of the context.
[0094] Optionally, to ensure the text completion model maintains a long-term ability to generate completion content in a specific language style, data related to that language style can be collected for supervised training / fine-tuning of the model. This allows the model to learn the vocabulary, sentence structure, and logic of that language style, improving the quality of the generated completion content. Alternatively, this application can train adapters (LoRAs) for different language styles (i.e., style plugins) without training the entire model. During inference, an adapter LoRA matching the language style of the input text or a pre-configured language style can be selected to generate completion content with that language style. This application does not limit the implementation method for generating completion content in different language styles. Furthermore, the generated candidate completion content or candidate completion content in different language styles can be based on a pre-configured completion content length, which can be, for example, […]. Figure 4 The visualization interface shown allows users to personalize their settings, or adaptively determine the length of the completed content based on information such as the context of the input text, in order to constrain the length of each completed content. This application does not limit the method for configuring the length of the completed content. Regarding the various variable parameters involved in implementing the method of this application, they can all be configured personally through the visualization interface, but are not limited to this method, to improve the flexibility of the completed content and adapt to the personal habits of different users.
[0095] In some embodiments, to improve the performance of the text completion model, incremental training can be performed based on user feedback on the completed content to adjust the model parameters. The text completion model with the adjusted parameters (i.e., the incrementally trained model) is then used to execute the next triggered completion task. This closed-loop optimization method ensures continuous improvement in the model's completion effect. Based on this, this application can respond to the user's acceptance or rejection of the completed content (target completed content) (which can be based on the output prompt indicating whether the completed content is accepted). If the user accepts the completed content, the feature information related to generating the completed content (such as context information, completed content, multimodal information, etc., or their feature representations) can be used as positive sample data. When the user rejects the completed content, the relevant feature information is recorded as negative sample data. Then, according to a preset incremental training period, or when the number of positive and negative samples reaches a preset number, these positive and negative sample data can be used to incrementally train the text completion model to improve the completion accuracy of the trained model.
[0096] Reference Figure 5 This is a flowchart illustrating the bidirectional context-aware text completion method proposed in Embodiment 4 of this application. This embodiment can describe an optional method for obtaining the aforementioned context feature representation and attention weight distribution, such as... Figure 5 As shown, the method of obtaining this information may include, but is not limited to:
[0097] Step S51: Encode the input identifier position to obtain a position-aware vector;
[0098] Step S52: Fuse the position-aware vectors into the feature representations of the preceding and following paragraphs respectively to obtain the preceding and following feature representations.
[0099] In this embodiment, to be applicable to the natural language domain, the FIM model can be pre-trained based on a large amount of natural language text data to learn the semantics, syntax, and common expression patterns of the text. This enables the trained FIM model to extract bidirectional contextual information from the input text in a natural language processing scenario. The implementation process is not detailed here. Based on the above description of the context encoding module, this application can use an improved Transformer architecture text encoder as the context encoding module. Utilizing the Transformer architecture's positional encoding mechanism, the input identifier position k is encoded into a position-aware vector, which is then fused into the elements of the feature representations (initial word embedding vectors) of the preceding and following fragments. For example, the position-aware vector is iterated over the word embedding vector of each token to obtain the corresponding preceding / following context feature representation, thereby enhancing the semantic meaning of the context through a self-attention mechanism.
[0100] Step S53: Capture the semantic dependencies between data units contained in the preceding and following context feature representations through a multi-head attention mechanism to generate a fused context feature representation; the capture of these semantic dependencies can be based on the distance weight between the data unit and the input identifier position.
[0101] Step S54: Temporal modeling and feature enhancement of the context feature representation are performed using a recurrent neural network to generate a position enhancement vector;
[0102] Since each token in the preceding and following feature representations contains the input identifier position k, the semantic dependency relationship between the n tokens before and after the input identifier position k can be captured using a multi-head attention mechanism based on the distance weight between the token and the input identifier position. That is, tokens closer to the input identifier position are assigned higher weights, making the text completion model pay more attention to tokens near the input identifier position and improving the relevance of the completion. Optionally, this application may adopt, but is not limited to, the following methods: This calculation formula achieves the allocation of the weights, where =0.02, etc., can be configured based on experience or prior historical knowledge. Let be the distance between the i-th token and the input identifier position. This represents the weight assigned to the i-th token. After generating a fused contextual feature representation based on captured semantic dependencies using a multi-head attention mechanism, a recurrent neural network (such as a gated recurrent unit, GRU) can be used to perform temporal modeling and feature enhancement on this contextual feature representation. This involves capturing the sequential information and semantic dependencies within the context window, and using a gating mechanism to select and retain contextual information that is more important (more relevant) to the completed content, while discarding or reducing the weight of less relevant contextual information. This dynamically focuses on text fragments related to the completed content, achieving refined modeling of contextual semantics and generating a position enhancement vector that fuses the input identifier position and key semantic information.
[0103] Therefore, it is evident that contextual relevance can be calculated using enhanced attention mechanisms. For example, by introducing a position enhancement vector P on top of the traditional attention mechanism, local semantic enhancement of bidirectional contextual information can be achieved, thereby improving the accuracy of contextual feature representation and enabling the text completion model to better adapt to different completion scenarios, thus improving the accuracy and flexibility of the completed content. Specifically, the position enhancement vector P can be the position embedding vector matching the input identifier position in the position embedding matrix updated through backpropagation during the training process of the text completion model, enhancing the influence of the input identifier position (completed position information) on the text semantics. During the training of the position embedding matrix, an initial position embedding matrix of size L×d can be initialized based on the maximum input length L and feature dimension d of the initial completion model, so that each position corresponds to a learnable position vector. During the backpropagation process of model training, this position vector is updated according to the method described above based on multi-head attention mechanism and gated recurrent units to obtain the position embedding vector corresponding to that position. During inference, the position embedding vector corresponding to the current input identifier position can be read directly from the trained position embedding matrix to perform subsequent steps, or the position embedding vector can be updated according to the method described above to adjust the context feature representation.
[0104] Step S55: Adjust the context feature representation based on the trained multimodal interaction weight matrix and position enhancement vector to obtain the adjusted context feature representation, and adjust the multimodal feature representation associated with the context feature representation based on the multimodal interaction weight matrix to obtain the adjusted multimodal feature representation;
[0105] Step S56: Based on the dot product attention mechanism, perform correlation analysis on the adjusted context feature representation and the adjusted multimodal feature representation to generate the attention weight distribution.
[0106] Based on the above analysis, compared with the traditional attention mechanism shown in formula (2), this application adopts the improved attention mechanism represented by formula (3) to implement steps S55 and S56.
[0107] (2)
[0108] (3)
[0109] In formula (2), , , , Represents the query matrix, such as contextual feature representation; Let V represent the key matrix, V represent the value matrix (e.g., a multimodal feature representation mapped to a feature space matching the context feature representation), M represent the pre-trained multimodal interaction weight matrix, and P be the position augmentation vector corresponding to the current input identifier position. This application obtains the new query matrix in this manner. and the new key matrix That is, the adjusted context feature representation and the adjusted multimodal feature representation are obtained. Among them, the multimodal interaction weight matrix M can calculate the similarity between various modal feature representations such as interactive behavior, text, image, and audio. Based on the obtained similarity matrix, an initial interaction weight matrix of size d×d (to ensure that it will not change its original size when multiplied with L×d QKV) is initialized, and then the weight values of each modality are iteratively optimized based on the dynamic programming algorithm, so as to use the optimized interaction weight matrix as the multimodal interaction weight matrix M. For example, when processing mixed text and image content, if the image is detected to contain key information, M will automatically increase the weight ratio of the image features, and update the weight value of the modal feature representation through the loss function and backpropagation algorithm to ensure that the fusion of multimodal information conforms to semantic logic and meets the real-time dynamic adjustment requirements. Afterwards, this application can calculate the context feature representation (i.e., the context feature representation) through the dot product attention mechanism, such as the scaled dot product attention shown in formula (4). From ) to multimodal feature representation (i.e. Association weights / attention weights The attention weight distribution is obtained.
[0110] (4)
[0111] In formula (3), and The meanings of each are as described above, with exp representing an exponential function. The numerator on the right side of the formula... The difference in the vector dot product is amplified by the exponential function, and the denominator part... This ensures that the sum of the attention weights for each line is 1, satisfying the probability distribution characteristics. The higher the attention weights calculated, the more the text completion model will focus on the corresponding feature representations in subsequent processing, adaptively focusing on the feature representations most relevant to the completed content, thereby improving the accuracy of the generated completed content.
[0112] Subsequently, this application can perform a weighted summation of the adjusted multimodal feature representations based on the attention weight distribution generated above (i.e., for...). The multimodal fusion feature representation is obtained by weighted summation. This multimodal fusion feature representation is then fused with the context feature representation to obtain a fused feature representation F. This fused feature representation F serves as an input vector to the completion decoding module. It, along with the context feature representation H output by the context encoding module, is then input to the completion decoder to generate the completed content. Therefore, in the natural language semantic completion scenario, this application's text completion model emphasizes multimodal interaction, pursues semantic understanding depth and interactive experience, improves the efficiency and quality of natural language interaction, and meets the actual needs of users in diverse scenarios when processing multimodal information (such as user behavior, real-time input, multimedia features, etc.).
[0113] Based on the description of the inference process of the text completion model performing the completion task described in the above embodiments, the following describes the implementation process of training an initial completion model, including an initial context encoding module, an initial multimodal encoding module, an initial multimodal feature fusion module, and an initial completion decoding module, to obtain a text completion model. Based on this, in some embodiments, this application can construct sample datasets for different scenarios; each sample data is labeled with an input identifier position, completion content label, and multimodal label; and obtain multimodal training information associated with the sample dataset, and obtain a multimodal sample training set after preprocessing; thereby, based on the sample dataset and the multimodal sample training set, the initial completion model is jointly trained for multiple tasks to obtain the task loss for each task; here, different task losses may include, but are not limited to, at least two combinations of completion prediction loss, multimodal matching loss, mask model prediction loss, and sentiment classification loss for the corresponding sample data; then, by minimizing the weighted sum of the task losses corresponding to different tasks, the model parameters of the initial completion model and the parameters used to generate the predicted content of the masked region in the sample data can be adjusted to obtain the text completion model.
[0114] To improve the quality of the sample dataset used for model training, this application obtains a large amount of natural language data (such as 500 social comments) from natural language corpora of various scenarios (which can be 12 known mainstream social scenarios, such as hotels, restaurants, movies, and commodities) to construct the sample dataset. During this construction process, a dynamic masking strategy can be used to process the original data to generate sample data labeled with input identifier positions and completion content tags. Optionally, the preceding and following sample fragments and completion content tags contained in each sample data can be obtained by masking the corresponding original data (such as the collected natural language corpus) based on a masking ratio using a masking model. This masking ratio is related to the length of the original data to improve the diversity of the sample data.
[0115] In one possible implementation, assuming the length of the original data is L, if L < 50, the mask ratio can be configured to 20%; if 50 ≤ L < 200, the mask ratio can be configured to 15%; and if L ≥ 20, the mask ratio can be configured to 10%. However, this method of adjusting the mask ratio is not limited to this, and the segmentation range of the original data length L is not limited to the method described in this embodiment; it can be dynamically adjusted according to actual needs. Then, the original data T can be divided into prefixes (as shown in the sample segment above) in an 8:1:1 ratio. , mask (Complete the content tags) , suffix (sample excerpt below) Construct the input triplet The sample data is then associated with multimodal tags (such as sentiment category tags, theme / scene category tags, etc.; this application only uses sentiment category tags as an example) to obtain a corresponding sample data.
[0116] The process of obtaining multimodal training information associated with the sample dataset is similar to the process of obtaining multimodal information associated with bidirectional context information described above, and will not be detailed here. Therefore, different multimodal training information may include interactive behavior data, text data, image data, audio data, and video data, etc. Corresponding to different types of multimodal training information, appropriate preprocessing methods can be used to obtain a multimodal training set. In some embodiments, the text data mentioned above may include user historical input text, comments from public social media platforms, etc., and each piece of text data may be associated with metadata such as publication time, domain / scene, etc. Optionally, this application performs multi-granular data augmentation operations on the text data to enrich it. At the lexical level, non-core words in the text data can be replaced with synonyms based on the word substitution probability (e.g., 15%) to obtain new text data. At the sentence level, the text data can be transformed into sentence structure through methods such as converting active sentences to passive sentences and affirmative sentences to double negative sentences to obtain new text data. At the discourse level, the text data can be partially reorganized, and the sentence order can be adjusted while maintaining semantic integrity to obtain new text data. The implementation process will not be detailed in this application.
[0117] For text data such as user historical input text, preprocessing can involve temporal processing, such as constructing a temporal sequence from the user's historical text input data in chronological order. This temporal sequence can then be modeled using, but is not limited to, an LSTM (Long Short-Term Memory) model to extract the user's language style evolution features and integrate them into the text completion model, making the completed content more consistent with the user's language habits. Behavioral data in multimodal training information can include, but is not limited to, the user's browsing trajectory on the corresponding platform (page dwell time). (redirect path), search keyword sequence () The data includes interactive operations (such as the frequency distribution of likes / favorites / shares), which are then converted into corresponding modal feature representations by an encoder. Image data can be crawled scene images related to the sample data content (such as hotel exterior images, food images, etc.), which can be manually labeled or labeled with entity labels and sentiment category labels (such as "clean" and "delicious") based on a large model. Audio data can be collected from a large number of voice reviews carrying sentiment category labels, which can be converted into text while retaining voiceprint features (such as fundamental frequency and energy), and then converted into corresponding modal feature representations by an audio encoder.
[0118] It should be understood that in the preprocessing of the above-mentioned training information and sample data for each modality, the following cleaning methods can also be used, but are not limited to: removing duplicate samples (samples with a duplication rate > 90%), filtering text containing sensitive words (which can be done using, but is not limited to, the AC (Aho-Corasick automaton, multimodal string matching) automaton algorithm to detect sensitive words), and correcting image recognition error labels (such as samples with an accuracy rate of less than 85% after manual review). The preprocessed multimodal training information constitutes the multimodal sample training set, i.e. T represents the preprocessed text data (i.e., the raw data such as the natural language corpus mentioned above) or its feature representation; B represents the preprocessed behavior sequence or its feature representation; I represents the preprocessed image data or its feature representation; and A represents the preprocessed audio data or its feature representation. It can represent the complete content tag of the T tag.
[0119] In the multi-task joint training process of the initial completion model based on the sample dataset and the multimodal sample training set, each sample data is equivalent to the bidirectional contextual information / contextual feature representation of the input text in the above inference process, and the multimodal sample is equivalent to the multimodal information / multimodal feature representation in the above inference process. Following the inference process described above, the predicted content of the masked region in the sample data is generated based on the initial completion model. The implementation process is not detailed in this application. Subsequently, the parameters of the initial completion model can be adjusted by collaboratively optimizing its performance on various tasks such as FIM mask completion, multimodal understanding, and sentiment analysis.
[0120] Among them, the completion prediction loss of the FIM mask completion task (such as the standard FIM task loss, denoted as...) The loss of the predicted content of the masked region in the corresponding sample data relative to the completed content label of the sample data can be calculated using the following formula (5). Multimodal matching loss This can represent the correlation between sample interaction behavior data and sample data in the multimodal sample training set. It can be calculated using the loss function shown in formula (6). The prediction loss of the above mask model... This can represent the prediction loss of the masking model for the non-masked regions in the corresponding sample data. These non-masked regions can include preceding and following sample segments from the sample data. (Sentence classification loss) This can be a sentiment category prediction of the sample data, where the predicted sentiment category (e.g., positive, negative, or neutral) is the prediction loss relative to the sentiment category label of the sample data. This application does not detail the implementation process of the sentiment analysis task. After each training iteration of the model, a weighted summation can be performed according to the method shown in formula (7) (not limited to the weighted values in the formula, but dynamically adjusted as needed) to minimize the total loss. The initial model parameters for the text completion model are determined, along with the parameters used to generate the predicted content of the masked regions in the sample data (such as the multimodal interaction weight matrix M and the position increment vector P described above). After multiple training iterations, the text completion model is obtained.
[0121] (5)
[0122] (6)
[0123] (7)
[0124] In formula (5), Represents the mask area The i-th real token in the middle, Indicates model parameters; Indicates that, given the above sample fragment The following sample fragment and model parameters In this case, the model predicts the probability that the next token is a real token. The length of the text in the masked region is represented by m. In formula (6), m can be the length of the action sequence B, such as the number of frames in a video clip or the number of time steps in an audio clip. This indicates that given sample data T and model parameters... In this case, the model predicts the i-th element in B. The probability, to reflect Correlation with T.
[0125] In some embodiments, to improve the training efficiency of the aforementioned models, reduce communication overhead, and enhance the adaptability of the models, this application proposes to introduce the DeepSpeed framework to achieve efficient distributed training. To this end, before training, this application configures a suitable optimizer and its hyperparameters, such as configuring the AdamW optimizer with an initial learning rate of 5e-5. Other hyperparameters can also be configured as needed, such as linear learning rate decay strategies, gradient pruning thresholds, etc., and the number of epoch batches can also be configured (e.g., training 100 epochs on 8 H20 GPUs). Then, based on the pre-configured optimizer, a mixed-precision training method can be used to train the initially completed model. During this training process, a zero-redundancy optimizer sharding parallel strategy (the three-level sharding strategy of the ZeRO-Infinity algorithm) can be used to shard the optimizer state, gradients, and model parameters and store them on different processors (such as GPUs) of the computer device, effectively reducing the memory usage of a single GPU. For example, when training a Transformer model, the parameters and gradients of each layer are divided according to the ZeRO-3 strategy, enabling 8 H20 GPUs to collaboratively process ultra-large-scale models.
[0126] The training process described above can employ a gradient accumulation strategy to update the optimizer state and synchronize model parameters. Based on an asynchronous checkpoint saving mechanism, when the gradient accumulation steps reach a pre-configured threshold, the initial state data of the completed model for this training session is stored. This state data includes the optimizer state, gradients, and model parameters after this training. For example, a checkpoint file of the model is saved every 2000 steps. In this implementation, DeepSpeed's asynchronous computation function can be used to continue training computation while saving the model's state data, reducing the performance loss caused by I / O operations and achieving efficient utilization of computing resources.
[0127] Furthermore, to further improve training efficiency, the different processors are each assigned different computational tasks to employ an interleaved execution mechanism during the training process, thereby reducing computational burden at different training stages of the model. In other words, this application, based on DeepSpeed's pipelined parallel technology, divides the Transformer layer into multiple stages, with each GPU processor responsible for handling the computational tasks of a specific stage. This interleaved execution mechanism allows different batches of data to undergo forward or backward propagation computations simultaneously on different GPUs, enabling overlapping computational tasks between different CPUs and hiding communication overhead. In summary, the various distributed training implementations described above rationally distribute the Transformer layers in the initial completed model across different GPUs, significantly improving parallelism and resource utilization during training, and ultimately enhancing model training efficiency.
[0128] In some embodiments, this application can also fine-tune the model obtained by the training method described above to improve the model's scene adaptability. Therefore, during the iterative training of the initial completion model, in response to the satisfaction of the training termination condition (such as the number of training iterations reaching a preset number, task loss convergence, etc.), the text completion model to be trained is determined. Then, based on the scene datasets constructed for different fine-tuning scenarios, the model parameters of the text completion model to be trained can be fine-tuned through reinforcement learning optimization to obtain the text completion model, thereby improving the model's rapid response and flexible adaptation to various user text input regions and improving the natural language completion effect.
[0129] The scenario datasets constructed for different fine-tuning scenarios can be implemented by referring to the sample dataset construction process described above. These fine-tuning scenarios can include, but are not limited to, multiple preset scenarios such as hotels and restaurants described above, or scenario clustering of the collected data to obtain scenario datasets corresponding to different scenarios. Each scenario data point can be labeled according to the method described above; the implementation process is not detailed in this application. Furthermore, during the reinforcement learning optimization process described above, reward feedback can be given based on the degree of matching between the model-generated completion content and the user's true intent, continuously adjusting the model parameters. In this implementation process, the PPO algorithm can be introduced to define states. ,in This refers to contextual text (such as bidirectional contextual information in the scene data). For user behavior features (refer to the process of obtaining the feature representation of behavior data above), referring to the reward function shown in formula (8), the reward value fed back can be based on the matching degree between the predicted completion content of the scene data and the labeled content (such as the manually labeled BLEU score, which will not be described in detail in this application). The semantic matching degree between the predicted completion content and the bidirectional contextual information in the scene data (such as cosine similarity or other similarity calculation methods). And the language style matching degree between the predicted completion content and the scene data (such as calculated based on n-grams, etc.). The weighted sum is obtained to fine-tune the model based on this reward value.
[0130] (8)
[0131] It should be noted that the weight coefficients in the reward function, such as 0.7, 0.2, and 0.1, can be dynamically adjusted according to the actual situation and are not limited to this. Preferably, in the above fine-tuning process, this application can also implement local model parameters of the text completion model based on an early stopping strategy. These local model parameters can include the parameters contained in the attention layer and output layer of the text completion model to achieve domain adaptability of the model. For example, when fine-tuning on various scene data, 60% of the parameters of the bottom layer of the model can be frozen, and only the upper attention layer and output layer can be trained. The model can be trained for 20 epochs with a learning rate of 1e-5, and the text completion model can be obtained by using an early stopping strategy (terminating if the BLEU score on the validation set does not improve for 3 consecutive epochs).
[0132] In conjunction with the method embodiments described above, this application also provides a text completion device based on bidirectional context awareness. This device may include: a bidirectional context information acquisition module, configured to acquire bidirectional context information in the input text in response to a completion trigger signal presented on an interactive interface; the bidirectional context information includes the preceding context segment before the input identifier position and the following context segment after the input identifier position; a context feature representation generation module, configured to generate a fused context feature representation based on the preceding context segment and the following context segment; and a multimodal feature representation acquisition module, configured to acquire a multimodal feature representation associated with the context feature representation; the multimodal feature representation includes feature representations of at least one multimodal information selected from interactive behavior data, text data, image data, audio data, and video data. An attention weight distribution acquisition module is used to perform correlation analysis on the context feature representation and the multimodal feature representation to obtain a corresponding attention weight distribution; the attention weight distribution can indicate the degree of contribution of different multimodal information to generating completion content for the input text; a completion content generation module is used to generate completion content for the input text based on the attention weight distribution, the multimodal feature representation, and the context feature representation; a completion content output module is used to present the completion content on the interactive interface.
[0133] Those skilled in the art will understand that the functions and technical effects of each module in the above device embodiments, as well as the units that implement the functions, are equivalent to the corresponding steps described in the foregoing method embodiments. For specific implementation details, please refer to the description in the method section. The device embodiments have not been described in detail.
[0134] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the bidirectional context-aware text completion methods provided in this application.
[0135] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the bidirectional context-aware text completion methods provided in this application.
[0136] Reference Figure 6 The present application provides a schematic diagram of the hardware structure of a computer device, as shown in the embodiment. Figure 6 As shown, the computer device may include, but is not limited to, at least one memory 61 and at least one processor 62, wherein: the memory 61 may be used to store multiple computer instructions for implementing the bidirectional context-aware text completion method proposed in the embodiments of this application; the processor 62 may load and execute the computer instructions stored in the memory 61 to implement the various steps of the bidirectional context-aware text completion method proposed in the embodiments of this application, and the implementation process can be referred to the description of the corresponding part of the method embodiments above. It should be understood that... Figure 6 The structure of the computer device shown does not constitute a limitation on the computer device in the embodiments of this application. In practical applications, the computer device may include more than Figure 6 The present application does not provide a detailed list of all the more or fewer components, or combinations of certain components, such as various communication elements, storage devices, displays, multiple sensors, audio acquisition devices, image acquisition devices, power management modules, or other input / output components.
[0137] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines. In the above embodiments, they can be implemented entirely or partially by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented entirely or partially in the form of a computer program product. The various embodiments in this specification are described in a progressive or parallel manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to mutually.
Claims
1. A text completion method based on bidirectional context awareness, characterized in that, The method includes: In response to a completion trigger signal for the input text presented on the interactive interface, bidirectional contextual information in the input text is obtained; the bidirectional contextual information includes the preceding text segment before the input marker position and the following text segment after the input marker position. Based on the preceding and following fragments, a fused contextual feature representation is generated; Obtain a multimodal feature representation associated with the context feature representation; the multimodal feature representation includes feature representations of at least one multimodal information from interactive behavior data, text data, image data, audio data, and video data; A correlation analysis is performed on the context feature representation and the multimodal feature representation to obtain the corresponding attention weight distribution; the attention weight distribution can indicate the degree of contribution of different multimodal information to the generation of the completion content for the input text; Based on the attention weight distribution, the multimodal feature representation, and the context feature representation, complete content is generated for the input text; The completed content is displayed on the interactive interface; The step of generating a fused contextual feature representation based on the preceding and following fragments includes: The context encoding module in the trained text completion model generates a fused context feature representation based on the semantic dependency between the preceding and following fragments; the length of each of the preceding and following fragments is associated with the input length of the input text; the text completion model can perform incremental optimization based on user feedback on historical completion content; The step of generating complete content for the input text based on the attention weight distribution, the multimodal feature representation, and the context feature representation includes: The multimodal feature fusion module in the text completion model fuses the context feature representation and the multimodal feature representation mapped to a feature space that matches the context feature representation based on the attention weight distribution to obtain a fused feature representation; the multimodal feature representation is obtained by the multimodal encoding module in the text completion model through feature extraction of multimodal information associated with bidirectional context information. The completion decoding module in the text completion model generates completed content for the input text based on the fused feature representation and the context feature representation; Specifically, generating complete content for the input text based on the fused feature representation and the contextual feature representation includes: Based on the fused feature representation and the context feature representation, multiple candidate completion contents are generated for the input text, as well as the completion confidence of each candidate completion content; In response to the existence of a target completion content among the plurality of candidate completion contents with a completion confidence greater than or equal to a confidence threshold, the target completion content is inserted into the input identifier position of the interactive interface to update the input text; In response to each of the completion confidence scores being less than the confidence threshold, a completion suggestion window is output on the interactive interface to present the multiple candidate completion contents; In response to the selection operation of the candidate completion content, the selected candidate completion content is inserted as the target completion content into the input identifier position to update the input text.
2. The method according to claim 1, characterized in that, The completion trigger signal is generated in response to at least one of the following conditions: The input identifier position was detected to remain unchanged within a preset time period; the preset time period is negatively correlated with the input speed of the input text. A semantic integrity identifier was detected in the input text; The input text was detected to contain at least one keyword related to a preset domain; The system detects that there is interaction data in the user's historical interaction data that meets the association conditions with the input text; In the case of different response priorities for different conditions, a corresponding completion trigger probability can be assigned to different conditions based on the response priority, so as to determine whether the completion trigger signal is generated when responding to the corresponding condition based on the completion trigger probability. The completion trigger probability can be dynamically adjusted based on user feedback on historical completion content; The variable parameters included in different conditions can be adjusted or configured through a visual interface.
3. The method according to claim 1 or 2, characterized in that, The generation of context feature representations based on the semantic dependency relationship between the preceding and following fragments includes: The input identifier position is encoded to obtain a position-aware vector; The position-aware vector is fused into the feature representations of the preceding and following paragraphs respectively to obtain the preceding and following feature representations. A multi-head attention mechanism is used to capture the semantic dependencies between the data units contained in the preceding and following context feature representations, generating a fused context feature representation; the capture of the semantic dependencies can be based on the distance weight between the data unit and the input identifier position. The context feature representation is temporally modeled and feature-enhanced using a recurrent neural network to generate a position enhancement vector, which is then used to adjust the context feature representation.
4. The method according to claim 1 or 2, characterized in that, in, The correlation analysis between the context feature representation and the multimodal feature representation to obtain the corresponding attention weight distribution is achieved through the multimodal feature fusion module in the text completion model. This process includes: Based on the trained multimodal interaction weight matrix and position enhancement vector, the context feature representation is adjusted to obtain the adjusted context feature representation; and based on the multimodal interaction weight matrix, the multimodal feature representation is further adjusted to obtain the adjusted multimodal feature representation. Based on the dot product attention mechanism, a correlation analysis is performed on the adjusted context feature representation and the adjusted multimodal feature representation to generate an attention weight distribution; The fused feature representation is generated based on the attention weight distribution, the multimodal feature representation, and the context feature representation, including: Based on the attention weight distribution, the adjusted multimodal feature representations are weighted and summed to obtain the multimodal fusion feature representation; The multimodal fusion feature representation and the context feature representation are fused to obtain the fusion feature representation.
5. The method according to claim 1 or 2, characterized in that, in, The multiple candidate completion contents can be generated based on the language style indicated by the input text or a pre-configured language style.
6. The method according to claim 1 or 2, characterized in that, The completion decoding module generates the completed content based on the fused feature representation and the context feature representation, including: The completion decoding module decodes the fused feature representation and the context feature representation to obtain the output feature representation; Based on the gating mechanism, the output feature representation and the knowledge graph embedding feature of the scene to which the input text belongs are weighted feature fusion to obtain the fused output feature; Based on the fused output features, the completed content of the input text is predicted to obtain the corresponding prediction probability distribution; The content to be completed is determined based on the predicted probability distribution; In the process of determining the supplementary content, the predicted probability distribution is adjusted according to the probability adjustment coefficient generated based on the entropy value of the predicted probability distribution, so as to determine the supplementary content based on the adjusted predicted probability distribution.
7. The method according to claim 1 or 2, characterized in that, The text completion model is obtained by training an initial completion model that includes an initial context encoding module, an initial multimodal encoding module, an initial multimodal feature fusion module, and an initial completion decoding module. The training process of the text completion model includes: Construct sample datasets for different scenarios; each sample data is labeled with input identifier position, completion content label and multimodal label. The preceding and following sample fragments and the completion content label contained in the sample data are obtained by masking the corresponding original data based on the mask ratio of the mask model. The mask ratio is related to the length of the original data; the multimodal label includes sentiment category label. Obtain multimodal training information associated with the sample dataset, and obtain a multimodal sample training set after preprocessing; Based on the sample dataset and the multimodal sample training set, the initial completion model is jointly trained for multiple tasks to obtain the task loss for each task. The text completion model is obtained by adjusting the model parameters of the initial completion model and the weights used to generate the predicted content of the masked region in the sample data by minimizing the weighted sum of the task losses of different tasks. The different task losses include at least two combinations of the following: completion prediction loss, multimodal matching loss, prediction loss of the mask model, and sentiment classification loss of the corresponding sample data. The completion prediction loss represents the loss of the predicted content of the masked region in the corresponding sample data relative to the completed content label of the sample data. The multimodal matching loss represents the correlation between the sample interaction behavior data in the multimodal sample training set and the sample data; The prediction loss of the masking model represents the prediction loss of the masking model for the non-masked regions in the corresponding sample data, where the non-masked regions include the preceding and following sample segments in the sample data.
8. The method according to claim 7, characterized in that, The training of the initial completion model is based on a pre-configured optimizer and is implemented using a mixed precision training method. During the training process, the optimizer state, gradient and model parameters are each sharded and stored on different processors of the computer device through a zero-redundancy optimizer sharding parallel strategy. The different computing tasks assigned to different processors are implemented using an interleaved execution mechanism during the training process. The training process employs a gradient accumulation strategy to update the optimizer state and synchronize model parameters. Based on an asynchronous checkpoint saving mechanism, when the gradient accumulation steps reach a pre-configured step threshold, the state data of the initial completed model for this training is stored. The state data includes the optimizer state, gradients, and model parameters after this training. The training process for the initial completion model also includes: In response to the satisfaction of the training termination condition, determine the text completion model to be determined in this training; Based on the scenario datasets constructed for different fine-tuning scenarios, the model parameters of the text completion model to be determined are fine-tuned through reinforcement learning optimization to obtain the text completion model. The fine-tuning is achieved by weighting the reward value obtained from the matching degree between the predicted and annotated content of the scene data, the semantic matching degree between the predicted and annotated content and the bidirectional contextual information in the scene data, and the linguistic style matching degree between the predicted and annotated content and the scene data. The fine-tuning can be implemented on the local model parameters of the text completion model based on an early stopping strategy. The local model parameters include the parameters contained in the attention layer and the output layer of the text completion model.
9. A computer device, characterized in that, The computer device includes: at least one memory and at least one processor, wherein: The memory is used to store multiple computer instructions; The processor is used to execute the computer instructions to perform the following steps: In response to a completion trigger signal for the input text presented on the interactive interface, bidirectional contextual information in the input text is obtained; the bidirectional contextual information includes the preceding text segment before the input marker position and the following text segment after the input marker position. Based on the preceding and following fragments, a fused contextual feature representation is generated; Obtain a multimodal feature representation associated with the context feature representation; the multimodal feature representation includes feature representations of at least one multimodal information from interactive behavior data, text data, image data, audio data, and video data; A correlation analysis is performed on the context feature representation and the multimodal feature representation to obtain the corresponding attention weight distribution; the attention weight distribution can indicate the degree of contribution of different multimodal information to the generation of the completion content for the input text; Based on the attention weight distribution, the multimodal feature representation, and the context feature representation, complete content is generated for the input text; The completed content is displayed on the interactive interface; The step of generating a fused contextual feature representation based on the preceding and following fragments includes: The context encoding module in the trained text completion model generates a fused context feature representation based on the semantic dependency between the preceding and following fragments; the length of each of the preceding and following fragments is associated with the input length of the input text; the text completion model can perform incremental optimization based on user feedback on historical completion content; The step of generating complete content for the input text based on the attention weight distribution, the multimodal feature representation, and the context feature representation includes: The multimodal feature fusion module in the text completion model fuses the context feature representation and the multimodal feature representation mapped to a feature space that matches the context feature representation based on the attention weight distribution to obtain a fused feature representation; the multimodal feature representation is obtained by the multimodal encoding module in the text completion model through feature extraction of multimodal information associated with bidirectional context information. The completion decoding module in the text completion model generates completed content for the input text based on the fused feature representation and the context feature representation; Specifically, generating complete content for the input text based on the fused feature representation and the contextual feature representation includes: Based on the fused feature representation and the context feature representation, multiple candidate completion contents are generated for the input text, as well as the completion confidence of each candidate completion content; In response to the existence of a target completion content among the plurality of candidate completion contents with a completion confidence greater than or equal to a confidence threshold, the target completion content is inserted into the input identifier position of the interactive interface to update the input text; In response to each of the completion confidence scores being less than the confidence threshold, a completion suggestion window is output on the interactive interface to present the multiple candidate completion contents; In response to the selection operation of the candidate completion content, the selected candidate completion content is inserted as the target completion content into the input identifier position to update the input text.
Citation Information
Patent Citations
Generative dialogue method and system based on multi-modal knowledge enhancement
CN116450787A
Multi-modal knowledge graph completion method based on modal hierarchical fusion
CN119089992A