Interaction processing method and device based on large language model, storage medium and equipment
Patent Information
- Application Number
- CN202610702784.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-09-29
AI Technical Summary
然而,当前大语言模型在进行角色扮演时,角色扮演时间遵循能力较差,经常会出现角色幻觉问题,即知道角色不应知晓的未来知识或事件
[0015]根据本申请的另一实施例,一种计算机程序产品或计算机程序,该计算机程序产品或计算机程序包括计算机指令,该计算机指令存储在计算机可读存储介质中。设备的处理器从计算机可读存储介质读取该计算机指令,处理器执行该计算机指令,使得该设备执行本申请实施例所述的各种可选实现方式中提供的方法。
Smart Images

Figure CN122838595A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to an interactive processing method, apparatus, storage medium, and device based on a large language model. Background Technology
[0002] With the development of Large Language Model (LLM) technology, role-playing has become an important application scenario for artificial intelligence, widely used in fields such as game NPCs, virtual assistants, and educational simulations. However, current large language models have poor time-following ability when performing role-playing, often exhibiting the role illusion problem, i.e., knowing future knowledge or events that the role should not know.
[0003] In response, related technologies guide large language models through system prompts and a few examples. However, due to the rich knowledge system of large language models, their spatiotemporal information is relatively disordered, and these methods have limited guiding effects on large language models. The time-following ability of large language models during role-playing is still poor, resulting in a poor user experience during role-playing. Summary of the Invention
[0004] This application provides a solution that can effectively improve the time-following ability of large language models in role-playing, thereby enhancing the user experience.
[0005] The embodiments of this application provide the following technical solutions: According to one embodiment of this application, an interactive processing method based on a large language model includes: segmenting input text according to temporal integrity to obtain multiple first segments; segmenting the input text using a large language model to obtain multiple second segments; performing time prediction processing on the multiple first segments to obtain time prediction results corresponding to each first segment; and assigning time information to the second segments corresponding to each first segment based on the time prediction results, so that the large language model performs content inference based on the multiple second segments and the corresponding time information.
[0006] In some embodiments of this application, the time prediction result includes a predetermined symbol or a timestamp; the time information includes a time embedding value; the step of assigning time information to the second segment corresponding to each first segment based on the time prediction result of each first segment includes: assigning the time prediction result of each first segment to the second segment corresponding to each first segment; if the time prediction result is a timestamp, then subtracting the timestamp from the anchor time to obtain a relative time; and binning the relative time or the predetermined symbol corresponding to each second segment to obtain the time embedding value corresponding to each second segment.
[0007] In some embodiments of this application, the time embedding value includes a time representation embedding value or a symbol representation embedding value; the step of binning the relative time or the predetermined symbol corresponding to each second word to obtain the time embedding value corresponding to each second word includes: querying the time representation embedding value corresponding to the bin to which the relative time belongs from a preset bin set, the preset bin set including the time representation embedding values corresponding to bins divided according to predetermined boundaries; and converting the predetermined symbol into the symbol representation embedding value corresponding to the symbol representation bin.
[0008] In some embodiments of this application, the content reasoning based on the plurality of second word segments and their corresponding time information includes: creating a time difference embedding matrix according to the number of word segments and the number of attention heads of the plurality of second word segments; embedding the time embedding value corresponding to each second word segment into the time difference embedding matrix; and performing content reasoning based on the time difference embedding matrix and the word embedding matrix of the plurality of second word segments to obtain the output text.
[0009] In some embodiments of this application, the step of performing content reasoning based on the temporal difference embedding matrix and the word embedding matrices of the plurality of second word segments to obtain output text includes: superimposing the word embedding matrices of the plurality of second word segments, the relative position bias matrix of the plurality of second word segments, and the temporal difference embedding matrix to obtain a superimposed matrix; performing attention mechanism processing based on the superimposed matrix to obtain an attention score matrix; and performing content reasoning based on the attention score matrix to obtain output text.
[0010] In some embodiments of this application, the training method of the large language model includes: obtaining a training set, the training set including training data corresponding to different roles, the training data including multiple input samples and output samples corresponding to each input sample that conform to the role's timeline and role style; training the large language model with the input samples as model input and the output samples as the model's expected output until it meets the preset training objective, thereby obtaining the trained large language model.
[0011] In some embodiments of this application, the method further includes: using a reward model to perform a role timeline consistency score on the output text of the large language model; and optimizing the large language model based on the consistency score using a proximal strategy optimization strategy.
[0012] According to one embodiment of this application, an interactive processing device based on a large language model is provided. The device includes: a first word segmentation module, configured to: segment input text according to temporal integrity to obtain multiple first words; a second word segmentation module, configured to: segment the input text using a large language model to obtain multiple second words; a time prediction module, configured to: perform time prediction processing on the multiple first words to obtain a time prediction result corresponding to each first word; and a time assignment module, configured to: assign time information to the second words corresponding to each first word based on the time prediction result corresponding to each first word, so that the large language model performs content reasoning based on the multiple second words and the corresponding time information.
[0013] According to another embodiment of this application, a storage medium stores a computer program thereon, which, when executed by a device's processor, causes the device to perform the methods described in the embodiments of this application.
[0014] According to another embodiment of this application, an apparatus may include: a memory storing a computer program; and a processor reading the computer program stored in the memory to execute the methods described in the embodiments of this application.
[0015] According to another embodiment of this application, a computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of the device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the device to perform the methods provided in the various optional implementations described in the embodiments of this application.
[0016] In this embodiment, the input text is segmented according to temporal integrity to obtain multiple first segments; the input text is segmented using a large language model to obtain multiple second segments; time prediction processing is performed on the multiple first segments to obtain time prediction results corresponding to each first segment; based on the time prediction results corresponding to each first segment, time information is assigned to the second segments corresponding to each first segment, so that the large language model can perform content reasoning based on the multiple second segments and the corresponding time information.
[0017] In this embodiment of the application, the input text is segmented according to temporal integrity, and the time prediction results corresponding to the multiple first segment words are obtained. The time prediction results corresponding to each first segment word are then assigned to the second segment words corresponding to each first segment word segmented by the large language model. This allows the large language model to accurately follow time for content reasoning based on multiple second segment words and their corresponding time prediction results. This transforms the large language model from an all-knowing but spatiotemporally disordered information database into an intelligent entity capable of thinking, communicating, and creating based on the role's temporal context, effectively improving the large language model's ability to follow time in role-playing and enhancing the user experience. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart of an interactive processing method based on a large language model according to an embodiment of this application is shown.
[0020] Figure 2 A flowchart illustrating the time information assignment process according to an embodiment of this application is shown.
[0021] Figure 3 A flowchart illustrating a model training process according to an embodiment of this application is shown.
[0022] Figure 4 A block diagram of an interactive processing apparatus based on a large language model according to an embodiment of this application is shown.
[0023] Figure 5 A block diagram of a device according to one embodiment of this application is shown. Detailed Implementation
[0024] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are merely illustrative of the present disclosure and are not intended to limit the present disclosure. Furthermore, the embodiments provided below are some embodiments for implementing the present disclosure, and not all embodiments for implementing the present disclosure. Unless otherwise specified, the technical solutions described in the embodiments of the present disclosure can be implemented in any combination.
[0025] It should be noted that, in the embodiments of this disclosure, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method or apparatus that includes a list of elements includes not only the elements expressly described, but also other elements not expressly listed, or elements inherent to implementing the method or apparatus. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other related elements (e.g., steps in the method or units in the apparatus; for example, a unit may be a portion of circuitry, a portion of a processor, a portion of a program or software, etc.) in the method or apparatus that includes that element.
[0026] For example, the interactive processing method based on a large language model provided in this disclosure includes a series of steps. However, the interactive processing method based on a large language model provided in this disclosure is not limited to the steps described. Similarly, the interactive processing device based on a large language model provided in this disclosure includes a series of units. However, the device provided in this disclosure is not limited to the units explicitly described, but may also include units that need to be set up for obtaining relevant information or processing based on information.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure.
[0028] It is understood that in the specific implementation of this application, relevant data is involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0029] With the development of Large Language Model (LLM) technology, role-playing has become an important application scenario for artificial intelligence, widely used in fields such as game NPCs, virtual assistants, and educational simulations. However, current large language models have poor time-following ability when performing role-playing, often exhibiting the role illusion problem, i.e., knowing future knowledge or events that the role should not know.
[0030] In response, related technologies guide large language models through system prompts and a few examples. However, due to the rich knowledge system of large language models, their spatiotemporal information is relatively disordered, and these methods have limited guiding effects on large language models. The time-following ability of large language models during role-playing is still poor, resulting in a poor user experience during role-playing.
[0031] To address these issues, this application provides an interaction processing scheme based on a large language model, which can effectively improve the time-following ability of the large language model during role-playing and enhance the user experience.
[0032] Based on the interactive processing scheme based on the large language model in the embodiments of this application, when the large language model plays the role of a game NPC, the NPC can possess knowledge and language habits that are completely consistent with its historical context. For example, an NPC from ancient Rome would not talk about the "Internet," and a medieval knight would be surprised by "muskets." When the large language model plays the role of a historical figure or a fictional character as a communication partner, for example, "Character 1" would not discuss the latest breakthroughs in quantum computing, and "Character 2" would reason using the language style of his era. When the large language model plays the role of educational virtual characters such as Character 3 and Character 4, for example, the model will answer questions in Character 4's voice that he cannot understand or speculates based on his own theories. When the large language model plays the role of a historical research and fact-checking tool, for example, when a description about the "Tang Dynasty" is input, the large language model can quickly identify words or events that do not conform to the characteristics of the era (such as "they used corn to fill their stomachs") and provide prompts. When the large language model plays the role of a content creation and media assistant, the assistant can ensure that the dialogue, behavior, and description of the surrounding environment of the characters are consistent with the historical setting and avoid common sense errors. The assistant can automatically generate marketing copy, short video scripts, or articles that conform to the style of a specific era.
[0033] The following describes in detail the relevant embodiments of the interactive processing scheme based on the large language model provided in this application.
[0034] Figure 1 The diagram illustrates a flowchart of an interaction processing method based on a large language model according to an embodiment of this application. The execution entity of this interaction processing method based on a large language model can be a device or a server. The device can be a mobile phone, computer, smartwatch, in-vehicle device, robot, speaker, and other home appliances, etc. The server can be a cloud server or a physical server, etc.
[0035] like Figure 1 As shown, the interactive processing method based on a large language model may include steps S110 to S140.
[0036] Step S110: The input text is segmented according to time integrity to obtain multiple first segments; Step S120: The input text is segmented using a large language model to obtain multiple second-level words; Step S130: Perform time prediction processing on multiple first word segments to obtain the time prediction results corresponding to each first word segment; Step S140: Based on the time prediction results corresponding to each first word, assign time information to the second word corresponding to each first word, so that the large language model can perform content reasoning based on multiple second words and their corresponding time information.
[0037] When using a large language model for role-playing, the content expressed by the user (such as content expressed by the user through voice, text input, or gestures) can be converted into input text.
[0038] For the input text, a pre-trained tokenizer (which can be described as a Time-Tokenizer) is used to segment the input text according to temporal integrity, resulting in multiple first-order tokens. The pre-trained tokenizer can completely segment words containing time information into first-order tokens, meaning the segmented first-order tokens possess temporal integrity. For example, it can segment "iphoneX" and "Harry Potter and the Philosopher's Stone" into first-order tokens with temporal integrity (i.e., words containing time information). Based on "iphoneX," we can determine its creation time as 2017. However, when "iphoneX" is segmented into the two words "iphone" and "X," the exact time corresponding to "iphone" and "X" cannot be accurately determined, thus "iphone" and "X" do not possess temporal integrity.
[0039] For example, a pre-trained word segmenter can be used to segment the input text “Do you think character 5 will use an iPhone X?” into several first-level words: [“you”, “think”, “character 5”, “will use”, “iphoneX”, “ma”].
[0040] In addition, the large language model performs word segmentation on the input text (that is, the large language model also performs word segmentation on the input text through the native word segmenter (that is, the word segmenter that comes with the large language model itself), and obtains multiple second words.
[0041] For example, the large language model can use the native word segmenter to segment the input text "Do you think character 5 will use an iPhone X?" into several second-level words: ["you", "think", "character 5", "will use", "iphone", "X", "ma"].
[0042] In addition, a pre-trained temporal inference module is used to perform temporal prediction processing on a plurality of first word segments. Since the segmented first word segments have temporal integrity, the temporal prediction results corresponding to each first word segment can be accurately predicted. Wherein, for the first word segments that have temporal integrity and contain time (i.e., temporal words) such as "person 5" and "iphoneX", the corresponding temporal prediction result is a specific time stamp. For example, the temporal prediction result corresponding to "person 5" can be time stamp 762, and the temporal prediction result corresponding to "iphoneX" can be time stamp 2017. For first word segments that do not contain time (i.e., non-temporal words) such as "you" and "think", the corresponding temporal prediction result can use a predetermined symbol (e.g., <ut>For example, the time prediction result corresponding to "you" could be... <ut>The time prediction result corresponding to "feeling" can be <ut>.
[0043] For example, the first segmentation words ["you", "feel", "character 5", "will use", "iphoneX", "ma"] can correspond to the time prediction results as [" <ut> ”," <ut> ”,"762”," <ut> ”,"2017”," <ut>”).
[0044] The pre-trained tokenizer (which can be described as a Time-Tokenizer) and the pre-trained time inference module can be set independently of the large language model, or they can be incorporated into the large language model. In one specific embodiment of this application, the pre-trained tokenizer (which can be described as a Time-Tokenizer) and the pre-trained time inference module can be incorporated into the large language model, facilitating direct role-playing using the large language model.
[0045] Then, based on the time prediction results corresponding to each first word segment, time information is assigned to the second word segment corresponding to each first word segment. The second word segment corresponding to the first word segment is the second word segment from the same source as the first word segment. For example, the second word segment corresponding to the first word segment "iphoneX" is "iphone" and "X", and the second word segment corresponding to the first word segment "you" is "you". Furthermore, based on [" <ut> ”," <ut> ”,"762”," <ut> ”,"2017”," <ut>The system can assign time information to the second segmentation words ["you", "feel", "character 5", "will use", "iphone", "X", "ma"]. Therefore, the second segmentation words such as "iphone" and "X" switched to by the large language model through the native word segmenter will also have accurate time information.
[0046] The large language model then performs content reasoning based on multiple second-segment words and the corresponding time information of each second-segment word. Since second-segment words such as "iphone" and "X" have accurate time information, the large language model can accurately follow the time of these second-segment words to perform content reasoning, thereby improving the large language model's time-following ability and avoiding role illusion in the large language model.
[0047] For example, when the large language model plays a role that occurred between Person 5 and the iPhone X, the large language model performs content reasoning based on the second segment words ["you", "think", "person 5", "will use", "iphone", "X", "ma"] and their corresponding time information. The output text obtained by the reasoning could be "I don't know when the iPhone X appeared, but it seems that there was no such thing as the iPhone X in Person 5's era, so Person 5 probably wouldn't use the iPhone X."
[0048] In summary, the method described in this embodiment of the application segments the input text according to temporal integrity and predicts the time prediction results corresponding to the multiple first segment words. The time prediction results corresponding to each first segment word are then assigned to the second segment words corresponding to each first segment word segmented by the large language model. This allows the large language model to accurately follow time for content reasoning based on multiple second segment words and their corresponding time prediction results. This transforms the large language model from an all-knowing but spatiotemporally disordered information database into an intelligent entity capable of thinking, communicating, and creating based on the role's temporal context, effectively improving the large language model's ability to follow time in role-playing and enhancing the user experience.
[0049] The following description Figure 1 Further optional specific embodiments for each step performed during interactive processing based on a large language model, as described in the example.
[0050] Step S110: The input text is segmented according to time integrity to obtain multiple first segments.
[0051] For the input text, a pre-trained tokenizer (which can be described as Time-Tokenizer) is used to perform tokenization processing on the input text according to temporal integrity, and multiple first tokens with temporal integrity can be obtained. For example, the pre-trained tokenizer can split the input text "Do you think character 5 will use iphoneX" into the following first tokens: ["you", "think", "character 5", "will use", "iphoneX", "?"]
[0052] The pre-trained tokenizer can be set independently of the large language model, or alternatively, the pre-trained tokenizer can be integrated inside the large language model. In a specific implementation of the present application, the pre-trained tokenizer is integrated inside the large language model, which facilitates directly using the large language model for role-playing.
[0053] Wherein, the specific training process of the pre-trained tokenizer may include: preparing a temporal integrity vocabulary, which includes words with temporal integrity (for example, words such as iphoneX and Harry Potter and the Philosopher's Stone); using the temporal integrity vocabulary to train a tokenizer (i.e., the pre-trained tokenizer) that performs tokenization according to the words in the temporal integrity vocabulary, and the pre-trained tokenizer can perform tokenization processing on the input text according to temporal integrity.
[0054] Step S120, performing tokenization processing on the input text by the large language model to obtain a plurality of second tokens.
[0055] For example, the large language model can split the input text "Do you think character 5 will use iphoneX" into the following second tokens via a native tokenizer: ["you", "think", "character 5", "will use", "iphone", "X", "?"]
[0056] Step S130, performing time prediction processing on the plurality of first tokens to obtain a time prediction result corresponding to each first token.
[0057] A pre-trained time inference module is used to perform time prediction processing on the plurality of first tokens. Since the obtained first tokens have temporal integrity, the time prediction result corresponding to each first token can be accurately predicted. For example, for the first tokens ["you", "think", "character 5", "will use", "iphoneX", "?"], the corresponding time prediction results can be [" <ut> ”," <ut> ”,"762”," <ut> ”,"2017”," <ut>”).
[0058] The pre-training time inference module can be set independently of the large language model, or it can be incorporated into the large language model. In one specific embodiment of this application, the pre-training time inference module is incorporated into the large language model, facilitating direct role-playing using the large language model.
[0059] The specific training process of the pre-trained time inference module may include: acquiring a large number of knowledge label pairs; training the neural network with "knowledge" as input and "label" as the expected output until the time prediction conditions are met; the trained neural network can then be used as the pre-trained time inference module. This neural network can be a recurrent neural network (RNN), a long short-term memory network (LSTM), or a temporal convolutional network (TCN), etc., and this application does not impose any special limitations on this. The time prediction conditions may include training reaching a predetermined number of iterations, accuracy reaching a predetermined threshold, or passing a test, etc., and this application does not impose any special limitations on this.
[0060] Here, "knowledge tag pairs" represent the correspondence between knowledge (such as knowledge of people, objects, or terms) and tags. Knowledge that implies time can be tagged with a specific timestamp, while knowledge that does not imply time (such as "you," "me," etc.) can be tagged with predetermined symbols (such as...). <ut>For example, the timestamp for a "person" could be the "year of death," and the timestamp for an "object or term" could be the "year of invention." Timestamps before Christ could be represented by a negative sign; your corresponding predefined symbol could be... <ut>The corresponding predefined symbol can be <ut>.
[0061] Step S140: Based on the time prediction results corresponding to each first word, assign time information to the second word corresponding to each first word, so that the large language model can perform content reasoning based on multiple second words and their corresponding time information.
[0062] For example, the second participle corresponding to the first participle "iphoneX" is "iphone" and "X", and the second participle corresponding to the first participle "you" is "you". Furthermore, according to [" <ut> ”," <ut> ”,"762”," <ut> ”,"2017”," <ut>The word "] can be assigned time information to the second participles ["you", "feel", "character 5", "will use", "iphone", "X", "ma"].
[0063] Time information is assigned to the second segment corresponding to each first segment. The second segment, such as "iphone" and "X," switched to by the large language model through the native word segmenter will also have accurate time information. The large language model then performs content inference based on multiple second segment and their corresponding time information. Since the second segment, such as "iphone" and "X," has accurate time information, the large language model can accurately follow the time of these second segment to perform content inference and obtain output text that conforms to the character's timeline.
[0064] See Figure 2 In one embodiment, the time prediction result may include a predetermined symbol or a timestamp; the time information includes a time embedding value; in step S140, according to the time prediction result corresponding to each first word, time information is assigned to the second word corresponding to each first word, including: step S210, assigning the time prediction result corresponding to each first word to the second word corresponding to each first word respectively; step S220, if the time prediction result is a timestamp, then the difference between the timestamp and the anchor time is calculated to obtain the relative time; step S230, the relative time or predetermined symbol corresponding to each second word is bucketed to obtain the time embedding value corresponding to each second word.
[0065] The time prediction results corresponding to each first word are assigned to their respective second words. For example, the time prediction results for the first words ["you", "feel", "character5", "will use", "iphoneX", "ma"] could be ["... " <ut> ”," <ut> ”,"762”," <ut> ”,"2017”," <ut>Assigning the corresponding second segment to each of the first segment words, the time prediction results for the second segment words ["you", "feel", "character 5", "will use", "iphone", "X", "ma"] can be [" <ut> ”," <ut> ”,"762”," <ut> ”,"2017”,"2017”," <ut>In other words, both "iphone" and "X" are associated with "2017".
[0066] Furthermore, if the predicted time for a certain second word is a timestamp t_j, then the difference between the timestamp t_j and the anchor time t_i is used to obtain the relative time Δt = t_j - t_i. Here, the anchor time t_i can be the character's birth time or the current time of processing the input text, etc. Calculating the relative time allows for the integration of relative time information into subsequent processing, making the large language model more accurately follow the character's timeline. For example, if the timestamp t_j for "character 5" is 762, and the anchor time t_i is 2026, then the relative time Δt = 762 - 2026 = -1264.
[0067] Furthermore, since the time span is large and time is usually independent of the input text sequence, and time is only related to the semantics expressed by the token, by binning the relative time or predetermined symbol corresponding to each second token, the number of final time embedding values can be reduced. Also, when some second tokens correspond to predetermined symbols, the corresponding time embedding values can be obtained through binning, which is easier for large language models to perform computational inference.
[0068] For example, in one example, "you" corresponds to the predefined symbol. <ut>"Bugling can convert it into temporal embedding values (embedding)." <ut>), the corresponding relative time for "character 5" is -1264, which can be converted into a time embedding value embedding(-1264) through bucketing. For example, the finally corresponding time embedding values of the second participles ["you", "think", "character 5", "will use", "iphone", "X", "right"] can be [embedding( <ut>),embedding( <ut>),embedding(-1264),embedding( <ut>),embedding(-9),embedding(-9),embedding( <ut>)].
[0069] Furthermore, in one embodiment, the time embedding value includes a time representation embedding value or a symbol representation embedding value; in step S230, the relative time or predetermined symbol corresponding to each second word is divided into buckets to obtain the time embedding value corresponding to each second word, which may specifically include: querying the time representation embedding value corresponding to the bucket to which the relative time belongs from a preset bucket set, the preset bucket set including the time representation embedding values corresponding to buckets divided according to predetermined boundaries; and converting the predetermined symbol into the symbol representation embedding value corresponding to the symbol representation bucket.
[0070] Set a preset bucket set and special symbolic representation buckets. The preset bucket set includes buckets divided according to predetermined boundaries, such as 0, 1, -1, 2, -2, (2, 10], (-10, -2], (10, 100], (-100, -10], >100, <-100, etc. Each bucket has a corresponding time representation embedding value (i.e., the embedding value used to represent time), for example, (-10, The time representation embedding value corresponding to the bucket with value <-2] can be 05, and the time representation embedding value corresponding to the bucket with value <-100 can be 09. When querying the time representation embedding value corresponding to the bucket to which the relative time belongs from the preset bucket set, for example, the relative time corresponding to "person 5" is -1264, then the bucket to which -1264 belongs is <-100, and therefore, the time representation embedding value of the relative time corresponding to "person 5" is -1264, embedding(-1264) = 09; similarly, the time representation embedding value of the relative time corresponding to "iphone" and "X" is -9, embedding(-9) = 05.
[0071] A symbolic representation bucket is a bucket corresponding to a predetermined symbol. A special symbolic representation bucket has a corresponding symbolic representation embedding value (an embedding value used to represent the symbol). For example, the symbolic representation embedding value corresponding to a symbolic representation bucket can be 00. When converting a predetermined symbol to the symbolic representation embedding value corresponding to a symbolic representation bucket, the predetermined symbols such as "you" and "feel" are... <ut>The corresponding symbolic representation embedding value ( <ut>)=00.
[0072] In one embodiment, the large language model performs content reasoning based on multiple second-segment words and their corresponding time information. Specifically, this may include: creating a time difference embedding matrix according to the number of segments of the multiple second-segment words and the number of attention heads; embedding the time embedding value corresponding to each second-segment word into the time difference embedding matrix; and performing content reasoning based on the time difference embedding matrix and the word embedding matrix of the multiple second-segment words to obtain the output text.
[0073] For example, the statement "Does Person 5 use an iPhone X?" includes the second word ["Person 5", "use", "iphone", "X", "?"], with a word count of 5. If the number of attention heads is n, then the temporal difference embedding matrix can be obtained by embedding the temporal embedding values corresponding to each second word into the temporal difference embedding matrix as shown in the following illustration.
[0074]
[0075] In addition, multiple second-segmentation word embedding matrices will be encoded within the large model. For example, the word embedding matrices of multiple second-segmentation words can be shown in the following matrix.
[0076]
[0077] Finally, by combining the time difference embedding matrix and the word embedding matrices of multiple second-segment words, the large language model can accurately follow the character's timeline to perform content reasoning and obtain output text that conforms to the character's historical background.
[0078] Furthermore, in one embodiment, content reasoning based on the temporal difference embedding matrix and the word embedding matrices of multiple second-segment words to obtain the output text may include: superimposing the word embedding matrices of multiple second-segment words, the relative position bias matrices of multiple second-segment words, and the temporal difference embedding matrix to obtain a superimposed matrix; performing attention mechanism processing based on the superimposed matrix to obtain an attention score matrix; and performing content reasoning based on the attention score matrix to obtain the output text.
[0079] For multiple second-word segmentation models, multiple relative position bias matrices for the second-word segments can also be encoded. For example, the relative position bias matrices for multiple second-word segments can be as shown in the following matrix. The relative position bias matrix includes the relative position embedding value of the relative position corresponding to each second-word segment.
[0080]
[0081] By superimposing the word embedding matrices, relative position bias matrices, and temporal difference embedding matrices of multiple second-segment words, a superimposed matrix can be obtained. This superimposed matrix includes the superimposed embedding value corresponding to each second-segment word. For example, the superimposed embedding value for character 5 is embedding(die1) = embedding(-1264) + embedding(character5) + embedding(P1), and the corresponding superimposed embedding value is embedding(die2) = embedding(... <ut>) + embedding (will be used) + embedding (P2), and so on.
[0082] Furthermore, by performing attention mechanism processing based on the superposition matrix, an attention score matrix can be obtained. Specifically, based on the submatrix corresponding to attention head 1 in the superposition matrix, the first attention score matrix Attention1 can be obtained through the attention mechanism; based on the submatrix corresponding to attention head i in the superposition matrix, the ith attention score matrix Attentioni can be obtained through the attention mechanism; and based on the submatrix corresponding to attention head n in the superposition matrix, the nth attention score matrix Attentionn can be obtained through the attention mechanism. Finally, superimposing the first attention score matrix Attention1 to the nth attention score matrix Attentionn yields the final attention score matrix.
[0083] Large language models can obtain output text by performing content reasoning based on the attention score matrix through subsequent inference layers and output layers. Different large language models can typically include corresponding subsequent inference layers.
[0084] For example, a subsequent inference layer may include, in sequence, a linear projection layer, a residual connection and normalization layer, a feedforward neural network layer, and a repeated residual connection and normalization layer. A large language model can have at least one subsequent inference layer set up sequentially. The attention score matrix is processed by at least one subsequent inference layer to obtain the hidden state. The hidden state is then processed by an output layer (which may include a linear layer and a softmax layer) to obtain the final output text. Specifically, the linear projection layer can linearly project the input matrix (such as the attention score matrix or the output matrix of the previous subsequent inference layer) to obtain a projection matrix; the residual connection and normalization layer can add and normalize the projection matrix and the input matrix to obtain a normalized matrix; the feedforward neural network layer can perform a nonlinear transformation on the normalized matrix to obtain a transformed matrix; and the repeated residual connection and normalization layer can add and normalize the transformed matrix and the normalized matrix to obtain the output matrix of the subsequent inference layer.
[0085] Furthermore, in one embodiment of this application, role-playing training can also be performed on the large language model, see [reference]. Figure 3 The specific training method for the large language model may include: Step S310, obtaining a training set, which includes training data corresponding to different roles, including multiple input samples and output samples corresponding to each input sample that conform to the role's timeline and style; Step S320, training the large language model with the input samples as model input and the output samples as the model's expected output until it meets the preset training objective, and obtaining the trained large language model.
[0086] The constructed training set includes training data for different characters (such as character 6, character 7, character 8, etc.). The training data for each character includes multiple input samples and output samples corresponding to each input sample that conform to the character's timeline and style.
[0087] The input sample for each character can be a question that includes a timeline challenge in the character design, such as "Is there a subway station near your home?" or "What do you usually order takeout?". Each character's input sample has a corresponding output sample that matches the character's timeline and style, such as "What is a subway station?"
[0088] Output samples that conform to the character's timeline can be output samples that "do not show that the character knows knowledge that they should not know" and "do not show that the character knows information beyond the character's cognition"; output samples that conform to the character's style can be output samples that "use a first-person perspective and avoid statements made by other persons such as character A", "do not include thought processes", and "do not contain unnatural precise time statements".
[0089] Using the above input samples as model input and the output samples as the model's expected output, the large language model is trained through role-playing until it meets the preset training objectives (such as all roles passing training or training reaching a predetermined number of times). The trained large language model is then obtained. Furthermore, the output text of the trained large language model further conforms to the role's timeline (i.e., the timeline before the role's death) and role style. Combined with this time-reasoning training process and the interactive processing flow of steps S110 to S140, it further ensures that typical problems such as "should not know, should know, different cognition, and different era context" are solved during role-playing, and further improves the large language model's time-following ability and user experience during role-playing.
[0090] Specifically, training a large language model by using input samples as model inputs and output samples as the model's expected outputs can include: segmenting the input samples according to temporal integrity to obtain multiple first-level segments; and segmenting the input samples using the large language model to obtain multiple second-level segments. Time prediction processing is performed on multiple first-segment words to obtain the time prediction results corresponding to each first-segment word; based on the time prediction results corresponding to each first-segment word, time information is assigned to the second-segment words corresponding to each first-segment word, so that the large language model can perform content inference based on multiple second-segment words and their corresponding time information to obtain prediction samples; the loss is calculated based on the prediction samples and output samples, and the model parameters are updated by gradient descent.
[0091] In addition, some methods involve testing large language models using a timeline consistency test evaluation system. This system can be used to set up a set of test roles (e.g., selecting 5 representative roles from different eras) and corresponding timeline challenge questions (e.g., what software do you usually use for payments). Once the model passes the timeline consistency test evaluation system, it is considered to meet the preset training objectives.
[0092] Furthermore, in one embodiment, it may also include: using a reward model to score the consistency of the role timeline of the output text of the large language model; and optimizing the large language model based on the consistency score using a proximate strategy optimization strategy.
[0093] In this embodiment, when using a large language model for role-playing, a pre-trained reward model is used to score the consistency of the output text of the large language model with the role timeline. The consistency score is used to reflect whether the output text conforms to the role timeline. If it does not conform to the role timeline, the consistency score can be 0, and if it conforms to the role timeline, the consistency score can be 1.
[0094] Based on the consistency score of the role timeline in the output text, the large language model can be optimized based on the Proximal Policy Optimization (PPO) strategy. In other words, the large language model is continuously trained by reinforcement learning based on the Proximal Policy Optimization (PPO) strategy to further improve the large language model's time compliance ability and user experience during role-playing.
[0095] To facilitate better implementation of the large language model-based interactive processing method provided in this application, this application also provides a large language model-based interactive processing device based on the aforementioned large language model-based interactive processing method. The meanings of the terms used are the same as in the aforementioned large language model-based interactive processing method, and specific implementation details can be found in the descriptions in the method embodiments. Figure 4 A block diagram of an interactive processing apparatus based on a large language model according to an embodiment of this application is shown.
[0096] like Figure 4 As shown, the interactive processing device 400 based on a large language model may include: a first word segmentation module 410, which can be used to: segment the input text according to temporal integrity to obtain multiple first words; a second word segmentation module 420, which can be used to: segment the input text using a large language model to obtain multiple second words; a time prediction module 430, which can be used to: perform time prediction processing on the multiple first words to obtain time prediction results corresponding to each first word; and a time assignment module 440, which can be used to: assign time information to the second words corresponding to each first word according to the time prediction results corresponding to each first word, so that the large language model can perform content reasoning based on the multiple second words and the corresponding time information.
[0097] In some embodiments of this application, the time prediction result includes a predetermined symbol or a timestamp; the time information includes a time embedding value; when assigning time information to the second segment corresponding to each first segment based on the time prediction result corresponding to each first segment, the time assignment module 440 can be used to: assign the time prediction result corresponding to each first segment to the second segment corresponding to each first segment respectively; if the time prediction result is a timestamp, then the difference between the timestamp and the anchor time is calculated to obtain the relative time; the relative time or the predetermined symbol corresponding to each second segment is bucketed to obtain the time embedding value corresponding to each second segment.
[0098] In some embodiments of this application, the time embedding value includes a time representation embedding value or a symbol representation embedding value; when the relative time or the predetermined symbol corresponding to each second word is divided into buckets to obtain the time embedding value corresponding to each second word, the time assignment module 440 can be used to: query the time representation embedding value corresponding to the bucket to which the relative time belongs from a preset bucket set, the preset bucket set including the time representation embedding value corresponding to the bucket divided according to a predetermined boundary; and convert the predetermined symbol into the symbol representation embedding value corresponding to the symbol representation bucket.
[0099] In some embodiments of this application, when performing content reasoning based on the plurality of second word segments and their corresponding time information, the device further includes a content reasoning module that can be used to: create a time difference embedding matrix according to the number of word segments and the number of attention heads of the plurality of second word segments; embed the time embedding value corresponding to each second word segment into the time difference embedding matrix; and perform content reasoning based on the time difference embedding matrix and the word embedding matrix of the plurality of second word segments to obtain output text.
[0100] In some embodiments of this application, when performing content reasoning based on the time difference embedding matrix and the word embedding matrices of the plurality of second word segments to obtain output text, the content reasoning module can be used to: superimpose the word embedding matrices of the plurality of second word segments, the relative position bias matrix of the plurality of second word segments, and the time difference embedding matrix to obtain a superimposed matrix; perform attention mechanism processing based on the superimposed matrix to obtain an attention score matrix; and perform content reasoning based on the attention score matrix to obtain output text.
[0101] In some embodiments of this application, the apparatus further includes a training module for a large language model, which can be used to: acquire a training set, the training set including training data corresponding to different roles, the training data including multiple input samples and output samples corresponding to each input sample that conform to the role's timeline and style; train the large language model using the input samples as model input and the output samples as the model's expected output until it meets a preset training objective, thereby obtaining the trained large language model.
[0102] In some embodiments of this application, the apparatus further includes a model optimization module that can be used to: use a reward model to perform a role timeline consistency score on the output text of the large language model; and optimize the large language model based on the consistency score using a proximate strategy optimization strategy.
[0103] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0104] Furthermore, embodiments of this application also provide a device, such as... Figure 5 As shown, Figure 5 A block diagram of a device according to an embodiment of this application is shown, specifically: The device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that... Figure 5 The device structure shown does not constitute a limitation on the device and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. Wherein: The processor 501 is the control center of the device, connecting various parts of the computer device via various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user page, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 501.
[0105] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0106] The device also includes a power supply 503 that supplies power to the various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0107] The device may also include an input unit 504, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0108] Although not shown, the device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the device can load the executable files corresponding to the processes of one or more computer programs into the memory 502 according to the following instructions, and the processor 501 runs the computer programs stored in the memory 502, thereby realizing the various functions in the foregoing embodiments of this application.
[0109] For example, processor 501 can perform the following: segment the input text according to temporal integrity to obtain multiple first segments; segment the input text using a large language model to obtain multiple second segments; perform temporal prediction processing on the multiple first segments to obtain temporal prediction results corresponding to each first segment; and assign temporal information to the second segments corresponding to each first segment based on the temporal prediction results, so that the large language model can perform content inference based on the multiple second segments and the corresponding temporal information.
[0110] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0111] Therefore, embodiments of this application also provide a storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the methods provided in embodiments of this application.
[0112] The storage medium can be a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0113] Since the computer program stored in the storage medium can execute the steps of any of the methods provided in the embodiments of this application, the beneficial effects that the methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0114] According to another embodiment of this application, a computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of the device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the device to perform the methods provided in the various optional implementations described in the embodiments of this application.
[0115] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0116] It should be understood that this application is not limited to the embodiments described above and shown in the accompanying drawings, but various modifications and changes can be made without departing from its scope.< / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut> < / ut>
Claims
1. An interactive processing method based on a large language model, characterized in that, include: The input text is segmented according to temporal integrity to obtain multiple first-level segments; The input text is segmented using a large language model to obtain multiple second-level segments. Perform time prediction processing on the plurality of first word segments to obtain the time prediction result corresponding to each first word segment; Based on the time prediction results corresponding to each of the first word segments, time information is assigned to the second word segments corresponding to each of the first word segments, so that the large language model can perform content reasoning based on the multiple second word segments and the corresponding time information.
2. The method according to claim 1, characterized in that, The time prediction result includes a predetermined symbol or timestamp; the time information includes a time embedding value; The step of assigning time information to the second segment corresponding to each of the first segment based on the time prediction results of each of the first segment includes: Assign the time prediction results corresponding to each of the first word segments to the corresponding second word segments; If the time prediction result is a timestamp, then the difference between the timestamp and the anchor time is calculated to obtain the relative time; The relative time or predetermined symbol corresponding to each second word is bucketed to obtain the time embedding value corresponding to each second word.
3. The method according to claim 2, characterized in that, The time embedding value includes a time representation embedding value or a symbolic representation embedding value; The step of binning the relative time or predetermined symbol corresponding to each second word to obtain the time embedding value corresponding to each second word includes: The time representation embedding value corresponding to the bucket to which the relative time belongs is queried from the preset bucket set, which includes the time representation embedding value corresponding to the buckets divided according to a predetermined boundary; The predetermined symbol is converted into the symbol representation embedding value corresponding to the symbol representation bucket.
4. The method according to claim 2, characterized in that, The content reasoning based on the multiple second word segments and corresponding time information includes: A temporal difference embedding matrix is created based on the number of segments in the multiple second word segments and the number of attention heads. The time embedding value corresponding to each of the second word segments is embedded in the time difference embedding matrix; Content reasoning is performed based on the time difference embedding matrix and the word embedding matrices of the multiple second-segment words to obtain the output text.
5. The method according to claim 4, characterized in that, The content reasoning based on the time difference embedding matrix and the word embedding matrices of the multiple second word segments yields the output text, including: The word embedding matrix of the plurality of second word segments, the relative position bias matrix of the plurality of second word segments, and the time difference embedding matrix are superimposed to obtain the superimposed matrix; Based on the superposition matrix, attention mechanism processing is performed to obtain the attention score matrix; Content reasoning is performed based on the attention score matrix to obtain the output text.
6. The method according to any one of claims 1 to 5, characterized in that, The training methods for the large language model include: Obtain a training set, which includes training data corresponding to different roles. The training data includes multiple input samples and output samples corresponding to each input sample that conform to the role's timeline and style. The large language model is trained by using the input sample as the model input and the output sample as the model expected output until it meets the preset training objective, thus obtaining the trained large language model.
7. The method according to any one of claims 1 to 5, characterized in that, The method further includes: A reward model is used to score the consistency of the role timeline in the output text of the large language model. The large language model is optimized based on the consistency score using a proximal strategy optimization strategy.
8. An interactive processing device based on a large language model, characterized in that, include: The first word segmentation module is used to: segment the input text according to time integrity to obtain multiple first words; The second word segmentation module is used to: segment the input text using a large language model to obtain multiple second word segments; The time prediction module is used to: perform time prediction processing on the plurality of first word segments to obtain the time prediction result corresponding to each first word segment; The time assignment module is used to: assign time information to the second segment corresponding to each of the first segment based on the time prediction results of each first segment, so that the large language model can perform content reasoning based on the multiple second segment and the corresponding time information.
9. A storage medium, characterized in that, It stores a computer program that, when executed by the device's processor, causes the device to perform the method described in any one of claims 1 to 7.
10. A device, characterized in that, include: Memory, which stores computer programs; A processor reads a computer program stored in memory to perform the method described in any one of claims 1 to 7.