Statement association method, device, equipment and product

By storing historical input data locally in the terminal and using a large language model to generate the following statements of personalized association, the problem of lack of personalization and accuracy of existing input methods is solved, efficient and accurate statement association is achieved, and user privacy is protected.

CN120336500APending Publication Date: 2025-07-18IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204935.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing input method lacks personalization of the association method, which leads to the generated association content that does not conform to the user's input intention, the accuracy needs to be improved, and the existing technology has shortcomings in the association effect, degree of personalization, association efficiency and privacy protection in the whole sentence.

Method used

The historical input data is stored locally in the terminal, matched with similar statements in the previous statements through a large language model, and generated personalized associative statements below. The matching area and record area are used to optimize the search for similar statements, and combined with the rough and fine placing algorithms to improve the search efficiency and accuracy. The overall process is completed at the terminal to protect user privacy.

Benefits of technology

It improves the association effect, personalization and association accuracy of the entire sentence of the input method association, reduces the amount of data processing, improves the efficiency of association, and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336500A_ABST
    Figure CN120336500A_ABST
Patent Text Reader

Abstract

The invention provides a statement association-based method, apparatus, device and product, and is applied to the technical field of natural language processing. The statement association-based method comprises the following steps: acquiring a first preceding statement, the first preceding statement comprising a statement input or sent on a terminal at the current time; in historical input data locally stored in the terminal, previous statements with the statement content similar to that of the first previous statement are searched for, a second previous statement is obtained, and the historical input data comprise context statements input on the terminal in the past time; according to the first preceding sentence and the second preceding sentence, generating an association following sentence of the first preceding sentence through a large language model; and outputting the associated next text statement. Therefore, personalized association of the next-text statements is realized through similar historical previous-text statements and a large language model, and the accuracy of association of the next-text statements is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application is applied to the field of natural language processing technology, and particularly relates to a sentence association method, device, equipment and product. Background Art

[0002] With the popularization of intelligent devices and the development of input method technology, the input method can provide association content for the user's next input based on the user's input content.

[0003] In the related technology of the input method, it mainly relies on a static thesaurus and a general association algorithm (such as a rule-based association algorithm) to generate association content for the user's next input.

[0004] However, the above-mentioned association method of the input method lacks personalization, resulting in the generated association content not conforming to the user's input intention, and the accuracy needs to be improved. Summary of the Invention

[0005] To solve the above problems, this application proposes a sentence association method, device, equipment and product based on, which can realize personalized association of sentences and improve association accuracy.

[0006] The first aspect of this application provides a sentence association method, including: obtaining a first previous sentence, where the first previous sentence includes a sentence input or sent on the terminal at the current time; in the historical input data stored locally on the terminal, searching for a previous sentence whose sentence content is similar to the first previous sentence to obtain a second previous sentence, where the historical input data includes context sentences input on the terminal in the past; according to the first previous sentence and the second previous sentence, generating an associated next sentence of the first previous sentence through a large language model; outputting the associated next sentence.

[0007] In some embodiments, the historical input data is distributed in a matching area and a recording area, and the input times of the context sentences in the matching area are greater than the input times of the context sentences in the recording area; the step of searching for a previous sentence whose sentence content is similar to the first previous sentence in the historical input data stored locally on the terminal to obtain a second previous sentence includes: searching for a previous sentence whose sentence content is similar to the first previous sentence in the previous sentences included in the matching area to obtain the second previous sentence.

[0008] In some embodiments, finding an upstream statement whose statement content is similar to the first upstream statement from the upstream statements included in the matching region to obtain the second upstream statement includes: matching the upstream statements included in the matching region with the first upstream statement to obtain a rough ranking similarity between the upstream statements included in the matching region and the first upstream statement; selecting, from the upstream statements included in the matching region, the upstream statements whose rough ranking similarity with the first upstream statement meets a first screening condition to obtain candidate upstream statements; matching the candidate upstream statements with the first upstream statement to obtain a fine ranking similarity between the candidate upstream statements and the first upstream statement; and selecting, from the candidate upstream statements, the upstream statements whose fine ranking similarity with the first upstream statement meets a second screening condition to obtain the second upstream statement.

[0009] In some embodiments, matching the upstream statements included in the matching region with the first upstream statement to obtain a rough ranking similarity between the upstream statements included in the matching region and the first upstream statement includes: segmenting the upstream statements included in the matching region to obtain a first vocabulary sequence; segmenting the first upstream statement to obtain a second vocabulary sequence; determining a Jaccard similarity between the upstream statements included in the matching region and the first upstream statement according to the first vocabulary sequence and the second vocabulary sequence; determining an edit distance between the upstream statements included in the matching region and the first upstream statement according to the upstream statements included in the matching region and the first upstream statement; and determining the rough ranking similarity between the upstream statements included in the matching region and the first upstream statement according to the Jaccard similarity and the edit distance.

[0010] In some embodiments, matching the candidate upstream statements with the first upstream statement to obtain a fine ranking similarity between the candidate upstream statements and the first upstream statement includes: encoding the candidate upstream statements and the first upstream statement to obtain a vector representation of the candidate upstream statements and a vector representation of the first upstream statement; and determining a cosine similarity between the vector representation of the candidate upstream statements and the vector representation of the first upstream statement according to the vector representation of the candidate upstream statements and the vector representation of the first upstream statement, where the fine ranking similarity between the candidate upstream statements and the first upstream statement is the cosine similarity.

[0011] In some embodiments, generating an associative downstream statement of the first upstream statement through a large language model according to the first upstream statement and the second upstream statement includes: generating a prompt according to the first upstream statement and the second upstream statement; and generating the associative downstream statement through the large language model according to the prompt.

[0012] In some embodiments, the obtaining of the first preceding sentence includes: receiving an input sentence at the current time; after receiving the input sentence, if a punctuation mark, a space, a line break instruction, or a send instruction is received again, determining the first preceding sentence as the input sentence.

[0013] A second aspect of the present application provides a sentence association device, including: an obtaining unit, configured to obtain a first preceding sentence, where the first preceding sentence includes a sentence input or sent on a terminal at the current time; a searching unit, configured to search, in historical input data locally stored on the terminal, for a preceding sentence whose sentence content is similar to the first preceding sentence to obtain a second preceding sentence, where the historical input data includes context sentences input on the terminal in the past; an associating unit, configured to generate an associated following sentence of the first preceding sentence through a large language model according to the first preceding sentence and the second preceding sentence; and an output unit, configured to output the associated following sentence.

[0014] A third aspect of the present application provides an electronic device, including a memory and a processor; the memory is connected to the processor and is configured to store a program; the processor is configured to implement the sentence association method as described in the first aspect or any embodiment of the first aspect by running the program in the memory.

[0015] A fourth aspect of the present application provides a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the sentence association method as described in the first aspect or any embodiment of the first aspect.

[0016] A fifth aspect of the present application provides a storage medium, on which a computer program is stored, and when the computer program is run by a processor, it implements the sentence association method as described in the first aspect or any embodiment of the first aspect.

[0017] A sentence association method, device, device, and product provided by the present application, for a sentence input or sent on a terminal at the current time, searches in historical input data locally stored on the terminal for a similar preceding sentence of the sentence, and generates an associated following sentence based on the sentence input or sent at the current time, the similar preceding sentence, and a large language model. Among them, the historical input data reflects the input habits of users on the terminal, and the large language model can improve the accuracy of sentence association. Through the above process, personalized associated following sentences are generated, improving the accuracy of the associated following sentences. Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.

[0019] Figure 1 It is a schematic diagram of the implementation environment related to the embodiments of the present application;

[0020] Figure 2 It is a flowchart of the sentence association method provided according to the embodiments of the present application Figure 1 ;

[0021] Figure 3 It is an example diagram of the storage strategy of historical input data provided according to the embodiments of the present application;

[0022] Figure 4 It is a flowchart of the sentence association method provided according to the embodiments of the present application Figure 2 ;

[0023] Figure 5 It is an example diagram of obtaining the first previous sentence provided according to the embodiments of the present application;

[0024] Figure 6 It is a structural schematic diagram of the sentence association device provided according to the embodiments of the present application;

[0025] Figure 7 It is a structural schematic diagram of the electronic device provided according to the embodiments of the present application. Specific implementation manners

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0027] In one way, the input method association relies on a static thesaurus and a rule-based association algorithm for word association. In this way, the historical input data of the user is not utilized, the degree of personalization is low, and it is also impossible to perform whole sentence association, and it is impossible to process long sentences and complex contexts.

[0028] In yet another way, an association model based on word frequency statistics is adopted for input method association. In this way, the frequency of words appearing in the input resources is calculated, and then a candidate word list is generated through statistical methods and according to set rules for word association. On the one hand, the input resources adopted are fixed and applied in batches to all users. When different users input the same content, the obtained association results are the same, and the differentiated association function cannot be realized. On the other hand, a large amount of input resources and a large-sized candidate list are required for processing long sentences and complex contexts, resulting in a large data volume and low efficiency. On the other hand, in the word frequency statistics-based association algorithm, the dictionary resources are relatively large and need to be deployed in a cloud environment to obtain better results. It is impossible to complete all association work locally and cannot meet the user's privacy protection requirements.

[0029] In yet another way, a sequence generation model based on deep learning is adopted for input method association. In this way, a recurrent neural network (RNN) and a long short-term memory (LSTM) are used to perform sentence association for the user's input content. On the one hand, the sequence generation model based on deep learning is pre-trained, and the characteristics of different users are not applied during the training process, resulting in the lack of personalization in the associated sentences. On the other hand, in actual applications, the generation effect of this method is not ideal. Especially for small-sized generation models, the generation effect in long sentence association is not ideal, and the computational complexity is relatively high.

[0030] It can be seen that the current input method association has deficiencies in aspects such as the whole sentence association effect, personalization degree, association efficiency, association accuracy, and privacy protection.

[0031] To solve the above problems, the embodiments of the present application provide a sentence association method, device, device, and product, which are applied to the field of natural language processing technology. In the embodiments of the present application, in the historical input data stored locally in the terminal, an upper sentence similar to the user's input or sent sentence is queried, and an associated lower sentence is generated based on the user's input or sent sentence, the similar upper sentence, and a large language model. The introduction of historical input data and the large language model realizes retrieval-augmented generation (RAG), enabling the large language model to generate high-quality personalized sentences, improving the whole sentence association effect, personalization degree, and association accuracy of the input method association; it does not require a large amount of user data, reduces the data processing volume, and improves the association efficiency; the overall method is completed locally on the terminal, protecting the user's privacy.

[0032] Exemplary implementation environment

[0033] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the implementation environment related to the embodiments of the present application. In the implementation environment involved in the present application, it may include: terminal 101.

[0034] On the terminal 101, historical input data is locally stored and a large language model is deployed. The user can input (which can be through text or voice input) or send a statement on the terminal 101. The terminal takes this statement as the previous statement and, based on the historical input data and the large language model, associates and outputs the next statement. For example, in Figure 1 , taking the conversation scenario as an example, the user inputs the previous statement "It will rain today" on the display screen of the terminal 101 and sends it to the conversation area. The terminal can output the associated next statement "Remember to bring an umbrella when going out" generated in the above manner.

[0035] It should be noted that the embodiments of the present application can be used in various input scenarios, such as search input scenarios, text editing scenarios, chat input scenarios, etc.

[0036] It should be noted that the terminal involved in the embodiments of the present application can be a personal digital assistant (PDA) device, a handheld device with wireless communication functions (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC)), a wearable device (such as a smart watch, a smart bracelet), and a smart home device (such as a smart speaker, a smart display device), etc.

[0037] Exemplary method

[0038] Please refer to Figure 2 , in an exemplary embodiment, a method for statement association is provided, which is applied to a terminal. The method includes the following steps:

[0039] S201, obtain a first previous statement, where the first previous statement includes the statement input or sent on the terminal at the current time.

[0040] In this embodiment, when it is detected that the user inputs a statement on the terminal, the statement input by the user is obtained, or when it is detected that the user sends a statement on the terminal, the statement sent by the user is obtained, so as to obtain the first previous statement.

[0041] For example, the terminal detects that the user types a statement in the input box and obtains the first preceding statement from the input box; or, the terminal detects a statement sending instruction in the chat application, obtains the first preceding statement from the statement sending instruction, or obtains the first preceding statement by obtaining the previous sent statement from the chat box; or, the terminal detects a statement input by the user in voice mode and obtains the first preceding statement from the collected user voice.

[0042] S202. In the historical input data stored locally in the terminal, search for a preceding statement whose statement content is similar to the first preceding statement to obtain a second preceding statement. The historical input data includes context statements input on the terminal in the past time.

[0043] Among them, in the historical input data, one context statement corresponds to one input record.

[0044] Among them, the context statement satisfies the context structure (that is, the input record in the historical input data satisfies the context structure), that is, the context statement includes multiple statements, and the relationship between the multiple statements is a context relationship. In other words, the context statement can be understood as a paragraph of text, which contains multiple statements. Among the multiple statements, there are preceding statements and following statements. Therefore, the context statement is expressed as a statement combination composed of the preceding statement and the following statement. The preceding statement and the following statement are relative concepts. In a paragraph of text, for the previous statement of a certain statement, this statement is the following statement; for the next statement of a certain statement, this statement is the preceding statement.

[0045] In this embodiment, the context statement in the historical input data is matched with the first preceding statement in terms of statement content to obtain the similarity between the preceding statement in the context statement and the first preceding statement. According to the similarity between the preceding statement in the context statement and the first preceding statement, search for a preceding statement whose statement content is similar to the first preceding statement in the historical input data to obtain a second preceding statement.

[0046] S203. According to the first preceding statement and the second preceding statement, generate an associative following statement of the first preceding statement through a large language model.

[0047] Among them, the large language model belongs to a generative large model. By pre-training the large language model in advance, the large language model is enabled to have the ability of statement association. After that, the large language model with the ability of statement association is deployed locally on the terminal to perform statement association locally on the terminal.

[0048] In this embodiment, the input data of the large language model can be determined according to the first preceding statement and the second preceding statement, and the input data is input into the large language model, and an associative following statement of the first preceding statement is generated in the large language model based on the input data.

[0049] S204, output the associated subsequent sentence.

[0050] In this embodiment, the associated subsequent sentence can be displayed in the input candidate area on the display screen of the terminal for the user to select whether to input or send the associated subsequent sentence, improving the user's input efficiency. Alternatively, a prompt message can be output in a voice manner, and the prompt message includes the associated subsequent sentence, and the prompt message is used to ask the user whether to input or send the associated subsequent sentence.

[0051] In the embodiment of the present application, the historical input data can reflect the user's expression habits, such as the subsequent sentences that the user habitually expresses after saying the previous text, the tone that the user likes in the expression, the words that the user likes to use in the expression, etc. By searching for similar previous text sentences in the historical input data, it provides a reference for the large language model to perform sentence association, enabling the large language model to generate personalized associated subsequent sentences that conform to the user's expression habits, improving the accuracy and quality of the associated subsequent sentences. The data used throughout the process is the input data of the user on the terminal in the past, and the amount of data processed is small, improving the efficiency of sentence association of the input method and thus improving the user input efficiency. In addition, the data used throughout the process is stored locally on the terminal, improving the privacy protection effect.

[0052] In some embodiments, the historical input data can be stored and / or updated according to a cache period. Thus, it avoids the historical input data occupying too much storage resources on the terminal and also controls the search range of similar previous text sentences within a reasonable range, improving the search efficiency.

[0053] Among them, the cache period can be a set fixed time period; or, the cache period can be a time period input by the user, which can be flexibly adjusted by the user according to the device situation of the terminal.

[0054] Among them, the cache period is, for example, one month, one week, etc.

[0055] In this embodiment, during the initial storage process of historical input data, the input records of the user on the terminal within the cache period can be obtained, and it is determined whether the input records satisfy the context structure. If the input records satisfy the context structure, the input records are stored as context statements in the historical input data. The context structure can refer to the description of the foregoing embodiment and will not be elaborated herein. During the update process of the historical input data, it includes the deletion process of old context statements and the storage process of new context statements. The deletion process of old context statements may include: obtaining the input time of the context statements in the historical input data, and determining whether the time interval between the input time of the context statements and the current time is greater than the cache period; if the time interval between the input time of the context statements and the current time is greater than the cache period, the context statements are deleted from the historical input data. The storage process of new context statements may include: when it is detected that the user inputs, obtaining the current input record of the user, and determining whether the input record satisfies the context structure; if the input record satisfies the context structure, the input record is stored as a context statement in the historical input statements.

[0056] In some embodiments, the historical input data is distributed in a matching area and a recording area, that is, the historical input data includes context statements belonging to the matching area and context statements belonging to the recording area. Among them, the number of input times of the context statements in the matching area is greater than the number of input times of the context statements in the recording area, and the number of context statements in the matching area is less than the number of context statements in the recording area. During the process of searching for the second previous context statement, a previous context statement whose statement content is similar to the first previous context statement can be searched in the matching area. Thus, by dividing the context statements in the historical input data into a matching area and a recording area and searching for similar previous context statements in the matching area, the search range of similar previous context statements is narrowed, and the search efficiency is improved; compared with the input records with fewer input times, the input records that satisfy the context structure and have more input times can reflect the input habits of the user. Storing such input records in the matching area for searching for similar context statements can improve the accuracy of generating personalized associated subsequent context statements.

[0057] In one example, during the initial storage process of the historical input data, the input records of the user on the terminal within the cache period can be obtained, and it is determined whether the input records satisfy the context structure. If the input records satisfy the context structure, the number of input times of the input records is counted, and according to the number of input times of the input records, the input records are stored as context statements in the matching area or the recording area. If the number of input times of the input record is greater than the number threshold, the input record is stored as a context statement in the matching area; if the number of input times of the input record is less than or equal to the number threshold, the input record is stored as a context statement in the recording area.

[0058] In another example, in the historical input data, the process of storing a new context statement may include: upon detecting a user input, obtaining the user's current input record, and determining whether the input record satisfies the context structure; if the input record satisfies the context result, checking whether the input record exists in the record area and the matching area; if the input record does not exist in both the record area and the matching area, storing the input record as a context statement in the record area and incrementing the input count of the input record by one; if the input record exists in the record area but does not exist in the matching area, incrementing the input count of the input record by one, and determining whether the input count of the input record is greater than the count threshold. If the input count of the input record is greater than the count threshold, updating the input record as a context statement in the matching area; otherwise, the input record remains as a context statement in the record area; if the input record does not exist in the record area but exists in the matching area, incrementing the input count of the input record by one, and the input record remains as a context statement in the matching area.

[0059] Optionally, the count threshold is 1.

[0060] In this optional method, after obtaining an input record that conforms to the context structure from the user input, check whether the input record exists in the record area; if not, a new input record can be added to the record area, and the input count of the input record is incremented by one; if the input record exists in the record area and the input count of the input record is 1, move the input record from the record area to the matching area, that is, update the input record as a context statement in the matching area, and increment the input count of the input record by one; if the input record exists in the matching area, increment the input count of the input record by one, and the input record remains as a context statement in the matching area.

[0061] Optionally, by adding a storage mark corresponding to the record area to the input record, the input record is stored as a context statement in the record area; by adding a storage mark corresponding to the matching area to the input record, the input record is stored as a context statement in the matching area.

[0062] Please refer to Figure 3 , Figure 3 , which is an example diagram of the storage strategy for historical input data provided according to an embodiment of the present application. The storage strategy includes a matching area and a record area. For an input record that conforms to the context structure input by the user, it can be newly added as a context statement in the record area during the first storage. As the input count of the input record increases, the input record will be updated from the record area to the matching area; for the above-mentioned statement input by the user, similar statements can be searched in the matching area.

[0063] Please refer toFigure 4 In yet another exemplary embodiment, a statement association method is provided. The method includes the following steps:

[0064] S401. Obtain a first previous statement, where the first previous statement includes a statement input or sent on the terminal at the current time.

[0065] Among them, the implementation principle and technical effect of S401 can be referred to the foregoing embodiments and will not be elaborated here.

[0066] S402. In the previous statements included in the matching area, search for a previous statement whose statement content is similar to the first previous statement to obtain a second previous statement.

[0067] Among them, the matching area can be referred to the description of the foregoing embodiments and will not be elaborated here.

[0068] In this embodiment, by matching the previous statements included in the matching area with the first previous statement, the similarity between the previous statements included in the matching area and the first previous statement can be obtained. According to the similarity between the previous statements included in the matching area and the first previous statement, in the previous statements included in the matching area, search for a previous statement whose statement content is similar to the first previous statement to obtain a second previous statement.

[0069] In one example, as Figure 4 shown, S402 includes: S4021. Match the previous statements included in the matching area with the first previous statement to obtain the rough sorting similarity between the previous statements included in the matching area and the first previous statement; S4022. In the previous statements included in the matching area, select the previous statements whose rough sorting similarity with the first previous statement meets the first screening condition to obtain candidate previous statements; S4023. Match the candidate previous statements with the first previous statement to obtain the fine sorting similarity between the candidate previous statements and the first previous statement; S4024. In the candidate previous statements, select the previous statements whose fine sorting similarity with the first previous statement meets the second screening condition to obtain a second previous statement. Thus, through the combination of rough sorting and fine sorting, the efficiency and accuracy of searching for similar previous statements are improved.

[0070] In S4021, through a rough sorting algorithm, match the previous statements included in the matching area with the first previous statement to obtain the rough sorting similarity between the previous statements included in the matching area and the first previous statement.

[0071] Among them, the rough ranking algorithm is, for example, the keyword matching algorithm, which determines the rough ranking similarity between the above-mentioned sentence included in the matching region and the first above-mentioned sentence through the keyword matching algorithm; the rough ranking algorithm is also, for example, the term frequency-inverse document frequency (TF-IDF) algorithm, which determines the rough ranking similarity between the above-mentioned sentence included in the matching region and the first above-mentioned sentence through the TF-IDF algorithm.

[0072] In a possible implementation manner, S4021 may include: performing word segmentation on the above-mentioned sentence included in the matching region to obtain a first vocabulary sequence; performing word segmentation on the first above-mentioned sentence to obtain a second vocabulary sequence; determining the Jaccard similarity between the above-mentioned sentence included in the matching region and the first above-mentioned sentence according to the first vocabulary sequence and the second vocabulary sequence; determining the edit distance between the above-mentioned sentence included in the matching region and the first above-mentioned sentence according to the above-mentioned sentence included in the matching region and the first above-mentioned sentence; determining the rough ranking similarity between the above-mentioned sentence included in the matching region and the first above-mentioned sentence according to the Jaccard similarity and the edit distance. Among them, the Jaccard similarity is a similarity calculation method based on sets, and the edit distance is the minimum number of edits required to convert one string into another string. The Jaccard similarity and the edit distance have their respective advantages. By combining the Jaccard similarity and the edit distance, the advantages of both are combined with each other, effectively improving the efficiency and accuracy of the rough ranking similarity calculation process.

[0073] In this implementation manner, the above-mentioned sentence included in the matching region is divided into multiple vocabularies (vocabularies composed of single characters or multiple characters), and the first vocabulary sequence is composed of the multiple vocabularies. The first above-mentioned sentence is divided into multiple vocabularies, and the second vocabulary sequence is composed of the multiple vocabularies. Determine the intersection of the first vocabulary sequence and the second vocabulary sequence, and determine the union of the first vocabulary sequence and the second vocabulary sequence; calculate the Jaccard similarity between the above-mentioned sentence included in the matching region and the first above-mentioned sentence according to the intersection and the union. Regarding the above-mentioned sentence included in the matching region and the first above-mentioned sentence as different strings, calculate the edit distance between the two strings through the edit distance calculation formula, that is, the edit distance between the above-mentioned sentence included in the matching region and the first above-mentioned sentence. Finally, combine the Jaccard similarity and the edit distance to determine the rough ranking similarity between the above-mentioned sentence included in the matching region and the first above-mentioned sentence.

[0074] In this implementation manner, the calculation formula of the Jaccard similarity is expressed as:

[0075]

[0076] Among them, A represents the first vocabulary sequence, and B represents the second vocabulary sequence. |A∩B| represents the size of the intersection of the first vocabulary sequence and the second vocabulary sequence, and |A∪B| represents the size of the intersection of the first vocabulary sequence and the second vocabulary sequence. J(A, B) represents the Jaccard similarity between A and B.

[0077] In this implementation, it can be assumed that the above - text statements included in the matching region are string S1, and the first above - text statement is string S2. The string length of S1 is n, and the string length of S2 is m. The calculation process of the edit distance between the above - text statements included in the matching region and the first above - text statement can be expressed as: First, initialize the distance matrix corresponding to the edit distance. The number of rows of this distance matrix is n + 1, and the number of columns is m + 1; then, fill the distance matrix corresponding to the edit distance according to the edit operations required for the conversion between the above - text statements included in the matching region and the first above - text statement, that is, the edit distance is obtained.

[0078] Optionally, the initialization process of the distance matrix is expressed as:

[0079] D[i][0] = i

[0080] D[0][j] = j

[0081] Among them, 0 ≤ i ≤ n, 0 ≤ j ≤ m, and D represents the distance matrix.

[0082] Optionally, the filling process of the edit matrix is expressed as:

[0083]

[0084] Among them, the meaning of the above formula is as follows: When the (i - 1) - th character A[i - 1] in the above - text statement included in the matching region is equal to the (j - 1) - th character B[j - 1] in the first above - text statement, it can be determined that D[i][j] = D[i - 1][j - 1]; when A[i - 1] is not equal to B[j - 1], if a deletion operation is performed, it can be determined that D[i][j] = D[i - 1][j]+1, if an insertion operation is performed, it can be determined that D[i][j] = D[i][j - 1]+1, and if a replacement operation is performed, it can be determined that D[i][j] = D[i - 1][j - 1]+1. Finally, D[n][m] is obtained, that is, the edit distance between the above - text statement included in the matching region and the first above - text statement.

[0085] Optionally, in the process of determining the Jaccard similarity, after segmenting the above - text statement included in the matching region and the first above - text statement respectively, pre - processing can be performed on the first vocabulary sequence and the second vocabulary sequence to improve the accuracy and efficiency of calculating the Jaccard similarity.

[0086] In this optional approach, the preprocessing of the first vocabulary sequence and the second vocabulary sequence includes: removing stop words and unifying the case of letters. By removing stop words, some common words that have little impact on similarity calculation, such as "de", "shi", "zai", are deleted, reducing the length of the vocabulary sequence and improving the accuracy and efficiency of calculating the Jaccard similarity. By unifying the case of letters, the adverse impact of different letter cases on Jaccard similarity calculation is avoided.

[0087] In S4022, according to the rough similarity between the above-mentioned sentences included in the matching region and the first above-mentioned sentence in accordance with the first screening condition, candidate above-mentioned sentences are selected from the above-mentioned sentences included in the matching region.

[0088] In a possible implementation, the first screening condition includes a first similarity threshold and a second similarity threshold, and S4022 includes: selecting, from the above-mentioned sentences included in the matching region, the above-mentioned sentences whose Jaccard similarity with the first above-mentioned sentence is greater than or equal to the first similarity threshold and whose edit distance is greater than or equal to the second similarity threshold to obtain candidate above-mentioned sentences. By setting two thresholds respectively for comparing the Jaccard similarity and the edit distance between the above-mentioned sentences included in the matching region and the first above-mentioned sentence, the screening accuracy of candidate above-mentioned sentences is improved.

[0089] In another possible implementation, the first screening condition may be to select the top M above-mentioned sentences arranged in descending order of rough similarity, and S4022 includes: sorting the above-mentioned sentences included in the matching region in descending order of the rough similarity (including Jaccard similarity and edit distance) between the above-mentioned sentences included in the matching region and the first above-mentioned sentence, and then selecting the top M above-mentioned sentences therefrom to obtain M candidate above-mentioned sentences, where M is greater than 1.

[0090] In S4023, through a fine-ranking algorithm, the candidate above-mentioned sentences are matched with the first above-mentioned sentence to obtain the rough similarity between the candidate above-mentioned sentences and the first above-mentioned sentence.

[0091] Among them, the fine-ranking algorithm is, for example, a sentence matching algorithm based on deep learning or reinforcement learning.

[0092] In a possible implementation, S4023 includes: encoding the candidate previous sentence and the first previous sentence to obtain the vector representations of the candidate previous sentence and the first previous sentence; determining the cosine similarity between the vector representation of the candidate previous sentence and the vector representation of the first previous sentence according to the vector representations of the candidate previous sentence and the first previous sentence, and the fine-ranking similarity between the candidate previous sentence and the first previous sentence is the cosine similarity. The cosine similarity can reflect the semantic similarity between the candidate previous sentence and the first previous sentence. By means of word vector encoding and calculating the cosine similarity, the accuracy of the fine-ranking similarity between the candidate previous sentence and the first previous sentence can be improved.

[0093] In this implementation, the candidate previous sentence and the first previous sentence can be encoded through a pre-trained word vector model to obtain the vector representations of the candidate previous sentence and the first previous sentence. Then, the cosine similarity between the vector representation of the candidate previous sentence and the vector representation of the first previous sentence is calculated through the cosine similarity calculation formula.

[0094] For example, the word vector model can adopt Word2Vec, global vectors for word representation (GloVe), bidirectional encoder representations from transformers (BERT).

[0095] Taking the word vector model as Word2Vec as an example, encoding the candidate previous sentence and the first previous sentence, the obtained corresponding vector representations can be expressed as:

[0096] V s1 =Word2Ver(s1)

[0097] V s2 =Word2Ver(s2)

[0098] Among them, s1 represents the candidate previous sentence, s2 represents the first previous sentence, V s1 represents the vector representation of the candidate previous sentence, and V s2 represents the vector representation of the first previous sentence. Word2Ver() represents the word vector model Word2Vec.

[0099] The cosine similarity calculation formula can be expressed as:

[0100]

[0101] Among them, Cosim(V s1 ,Vs2 ) represents the cosine similarity between the vector representation of the candidate previous context sentence and the vector representation of the first previous context sentence, and k represents the vector lengths of the vector representations of the candidate previous context sentence and the first previous context sentence. V s1 [i] represents the i-th element in the vector representation of the candidate previous context sentence, V s2 [i] represents the i-th element in the vector representation of the first previous context sentence.

[0102] In S4024, according to the second screening condition and the fine-rank similarity between the candidate previous context sentence and the first previous context sentence, select the second previous context sentence from the candidate previous context sentences.

[0103] In one possible implementation, the second screening condition includes a third similarity threshold, and S4024 includes: selecting, from the candidate previous context sentences, the previous context sentences whose fine-rank similarity with the first previous context sentence is greater than or equal to the third similarity threshold to obtain the second previous context sentence.

[0104] In another possible implementation, the second screening condition may be the candidate previous context sentences ranked in the top N positions according to the fine-rank similarity from large to small. S4024 includes: selecting, from the candidate previous context sentences in the order of the fine-rank similarity from large to small with the first previous context sentence, the candidate previous context sentences ranked in the top N positions to obtain the second previous context sentence, where N is greater than or equal to 1. When N is equal to 1, the second screening condition can be expressed as selecting the candidate previous context sentence with the highest fine-rank similarity with the first previous context sentence as the second previous context sentence.

[0105] S403, according to the first previous context sentence and the second previous context sentence, generate an associated next context sentence of the first previous context sentence through a large language model.

[0106] S404, output the associated next context sentence.

[0107] Among them, the implementation principles and technical effects of S403 to S404 can be referred to the foregoing embodiments and will not be elaborated here.

[0108] In the embodiments of the present application, when realizing the association of the next context sentence through the historical input data stored locally in the terminal and the large language model deployed locally, improving the personalization degree, quality and accuracy of the associated next context sentence, improving the association efficiency and protecting the user privacy, by dividing the historical input data into a matching area and a recording area, the search efficiency of similar previous context sentences is improved, and through the combination of rough ranking and fine ranking, the search efficiency and search accuracy of similar previous context sentences are improved, thereby further improving the sentence association efficiency and accuracy.

[0109] Next, based on any of the foregoing embodiments, some extended embodiments are provided.

[0110] In some embodiments, according to the first preceding statement and the second preceding statement, an associative following statement of the first preceding statement is generated through a large language model, including: generating a prompt according to the first preceding statement and the second preceding statement; generating an associative following statement through the large language model according to the prompt. For the large language model, the prompt is equivalent to a task instruction, and the large language model will complete the task of generating the associative following statement of the first following statement according to the prompt.

[0111] In this embodiment, the following statement of the second preceding statement can be obtained from the historical input data, the second preceding statement and the following statement of the second preceding statement are combined into the reference information in the prompt, and the first preceding statement is used as the task premise in the prompt. The finally formed prompt can be expressed as: Please refer to the second preceding statement and the following statement of the second preceding statement to generate the following statement of the first preceding statement. Input this prompt to the large language model, and the large language model will perform the following statement association according to this prompt. Thus, an accurate task instruction is provided to the large language model through the prompt formed by the first preceding statement and the second preceding statement, improving the accuracy of the following statement association.

[0112] In some embodiments, obtaining the first preceding statement includes: receiving an input statement at the current time; after receiving the input statement, if a punctuation mark, space, line break instruction or send instruction is received again, it is determined that the first preceding statement is the input statement.

[0113] In this embodiment, after receiving the input statement, that is, when there is already input, if the user inputs a punctuation mark, space, line break instruction or send instruction again, it can be determined that the user has completed the input of the preceding statement, and this input statement is the first preceding statement. Thus, by identifying the content input again by the user, it is judged whether the previous input of the user is a complete preceding statement, improving the integrity and accuracy of the obtained first preceding statement.

[0114] Please refer to Figure 5 , Figure 5 FIG. is an example diagram for obtaining the first preceding statement according to an embodiment of the present application. When there is already input, if the user sends this existing input, this existing input is recorded as the first preceding statement. If the user inputs a punctuation mark, space or line break instruction again, this existing input is also recorded as the first preceding statement. If the content input again is other than punctuation marks, spaces and line break instructions, there is no need to record this existing input as the first preceding statement.

[0115] In some embodiments, the volume of the large language model can be reduced by cropping and / or quantization, and then the large language model with reduced volume can be deployed to the terminal, so that the large language model can be flexibly deployed to the terminal and the terminal has sufficient capabilities to support the operation of the large language model.

[0116] It should be noted that in the embodiments of the present application, the specific structure of the large language model is not limited.

[0117] Exemplary device

[0118] Correspondingly, the embodiments of the present application also provide a statement association device.

[0119] Please refer to Figure 6 , in an exemplary embodiment, a statement association device 600 is provided. The statement association device 600 includes: an acquisition unit 601, a search unit 602, an association unit 603, and an output unit 604, where:

[0120] The acquisition unit 601 is configured to acquire a first previous statement, where the first previous statement includes a statement input or sent on the terminal at the current time; the search unit 602 is configured to search for a previous statement whose statement content is similar to the first previous statement in the historical input data stored locally on the terminal to obtain a second previous statement, and the historical input data includes context statements input on the terminal in the past; the association unit 603 is configured to generate an associated subsequent statement of the first previous statement through a large language model according to the first previous statement and the second previous statement; the output unit 604 is configured to output the associated subsequent statement.

[0121] In some embodiments, the historical input data is distributed in a matching area and a recording area, and the input frequency of the context statements in the matching area is greater than that of the context statements in the recording area; specifically, the search unit 602 is configured to: search for a previous statement whose statement content is similar to the first previous statement among the previous statements included in the matching area to obtain a second previous statement.

[0122] In some embodiments, the search unit 602 is specifically configured to: match the previous statements included in the matching area with the first previous statement to obtain a rough sorting similarity between the previous statements included in the matching area and the first previous statement; among the previous statements included in the matching area, select the previous statements whose rough sorting similarity with the first previous statement meets the first screening condition to obtain candidate previous statements; match the candidate previous statements with the first previous statement to obtain a fine sorting similarity between the candidate previous statements and the first previous statement; among the candidate previous statements, select the previous statements whose fine sorting similarity with the first previous statement meets the second screening condition to obtain a second previous statement.

[0123] In some embodiments, the search unit 602 is specifically configured to: segment the previous statement included in the matching region to obtain a first vocabulary sequence; segment the first previous statement to obtain a second vocabulary sequence; determine the Jaccard similarity between the previous statement included in the matching region and the first previous statement according to the first vocabulary sequence and the second vocabulary sequence; determine the edit distance between the previous statement included in the matching region and the first previous statement according to the previous statement included in the matching region and the first previous statement; determine the rough ranking similarity between the previous statement included in the matching region and the first previous statement according to the Jaccard similarity and the edit distance.

[0124] In some embodiments, the search unit 602 is specifically configured to: encode the candidate previous statement and the first previous statement to obtain a vector representation of the candidate previous statement and a vector representation of the first previous statement; determine the cosine similarity between the vector representation of the candidate previous statement and the vector representation of the first previous statement according to the vector representation of the candidate previous statement and the vector representation of the first previous statement, and the fine ranking similarity between the candidate previous statement and the first previous statement is the cosine similarity.

[0125] In some embodiments, the association unit 603 is specifically configured to: generate a prompt word according to the first previous statement and the second previous statement; generate an associated subsequent statement through a large language model according to the prompt word.

[0126] In some embodiments, the acquisition unit 601 is specifically configured to: receive an input statement at the current time; after receiving the input statement, if a punctuation mark, a space, a line break instruction, or a send instruction is received again, determine the first previous statement as the input statement.

[0127] The statement association device provided in this embodiment belongs to the same inventive concept as the statement association method provided in the foregoing embodiments of the present application, can execute the statement association method provided in any of the foregoing embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the statement association method provided in the foregoing embodiments of the present application, which will not be elaborated here.

[0128] The functions implemented by each unit in the device (statement association device) provided in the foregoing embodiments may be implemented by the same or different processors respectively, which is not limited in the embodiments of the present application.

[0129] It should be understood that each unit in the device provided in the above embodiments can be implemented in the form of a processor invoking software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor invokes the instructions stored in the memory to implement any of the above methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits. By designing the hardware circuits, the functions of some or all of the units can be implemented. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and by designing the logical relationships of the components in the circuit, the functions of some or all of the above units are implemented. Again, for example, in another implementation, the hardware circuit can be implemented by a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationships between the logic gate circuits are configured through a configuration file, so as to implement the functions of some or all of the above units. All units of the above device can be all implemented in the form of a processor invoking software, or all implemented in the form of hardware circuits, or part implemented in the form of a processor invoking software, and the remaining part implemented in the form of hardware circuits.

[0130] In the embodiments of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor can be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a GPU, or a DSP, etc. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the hardware circuits are fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a kind of ASIC, such as an NPU, a TPU, a DPU, etc.

[0131] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.

[0132] In addition, each unit in the above device can be integrated in whole or in part, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of an SOC. The SOC may include at least one processor for implementing any of the above methods or the functions of each unit of the device. The types of the at least one processor can be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.

[0133] Exemplary electronic device

[0134] Another embodiment of the present application also provides an electronic device. Refer to Figure 7 As shown, the electronic device includes: a memory 700 and a processor 710; wherein, the memory 700 is connected to the processor 710 for storing programs; the processor 710 is configured to implement the sentence association method disclosed in any of the above embodiments by running the programs stored in the memory 700.

[0135] Specifically, the above electronic device may further include: a bus, a communication interface 720, an input device 730, and an output device 740.

[0136] The processor 710, the memory 700, the communication interface 720, the input device 730, and the output device 740 are interconnected via the bus. Among them:

[0137] The bus may include a path for transmitting information between various components of the computer system.

[0138] The processor 710 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or may be an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0139] The processor 710 may include a main processor, and may also include a baseband chip, a modem, etc.

[0140] The program for implementing the technical solution of this application is stored in the memory 700, and the operating system and other key services may also be stored. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 700 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.

[0141] The input device 730 may include devices for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.

[0142] The output device 740 may include devices for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0143] The communication interface 720 may include devices of any transceiver type for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0144] The processor 710 executes the program stored in the memory 700 and calls other devices, and can be used to implement each step of any one of the statement association methods provided in the above embodiments of this application.

[0145] An embodiment of this application also provides a chip, which includes a processor and a data interface. The processor reads and runs the program stored on the memory through the data interface to execute the statement association method introduced in any of the above embodiments. The specific processing process and its beneficial effects can be referred to the embodiments of the statement association method described above.

[0146] Exemplary computer program product and storage medium

[0147] In addition to the above methods and devices, an embodiment of this application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the statement association method according to various embodiments of this application described in any of the above embodiments of this specification.

[0148] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0149] In addition, an embodiment of the present application can also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to perform the steps in the sentence association method according to various embodiments of the present application described in any of the above embodiments of this specification.

[0150] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0151] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0152] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs, and the technical features recorded in the embodiments can be replaced or combined.

[0153] The units in the devices in the embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0154] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0155] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0156] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware, or in the form of software functional modules or sub-modules.

[0157] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0158] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0159] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0160] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A sentence association method, characterized in that, Including: Obtain a first previous sentence, where the first previous sentence includes a sentence input or sent on the terminal at the current time; In the historical input data stored locally on the terminal, search for a previous sentence whose sentence content is similar to the first previous sentence to obtain a second previous sentence, where the historical input data includes context sentences input on the terminal in the past; According to the first previous sentence and the second previous sentence, generate an associated next sentence of the first previous sentence through a large language model; Output the associated next sentence.

2. The statement association method according to claim 1, wherein The historical input data is distributed in a matching area and a recording area, and the input frequency of the context sentences in the matching area is greater than that of the context sentences in the recording area; The step of searching for a previous sentence whose sentence content is similar to the first previous sentence in the historical input data stored locally on the terminal to obtain a second previous sentence includes: Search for a previous sentence whose sentence content is similar to the first previous sentence among the previous sentences included in the matching area to obtain the second previous sentence.

3. The statement association method according to claim 2, wherein The step of searching for a previous sentence whose sentence content is similar to the first previous sentence among the previous sentences included in the matching area to obtain the second previous sentence includes: Match the previous sentences included in the matching area with the first previous sentence to obtain the rough ranking similarity between the previous sentences included in the matching area and the first previous sentence; Among the previous sentences included in the matching area, select the previous sentences whose rough ranking similarity with the first previous sentence meets the first screening condition to obtain candidate previous sentences; Match the candidate previous sentences with the first previous sentence to obtain the fine ranking similarity between the candidate previous sentences and the first previous sentence; Among the candidate previous sentences, select the previous sentences whose fine ranking similarity with the first previous sentence meets the second screening condition to obtain the second previous sentence.

4. The statement association method according to claim 3, characterized in that The step of matching the previous sentences included in the matching area with the first previous sentence to obtain the rough ranking similarity between the previous sentences included in the matching area and the first previous sentence includes: Segment the previous sentences included in the matching area to obtain a first vocabulary sequence; Segment the first previous sentence to obtain a second vocabulary sequence; According to the first vocabulary sequence and the second vocabulary sequence, determine the Jaccard similarity between the previous sentences included in the matching area and the first previous sentence; According to the previous sentences included in the matching area and the first previous sentence, determine the edit distance between the previous sentences included in the matching area and the first previous sentence; According to the Jaccard similarity and the edit distance, determine the rough ranking similarity between the previous sentences included in the matching area and the first previous sentence.

5. The statement association method according to claim 3, wherein The step of matching the candidate previous sentences with the first previous sentence to obtain the fine ranking similarity between the candidate previous sentences and the first previous sentence includes: Encode the candidate previous sentences and the first previous sentence to obtain the vector representation of the candidate previous sentences and the vector representation of the first previous sentence; Based on the vector representation of the candidate previous sentence and the vector representation of the first previous sentence, determine the cosine similarity between the vector representation of the candidate previous sentence and the vector representation of the first previous sentence, and the fine-rank similarity between the candidate previous sentence and the first previous sentence is the cosine similarity.

6. The statement association method according to any one of claims 1 to 5, characterized in that The generating, by a large language model, an associated next sentence of the first previous sentence according to the first previous sentence and the second previous sentence includes: Generating a prompt word according to the first previous sentence and the second previous sentence; Generating the associated next sentence through the large language model according to the prompt word.

7. The method for sentence association according to any one of claims 1 to 5, characterized in that The obtaining of the first previous sentence includes: Receiving an input sentence at the current time; After receiving the input sentence, if a punctuation mark, a space, a line break instruction or a send instruction is received again, determining the first previous sentence as the input sentence.

8. A sentence association device, characterized in that, Including: An obtaining unit, configured to obtain a first previous sentence, where the first previous sentence includes a sentence input or sent on a terminal at the current time; A searching unit, configured to search, in historical input data locally stored on the terminal, for a previous sentence whose sentence content is similar to the first previous sentence to obtain a second previous sentence, where the historical input data includes context sentences input on the terminal in the past; An associating unit, configured to generate an associated next sentence of the first previous sentence through a large language model according to the first previous sentence and the second previous sentence; An output unit, configured to output the associated next sentence.

9. An electronic device, characterized in that, Including a memory and a processor; The memory is connected to the processor and is used for storing programs; The processor is configured to implement the sentence association method according to any one of claims 1 to 7 by running the programs in the memory.

10. A computer program product, characterized in that, Including a computer program, where when the computer program is executed by a processor, it implements the sentence association method according to any one of claims 1 to 7.