Information updating method, task execution method, equipment, medium and program product
By updating dynamic weights and association sequences in large language models, combining dynamic information sets and reference information sets, and optimizing cache strategy, the cache efficiency problems caused by semantic drift and cold start are solved, and the model inference speed and user experience are improved.
Patent Information
- Application Number
- CN202510725289.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The prior art cannot effectively deal with semantic drift and cold start delay in large language models, resulting in low cache utilization and low cache hit rate.
By determining the similarity between the input sequence in the input information and the dynamic information set, update the dynamic weight and association sequence, the target sequence in the dynamic information set is combined with the reference information set, the candidate information set is updated in real time, and the cache strategy is optimized.
Improves cache hit rate in cold start and semantic drift cases, improving model inference speed and user experience.
Smart Images

Figure CN120234401A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of information storage technology and model inference technology, and particularly relates to an information update method, a task execution method, a device, a medium, and a program product. Background Art
[0002] Large language models usually have a huge number of parameters and intermediate results, resulting in limited system storage resources. The key-value caching technology can reduce redundant calculations by storing intermediate keys and value vectors, which can improve the inference efficiency of the model, reduce the occupation of system resources, and support efficient decoding during the inference process. In related technologies, key-value caches are mainly managed through hardware co-optimization, cache compression, and adjusting cache priorities based on word frequencies. However, in actual use, related technologies cannot be applied to situations such as semantic drift and cold start latency during the model inference process, resulting in problems such as low cache utilization and low cache hit rate. Summary of the Invention
[0003] In view of the above problems, the present invention provides an information update method, a task execution method, a device, a device, a medium, and a program product.
[0004] According to the first aspect of the present invention, an information update method is provided, including: in response to receiving input information, determining the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set, where the dynamic sequence has an initial dynamic weight, and the dynamic information set includes a plurality of information pairs composed of the initial dynamic weight and the dynamic sequence; updating the dynamic information set based on the updated dynamic weight obtained by using the similarity to update the initial dynamic weight and the associated sequence to obtain an updated dynamic information set, where the associated sequence is the sequence in the dynamic sequence that is semantically associated with the input sequence; updating the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, so as to process the target task by using the candidate information set, where the initial candidate information set is the candidate information set before receiving the input information.
[0005] The second aspect of the present invention provides a task execution method, including: in response to receiving a task execution instruction, processing the target task by using the candidate information set; where the target task includes any one of text generation, speech recognition, and financial analysis, and the candidate set is obtained according to the above information update method.
[0006] A third aspect of the present invention provides an information update device, including: a similarity determination module, configured to determine the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set in response to receiving the input information, wherein the dynamic sequence has an initial dynamic weight, and the dynamic information set includes a plurality of information pairs composed of the initial dynamic weight and the dynamic sequence; a dynamic information set update module, configured to update the dynamic information set based on the updated dynamic weight obtained by updating the initial dynamic weight using the similarity and the associated sequence to obtain an updated dynamic information set, wherein the associated sequence is obtained from a sequence semantically associated with the dynamic sequence; a candidate information set update module, configured to update the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set for processing the target task using the candidate information set, wherein the initial candidate information set is the candidate information set before receiving the input information.
[0007] A fourth aspect of the present invention provides a task execution device, including: a task processing module, configured to process the target task using the candidate information set in response to receiving a task execution instruction; wherein the target task includes any one of text generation, speech recognition, and financial analysis, and the candidate information set is obtained according to the above task execution method.
[0008] A fifth aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0009] A sixth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and the above computer program or instruction implements the steps of the above method when executed by a processor.
[0010] A seventh aspect of the present invention further provides a computer program product, including a computer program or instruction, and the above computer program or instruction implements the steps of the above method when executed by a processor. Description of the Drawings
[0011] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer.
[0012] Figure 1 The application scenario diagram of the information update method, task execution method, device, medium, and program product according to the embodiments of the present invention is shown.
[0013] Figure 2A The flowchart of the information update method according to the embodiments of the present invention is shown.
[0014] Figure 2B The flowchart of the sequence weight update method according to an embodiment of the present invention is shown.
[0015] Figure 2C The flowchart of the initial candidate information set update method according to an embodiment of the present invention is shown.
[0016] Figure 3 The flowchart of the task execution method according to an embodiment of the present invention is shown.
[0017] Figure 4 The structural block diagram of the information update device according to an embodiment of the present invention is shown.
[0018] Figure 5 The structural block diagram of the task execution device according to an embodiment of the present invention is shown.
[0019] Figure 6 The block diagram of the electronic device according to an embodiment of the present invention is shown.
[0020] Figure 7 The block diagram of the electronic device suitable for implementing the information update method according to an embodiment of the present invention is shown. Detailed implementation manners
[0021] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0022] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0024] In cases where expressions such as "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning that those skilled in the art usually understand for such expressions (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0025] In some examples, different hardware architectures are utilized to work together. For example, by offloading cyclic redundancy check and data compression through a Data Processing Unit (DPU), the video memory occupancy of a Graphics Processing Unit (GPU) is reduced.
[0026] However, the DPU and the GPU are different hardware architectures, and their cooperation requires good adaptation and compatibility. In practical applications, problems such as high communication latency and limited data transmission bandwidth between the DPU and the GPU may occur, affecting the optimization effect of the overall model inference performance.
[0027] In some examples, based on word frequencies such as term frequency-inverse document frequency (TF-IDF), the priorities are adjusted, and high-frequency hot words are preferentially stored in the cache. Placing high-frequency hot words (words with high TF-IDF values) in the cache speeds up the access to these words to a certain extent.
[0028] However, the method based only on word frequencies such as TF-IDF relies on historical data to determine the priorities of words. In the cold start phase of model inference, the system does not have enough historical data to calculate the word frequencies and inverse document frequencies, so it is unable to accurately identify high-frequency hot words.
[0029] In some examples, according to the analysis results of the attention module, a selective storage strategy for large language models is improved through lightweight model analysis and adaptive key-value caching. This reduces the occupied space of the cache to a certain extent. However, as the system runs for a longer time, the meaning of the data may change, and semantic drift occurs due to the change in the theme of the input information. If the data in the cache cannot be updated in a timely manner to reflect this semantic change, then the cached data may become irrelevant or inaccurate and cannot meet the current needs of users.
[0030] Based on the above problems, the present invention provides an information update method, including: in response to receiving input information, determining the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set, where the dynamic sequence has an initial dynamic weight, and the dynamic information set includes multiple information pairs composed of the initial dynamic weight and the dynamic sequence; updating the dynamic information set based on the updated dynamic weight obtained by using the similarity to update the initial dynamic weight and the associated sequence to obtain an updated dynamic information set, where the associated sequence is the sequence in the dynamic sequence that is semantically associated with the input sequence; updating the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, so as to process the target task by using the candidate information set, where the initial candidate information set is the candidate information set before receiving the input information.
[0031] According to an embodiment of the present invention, since the updated dynamic weight is flexibly determined based on the similarity between the input feature vector and the dynamic feature vector, the updated dynamic information set obtained based on the updated dynamic weight takes into account the dual influence of real-time input information and the associated sequence related thereto. Therefore, the initial candidate information set is updated in real time by using the target sequences of the dynamic updated dynamic information set and the static reference information set respectively, improving the hit rate of the candidate information set in the case of cold start and semantic drift, accelerating the inference speed of the model, and further improving the user experience.
[0032] Figure 1 The application scenario diagram of the information update method, task execution method, device, medium and program product according to an embodiment of the present invention is shown.
[0033] As Figure 1 shown, the application scenario according to this embodiment may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0034] The user can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. The terminal device 101 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0035] For example, applications on the terminal device 101 (such as chat software, intelligent assistant applications, etc.) can receive user input information, perform basic preprocessing on the input content, and convert it into a format suitable for transmission over the network. For example, the question is encapsulated into a data packet according to a certain protocol format. For example, the user enters a text question in a specific dialogue application interface on the terminal device 101, such as "Hello, what's the weather like today?"
[0036] The server 103 can be a server that provides various services. For example, the corresponding service program on the server 103 (such as a dialogue processing service) will receive the data packet sent by the terminal device 101 transmitted from the network 102, and then parse it to extract the user's question content "What's the weather like today?".
[0037] For example, the large language model on the server 103 will receive this question. It can start reasoning operations based on its internal parameters and the knowledge obtained through training, analyze the semantics of the question, determine that this is a request for asking about the weather, and also generate a suitable answer by combining information such as the current time and the user's geographical location. The server 103 then encapsulates the generated answer content into a new data packet according to the corresponding protocol format as a reply to the request initiated by the terminal device 101.
[0038] It should be noted that the information update method or task execution method provided by the embodiments of the present invention can generally be executed by the server 103. Correspondingly, the information update device or task execution device provided by the embodiments of the present invention can generally be set in the server 103. The information update method or task execution method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103. Correspondingly, the information update device or task execution device provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103.
[0039] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0040] Figure 2A are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0041] As Figure 2A shown, the information update method of this embodiment includes operation S210 to operation S230.
[0042] In operation S210, in response to receiving input information, determine the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set, where the dynamic sequence has an initial dynamic weight, and the dynamic information set includes multiple information pairs composed of the initial dynamic weight and the dynamic sequence.
[0043] In an embodiment of the present invention, the input information may be text input by the user to the corresponding service application in real time according to actual needs, and the input information may include at least one of text information, image information, audio information, and video information. The input sequence may be a sequence obtained by performing word segmentation processing on the input information, and may be a single sequence or a sequence group. The input feature vector may be a vector obtained by performing feature extraction and processing on the input sequence. The dynamic information set may be used to store and update the importance value or weight of the dynamic sequence in real time. The dynamic sequence may be a sequence stored in the dynamic information set, and the dynamic feature vector may be a vector obtained by performing feature extraction and processing on the dynamic sequence. Multiple dynamic sequences and the weights corresponding to the multiple dynamic sequences may be combined into multiple information pairs and stored in the dynamic information set. It can be understood that the information pair may be a key-value pair.
[0044] In an embodiment of the present invention, the similarity may characterize the degree of direction difference between two different vectors and may reflect the direction similarity of the two vectors in space. The determination method of the similarity may include any one of cosine distance, Euclidean distance, or correlation coefficient, and specific details are not limited herein. The initial dynamic weight may be dynamically updated in real time according to the actual situation.
[0045] In operation S220, update the dynamic information set based on the updated dynamic weight obtained by updating the initial dynamic weight using the similarity and the associated sequence to obtain an updated dynamic information set, where the associated sequence is a sequence in the dynamic sequence that is semantically associated with the input sequence.
[0046] In an embodiment of the present invention, the updated dynamic weight may be a weight obtained by performing an update process on the initial dynamic weight using the similarity. The associated sequence may be a sequence whose degree of association or similarity with the dynamic information set in terms of semantics and / or topic satisfies a preset condition, and multiple associated sequences may form an associated information set.
[0047] For example, determine the associated sequence from the dynamic sequence based on the preset condition of the degree of association, so as to update the sorted dynamic sequence and the associated sequence using the determined updated dynamic weight and the updated associated weight to obtain multiple updated sequences, and further obtain an updated dynamic information set.
[0048] In operation S230, the initial candidate information set is updated based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set for processing the target task, where the initial candidate information set is the candidate information set before receiving the input information.
[0049] In an embodiment of the present invention, the reference information set may correspond to the dynamic information set and includes a plurality of information pairs composed of a reference sequence and a reference weight of the reference sequence, and is an information set for comparing with the dynamic sequence and the sequence weight in the dynamic information set at a later time. The reference sequence and the weight value of the reference sequence in the reference information set are static and unchanged. The first target sequence may be a sequence in the updated dynamic information set that satisfies a preset selection rule, and the second target sequence may be a sequence in the reference information set that satisfies a preset selection rule. The initial candidate information set may be an information set stored in the system before the target task is started and may be determined according to the weight of the initial sequence in the initial information set. It can be understood that the information update method in the present invention may be an information update method for the key-value cache (KV cache) technology.
[0050] For example, two weight key-value tables may be maintained in the system, including a dynamic key-value table A and a static key-value table B. It can be understood that the dynamic key-value table A corresponds to the dynamic information set, and the static key-value table B corresponds to the reference information set. The dynamic key-value table A can be used to update the importance of the sequence in real time, and the static key-value table B can save the weights of all sequences in the initial candidate information set at the beginning of each update cycle for comparing the weight changes of the sequences with the dynamic key-value table A.
[0051] For example, by comparing the difference information between the first target sequence and the second target sequence, the initial candidate information set is updated based on the difference information to obtain a candidate information set.
[0052] According to an embodiment of the present invention, since the updated dynamic weight is flexibly determined based on the similarity between the input feature vector and the dynamic feature vector, the updated dynamic information set obtained based on the updated dynamic weight takes into account the dual effects of real-time input information and its related associated sequences. Therefore, the initial candidate information set is updated in real time using the target sequences of the dynamic updated dynamic information set and the static reference information set respectively, improving the hit rate of the candidate information set in the case of cold start and semantic drift, accelerating the inference speed of the model, and further improving the user experience.
[0053] It can be understood that how to update the initial candidate information set has been described above, and below will describe how to determine the associated sequence.
[0054] According to an embodiment of the present invention, the method further includes: determining sequences in the dynamic sequence whose similarity is greater than or equal to the similarity threshold as associated sequences, so as to obtain an associated information set by using the associated sequences.
[0055] According to an embodiment of the present invention, the similarity threshold can be used to determine sequences in the dynamic sequence that have a relatively high degree of association with the input sequence at the current moment. The associated sequences can be sequences that have a relatively high degree of association with the input sequence in terms of semantics and theme. The associated information set can be constructed by using multiple associated sequences.
[0056] For example, traverse all sequence pairs in the dynamic sequence, calculate the similarity value for each sequence using a selected similarity calculation method, and store the similarity results in a data structure such as a similarity matrix or dictionary; thus, determine the similarity threshold according to statistical analysis or experimental methods. For example, in the text similarity task, if it is determined that two short texts are semantically similar, the cosine similarity threshold may be set around 0.8. It can be understood that the determination of the similarity threshold can be determined according to the actual target task, and specific details are not limited herein.
[0057] For example, according to the similarity matrix and the similarity threshold, determine sequence pairs whose similarity is greater than or equal to the similarity threshold as associated sequences, and these associated sequences can be stored in a list, where each element contains the identifier of the associated sequence and the similarity value.
[0058] According to an embodiment of the present invention, when the user inputs the latest text, by calculating the similarity between the input sequence and the dynamic sequence, determine the sequences associated with the input sequence according to the similarity threshold, and then can quickly identify the sequences related to the latest text, and update these associated sequences to the cache (candidate information set) in a timely manner, so that the cache can more quickly reflect the sequence association relationship in the current data environment, avoiding the problem of cache content lag caused by the appearance of new words or sequences, and ensuring the timeliness and effectiveness of the cache content.
[0059] According to an embodiment of the present invention, the similarity includes a first similarity between the associated sequence and the input sequence, and a second similarity between the remaining sequences in the dynamic sequence and the input sequence; based on the updated dynamic weight obtained by using the similarity to update the initial dynamic weight and the associated sequence, update the dynamic information set to obtain an updated dynamic information set, including: updating the initial dynamic weight by using the first similarity and the second similarity to obtain an updated dynamic weight; updating the dynamic information set by re - sorting the dynamic sequence based on the updated dynamic weight.
[0060] In an embodiment of the present invention, the first similarity can be determined according to the similarity between the associated sequence and the input sequence, the second similarity can be determined according to the similarity between the remaining sequence other than the associated sequence in the dynamic sequence and the input sequence, and the similarity between different sequences can be determined by the cosine distance between the respective feature vectors of different sequences. It can be understood that the initial dynamic weight can include the associated weight corresponding to the associated sequence and the remaining weight corresponding to the remaining sequence, and the initial dynamic weight can be determined based on experience, statistical analysis or experimental methods. The sequences can be sorted in descending order based on the weights of the dynamic sequence respectively, and the magnitude of the weight is proportional to the order of sorting. The greater the weight, the higher the ranking, and then multiple updated sequences can be obtained.
[0061] In a feasible embodiment, the dynamic information set can be a dynamic key-value table. The key in the dynamic key-value table is the importance degree (i.e., weight) of the sequence, and the weight can be initialized to the corresponding TF-IDF score; the value can be the embedding vector of the current sequence (or sequence group). The embedding vector can contain the semantic information of the sequence and can be used for semantic similarity and correlation retrieval.
[0062] It should be noted that, for the convenience of understanding, the embodiments of the present invention are described based on sequences, and the sequences in the relevant embodiments can be equivalently replaced by sequence groups.
[0063] For each sequence in the dynamic key-value table, its weight can decay exponentially according to the round index called by the user. Specifically, as shown in the following formula (1):
[0064] (1);
[0065] Among them, q0 can be the initial weight (importance degree) of the sequence, that is, the corresponding TF-IDF score; λ>0, λ can be the decay rate, which is used to control the speed of weight decay; t is the current round called by the user. q t can be the weight of the sequence in the current round, and can also be called the periodic weight, q t+1 can be the weight of the sequence in the next round of the current round.
[0066] It can be understood that through the exponential decay mechanism, the importance of the sequences that have not been used for a long time gradually decreases, so as to make room for new hot words (words with high TF-IDF scores) in the cache.
[0067] Figure 2B The flowchart of the sequence weight update method according to an embodiment of the present invention is shown.
[0068] As Figure 2B shown, the sequence weight update method includes operation S21~operation S29.
[0069] In operation S21, the weights of the sequences are initialized. It can be determined according to the TF-IDF scores of the sequences.
[0070] In operation S22, a dynamic key-value table A and a static key-value table B are constructed. The dynamic key-value table A can be used to update the importance of sequences in real time, and the static key-value table B can store the weights of all sequences at the start of each update cycle to compare the changes in sequence weights with the dynamic key-value table A.
[0071] In operation S23, according to the user input sequence, the cosine distance between the eigenvectors of the dynamic sequences and the input sequence in the dynamic key-value table A is calculated.
[0072] In operation S24, the weights of the sequences in the dynamic key-value table A with a sampling radius less than or equal to the distance threshold are updated. It can be understood that the distance threshold can also be referred to as the similarity threshold.
[0073] In operation S25, the distance threshold is updated according to the frequency of occurrence of the input sequence.
[0074] In operation S26, it is detected whether the current round meets the preset sampling period. If the preset period is not met, return to execute operation S23.
[0075] In operation S27, when the preset period is met, the weights of the dynamic sequences in the dynamic key-value table A are updated.
[0076] In operation S28, the top M sequences of the dynamic key-value table A and the static key-value table B are compared to determine the KV cache pairs that need to be added and replaced, and the distance threshold is reset.
[0077] In operation S29, a candidate information set is determined.
[0078] According to the embodiments of the present invention, in scenarios such as text processing, the usage frequency of sequences changes over time. In the case where the topic of the input information changes, sequences that were commonly used in the previous period but have not been used for a long time in the current period may gradually decrease in value in the current task or application. Based on the exponential decay mechanism, the importance weights of sequences are dynamically adjusted according to the length of time they have not been used. The cache space is a limited resource. By reducing the importance of sequences that have not been used for a long time, when new hot words with high TF-IDF scores need to be incorporated into the cache, the cache space can be more reasonably allocated, ensuring that the cache stores sequences that are currently more valuable and more in line with the needs of users and target tasks, and improving the utilization rate and effectiveness of the cache.
[0079] According to an embodiment of the present invention, updating the initial dynamic weight using the first similarity and the second similarity to obtain an updated dynamic weight includes: determining a first update value corresponding to the associated sequence and a second update value corresponding to the remaining sequences using the first similarity, the second similarity, and a similarity threshold; determining the updated dynamic weight based on the periodic weight of the current sequence in a preset period, the first update value, the second update value, the weight peak, and the weight valley.
[0080] In an embodiment of the present invention, the first update value may be a weight increase value corresponding to the associated sequence, and the second update value may be a weight increase value corresponding to the remaining sequences. The preset period may represent a period set based on the current model inference requirement. The weight value of the sequence may have a preset weight range, the weight peak is the maximum value in the weight range, and the weight valley may be the minimum value in the weight range.
[0081] For example, when a large language model receives a sequence of prompt words input by a user, since the model processes the sequences one by one. Taking the case where the user inputs a single sequence as an example, first calculate the embedding vector of the sequence, and then calculate the cosine distance between this vector and the embedding vectors of all sequences in the dynamic information set A; thus, the sequences (tokens) within the sampling distance threshold of the cosine distance form the associated information set A'. For each sequence in the set A', update its weight, and the degree of increase is negatively correlated with the distance between the word and the embedding vector of the user input word. Specifically, as shown in the following formula (2):
[0082] (2);
[0083] where Δq may be the weight increase value (update value) of the sequence, α is the learning rate, which can be used to control the amplitude of importance increase; d may be the cosine distance between the embedding vectors of this sequence and the user input sequence. D may be the similarity threshold, and the initial value of the similarity threshold may be set based on experience. For example, the initial value of D is D max 。
[0084] Furthermore, in combination with the above formula (2), the method for determining the weight of the sequence can be updated to obtain the updated sequence weight, which can be denoted as q' t+1 , specifically as shown in the following formula (3):
[0085] (3);
[0086] It should be noted that to make the value of the weight within a reasonable range, the weight range corresponding to the weight of the sequence can be set to be between.
[0087] According to an embodiment of the present invention, determining the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set includes: determining the direction relationship feature and the respective scale features between the input feature vector and the dynamic feature vector; determining the similarity based on the direction relationship feature and the scale feature.
[0088] In an embodiment of the present invention, the direction relationship feature can characterize the included angle size and direction relationship between two vectors. The direction relationship feature value is directly proportional to the similarity in direction between the two vectors. The scale feature can characterize the absolute size of the vector in space. Based on the square root relationship feature and the size feature, the similarity between the two vectors in terms of direction and magnitude can be comprehensively measured.
[0089] Taking the direction relationship feature as the dot product of vectors and the size feature as the modulus length of the vector as an example. The cosine similarity formed by combining the dot product of vectors and the modulus length can comprehensively measure the similarity between the two vectors in terms of direction and magnitude.
[0090] For example, in semantic analysis, for the embedding vectors of all dynamic sequences in the dynamic information set A, the cosine distance is calculated successively with the embedding vector of the input sequence, so as to obtain a set of distance values. These distance values can reflect the similarity degree of the input sequence and the dynamic sequences in the dynamic information set in the semantic space, and can more accurately reflect the semantic similarity between texts or word vectors.
[0091] According to an embodiment of the present invention, the method further includes: determining the initial weights of multiple initial sequences in the initial information set; constructing an initial candidate information set with the sequences corresponding to the initial weights greater than or equal to the weight threshold.
[0092] In an embodiment of the present invention, the initial information set can be the basic information set stored in the system before performing the target task for the model inference process, and can also be called a corpus. The initial weights of the initial sequences can be determined according to calculation rules.
[0093] For example, preprocess the initial sequence (including operations such as word segmentation, stop word removal, stemming, or lemmatization), and then obtain the processed sequence; thus, count the frequency of the sequence appearing in each document and the number of documents containing each sequence, and then obtain the TF-IDF score; then sort all the sequences in descending order according to the calculated TF-IDF score.
[0094] For example, in the cold start initialization stage of the key-value cache, the sequences in the sequence vocabulary of the corpus can be sorted according to the TF-IDF scores. The initial information set can include sequence combinations composed of a series of words that often appear simultaneously. The sequence group is, for example, {"large model", "is", "artificial intelligence", "field", "of", "popular", "technology"}. These sequences in the corpus together form the initial information set, as shown in the following formula (4):
[0095] (4);
[0096] Among them, TF(t’, d1) can represent the frequency of the term t’ in the document d1, and IDF(t’) can represent the inverse document frequency of the term t’. By calculating the TF-IDF scores of each sequence, the top M sequences with the highest scores are selected as the key-value pairs for cold start and loaded into the KV cache in the video memory. The value of M can be dynamically adjusted according to the scale of the model and the video memory capacity. For example, in some medium-scale models, M can be set to 1000.
[0097] According to an embodiment of the present invention, considering the problem of high initial inference latency caused by an empty cache in the cold start stage of the model, through the method of preloading the cache based on the TF-IDF scores, in the absence of sufficient historical data, the statistical information of the corpus can be used to reasonably select key data and load it into the cache, effectively alleviating the performance bottleneck caused by an empty cache in the cold start stage and providing a good foundation for the initial inference of the model. At the same time, during cold start, the top M sequences are pre-loaded into the cache, so that the model can directly obtain the relevant data of these high-frequency words from the cache during the first inference, avoiding the real-time calculation and loading of these data during the first inference, thereby significantly reducing the latency of the first inference.
[0098] According to an embodiment of the present invention, the initial information set includes initial sequences; the method further includes: processing the initial information set based on a decomposition strategy to obtain topic information corresponding to the initial information set and a weight matrix corresponding to multiple initial sequences, where the weight matrix represents the weights of multiple initial sequences corresponding to the topic information; determining multiple feature sequences from the multiple initial sequences based on the weight matrix, and constructing an initial candidate information set based on the multiple feature sequences.
[0099] In an embodiment of the present invention, the decomposition strategy can be a method of processing the initial information set through non-negative matrix factorization. The feature sequence can be a key sequence determined based on each row information in the first weight matrix. The initial information set can further include initial text, and the initial text can be a document in the corpus, and each document can correspond to known topic information or category information. The initial sequence can be a sequence in the initial text.
[0100] For example, by counting the occurrence frequency of each sequence in each document, a vocabulary-document matrix can be constructed. The rows of the matrix can represent the vocabulary, the columns represent the documents, and the elements represent the number of occurrences of the vocabulary in the document or the value after weighting by TF-IDF, etc. Then, the number of topics to be extracted can be determined according to prior knowledge or the elbow method. Next, the non-negative matrix is initialized. By setting the number of topics and other parameters (such as the number of iterations, initialization method, etc.), and using the constructed vocabulary-document matrix as the input, the non-negative matrix factorization algorithm is run to obtain the decomposed topic-vocabulary matrix (W) and document-topic matrix (H). For each topic (row) in the topic-vocabulary matrix (W), the vocabulary can be sorted in descending order according to the weight value, and the top N vocabulary with the highest weights are selected as the multiple feature sequences of the topic, and an initial candidate information set is constructed using the multiple feature sequences.
[0101] According to an embodiment of the present invention, through the decomposition strategy, important feature vocabulary for different categories (topics) can be extracted from a large amount of sequences, mapping the high-dimensional text data to a low-dimensional feature space, which helps to reduce the complexity of the data, highlight key information, and make subsequent data analysis and processing more efficient.
[0102] According to an embodiment of the present invention, the method further includes: updating the similarity threshold based on the frequency of the input sequence appearing within a preset period and the threshold interval corresponding to the similarity threshold to obtain the updated similarity threshold.
[0103] In an embodiment of the present invention, within a preset period (such as 10 rounds of conversations), the similarity threshold for determining associated sequences can be updated according to the frequency of the sequence appearing in the user input information, and the value is within [D min , D max , where D min and D max can be preset values. The update method of the similarity threshold is shown in the following formula (5):
[0104] (5);
[0105] where D t+1 can be the updated similarity threshold; D min and D max can be the minimum and maximum values of the similarity threshold respectively; D t can be the similarity threshold at the previous moment. f can be the frequency of the sequence, and β can be a parameter controlling the adjustment sensitivity. When the frequency of the sequence is lower, the value of D is larger, which means the search range of associated words will be expanded; when the frequency of the sequence is higher, the value of D is smaller.
[0106] According to an embodiment of the present invention, by dynamically updating the similarity threshold based on the occurrence frequency of the input sequence, it is possible to avoid storing too many unnecessary similar fields and semantic words in the cache, and improve the utilization rate of system resources.
[0107] According to an embodiment of the present invention, the initial candidate information set is updated based on the first target sequence in the dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, including: determining a difference sequence between the first target sequence and the second target sequence, where the difference sequence includes a first difference sequence and a second difference sequence, the first difference sequence is a sequence included only in the first target sequence, and the second difference sequence is a sequence included only in the second target sequence; using the first difference sequence and the second difference sequence to update the initial candidate information set to obtain a candidate information set.
[0108] In an embodiment of the present invention, the difference sequence may be a sequence that is different from each other in the first target sequence and the second target sequence.
[0109] Taking the dynamic information set as the dynamic key-value table A and the reference information set as the static key-value table B as an example, after every T rounds of conversations, the sequences in the dynamic key-value table A can be sorted in descending order according to the weight, and then the first M words (i.e., sequences) in the dynamic key-value table A and the static key-value table B are compared. Assume X add is in the dynamic key-value table A (which can be denoted as A M ), and not in the static key-value table B (which can be denoted as ), that is, the set of sequences , that is, X add is the first difference sequence; X del is in the static key-value table B (which can be denoted as B M ), and not in the dynamic key-value table A (which can be denoted as ), that is, the set of sequences , that is, X del is the second difference sequence. Obviously, , that is, the number of newly added sequences is equal to the number of sequences to be removed; thus, after determining the sequences X add and X del to be updated and removed, the static key-value table B can be updated by overwriting it with the dynamic key-value table A.
[0110] Figure 2C shows a flowchart of a method for updating an initial candidate information set according to an embodiment of the present invention.
[0111] As Figure 2C shown, the method for updating the initial candidate information set includes operations S201 to S206.
[0112] In operation S201, initialize the initial sequence in the initial information set. The initial information set may include a sequence combination composed of a series of words that often appear simultaneously.
[0113] In operation S202, determine the TF-IDF score corresponding to the initial sequence.
[0114] In operation S203, load it into the KV cache of the graphics card in descending order according to the TF-IDF score. For example, select the M sequences with the highest scores as key-value pairs for cold start and load them into the KV cache of the video memory.
[0115] In operation S204, update the KV cache in the video memory according to the weight of the sequence.
[0116] In operation S205, determine whether the update round meets the preset cycle requirement. If the current round does not meet the preset cycle, return to execute operation S204 to continue updating the KV cache.
[0117] In operation S206, when the current round meets the preset cycle, update the KV cache in the system to obtain the target cache (candidate information set).
[0118] According to an embodiment of the present invention, updating the initial candidate information set using the first difference sequence and the second difference sequence to obtain a candidate information set includes: obtaining address information for the second difference sequence based on the second difference sequence and the current candidate information set at the current moment; using the address information to store the first difference sequence into the current candidate information set to obtain the candidate information set.
[0119] In an embodiment of the present invention, the address information may be the storage address information corresponding to the second difference sequence in the initial candidate information set. The X to be eliminated del is sent into the inference model to determine the addresses corresponding to X in the KV caches of each layer in the model; thus, X del can be sent into the model to calculate the KV cache values of each layer and save them at the addresses corresponding to X add to complete the update of the cache; at the same time, the similarity threshold D of the associated sequence can be updated to D del . max .
[0120] In another feasible embodiment, regarding updating the video memory, it is also possible to, based on the load information in the system, when the load of the inference model is relatively low or relatively idle, reload all the first M sequences in the dynamic key-value table A into the KV cache of the video memory to replace the original cache.
[0121] According to an embodiment of the present invention, obtaining address information for a second difference sequence based on the second difference sequence and the current candidate information set at the current moment includes: matching the sequence identification information or sequence position information of the second difference sequence with the candidate identification information or candidate position information in the current candidate information set, and obtaining the address information based on the matching result; or obtaining the address information based on the mapping relationship between the sequence identification information and the candidate position information.
[0122] In an embodiment of the present invention, the sequence identification information and the sequence position information can be used to uniquely identify the second difference sequence. The candidate identification information and the candidate position information can be used to uniquely identify a candidate sequence in the candidate information set. The address information of the second difference sequence in the model can be obtained by matching the sequence identification information with the candidate identification information or matching the sequence position information with the candidate position information, when the matching result is a correct match.
[0123] For example, X del can be used as an input and sent to the model, and the model will search for the address corresponding to X del in the KV caches of each layer according to the content of X del and the current cache state. The matching method can include: matching the sequence identification or position information in the cache to determine the specific position of X del in the cache; or, each sequence in the cache may have a unique identifier, and the corresponding address can be found by comparing the identifier of X del with the identifiers stored in the cache.
[0124] For example, a dictionary (or a similar data structure) can be created for each layer of the model to store the mapping relationship between the sequence identification information and the candidate position information; during the model inference process, when a new sequence is loaded into the cache, the mapping table can be updated to associate the sequence identification information with the actual storage position information in the cache, so as to realize finding the corresponding cache address information in the mapping table through the sequence identifier.
[0125] Figure 3 The flowchart of the task execution method according to an embodiment of the present invention is shown.
[0126] As Figure 3 shown, the task execution method includes operation S310.
[0127] In operation S310, in response to receiving a task execution instruction, the candidate information set is used to process the target task; wherein, the target task includes any one of text generation, speech recognition, and financial analysis, and the candidate information set is obtained according to the above information update method.
[0128] In an embodiment of the present invention, the task execution instruction may be an instruction generated according to actual application requirements. For example, in a speech recognition task, the target instructions may include instructions such as speech-to-text, speech translation, and speech instruction recognition. For example, in a text generation task, the target instructions may include instructions such as text creation, text editing, and text storage. It should be noted that the above application scenarios are not limited to text generation, speech recognition, and financial analysis, and are not specifically limited herein.
[0129] Taking a medical Q&A system as an example, a large language model is used as the dialogue model. Suppose there are 10,000 sequence groups in the initial information set (corpus), and M = 1000 is set, that is, 1000 sequences with the highest TF-IDF scores are loaded into the KV cache during cold start. At the same time, a similarity threshold D min = 0.1, D max = 0.5, λ = 0.01, α = 0.1, β = 0.05, and a preset period T = 10 rounds of dialogue are set. It should be noted that the specific type of the large prediction model in the present invention can be determined according to the actual situation and is not limited herein.
[0130] Suppose there are 1000 medical documents. Taking the sequence "chest pain" as an example, this sequence appears 10 times in document d1, and document d1 has a total of 1000 words. Then the term frequency of this sequence in d1 . Suppose "chest pain" appears in 200 documents. Then the inverse document frequency of this sequence ; thus, the TF-IDF score of "chest pain" can be obtained . By calculating the TF-IDF of all sequences in the vocabulary, the 1000 sequences with the highest scores are selected as the key-value pairs for cold start and loaded into the KV cache in the video memory. At the same time, the TF-IDF scores of these sequences are stored as the initial weights in the dynamic key-value table A and the static key-value table B of the weight key-value table, and the embedding vectors of these sequences are recorded.
[0131] Suppose after cold start initialization, the dynamic key-value table A and the static key-value table B have stored the information of 1000 sequence groups. Taking the "chest pain" sequence as an example, its initial weight .
[0132] Before the user's first round of dialogue, the weights of all sequences remain at their initial values. When the user has the first round of dialogue, the weight of "chest pain" begins to decay. According to the formula , at this time t = 1, , then the weight of the sequence "chest pain" can be updated to .
[0133] Suppose in the second round of conversation, the user inputs "chest pain". First, the cosine distances between the embedding vector (feature vector) of "chest pain" and the embedding vectors of all sequences in the dynamic key-value table A can be calculated. Suppose there is a sequence "angina pectoris" in the vocabulary, and the cosine distance d between its embedding vector and the embedding vector of "chest pain" is 0.2. Thus, according to the semantic association promotion mechanism, for "chest pain" itself, the distance d = 0.2, D = 0.3, α = 0.1, then the weight is promoted . Meanwhile, for the associated sequence "angina pectoris" (because d = 0.2 is less than the current cosine distance), its weight is promoted. According to the formula, . Suppose the original weight of "angina pectoris" is , then the promoted weight is .
[0134] Furthermore, the similarity threshold can be updated according to the frequency of the user's input "chest pain". Suppose "chest pain" has appeared a total of 2 times in the previous conversation. According to formula (5), the associated sampling radius for "chest pain" can be obtained as . It should be noted that within one cycle, the associated sampling radii of sequences not input by other users do not change.
[0135] In each subsequent round of conversation within the preset cycle, the above steps can be repeated to update the weights, promote or decay the weights according to the user input, and update the sampling distance. For example, in the third round of conversation, the weights of all sequences continue to decay. If the user inputs a new word, then calculate the weight promotion and update the sampling distance of the associated words according to the above steps.
[0136] After reaching the tenth round of conversation, the sequences in the dynamic key-value table A in the weight key-value table are sorted according to the weights. Suppose among the top 1000 sequences in the sorted dynamic key-value table A, 100 sequences are not among the top 1000 words in the static key-value table B, that is, X add = 100; meanwhile, among the top 1000 words in the static key-value table B, 100 words are not among the top 1000 words in the dynamic key-value table A, that is, X del = 100.
[0137] The X del to be excluded can be sent into the model to determine the corresponding address information in the KV caches of each layer of the model; thus, the X del can be sent into the model, calculate the KV cache values of each layer, and save them in X add add The corresponding address information is used to complete the update of the cache. After that, the content of the dynamic key-value table A can be copied to the static key-value table B to prepare for the update of the next 10 rounds of conversations.
[0138] Based on the above information update method, the present invention also provides an information update device. The following will be combined with Figure 4 to describe this device in detail.
[0139] Figure 4 The structural block diagram of the information update device according to an embodiment of the present invention is shown.
[0140] As Figure 4 shown, the information update device of this embodiment includes a similarity determination module 410, a dynamic information set update module 420, and a candidate information set update module 430.
[0141] The similarity determination module 410 is configured to determine the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set in response to receiving the input information, where the dynamic sequence has an initial dynamic weight, and the dynamic information set includes a plurality of information pairs composed of the initial dynamic weight and the dynamic sequence. In one embodiment, the similarity determination module 410 can be used to perform the operation S210 described above, which will not be elaborated here.
[0142] The dynamic information set update module 420 is configured to update the dynamic information set based on the updated dynamic weight obtained by updating the initial dynamic weight using the similarity and the associated sequence to obtain an updated dynamic information set, where the associated sequence is the sequence in the dynamic sequence that is semantically associated with the input sequence. In one embodiment, the dynamic information set update module 420 can be used to perform the operation S220 described above, which will not be elaborated here.
[0143] The candidate information set update module 430 is configured to update the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, so as to process the target task using the candidate information set, where the initial candidate information set is the candidate information set before receiving the input information. In one embodiment, the candidate information set update module 430 can be used to perform the operation S230 described above, which will not be elaborated here.
[0144] According to an embodiment of the present invention, through the similarity determination module 410, the dynamic information set update module 420, and the candidate information set update module 430 in the information update device, since the updated dynamic weight is flexibly determined based on the similarity between the input feature vector and the dynamic feature vector, the updated dynamic information set obtained based on the updated dynamic weight takes into account the dual effects of real-time input information and its related associated sequences. Thus, the initial candidate information set is updated in real time using the target sequences of the dynamic updated dynamic information set and the static reference information set respectively, improving the hit rate of the candidate information set in cold start and semantic drift situations, accelerating the inference speed of the model, and further improving the user experience.
[0145] According to an embodiment of the present invention, the device further includes: a sequence determination module, configured to determine a sequence in the dynamic sequence whose similarity is greater than or equal to a similarity threshold as an associated sequence, so as to obtain an associated information set using the associated sequence.
[0146] According to an embodiment of the present invention, the similarity includes a first similarity between the associated sequence and the input sequence, and a second similarity between the remaining sequences in the dynamic sequence and the input sequence; the dynamic information set update module 420 includes: a dynamic weight update sub-module and a sequence update sub-module. The dynamic weight update sub-module is configured to update the initial dynamic weight using the first similarity and the second similarity to obtain an updated dynamic weight; the sequence update sub-module is configured to perform an updated sorting on the dynamic sequence based on the updated dynamic weight to update the dynamic information set.
[0147] According to an embodiment of the present invention, the dynamic weight update sub-module includes: an update value determination unit and a dynamic weight determination unit. The update value determination unit is configured to use the first similarity, the second similarity, and the similarity threshold to determine a first update value corresponding to the associated sequence and a second update value corresponding to the remaining sequences; the dynamic weight determination unit is configured to determine the updated dynamic weight based on the periodic weight of the current sequence in a preset period, the first update value and the second update value, the weight peak value, and the weight valley value.
[0148] According to an embodiment of the present invention, the similarity determination module 410 includes: a feature determination sub-module and a similarity determination sub-module. The feature determination sub-module is configured to determine the direction relationship feature and the respective scale features between the input feature vector and the dynamic feature vector; the similarity determination sub-module is configured to determine the similarity based on the direction relationship feature and the scale feature.
[0149] According to an embodiment of the present invention, the device further includes: an initial weight determination module and an information set construction module. The initial weight determination module is configured to determine the initial weights of multiple initial sequences in the initial information set respectively; the information set construction module is configured to construct an initial candidate information set with sequences whose initial weights are greater than or equal to a weight threshold.
[0150] According to an embodiment of the present invention, the initial information set includes an initial sequence; the apparatus further includes: an information processing module and a feature sequence determination module. The information processing module is configured to process the initial information set based on a decomposition strategy to obtain topic information corresponding to the initial information set and a weight matrix corresponding to a plurality of initial sequences, where the weight matrix represents the weights of the plurality of initial sequences corresponding to the topic information; the feature sequence determination module is configured to determine a plurality of feature sequences from the plurality of initial sequences based on the weight matrix and construct an initial candidate information set based on the plurality of feature sequences.
[0151] According to an embodiment of the present invention, the apparatus further includes: a threshold refinement module, configured to update a similarity threshold based on the frequency of occurrence of an input sequence within a preset period and a threshold interval corresponding to the similarity threshold to obtain an updated similarity threshold.
[0152] According to an embodiment of the present invention, the candidate information set update module 430 includes: a difference sequence determination module and a candidate information set update module. The difference sequence determination module is configured to determine a difference sequence between a first target sequence and a second target sequence; the candidate information set update module is configured to update the initial candidate information set using the first difference sequence and the second difference sequence to obtain a candidate information set.
[0153] According to an embodiment of the present invention, the candidate information set update module includes: an address information determination sub-module and a storage sub-module. The address information determination sub-module is configured to obtain address information for the second difference sequence based on the second difference sequence and the current candidate information set at the current moment; the storage sub-module is configured to store the first difference sequence into the current candidate information set using the address information to obtain a candidate information set.
[0154] According to an embodiment of the present invention, the address information determination sub-module includes: a matching unit and a mapping unit. The matching unit is configured to match the sequence identification information or sequence position information of the second difference sequence with the candidate identification information and candidate position information in the current candidate information set and obtain address information based on the matching result; or the mapping unit is configured to obtain address information based on the mapping relationship between the sequence identification information and the candidate position information.
[0155] Figure 5 The structural block diagram of a task execution apparatus according to an embodiment of the present invention is shown.
[0156] As Figure 5As shown, the task execution device of this embodiment includes a task processing module 510. The task processing module is configured to process a target task by using a candidate information set in response to receiving a task execution instruction; wherein, the target task includes any one of text generation, speech recognition, and financial analysis, and the candidate information set is obtained according to the above-mentioned task execution method. In one embodiment, the task processing module 510 may be used to execute the operation S310 described above, which will not be elaborated here.
[0157] According to an embodiment of the present invention, any multiple modules among the similarity determination module 410, the dynamic information set update module 420, and the candidate information set update module 430 may be combined and implemented in one module, or any one of them may be split into multiple modules. Or, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the similarity determination module 410, the dynamic information set update module 420, and the candidate information set update module 430 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Or, at least one of the similarity determination module 410, the dynamic information set update module 420, and the candidate information set update module 430 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions. Similarly, this embodiment is also applicable to the task processing module 510, which will not be elaborated here specifically.
[0158] Figure 6 The block diagram of an electronic device according to an embodiment of the present invention is shown.
[0159] As Figure 6 shown, the electronic device includes: a memory 610 and a processor 620. The processor 620 is configured to execute the above-mentioned information update method and / or task execution method according to the instructions and data stored in the memory.
[0160] In an embodiment of the present invention, the processor 620 may be a server that provides various services. The instructions include but are not limited to instructions related to text generation tasks, speech recognition tasks, and financial analysis tasks. The data may include but are not limited to a dynamic information set, an association information set, and a reference information set.
[0161] Figure 7A block diagram of an electronic device suitable for implementing an information update method according to an embodiment of the present invention is shown.
[0162] As Figure 7 shown, the electronic device according to an embodiment of the present invention includes a first processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The first processor 701 may include, for example, a general-purpose micro-first processor (e.g., CPU), an instruction set first processor, and / or a dedicated micro-first processor (e.g., an application specific integrated circuit (ASIC)), etc. The first processor 701 may also include on-board memory for caching purposes. The first processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0163] In the RAM 703, various programs and data required for the operation of the electronic device are stored. The first processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The first processor 701 performs various operations of the method flow according to an embodiment of the present invention by executing the program in the ROM 702 and / or the RAM 703. It should be noted that the program may also be stored in one or more memories other than the ROM 702 and the RAM 703. The first processor 701 may also perform various operations of the method flow according to an embodiment of the present invention by executing the program stored in the one or more memories.
[0164] According to an embodiment of the present invention, the electronic device may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device may further include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.
[0165] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0166] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.
[0167] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the information update method provided by the embodiments of the present invention.
[0168] When the computer program is executed by the first processor 701, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0169] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0170] In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the first processor 701, the above functions defined in the system of the embodiments of the present invention are executed. According to the embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.
[0171] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0173] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0174] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. An information update method, characterized in that, The method includes: In response to receiving input information, determining a similarity between an input feature vector of an input sequence in the input information and a dynamic feature vector of a dynamic sequence in a dynamic information set, where the dynamic sequence has an initial dynamic weight, and the dynamic information set includes a plurality of information pairs composed of the initial dynamic weight and the dynamic sequence; Updating the dynamic information set based on an updated dynamic weight obtained by updating the initial dynamic weight using the similarity and an associated sequence to obtain an updated dynamic information set, where the associated sequence is a sequence in the dynamic sequence that is semantically associated with the input sequence; Updating an initial candidate information set based on a first target sequence in the updated dynamic information set and a second target sequence in a reference information set to obtain a candidate information set for processing a target task using the candidate information set, where the initial candidate information set is the candidate information set before receiving the input information.
2. The method according to claim 1, characterized in that The method further includes: Determining a sequence in the dynamic sequence with a similarity greater than or equal to a similarity threshold as the associated sequence to obtain an associated information set using the associated sequence.
3. The method according to claim 2, wherein The similarity includes a first similarity between the associated sequence and the input sequence, and a second similarity between the remaining sequences in the dynamic sequence and the input sequence; Updating the dynamic information set based on an updated dynamic weight obtained by updating the initial dynamic weight using the similarity and an associated sequence to obtain an updated dynamic information set, including: Updating the initial dynamic weight using the first similarity and the second similarity to obtain an updated dynamic weight; Updating and sorting the dynamic sequence based on the updated dynamic weight to obtain the updated dynamic information set.
4. The method according to claim 3, characterized in that, Updating the initial dynamic weight using the first similarity and the second similarity to obtain an updated dynamic weight, including: Determining a first update value corresponding to the associated sequence and a second update value corresponding to the remaining sequences using the first similarity, the second similarity, and the similarity threshold; Determining the updated dynamic weight based on a periodic weight of the current sequence in a preset period, the first update value, the second update value, a weight peak value, and a weight valley value.
5. The method according to claim 1, wherein Determining a similarity between an input feature vector of an input sequence in the input information and a dynamic feature vector of a dynamic sequence in a dynamic information set, including: Determining a direction relationship feature and respective scale features between the input feature vector and the dynamic feature vector; Determining the similarity based on the direction relationship feature and the scale feature.
6. The method according to claim 1, wherein The method further includes: Determining initial weights of a plurality of initial sequences in an initial information set; Constructing the initial candidate information set with sequences corresponding to initial weights greater than or equal to a weight threshold.
7. The method according to claim 6, characterized in that, The initial information set includes a plurality of initial sequences; the method further includes: Processing the initial information set based on a decomposition strategy to obtain topic information corresponding to the initial information set and a weight matrix corresponding to the plurality of initial sequences, where the weight matrix represents weights of the plurality of initial sequences corresponding to the topic information. Determine a plurality of feature sequences from a plurality of initial sequences based on the weight matrix, and construct the initial candidate information set based on the plurality of feature sequences.
8. The method according to claim 4, wherein The method further includes: Updating the similarity threshold based on the frequency of occurrence of the input sequence within the preset period and the threshold interval corresponding to the similarity threshold to obtain an updated similarity threshold.
9. The method according to any one of claims 1 to 8, characterized in that, Updating the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, including: Determine the difference sequence between the first target sequence and the second target sequence, where the difference sequence includes a first difference sequence and a second difference sequence, the first difference sequence is a sequence included only in the first target sequence, and the second difference sequence is a sequence included only in the second target sequence; Update the initial candidate information set using the first difference sequence and the second difference sequence to obtain the candidate information set.
10. The method according to claim 9, characterized in that Updating the initial candidate information set using the first difference sequence and the second difference sequence to obtain the candidate information set, including: Based on the second difference sequence and the current candidate information set at the current moment, obtain the address information for the second difference sequence; Using the address information, store the first difference sequence into the current candidate information set to obtain the candidate information set.
11. The method according to claim 9, wherein Based on the second difference sequence and the current candidate information set at the current moment, obtaining the address information for the second difference sequence, including: Match the sequence identification information or sequence position information of the second difference sequence with the candidate identification information or candidate position information in the current candidate information set, and obtain the address information based on the matching result; or Based on the mapping relationship between the sequence identification information and the candidate position information, obtain the address information.
12. A task execution method, characterized in that The method includes: In response to receiving a task execution instruction, process the target task using the candidate information set; Wherein the target task includes any one of text generation, speech recognition, and financial analysis, and the candidate information set is obtained according to the method described in any one of claims 1 to 11.
13. An electronic device, characterized in that, Includes: A memory; A processor configured to execute the method according to any one of claims 1 to 12 based on the instructions and data stored in the memory.
14. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Chinese academic keyword extraction method and device and storage medium
CN113268995A
Medical dialogue text matching method and device, equipment and storage medium
CN116069918A
Calculation method and system for unstructured text data
CN119474383A
Multi-task parallel processing method and system based on AI target identification
CN119960946A
Representation and visualization of multivariate sensory time series data
US20220188343A1