Information updating method, task execution method, device, medium and program product

By updating dynamic weights and association sequences in large language models and optimizing cache management, the problems of semantic drift and cold start delay are solved, and the cache hit rate and model inference speed are improved.

CN120234401BActive Publication Date: 2025-08-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510725289.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-12
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The prior art cannot effectively deal with semantic drift and cold start delay in large language models, resulting in low cache utilization and low cache hit rate.

Method used

By determining the similarity between the input sequence in the input information and the dynamic information set, update the dynamic weight and association sequence, and update the candidate information set in combination with the reference information set, and optimize cache management.

Benefits of technology

Improve the cache hit rate in cold start and semantic drift cases, speed up model inference speed, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234401B_ABST
    Figure CN120234401B_ABST
Patent Text Reader

Abstract

The present invention provides an information updating method that can be applied to the fields of information storage technology and model reasoning technology. The method includes: in response to receiving input information, determining the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set, wherein the dynamic sequence has an initial dynamic weight; updating the dynamic information set based on the updated dynamic weight and the associated sequence obtained by updating the initial dynamic weight using the similarity, to obtain an updated dynamic information set, wherein the associated sequence is a sequence in the dynamic sequence that is semantically associated with the input sequence; updating the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set, to obtain a candidate information set, so as to process the target task using the candidate information set, wherein the initial candidate information set is the candidate information set before receiving the input information. The present invention also provides a task execution method, device, medium, and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to information storage technology and model reasoning technology, and in particular to an information updating method, a task execution method, a device, a medium and a program product. Background Art

[0002] Large language models typically have a massive number of parameters and intermediate results, resulting in limited system storage resources. Key-value caching technology reduces redundant computations by storing intermediate key and value vectors, which can improve model inference efficiency, reduce system resource usage, and support efficient decoding during inference. Related technologies primarily manage key-value caches through hardware co-optimization, cache compression, and adjusting cache priorities based on word frequency. However, in actual use, these technologies are not applicable to situations such as semantic drift and cold start delays that exist during model inference, resulting in low cache utilization and low cache hit rates. Summary of the Invention

[0003] In view of the above problems, the present invention provides an information updating method, a task execution method, an apparatus, a device, a medium and a program product.

[0004] According to a first aspect of the present invention, there is provided an information updating method, comprising: in response to receiving input information, determining the similarity between an input feature vector of an input sequence in the input information and a dynamic feature vector of a dynamic sequence in a dynamic information set, wherein the dynamic sequence has an initial dynamic weight, and the dynamic information set comprises a plurality of information pairs consisting of the initial dynamic weight and the dynamic sequence; updating the dynamic information set based on updated dynamic weights and associated sequences obtained by updating the initial dynamic weights using the similarity to obtain an updated dynamic information set, wherein the associated sequence is a sequence in the dynamic sequence that is semantically associated with the input sequence; updating the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, so as to process a target task using the candidate information set, wherein the initial candidate information set is the candidate information set before receiving the input information.

[0005] The second aspect of the present invention provides a task execution method, comprising: in response to receiving a task execution instruction, processing a target task using a candidate information set; wherein the target task includes any one of text generation, speech recognition and financial analysis, and the candidate set is obtained according to the above-mentioned information updating method.

[0006] The third aspect of the present invention provides an information updating device, including: a similarity determination module, used to determine the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set in response to receiving input information, wherein the dynamic sequence has an initial dynamic weight, and the dynamic information set includes multiple information pairs consisting of the initial dynamic weight and the dynamic sequence; a dynamic information set updating module, used to update the dynamic information set based on the updated dynamic weight and the associated sequence obtained by updating the initial dynamic weight using the similarity, to obtain an updated dynamic information set, wherein the associated sequence is obtained by a sequence associated with the semantics of the dynamic sequence; a candidate information set updating module, used to update the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set, to obtain a candidate information set, so as to use the candidate information set to process the target task, wherein the initial candidate information set is the candidate information set before receiving the input information.

[0007] The fourth aspect of the present invention provides a task execution device, which includes: a task processing module for processing a target task using a candidate information set in response to receiving a task execution instruction; wherein the target task includes any one of text generation, speech recognition and financial analysis, and the candidate information set is obtained according to the above-mentioned task execution method.

[0008] A fifth aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0009] The sixth aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0010] The seventh aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings.

[0012] Figure 1 An application scenario diagram of the information updating method, task execution method, device, medium, and program product according to an embodiment of the present invention is shown.

[0013] Figure 2A A flow chart of an information updating method according to an embodiment of the present invention is shown.

[0014] Figure 2B A flow chart of a sequence weight updating method according to an embodiment of the present invention is shown.

[0015] Figure 2C A flowchart of a method for updating an initial candidate information set according to an embodiment of the present invention is shown.

[0016] Figure 3 A flowchart of a task execution method according to an embodiment of the present invention is shown.

[0017] Figure 4 A structural block diagram of an information updating device according to an embodiment of the present invention is shown.

[0018] Figure 5 A structural block diagram of a task execution device according to an embodiment of the present invention is shown.

[0019] Figure 6 A block diagram of an electronic device according to an embodiment of the present invention is shown.

[0020] Figure 7 A block diagram of an electronic device suitable for implementing an information updating method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0021] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0022] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0024] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0025] In some examples, different hardware architectures are used to work together, for example, by offloading cyclic redundancy checking and data compression from the Data Processing Unit (DPU) to reduce the memory usage of the Graphics Processing Unit (GPU).

[0026] However, DPUs and GPUs have different hardware architectures, and their collaborative operation requires good adaptation and compatibility. In practical applications, problems such as high communication latency and limited data transmission bandwidth between the DPU and GPU may arise, affecting the optimization of overall model inference performance.

[0027] In some examples, we prioritize high-frequency hot words (words with high TF-IDF values) by adjusting the priority based on word frequency, such as term frequency-inverse document frequency (TF-IDF). This cached high-frequency hot words (words with high TF-IDF values) speeds up access to these words to a certain extent.

[0028] However, methods based solely on word frequency, such as TF-IDF, rely on historical data to prioritize words. During the cold start phase of model inference, the system lacks sufficient historical data to calculate word frequency and inverse document frequency, making it difficult to accurately identify high-frequency hot words.

[0029] In some examples, based on the analysis of the attention module, a selective storage strategy is used to improve the performance of large language models through lightweight model analysis and adaptive key-value caching. This reduces cache space usage to a certain extent. However, as the system runs longer, the meaning of the data may change, and changes in the input information topic may cause semantic drift. If the cached data is not updated in a timely manner to reflect this semantic change, it may become irrelevant or inaccurate, and fail to meet the user's current needs.

[0030] Based on the above problems, the present invention provides an information updating method, comprising: in response to receiving input information, determining the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set, wherein the dynamic sequence has an initial dynamic weight, and the dynamic information set includes a plurality of information pairs consisting of the initial dynamic weight and the dynamic sequence; updating the dynamic information set based on the updated dynamic weight and the associated sequence obtained by updating the initial dynamic weight using the similarity to obtain an updated dynamic information set, wherein the associated sequence is a sequence in the dynamic sequence that is semantically associated with the input sequence; updating the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, so as to use the candidate information set to process the target task, wherein the initial candidate information set is the candidate information set before receiving the input information.

[0031] According to an embodiment of the present invention, since the updated dynamic weight is flexibly determined based on the similarity between the input feature vector and the dynamic feature vector, the updated dynamic information set obtained based on the updated dynamic weight simultaneously takes into account the dual influence of the real-time input information and the associated sequence related thereto, thereby utilizing the respective target sequences of the dynamic updated dynamic information set and the static reference information set to update the initial candidate information set in real time, thereby improving the hit rate of the candidate information set in the case of cold start and semantic drift, accelerating the inference speed of the model, and further improving the user experience.

[0032] Figure 1 An application scenario diagram of the information updating method, task execution method, device, medium, and program product according to an embodiment of the present invention is shown.

[0033] like Figure 1 As shown, the application scenario according to this embodiment may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0034] A user can use a terminal device 101 to interact with a server 103 via a network 102 to receive or send messages, etc. The terminal device 101 can be any electronic device with a display screen and supporting web browsing, including but not limited to a smartphone, a tablet computer, a laptop computer, a desktop computer, and the like.

[0035] For example, an application (such as a chat software or intelligent assistant application) on the terminal device 101 can receive user input information, perform basic pre-processing on the input content, and convert it into a format suitable for transmission over the network, such as encapsulating the question into a data packet according to a certain protocol format. For example, a user enters a text question into a specific dialogue application interface on the terminal device 101, such as "Hello, how is the weather today?"

[0036] Server 103 can be a server that provides various services. For example, a corresponding service program (such as a conversation processing service) on server 103 receives a data packet sent by terminal device 101 from network 102, parses it, and extracts the user's question content, "What's the weather like today?"

[0037] For example, the large language model on server 103 receives the question. Based on its internal parameters and trained knowledge, it begins reasoning and analyzing the semantics of the question, determining that it is a request about the weather. It also generates an appropriate response based on information such as the current time and the user's location. Server 103 then encapsulates this response into a new data packet according to the corresponding protocol format, which serves as a response to the request initiated by terminal device 101.

[0038] It should be noted that the information updating method or task execution method provided in the embodiments of the present invention can generally be executed by the server 103. Accordingly, the information updating device or task execution device provided in the embodiments of the present invention can generally be set in the server 103. The information updating method or task execution method provided in the embodiments of the present invention can also be executed by a server or server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103. Accordingly, the information updating device or task execution device provided in the embodiments of the present invention can also be set in a server or server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103.

[0039] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0040] Figure 2A A flow chart of an information updating method according to an embodiment of the present invention is shown.

[0041] like Figure 2A As shown, the information updating method of this embodiment includes operations S210 to S230.

[0042] In operation S210, in response to receiving input information, a similarity between an input feature vector of an input sequence in the input information and a dynamic feature vector of a dynamic sequence in a dynamic information set is determined, wherein the dynamic sequence has an initial dynamic weight and the dynamic information set includes a plurality of information pairs consisting of the initial dynamic weight and the dynamic sequence.

[0043] In an embodiment of the present invention, the input information may be text that a user inputs into a corresponding service application in real time according to actual needs. The input information may include at least one of text information, image information, audio information, and video information. The input sequence may be a sequence obtained by performing word segmentation processing on the input information, and may be a single sequence or a sequence group. The input feature vector may be a vector obtained by extracting and processing features of the input sequence. The dynamic information set may be used to store and update the importance value or weight of the dynamic sequence in real time. The dynamic sequence may be a sequence stored in the dynamic information set, and the dynamic feature vector may be a vector obtained by extracting and processing features of the dynamic sequence. Multiple dynamic sequences and the weights corresponding to the multiple dynamic sequences may be combined into multiple information pairs and stored in the dynamic information set. It can be understood that the information pair may be a key-value pair.

[0044] In embodiments of the present invention, similarity can represent the degree of directional difference between two different vectors and can reflect the directional similarity of the two vectors in space. Similarity can be determined using any of the following methods: cosine distance, Euclidean distance, or correlation coefficient, without limitation herein. The initial dynamic weight can be dynamically updated in real time based on actual conditions.

[0045] In operation S220 , the dynamic information set is updated based on the updated dynamic weight obtained by updating the initial dynamic weight using the similarity and the associated sequence to obtain an updated dynamic information set, wherein the associated sequence is a sequence in the dynamic sequence that is semantically associated with the input sequence.

[0046] In embodiments of the present invention, the updated dynamic weight may be a weight obtained by updating the initial dynamic weight using similarity. An associated sequence may be a sequence whose semantic and / or topical association or similarity with a dynamic information set satisfies preset conditions. Multiple associated sequences may constitute an associated information set.

[0047] For example, based on the preset correlation condition, an associated sequence is determined from the dynamic sequence, and then the determined updated dynamic weight and updated associated weight are used to update the sorted dynamic sequence and the associated sequence to obtain multiple updated sequences, and then obtain an updated dynamic information set.

[0048] In operation S230, the initial candidate information set is updated based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set to process the target task using the candidate information set, wherein the initial candidate information set is the candidate information set before receiving the input information.

[0049] In an embodiment of the present invention, the reference information set may be an information set corresponding to a dynamic information set, including a plurality of information pairs consisting of a reference sequence and a reference weight of the reference sequence, and used to compare with the dynamic sequence and the sequence weight in the dynamic information set at a later moment. The reference sequence and the weight value of the reference sequence in the reference information set are static and unchanged. The first target sequence may be a sequence in the updated dynamic information set that meets the preset selection rules, and the second target sequence may be a sequence in the reference information set that meets the preset selection rules. The initial candidate information set may be an information set stored in the system before the target task is started, and may be determined based on the weight of the initial sequence in the initial information set. It is understandable that the information update method in the present invention may be an information update method for key-value cache (KV cache) technology.

[0050] For example, two weight key tables can be maintained in the system, including a dynamic key table A and a static key table B. It can be understood that dynamic key table A corresponds to the dynamic information set, and static key table B corresponds to the reference information set. Dynamic key table A can be used to update the importance of sequences in real time, while static key table B can store the weights of all sequences in the initial candidate information set at the beginning of each update cycle, and compare the changes in sequence weights with dynamic key table A.

[0051] For example, by comparing difference information between the first target sequence and the second target sequence, the initial candidate information set is updated based on the difference information to obtain the candidate information set.

[0052] According to an embodiment of the present invention, since the updated dynamic weight is flexibly determined based on the similarity between the input feature vector and the dynamic feature vector, the updated dynamic information set obtained based on the updated dynamic weight simultaneously considers the dual influence of the real-time input information and the associated sequence related thereto, thereby utilizing the respective target sequences of the dynamic updated dynamic information set and the static reference information set to update the initial candidate information set in real time, thereby improving the hit rate of the candidate information set in the case of cold start and semantic drift, accelerating the reasoning speed of the model, and further improving the user experience.

[0053] It can be understood that how to update the initial candidate information set has been described above, and how to determine the association sequence will be described below.

[0054] According to an embodiment of the present invention, the method further includes: determining sequences in the dynamic sequence whose similarity is greater than or equal to a similarity threshold as associated sequences, so as to obtain an associated information set using the associated sequences.

[0055] According to an embodiment of the present invention, a similarity threshold can be used to determine, from a dynamic sequence, sequences that are currently highly correlated with the input sequence. A correlated sequence can be a sequence that is highly correlated with the input sequence in terms of semantics and themes. A correlated information set can be constructed using multiple correlated sequences.

[0056] For example, all sequence pairs in a dynamic sequence are traversed, and a selected similarity calculation method is applied to each sequence to obtain a similarity value. The similarity results are stored in a data structure such as a similarity matrix or dictionary. A similarity threshold is then determined based on statistical analysis or experimental methods. For example, in a text similarity task, if two short articles are semantically similar, the cosine similarity threshold might be set around 0.8. It is understood that the determination of the similarity threshold can be determined based on the actual target task and is not specifically limited here.

[0057] For example, based on the similarity matrix and the similarity threshold, sequence pairs with similarities greater than or equal to the similarity threshold are determined as related sequences. These related sequences can be stored in a list, with each element containing an identifier of the related sequence and a similarity value.

[0058] According to an embodiment of the present invention, when a user inputs the latest text, the similarity between the input sequence and the dynamic sequence is calculated, and sequences associated with the input sequence are determined based on a similarity threshold. This allows for rapid identification of sequences related to the latest text, and timely updates of these associated sequences to the cache (candidate information set). This allows the cache to more quickly reflect sequence associations in the current data environment, avoiding cache content lags caused by the emergence of new words or sequences, and ensuring the timeliness and effectiveness of the cache content.

[0059] According to an embodiment of the present invention, the similarity includes a first similarity between the associated sequence and the input sequence, and a second similarity between the remaining sequences in the dynamic sequence and the input sequence; based on the updated dynamic weight obtained by updating the initial dynamic weight using the similarity and the associated sequence, the dynamic information set is updated to obtain an updated dynamic information set, including: updating the initial dynamic weight using the first similarity and the second similarity to obtain an updated dynamic weight; updating and sorting the dynamic sequence based on the updated dynamic weight, and updating the dynamic information set.

[0060] In an embodiment of the present invention, the first similarity can be determined based on the similarity between the associated sequence and the input sequence, and the second similarity can be determined based on the similarity between the remaining sequences in the dynamic sequence, excluding the associated sequence, and the input sequence. The similarity between different sequences can be determined by the cosine distance between their respective feature vectors. It is understood that the initial dynamic weights may include the associated weight corresponding to the associated sequence and the remaining weights corresponding to the remaining sequences. The initial dynamic weights may be determined based on experience, statistical analysis, or experimental methods. The sequences can be sorted in descending order based on the weights of the dynamic sequences. The weights are proportional to the order of sorting; larger weights indicate higher rankings, thereby obtaining multiple updated sequences.

[0061] In one feasible embodiment, the dynamic information set can be a dynamic key-value table. The key in the dynamic key-value table represents the importance (i.e., weight) of the sequence, which can be initialized to the corresponding TF-IDF score. The value can be the embedding vector of the current sequence (or sequence group). The embedding vector can contain semantic information about the sequence and can be used for semantic similarity and association retrieval.

[0062] It should be noted that, for ease of understanding, the embodiments of the present invention are explained based on sequences, and the sequences in the relevant embodiments can be equally replaced by sequence groups.

[0063] For each sequence in the dynamic key-value table, its weight can be exponentially decayed according to the number of rounds called by the user. The specific details are shown in the following formula (1):

[0064] (1);

[0065] Where q0 is the initial weight (importance) of the sequence, i.e. the corresponding TF-IDF score; λ>0, λ is the decay rate, which is used to control the speed of weight decay; t is the current round called by the user. t It can be the weight of the sequence in the current round, or it can be called the cycle weight, q t+1 It can be the weight of the sequence in the next round of the current round.

[0066] It can be understood that through the exponential decay mechanism, the importance of sequences that have not been used for a long time gradually decreases, making cache space for new hot words (words with high TF-IDF scores).

[0067] Figure 2B A flow chart of a sequence weight updating method according to an embodiment of the present invention is shown.

[0068] like Figure 2B As shown, the sequence weight updating method includes operations S21 to S29.

[0069] In operation S21, the weight of the sequence is initialized, which can be determined based on the TF-IDF score of the sequence.

[0070] In operation S22 , a dynamic key table A and a static key table B are constructed. Dynamic key table A can be used to update the importance of sequences in real time, and static key table B can store the weights of all sequences at the beginning of each update cycle to compare the changes in sequence weights with dynamic key table A.

[0071] In operation S23 , based on the user input sequence, the cosine distance between the feature vectors of the dynamic sequence in the dynamic key-value table A and the input sequence is calculated.

[0072] In operation S24, the weights of the sequences whose sampling radius is less than or equal to the distance threshold in the dynamic key-value table A are updated. It can be understood that the distance threshold can also be called a similarity threshold.

[0073] In operation S25 , the distance threshold is updated according to the frequency of occurrence of the input sequence.

[0074] In operation S26, it is detected whether the current round meets the preset sampling period. If it does not meet the preset period, the process returns to operation S23.

[0075] In operation S27 , when the preset period is satisfied, the weight of the dynamic sequence in the dynamic key value table A is updated.

[0076] In operation S28 , the first M sequences of the dynamic key-value table A and the static key-value table B are compared to determine the KV cache pairs that need to be added and replaced, and the distance threshold is reset.

[0077] In operation S29, a candidate information set is determined.

[0078] According to an embodiment of the present invention, in scenarios such as text processing, the frequency of use of sequences will change over time. When the topic of the input information shifts, the value of sequences that were frequently used in the previous period but have not been used for a long time in the current period in the current task or application may gradually decrease. The importance weight of a sequence is dynamically adjusted based on the length of time the sequence has not been used based on an exponential decay mechanism. Cache space is a limited resource. By reducing the importance of sequences that have not been used for a long time, when new hot words with high TF-IDF scores need to be included in the cache, cache space can be allocated more reasonably, ensuring that the cache stores currently more valuable words that are more in line with the needs of users and target tasks, thereby improving cache utilization and effectiveness.

[0079] According to an embodiment of the present invention, the initial dynamic weight is updated using the first similarity and the second similarity to obtain the updated dynamic weight, including: using the first similarity, the second similarity and the similarity threshold to determine the first update value corresponding to the associated sequence, and the second update value corresponding to the remaining sequences; based on the periodic weight of the current sequence in a preset period, the first update value and the second update value, the weight peak value and the weight valley value, determining the updated dynamic weight.

[0080] In an embodiment of the present invention, the first updated value may be a weight increase value corresponding to the associated sequence, and the second updated value may be a weight increase value corresponding to the remaining sequences. The preset period may represent a period set based on the current model inference requirements. The weight values of the sequence may have a preset weight range, where the weight peak value is the maximum value in the weight range, and the weight valley value may be the minimum value in the weight range.

[0081] For example, when a large language model receives a sequence of prompt words from a user, the model processes each sequence one by one. For example, if a user inputs a single sequence, the embedding vector of the sequence is first calculated. Then, the cosine distance between the vector and the embedding vectors of all sequences in the dynamic information set A is calculated. Sequences (tokens) whose cosine distances are within the sampling distance threshold are then combined into an associated information set A'. For each sequence in set A', its weight is updated, and the degree of improvement is negatively correlated with the distance between the embedding vector of the word and the user-input word. This is shown in the following formula (2):

[0082] (2);

[0083] Where Δq is the weight increase (update value) of the sequence, α is the learning rate, which can be used to control the magnitude of the importance increase; d is the cosine distance between the embedding vector of the sequence and the user input sequence. D is the similarity threshold, and the initial value of the similarity threshold can be set based on experience. For example, the initial value of D is D max .

[0084] Furthermore, combined with the above formula (2), the sequence weight determination method can be updated to obtain the updated sequence weight, which can be recorded as q' t+1 , as shown in the following formula (3):

[0085] (3);

[0086] It should be noted that in order to make the weight value within a reasonable range, the weight interval corresponding to the sequence weight can be set to between.

[0087] According to an embodiment of the present invention, determining the similarity between an input feature vector of an input sequence in input information and a dynamic feature vector of a dynamic sequence in a dynamic information set includes: determining a directional relationship feature and respective scale features between the input feature vector and the dynamic feature vector; and determining the similarity based on the directional relationship feature and the scale feature.

[0088] In embodiments of the present invention, the directional relationship feature can characterize the angle and directional relationship between two vectors. The directional relationship feature value is proportional to the directional similarity between the two vectors. The scale feature can characterize the absolute size of the vectors in space. Based on the square root relationship feature and the size feature, a comprehensive measure of the similarity between two vectors in terms of direction and magnitude can be used.

[0089] For example, if the directional feature is the dot product of vectors and the size feature is the modulus of vectors, the cosine similarity formed by combining the dot product of vectors and the modulus can comprehensively measure the similarity of two vectors in terms of direction and size.

[0090] For example, in semantic analysis, the cosine distance between the embedding vectors of all dynamic sequences in the dynamic information set A and the embedding vector of the input sequence is calculated in turn to obtain a set of distance values. These distance values can reflect the degree of similarity between the input sequence and the dynamic sequences in the dynamic information set in the semantic space, and can more accurately reflect the semantic similarity between text or word vectors.

[0091] According to an embodiment of the present invention, the method further includes: determining initial weights of each of a plurality of initial sequences in the initial information set; and constructing an initial candidate information set from sequences corresponding to initial weights greater than or equal to a weight threshold.

[0092] In an embodiment of the present invention, the initial information set may be a basic information set used for the model inference process stored in the system before executing the target task, which may also be referred to as a corpus. The initial weight of the initial sequence may be determined according to a calculation rule.

[0093] For example, the initial sequence is preprocessed (including operations such as word segmentation, stop word removal, stemming or word form merging) to obtain the processed sequence; the frequency of the sequence in each document and the number of documents containing each sequence are counted to obtain the TF-IDF score; and all sequences are sorted in descending order according to the calculated TF-IDF score.

[0094] For example, during the cold start initialization phase of the key-value cache, the sequences in the sequence vocabulary in the corpus can be sorted according to the TF-IDF score. The initial information set can contain a sequence combination consisting of a series of frequently co-occurring words, such as {"big model", "is", "artificial intelligence", "field", "of", "hot", "technology"}. These sequences in the corpus together constitute the initial information set, as shown in the following formula (4):

[0095] (4);

[0096] TF(t', d1) represents the frequency of term t' in document d1, and IDF(t') represents the inverse document frequency of term t'. By calculating the TF-IDF score for each sequence, the M sequences with the highest scores are selected as cold-start key-value pairs and loaded into the KV cache of the graphics memory. The value of M can be dynamically adjusted based on the model scale and graphics memory capacity. For example, in some medium-sized models, M can be set to 1000.

[0097] According to embodiments of the present invention, addressing the high initial inference latency caused by an empty cache during the cold start phase, a method for preloading the cache based on TF-IDF scores can be used to rationally select key data and load it into the cache, even when insufficient historical data is available. This effectively alleviates the performance bottleneck caused by an empty cache during the cold start phase and provides a solid foundation for the model's initial inference. Furthermore, during the cold start phase, the first M sequences are preloaded into the cache, allowing the model to directly retrieve the relevant data for these high-frequency words from the cache during the initial inference, avoiding the need for real-time calculation and loading of this data during the initial inference, thereby significantly reducing the initial inference latency.

[0098] According to an embodiment of the present invention, the initial information set includes an initial sequence; the method also includes: processing the initial information set based on a decomposition strategy to obtain topic information corresponding to the initial information set and a weight matrix corresponding to multiple initial sequences, the weight matrix representing the weights of the multiple initial sequences corresponding to the topic information; determining multiple feature sequences from the multiple initial sequences based on the weight matrix, and constructing an initial candidate information set based on the multiple feature sequences.

[0099] In an embodiment of the present invention, the decomposition strategy may be a method for processing the initial information set through non-negative matrix decomposition. The feature sequence may be a key sequence determined based on each row of information in the first weight matrix. The initial information set may also include initial text, which may be a document in a corpus, each document corresponding to known topic information or category information. The initial sequence may be a sequence in the initial text.

[0100] For example, a word-document matrix can be constructed by counting the frequency of occurrence of each sequence in each document. The rows of the matrix can represent words, the columns represent documents, and the elements represent the number of times a word appears in the document or the value after weighting by TF-IDF, etc.; the number of topics to be extracted can be determined based on prior knowledge or the elbow rule; the non-negative matrix can be initialized, and by setting the number of topics and other parameters (such as the number of iterations, initialization method, etc.), the constructed word-document matrix is used as input, and the non-negative matrix algorithm is run to obtain the decomposed topic-word matrix (W) and document-topic matrix (H); for each topic (row) in the topic-word matrix (W), the words can be sorted from high to low according to their weight values, and the top N words with the highest weights can be selected as multiple feature sequences for the topic, and the initial candidate information set can be constructed using multiple feature sequences.

[0101] According to an embodiment of the present invention, a decomposition strategy can be used to extract feature words important to different categories (topics) from massive sequences, mapping high-dimensional text data to a low-dimensional feature space, which helps to reduce the complexity of the data, highlight key information, and make subsequent data analysis and processing more efficient.

[0102] According to an embodiment of the present invention, the method further includes: updating the similarity threshold based on the frequency of occurrence of the input sequence within a preset period and a threshold interval corresponding to the similarity threshold to obtain an updated similarity threshold.

[0103] In an embodiment of the present invention, within a preset period (e.g., 10 rounds of dialogue), the similarity threshold for determining the associated sequence can be updated according to the frequency of occurrence of the sequence in the user input information, and the value is between [D min , D max ], D min and D max It can be a preset value. The update method of the similarity threshold is shown in the following formula (5):

[0104] (5);

[0105] Among them, D t+1 It can be the updated similarity threshold; D min and D max can be the minimum and maximum values of the similarity threshold respectively; D t It can be the similarity threshold at the previous moment. f can be the frequency of the sequence, and β can be a parameter that controls the sensitivity of the adjustment. When the frequency of the sequence is low, the D value is large, which means that the search range of related words will be expanded; when the frequency of the sequence is high, the D value is small.

[0106] According to an embodiment of the present invention, by dynamically updating the similarity threshold based on the frequency of occurrence of the input sequence, it is possible to avoid storing too many unnecessary similar domains and semantic words in the cache, thereby improving the utilization of system resources.

[0107] According to an embodiment of the present invention, an initial candidate information set is updated based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, including: determining a difference sequence between the first target sequence and the second target sequence, wherein the difference sequence includes a first difference sequence and a second difference sequence, the first difference sequence being a sequence included only in the first target sequence, and the second difference sequence being a sequence included only in the second target sequence; and using the first difference sequence and the second difference sequence to update the initial candidate information set to obtain a candidate information set.

[0108] In an embodiment of the present invention, the difference sequence may be a sequence that is different from the first target sequence and the second target sequence.

[0109] For example, if the dynamic information set is dynamic key table A and the reference information set is static key table B, after every T rounds of dialogue, the sequences in dynamic key table A can be sorted in descending order according to weight, and then the first M words (i.e., sequences) in dynamic key table A and static key table B can be compared. Assume X add In the dynamic key value table A (which can be recorded as A M ), which is not in the static key value table B (which can be recorded as ), that is, , that is, X add is the first difference sequence; X del In the static key value table B (which can be recorded as B M ), which is not in the dynamic key value table A (which can be recorded as ), that is, , that is, X del is the second difference sequence. Obviously, , that is, the number of newly added sequences is equal to the number of sequences to be removed; thus determining the sequence X to be updated and removed add and X del Finally, the dynamic key value table A can be used to overwrite and update the static key value table B.

[0110] Figure 2C A flowchart of a method for updating an initial candidate information set according to an embodiment of the present invention is shown.

[0111] like Figure 2C As shown, the method for updating the initial candidate information set includes operations S201 to S206.

[0112] In operation S201, an initial sequence in an initial information set is initialized. The initial information set may include a sequence combination consisting of a series of words that often appear at the same time.

[0113] In operation S202 , the TF-IDF score corresponding to the initial sequence is determined.

[0114] In operation S203, the sequences are loaded into the KV cache of the graphics card in descending order of TF-IDF scores. For example, the M sequences with the highest scores are selected as the key-value pairs for cold start and loaded into the KV cache of the graphics card.

[0115] In operation S204 , the video memory KV cache is updated according to the weight of the sequence.

[0116] In operation S205, it is determined whether the update round meets the preset period requirement. If the current round does not meet the preset period, the process returns to operation S204 to continue updating the KV cache.

[0117] In operation S206 , when the current round meets the preset period, the KV cache in the system is updated to obtain a target cache (candidate information set).

[0118] According to an embodiment of the present invention, the initial candidate information set is updated using the first difference sequence and the second difference sequence to obtain a candidate information set, including: obtaining address information for the second difference sequence based on the second difference sequence and the current candidate information set at the current moment; and using the address information, storing the first difference sequence in the current candidate information set to obtain the candidate information set.

[0119] In an embodiment of the present invention, the address information may be the storage address information corresponding to the second difference sequence in the initial candidate information set. del Send it to the inference model to determine the X in the KV cache of each layer in the model del The corresponding address; thus X add Enter the model, calculate the KV cache value of each layer, and save it in X del The corresponding address completes the cache update; at the same time, the similarity threshold D of the associated sequence can be updated to D max .

[0120] In another feasible embodiment, regarding updating the video memory, based on the load information in the system, when the load of the inference model is relatively low or relatively idle, the first M sequences in the dynamic key-value table A can all be reloaded into the KV cache of the video memory to replace the original cache.

[0121] According to an embodiment of the present invention, address information for the second difference sequence is obtained based on the second difference sequence and the current candidate information set at the current moment, including: matching the sequence identification information or sequence position information of the second difference sequence with the candidate identification information or candidate position information in the current candidate information set, and obtaining the address information based on the matching result; or obtaining the address information based on a mapping relationship between the sequence identification information and the candidate position information.

[0122] In embodiments of the present invention, the sequence identification information and sequence position information can be used to uniquely identify the second difference sequence. The candidate identification information and candidate position information can be used to uniquely identify the candidate sequence in the candidate information set. The sequence identification information can be matched with the candidate identification information, or the sequence position information can be matched with the candidate position information. If the match result is a correct match, the address information of the second difference sequence in the model can be obtained.

[0123] For example, you can del As input to the model, the model will del The content and current cache status of each layer are searched in the KV cache of each layer. del The matching method may include: matching the sequence identifier or location information in the cache to determine X del A specific location in the cache; alternatively, each sequence in the cache may have a unique identifier that can be compared to X del The identifier is compared with the identifier stored in the cache to find the corresponding address.

[0124] For example, a dictionary (or similar data structure) can be created for each layer of the model to store the mapping relationship between sequence identification information and candidate location information. During model inference, when a new sequence is loaded into the cache, the mapping table can be updated to associate the sequence identification information with the actual storage location information in the cache, thereby enabling the corresponding cache address information to be found in the mapping table through the sequence identifier.

[0125] Figure 3 A flowchart of a task execution method according to an embodiment of the present invention is shown.

[0126] like Figure 3 As shown, the task execution method includes operation S310.

[0127] In operation S310 , in response to receiving a task execution instruction, a target task is processed using a candidate information set; wherein the target task includes any one of text generation, speech recognition, and financial analysis, and the candidate information set is obtained according to the above-mentioned information updating method.

[0128] In an embodiment of the present invention, the task execution instruction may be an instruction generated according to actual application requirements. For example, in a speech recognition task, the target instruction may include speech-to-text instructions, speech translation instructions, and speech instruction recognition instructions. For example, in a text generation task, the target instruction may include text creation instructions, text editing instructions, and text storage instructions. It should be noted that the above application scenarios are not limited to text generation, speech recognition, and financial analysis, and are not specifically limited here.

[0129] Taking the medical question answering system as an example, a large language model is used as the dialogue model. Assuming that there are 10,000 sequence groups in the initial information set (corpus), set M = 1000, that is, the 1000 sequences with the highest TF-IDF scores are loaded into the KV cache during the cold start. At the same time, the similarity threshold D is preset. min =0.1, D max =0.5, λ=0.01, α=0.1, β=0.05, and a preset period T=10 rounds of dialogue. It should be noted that the specific type of the large prediction model in the present invention can be determined according to actual conditions and is not limited here.

[0130] Assume there are 1000 medical documents. Take the sequence "chest pain" as an example. This sequence appears 10 times in document d1. Document d1 has a total of 1000 words. Then the word frequency of this sequence in d1 is Assuming that “chest pain” appears in 200 documents, the inverse document frequency of the sequence is ; Thus, we can get the TF-IDF score of "chest pain" By calculating the TF-IDF of all sequences in the vocabulary, the 1000 sequences with the highest scores are selected as cold-start key-value pairs and loaded into the KV cache of the video memory. At the same time, the TF-IDF scores of these sequences are stored as initial weights in the dynamic key-value table A and the static key-value table B of the weight key-value table, and the embedding vectors of these sequences are recorded.

[0131] Assume that after cold start initialization, the dynamic key value table A and the static key value table B have stored information of 1000 sequence groups. Taking the "chest pain" sequence as an example, its initial weight is .

[0132] Before the user has the first round of conversation, the weights of all sequences remain at their initial values. After the user has the first round of conversation, the weight of “chest pain” begins to decay. According to the formula , at this time t=1, , then the weight of the sequence "chest pain" can be updated to .

[0133] Suppose that in the second round of conversation, the user enters "chest pain." First, we can calculate the cosine distance between the embedding vector (feature vector) of "chest pain" and the embedding vectors of all sequences in the dynamic key-value table A. Suppose there is a sequence in the vocabulary "angina pectoris" whose embedding vector has a cosine distance d = 0.2 with the embedding vector of "chest pain." Therefore, according to the semantic association promotion mechanism, the distance d = 0.2, D = 0.3, and α = 0.1 for "chest pain" itself can be weighted. At the same time, for the associated sequence "angina pectoris" (because d=0.2 is less than the current cosine distance), its weight is increased. According to the formula, . Assume that the original weight of "angina pectoris" is , then the increased weight is .

[0134] Furthermore, the similarity threshold can be updated based on the frequency of the user inputting “chest pain”. Assuming that “chest pain” appeared twice in the previous conversation, according to formula (5), the associated sampling radius for “chest pain” can be obtained as It should be noted that within a cycle, the correlation sampling radius of sequences not input by other users does not change.

[0135] In each subsequent round of conversation within the preset period, the above steps can be repeated to update the weights, increase or decrease the weights based on user input, and update the sampling distance. For example, in the third round of conversation, the weights of all sequences continue to decrease. If the user enters a new word, the above steps are repeated to calculate the weight of the related words and update the sampling distance.

[0136] After the 10th round of dialogue, the sequences in the dynamic key value table A in the weight key value table are sorted according to the weight. Assume that after sorting, 100 of the first 1000 sequences in the dynamic key value table A are not in the first 1000 words in the static key value table B, that is, X add =100; At the same time, 100 of the first 1000 words in the static key-value table B are not in the first 1000 words in the dynamic key-value table A, that is, X del =100.

[0137] You can remove the X del Enter the model and determine the X in the KV cache of each layer of the model del Corresponding address information; thus X add Enter the model, calculate the KV cache value of each layer, and save it in X addThe corresponding address information is used to complete the cache update. After that, the content of the dynamic key-value table A can be copied to the static key-value table B to prepare for the update of the next 10 rounds of dialogue.

[0138] Based on the above information updating method, the present invention also provides an information updating device. Figure 4 The device is described in detail.

[0139] Figure 4 A structural block diagram of an information updating device according to an embodiment of the present invention is shown.

[0140] like Figure 4 As shown, the information updating apparatus of this embodiment includes a similarity determination module 410 , a dynamic information set updating module 420 and a candidate information set updating module 430 .

[0141] Similarity determination module 410 is configured to, in response to receiving input information, determine a similarity between an input feature vector of an input sequence in the input information and a dynamic feature vector of a dynamic sequence in a dynamic information set, where the dynamic sequence has an initial dynamic weight, and the dynamic information set includes multiple information pairs consisting of the initial dynamic weight and the dynamic sequence. In one embodiment, similarity determination module 410 may be configured to perform operation S210 described above, and will not be further described herein.

[0142] Dynamic information set updating module 420 is configured to update the dynamic information set based on the updated dynamic weights obtained by updating the initial dynamic weights using the similarity and the associated sequences, thereby obtaining an updated dynamic information set. The associated sequences are sequences within the dynamic sequence that are semantically associated with the input sequence. In one embodiment, dynamic information set updating module 420 may be configured to perform operation S220 described above, and will not be further described herein.

[0143] Candidate information set updating module 430 is configured to update the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set, thereby obtaining a candidate information set for processing the target task using the candidate information set. The initial candidate information set is the candidate information set before receiving the input information. In one embodiment, candidate information set updating module 430 may be configured to perform operation S230 described above, and will not be further described herein.

[0144] According to an embodiment of the present invention, through the similarity determination module 410, the dynamic information set update module 420 and the candidate information set update module 430 in the information update device, since the updated dynamic weight is flexibly determined based on the similarity between the input feature vector and the dynamic feature vector, the updated dynamic information set obtained based on the updated dynamic weight takes into account the dual influence of the real-time input information and the associated sequence related thereto, thereby utilizing the respective target sequences of the dynamically updated dynamic information set and the static reference information set to update the initial candidate information set in real time, thereby improving the hit rate of the candidate information set in the case of cold start and semantic drift, accelerating the inference speed of the model, and further improving the user experience.

[0145] According to an embodiment of the present invention, the apparatus further comprises: a sequence determination module configured to determine sequences in the dynamic sequence whose similarity is greater than or equal to a similarity threshold as associated sequences, so as to obtain an associated information set using the associated sequences.

[0146] According to an embodiment of the present invention, the similarity includes a first similarity between the associated sequence and the input sequence, and a second similarity between the remaining sequences in the dynamic sequence and the input sequence. The dynamic information set updating module 420 includes a dynamic weight updating submodule and a sequence updating submodule. The dynamic weight updating submodule is configured to update the initial dynamic weight using the first similarity and the second similarity to obtain an updated dynamic weight. The sequence updating submodule is configured to update and sort the dynamic sequence based on the updated dynamic weight, thereby updating the dynamic information set.

[0147] According to an embodiment of the present invention, the dynamic weight update submodule includes: an update value determination unit and a dynamic weight determination unit. The update value determination unit is configured to determine a first update value corresponding to the associated sequence and a second update value corresponding to the remaining sequences using a first similarity, a second similarity, and a similarity threshold; and the dynamic weight determination unit is configured to determine and update the dynamic weight based on the periodic weight of the current sequence in a preset period, the first update value, the second update value, the weight peak value, and the weight valley value.

[0148] According to an embodiment of the present invention, similarity determination module 410 includes a feature determination submodule and a similarity determination submodule. The feature determination submodule is configured to determine the directional relationship features and the scale features of the input feature vector and the dynamic feature vector, and the similarity determination submodule is configured to determine the similarity based on the directional relationship features and the scale features.

[0149] According to an embodiment of the present invention, the apparatus further includes: an initial weight determination module and an information set construction module. The initial weight determination module is configured to determine the initial weights of each of a plurality of initial sequences in the initial information set; and the information set construction module is configured to construct an initial candidate information set from sequences corresponding to sequences whose initial weights are greater than or equal to a weight threshold.

[0150] According to an embodiment of the present invention, an initial information set includes an initial sequence; the apparatus further includes an information processing module and a feature sequence determination module. The information processing module is configured to process the initial information set based on a decomposition strategy to obtain topic information corresponding to the initial information set and a weight matrix corresponding to multiple initial sequences, wherein the weight matrix represents the weights of the multiple initial sequences relative to the topic information; and the feature sequence determination module is configured to determine multiple feature sequences from the multiple initial sequences based on the weight matrix and construct an initial candidate information set based on the multiple feature sequences.

[0151] According to an embodiment of the present invention, the apparatus further comprises: a threshold refinement module configured to update the similarity threshold based on the frequency of occurrence of the input sequence within a preset period and a threshold interval corresponding to the similarity threshold to obtain an updated similarity threshold.

[0152] According to an embodiment of the present invention, the candidate information set updating module 430 includes: a difference sequence determining module and a candidate information set updating module. The difference sequence determining module is configured to determine a difference sequence between a first target sequence and a second target sequence; and the candidate information set updating module is configured to update an initial candidate information set using the first difference sequence and the second difference sequence to obtain a candidate information set.

[0153] According to an embodiment of the present invention, the candidate information set update module includes an address information determination submodule and a storage submodule. The address information determination submodule is configured to obtain address information for the second difference sequence based on the second difference sequence and the current candidate information set at the current moment; the storage submodule is configured to use the address information to store the first difference sequence in the current candidate information set to obtain the candidate information set.

[0154] According to an embodiment of the present invention, the address information determination submodule includes: a matching unit and a mapping unit. The matching unit is configured to match the sequence identification information or sequence position information of the second difference sequence with the candidate identification information and candidate position information in the current candidate information set, and obtain the address information based on the matching result; or the mapping unit is configured to obtain the address information based on the mapping relationship between the sequence identification information and the candidate position information.

[0155] Figure 5 A structural block diagram of a task execution device according to an embodiment of the present invention is shown.

[0156] like Figure 5As shown, the task execution apparatus of this embodiment includes a task processing module 510. The task processing module is configured to, in response to receiving a task execution instruction, process a target task using a candidate information set; the target task includes any one of text generation, speech recognition, and financial analysis, and the candidate information set is obtained according to the task execution method described above. In one embodiment, task processing module 510 can be configured to perform operation S310 described above, and will not be further described here.

[0157] According to an embodiment of the present invention, any multiple modules among the similarity determination module 410, the dynamic information set update module 420, and the candidate information set update module 430 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present invention, at least one of the similarity determination module 410, the dynamic information set update module 420, and the candidate information set update module 430 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the similarity determination module 410, the dynamic information set update module 420, and the candidate information set update module 430 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality. Similarly, this embodiment can also be applied to the task processing module 510, and the details will not be repeated here.

[0158] Figure 6 A block diagram of an electronic device according to an embodiment of the present invention is shown.

[0159] like Figure 6 As shown, the electronic device includes: a memory 610 and a processor 620. The processor 620 is configured to execute the above-mentioned information updating method and / or task execution method according to the instructions and data stored in the memory.

[0160] In an embodiment of the present invention, processor 620 may be a server that provides various services. Instructions include, but are not limited to, instructions related to text generation tasks, speech recognition tasks, and financial analysis tasks. Data may include, but is not limited to, dynamic information sets, associated information sets, and reference information sets.

[0161] Figure 7A block diagram of an electronic device suitable for implementing an information updating method according to an embodiment of the present invention is shown.

[0162] like Figure 7 As shown, an electronic device according to an embodiment of the present invention includes a first processor 701, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 702 or a program loaded from a storage unit 708 into a random access memory (RAM) 703. The first processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)). The first processor 701 may also include onboard memory for caching purposes. The first processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0163] Various programs and data required for the operation of the electronic device are stored in RAM 703. The first processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The first processor 701 executes the programs in ROM 702 and / or RAM 703 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than ROM 702 and RAM 703. The first processor 701 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.

[0164] According to an embodiment of the present invention, the electronic device may further include an input / output (I / O) interface 705, which is also connected to the bus 704. The electronic device may further include one or more of the following components connected to the I / O interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage portion 708 including a hard disk; and a communication portion 709 including a network interface card such as a LAN card or a modem. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. Removable media 711, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 710 as needed, so that computer programs read from the removable media can be installed in the storage portion 708 as needed.

[0165] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0166] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above, and / or one or more memories other than ROM 702 and RAM 703.

[0167] The embodiments of the present invention further include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to enable the computer system to implement the information update method provided by the embodiments of the present invention.

[0168] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when the first processor 701 executes the computer program. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0169] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 709, and / or installed from a removable medium 711. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0170] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709 and / or installed from the removable medium 711. When the computer program is executed by the first processor 701, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0171] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0173] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.

[0174] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. An information updating method, characterized in that: The method comprises: In response to receiving input information, determining a similarity between an input feature vector of an input sequence in the input information and a dynamic feature vector of a dynamic sequence in a dynamic information set, wherein the dynamic sequence has an initial dynamic weight, and the dynamic information set includes a plurality of information pairs consisting of the initial dynamic weight and the dynamic sequence; updating the dynamic information set based on an updated dynamic weight and an associated sequence obtained by updating the initial dynamic weight using the similarity, to obtain an updated dynamic information set, wherein the associated sequence is a sequence in the dynamic sequence that is semantically associated with the input sequence, and the similarity includes a first similarity between the associated sequence and the input sequence, and a second similarity between a remaining sequence in the dynamic sequence and the input sequence; updating an initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain a candidate information set, so as to process a target task using the candidate information set, wherein the initial candidate information set is the candidate information set before receiving the input information; The updating of the dynamic information set based on the updated dynamic weight and the association sequence obtained by updating the initial dynamic weight using the similarity to obtain the updated dynamic information set includes: Updating the initial dynamic weight using the first similarity and the second similarity to obtain an updated dynamic weight; updating and sorting the dynamic sequence based on the updated dynamic weight to obtain the updated dynamic information set; Updating the initial candidate information set based on the first target sequence in the updated dynamic information set and the second target sequence in the reference information set to obtain the candidate information set includes: Determining a difference sequence between the first target sequence and the second target sequence, wherein the difference sequence includes a first difference sequence and a second difference sequence, the first difference sequence is a sequence included only in the first target sequence, and the second difference sequence is a sequence included only in the second target sequence; The initial candidate information set is updated using the first difference sequence and the second difference sequence to obtain the candidate information set.

2. The method according to claim 1, characterized in that The method further comprises: The sequences in the dynamic sequence whose similarity is greater than or equal to a similarity threshold are determined as the associated sequences, so as to obtain an associated information set using the associated sequences.

3. The method according to claim 1, characterized in that Updating the initial dynamic weight using the first similarity and the second similarity to obtain an updated dynamic weight includes: Determine a first update value corresponding to the associated sequence and a second update value corresponding to the remaining sequence by using the first similarity, the second similarity, and the similarity threshold; The updated dynamic weight is determined based on the period weight of the current sequence in a preset period, the first updated value, the second updated value, the weight peak value, and the weight valley value.

4. The method according to claim 1, wherein Determining the similarity between the input feature vector of the input sequence in the input information and the dynamic feature vector of the dynamic sequence in the dynamic information set includes: Determining a directional relationship feature between the input feature vector and the dynamic feature vector and respective scale features; The similarity is determined based on the directional relationship feature and the scale feature.

5. The method according to claim 1, wherein The method further comprises: Determine the initial weights of the multiple initial sequences in the initial information set; The initial candidate information set is constructed by using sequences corresponding to initial weights greater than or equal to a weight threshold.

6. The method according to claim 5, characterized in that The initial information set includes a plurality of initial sequences; the method further includes: Processing the initial information set based on a decomposition strategy to obtain topic information corresponding to the initial information set and a weight matrix corresponding to the multiple initial sequences, wherein the weight matrix represents weights corresponding to the multiple initial sequences and the topic information; A plurality of feature sequences are determined from a plurality of initial sequences based on the weight matrix, and the initial candidate information set is constructed based on the plurality of feature sequences.

7. The method according to claim 3, characterized in that The method further comprises: The similarity threshold is updated based on the frequency of occurrence of the input sequence within the preset period and a threshold interval corresponding to the similarity threshold to obtain an updated similarity threshold.

8. The method according to claim 1, characterized in that Updating the initial candidate information set using the first difference sequence and the second difference sequence to obtain the candidate information set includes: Obtaining address information for the second difference sequence based on the second difference sequence and a current candidate information set at a current moment; The first difference sequence is stored in the current candidate information set using the address information to obtain the candidate information set.

9. The method according to claim 1, characterized in that Obtaining address information for the second difference sequence based on the second difference sequence and a current candidate information set at a current moment, including: Matching the sequence identification information or sequence position information of the second difference sequence with the candidate identification information or candidate position information in the current candidate information set, and obtaining the address information based on the matching result; or The address information is obtained based on a mapping relationship between the sequence identification information and the candidate location information.

10. A task execution method, characterized in that: The method comprises: In response to receiving the task execution instruction, processing the target task using the candidate information set; The target task includes any one of text generation, speech recognition and financial analysis, and the candidate information set is obtained according to the method of any one of claims 1 to 9.

11. An electronic device, characterized in that: include: Memory; A processor configured to execute the method according to any one of claims 1 to 10 according to the instructions and data stored in the memory.

12. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

13. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Chinese academic keyword extraction method and device and storage medium

    CN113268995A

  • Multi-task parallel processing method and system based on AI target identification

    CN119960946A