Data processing method and electronic equipment

By pre-storing and obtaining embedded vectors of prefix information in the large language model, the problem of high computational cost in processing duplicate prefix information is solved, and more efficient processing is achieved.

CN120011412AActive Publication Date: 2025-05-16BEIJING FEISHU TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510487204.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

When processing prompt information containing duplicate prefix information, the generative large language model has too high calculation cost, resulting in inefficient processing.

Method used

The embedding vector that pre-stores prefix information, and directly obtains the stored embedding vector when processing prompt information, avoiding repeated processing of prefix information.

Benefits of technology

By avoiding repeated processing of prefix information, the calculation cost is reduced and the processing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011412A_ABST
    Figure CN120011412A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing method and electronic equipment. The method comprises the following steps: receiving a cache request from a client, wherein the cache request comprises prefix information; providing a first response to the client, the first response indicating that the embedded vector of the prefix information has been stored in the vector database; receiving a processing request from the client, wherein the processing request comprises prompt information and the prompt information has the prefix information; obtaining an embedded vector of prefix information pre-stored in a vector database; and based on the embedded vector of the prefix information and other information except the prefix information in the prompt information, utilizing the large model to obtain model output. Through the scheme, when the large model processes the prompt information, repeated processing of the prefix information can be avoided, so that the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of computers, and more particularly to data processing methods and electronic devices. Background Art

[0002] Generative large language models are deep learning models that can generate natural language text. These models are based on complex neural network architectures, such as transformers, and are pre-trained on large amounts of text data to learn the statistical properties and patterns of language.

[0003] Users can call the large language model through prompt information, such as prompt words. Users can input requests including prompt words through the client application interface (API) to guide the large language model to reply. Prompt words in different requests may have repeated sequences, and the large language model's processing of prompt words each time will result in excessive computational costs. Summary of the invention

[0004] According to an example embodiment of the present disclosure, a data processing scheme is provided, in which an embedding vector of prefix information can be pre-stored, so that when a large language model processes prompt information, repeated processing of prefix information can be avoided, thereby reducing costs.

[0005] In a first aspect of the present disclosure, a data processing method is provided. The method includes: receiving a cache request from a client, the cache request including prefix information; providing a first response to the client, the first response indicating that an embedding vector of the prefix information has been stored in a vector database; receiving a processing request from the client, the processing request including prompt information and the prompt information having the prefix information; obtaining an embedding vector of the prefix information pre-stored in the vector database; and obtaining a model output using a large model based on the embedding vector of the prefix information and other information in the prompt information except the prefix information.

[0006] In a second aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method described in the first aspect of the present disclosure.

[0007] In a third aspect of the present disclosure, a data processing device is provided. The device includes: a receiving unit configured to receive a cache request from a client, the cache request including prefix information; a providing unit configured to provide a first response to the client, the first response indicating that the embedding vector of the prefix information has been stored in a vector database; the receiving unit is also configured to receive a processing request from the client, the processing request including prompt information, and the prompt information has prefix information; a vector acquisition unit configured to acquire the embedding vector of the prefix information pre-stored in the vector database; and an output unit configured to obtain a model output using a large model based on the embedding vector of the prefix information and other information in the prompt information except the prefix information.

[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has machine-executable instructions stored thereon, and when the machine-executable instructions are executed by a device, the device executes the method described in the first aspect of the present disclosure.

[0009] In a fifth aspect of the present disclosure, a computer program product is provided, comprising computer executable instructions, wherein the computer executable instructions implement the method described according to the first aspect of the present disclosure when executed by a processor.

[0010] In a sixth aspect of the present disclosure, an electronic device is provided, comprising: a processing circuit configured to execute the method described according to the first aspect of the present disclosure.

[0011] The invention summary is provided to introduce a series of concepts in a simplified form, which will be further described in the specific embodiments below. The invention summary is not intended to identify the key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein: Figure 1 A schematic block diagram of a system according to some embodiments of the present disclosure is shown; Figure 2 A flowchart illustrating an example process according to some embodiments of the present disclosure; Figure 3 A schematic diagram showing a communication process according to some embodiments of the present disclosure is shown; Figure 4 A block diagram showing an example apparatus according to some embodiments of the present disclosure; and Figure 5 A block diagram of an example device that may be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0014] In the embodiments of the present disclosure, the large model may be referred to as a large language model (LLM), a generative large language model, a large language model based on a transformer architecture, etc., and the present disclosure is not limited thereto. The large model in the embodiments of the present disclosure may be applied to a variety of application scenarios, such as writing articles, generating dialogues, creating poems, generating content for social media / news, etc., generating code / annotations / documents, generating teaching materials, assisting language learning, etc.

[0015] Large language models can be trained with large amounts of text data to be able to understand and generate natural language text. Large language models can be based on transformer architectures, which can effectively handle long-distance dependencies through self-attention mechanisms and have excellent performance in parallel computing. Large language models can include multiple transformation layers, for example, a transformation layer can be an encoder layer or a decoder layer. Transformation layers can include self-attention layers, feedforward neural network (FNN) layers, and layer normalization (LN) layers.

[0016] Users can call the big model by inputting prompt information through the API of client devices (such as mobile phones, tablets, desktops, wearable devices, etc.). API is a set of predefined rules and protocols for building and interacting with interfaces between software applications. API defines the data format of requests and the structure of responses, so that different applications can understand and exchange data with each other. Through API, applications can provide their functions or services for other applications to use without sharing their source code. Developers can use API to integrate third-party services into their own applications or expand the functionality of applications.

[0017] Exemplarily, the components of an API may include endpoints, requests, and responses. An endpoint may be a URL address that can be accessed in an API. A request may include data sent by a user to an API, including methods (such as GET, POST, PUT, DELETE), paths, parameters, and a body. A response may be data returned by an API to a client, including a status code, header information, and a body.

[0018] Specifically, the user can input prompt information through the API, such as a string that the user wants to ask. The API can also implement the maximum number and format of reply word units, etc. Further, the large model can generate a reply string based on the prompt information and provide it to the client device.

[0019] The running cost of large models is related to the number of word-units. Large models based on the transformer architecture can use the full attention mechanism to consider the relevance of all other word-units when processing each word-unit. For example, each word-unit needs to interact with other word-units to calculate the attention weight. The computational complexity of this self-attention mechanism is the square of the number of word-units. When the input sequence is long, the computational cost of full attention is too high. The more word-units there are, the greater the computational workload and the higher the running cost. In addition, each word-unit has its own vector representation in the model, and these representations need to be stored. In fact, the number of word-units is one of the important indicators for measuring the processing workload of the model. Therefore, when processing more word-units, more storage resources are required, more computing resources are consumed, and it takes longer time.

[0020] There may be recurring text sequences in the user requests input into the big model. For example, in the disclosed embodiments, the recurring text sequences may be referred to as "common prefixes" or "prefix information". For example, in a question-and-answer system, the questions that a user may input each time may begin with phrases such as "Excuse me", "Please tell me", "I want to know", etc. For another example, the context of some questions may be the same long text sequence, such as a question-and-answer system for novel A, then all the text content of the novel A can be considered as a common prefix. In this case, if the big model processes the prompt information completely each time, it will lead to repeated processing of the common prefix, resulting in problems such as high computational cost and low processing efficiency.

[0021] In order to solve the above problems and other potential problems, the embodiments of the present disclosure provide a data processing solution for a large model. In this solution, the embedding vector of the prefix information can be stored, so when processing the prompt information, the embedded vector of the stored prefix information can be obtained to avoid repeated processing of the prefix information, which can improve processing efficiency.

[0022] Figure 1 1 shows a schematic block diagram of a system 100 according to some embodiments of the present disclosure. Figure 1 As shown, the system 100 includes a client 110 and a server 120 , wherein the server 120 may include a large model 122 and a vector database 124 .

[0023] In an embodiment of the present disclosure, a user may use the client 110 to access the vector database 124 via an API interface between the client 110 and the server 120. For example, a user may obtain permission to the vector database 124 through subscription, etc. For example, the size, period, access rate, etc. of the vector database 124 that a user can access depends on the user's permission. For example, a user may apply for access to the vector database 124 based on his own cost control strategy, etc.

[0024] Exemplarily, the client 110 may be implemented as an electronic device on the user side, such as a mobile phone, a tablet computer, a desktop computer, a wearable device, a vehicle-mounted device, etc.

[0025] For example, the server 120 may be implemented as a remote device or may be located in the cloud, for example, the server 120 may be implemented as an electronic device or a distributed cluster, etc. It should be understood that although Figure 1 The server 120 is shown to include a large model 122 and a vector database 124, but in some embodiments, the server 120 may be implemented in a distributed manner, for example, the server 120, the large model 122, and the vector database 124 may be physically located in different locations.

[0026] The embedding vectors of the prefix information may be stored in the vector database 124. It is understandable that the vector database 124 may store multiple embedding vectors of multiple prefix information, for example, the embedding vectors of "Excuse me", "Please tell me", "I want to know" and the like in the question-answering system are all stored in the vector database 124. Exemplarily, the multiple embedding vectors of the multiple prefix information may be stored based on a cache request from a user (from the client 110).

[0027] The prefix information and embedded vectors stored in the vector database 124 may also have corresponding cache priority information. As an example, the stored information may be as shown in the following Table 1. It should be noted that different prefix information may have the same or different cache priorities, depending on the user's settings.

[0028]

[0029] Exemplarily, different indexes may be assigned to different prefix information and stored together with the embedded vector. As an example, the stored information may be as shown in Table 2 below. It should be noted that the indexes of different prefix information are different, and the assignment of indexes is not necessarily continuous.

[0030]

[0031] It is understandable that the content and format of the stored information shown in Table 1 or Table 2 above are only for illustration and do not constitute a limitation on the embodiments of the present disclosure. Optionally, the index information and cache priority information can be combined into one field for storage. For example, "001010" is stored corresponding to prefix information 1, the first three bits "001" indicate that the index is 1, and the last three bits "010" indicate that the cache priority is level 2. The present disclosure does not limit this.

[0032] It is understandable that the vector database may also be referred to as other names, such as storage module, memory, cache area, or similar terms, and the present disclosure is not limited to this.

[0033] The user can input prompt information through the client 110. The prompt information input by the user may include prefix information and other information. Then the large model 122 can obtain the embedding vector of the prefix information from the vector database 124, and then perform subsequent processing together with other information to obtain the output of the model. Exemplarily, the model output can be provided or displayed to the user via the client 110. In this way, at least the process of determining the embedding vector of the prefix information can be omitted, saving the calculation cost.

[0034] Figure 2 A flow chart of a processing process 200 according to an example embodiment of the present disclosure is shown. In box 210, a cache request is received from a client, the cache request includes prefix information. In box 220, a first response is provided to the client, the first response indicating that an embedding vector of the prefix information has been stored in a vector database. In box 230, a processing request is received from a client, the processing request includes prompt information and the prompt information has the prefix information. In box 240, an embedding vector of the prefix information pre-stored in the vector database is obtained. In box 250, a model output is obtained using a large model based on the embedding vector of the prefix information and other information in the prompt information except the prefix information.

[0035] Exemplarily, the prompt information includes a plurality of tokens. A token may be the smallest processing unit of input text and output text in natural language processing and a large language model. For example, a token may be a character, a word, a subword, a special symbol, etc. Exemplarily, when processing text, the large model may perform a tokenization process on the text, thereby segmenting a string in the text into a series of tokens. For example, the prompt information is input into the large model, and the large model may segment the prompt information into a plurality of tokens through a tokenization process.

[0036] The prefix information may include one or more than one word element. The prefix information may be part of the prompt information, and different prompt information may have the same or different prefix information. Exemplarily, different prompt information may include the same prefix information, i.e., a common prefix. The common prefix may be used to set the context of the conversation, and the large model may understand the intention of the following conversation or text based on the common prefix. Based on the common prefix, the large model may be guided to provide a specific type of output (i.e., answer), such as a question-like, imperative, or explanatory one. In the embodiments of the present disclosure, as combined with Figure 1 As shown, if different prompt information has a common prefix, it is possible to avoid repeatedly processing the entire prompt information, thereby saving computing resources.

[0037] In some implementations, the large model is a large language model based on the transformer architecture, and the embedded vector of the stored prefix information can be calculated using a key-value cache (KV cache). The key-value pair can include a key (Key) and a value (Value) of the attention mechanism, and the cache can refer to storing two mathematical vectors of the key and the value in memory or hard disk.

[0038] In an embodiment of the present disclosure, the server 120 may provide an API of a large model so that a user may obtain permission to access the large model, such as obtaining permission to use and control the vector database 124. Exemplarily, a user may send an API request through the client 110. The server 120 may determine the vector database accessible to the client 110 based on the API request. Optionally, the API request may include an API key, an access method, a request body, and the like. For example, the access method may be a Hypertext Transfer Protocol (HTTP) method, and the like. Exemplarily, thereafter, the user may use the client to use the large model through the API. Optionally, the API may also be referred to as an API endpoint, such as a URL address.

[0039] The process of the embodiment of the present disclosure may include two stages: a first stage and a second stage. In some implementations, the first stage may be used to confirm whether the embedding vector of the prefix information has been cached, and the second stage may be used to request the large model to provide a model output corresponding to the prompt information.

[0040] In the first stage, a cache request from a client may be received, and the cache request may include cache control parameters. Exemplarily, the cache control parameters may include prefix information. Optionally, the cache control parameters may further include cache priority information. For example, the cache request may include prefix information (e.g., text XXXXX) requested to be stored and a cache priority (e.g., Lx) corresponding to the prefix information. Optionally, the cache request may include a first execution code to indicate the type of the cache request. Optionally, the first execution code may indicate that the request is a request for the first stage. For example, the first execution code may be 0, indicating that the request is a cache request for the first stage.

[0041] In some examples, based on the cache request, the large model can be used to determine the embedding vector of the prefix information, and the embedding vector of the prefix information can be stored in the vector database. Alternatively, a key-value pair cache based on the transformer large model can be used, for example, a forward propagation algorithm of the transformer neural network can be called to determine the embedding vector of the prefix information.

[0042] In the disclosed embodiments, cache priority information may indicate the length of time that prefix information is stored. Exemplarily, assuming that N levels of cache priority are preset, N different storage duration thresholds may correspond. The correspondence or mapping relationship between cache priority and storage duration threshold may be predefined. For prefix information of a certain cache priority, if the storage time length of the prefix information has reached the storage duration threshold corresponding to the cache priority, the embedded vector of the prefix information may be cleared from the cache. In this way, long-term invalid occupation of the cache can be avoided, and sufficient storage space can be ensured for the embedded vectors of other prefix information, thereby improving cache utilization.

[0043] In an embodiment of the present disclosure, cache priority information may indicate the order of removal from the cache. Exemplarily, if the occupancy rate of the cache has exceeded a threshold value (e.g., 90% or other value) or the idle rate is lower than a threshold value (e.g., 10% or other value), the embedded vector of the prefix information with a lower cache priority may be cleared from the cache. In this way, it is possible to ensure that the prefix information with a high cache priority is retained in the cache for as long as possible, avoiding repeated processing of the prefix information with a high cache priority.

[0044] In an embodiment of the present disclosure, the embedded vector of one or more prefix information may be cleared in combination with one or more of the following factors: available storage space of the vector database, length of time stored, cache priority information, size of the occupied storage space, frequency of being called or accessed within a historical predetermined period, etc. For example, among multiple prefix information with the lowest cache priority, the prefix information with the earliest last access time may be removed first.

[0045] In other examples, after receiving a cache request, it can be determined whether the embedded vector of the prefix information has been stored, that is, whether the embedded vector of the prefix information is stored (that is, whether it exists) in the vector database. If the embedded vector of the prefix information does not exist in the vector database, the embedded vector can be determined and stored through the above process. If the embedded vector of the prefix information already exists in the vector database, there is no need to repeat the above process, which can save processing power and improve efficiency. For example, based on historical cache requests, the vector database has already stored the embedded vector of the prefix information; but the user does not know whether it has been cleared, so the user can initiate a cache request again to ensure that the embedded vector of the prefix information is included in the vector database.

[0046] In other examples, if the embedded vector of the prefix information does not exist in the vector database, and the vector database does not have enough space to store the embedded vector of the prefix information, then it can be further determined whether to store the embedded vector of the prefix information. Exemplarily, if the cache priority of another prefix information already in the vector database is lower than the cache priority of the prefix information, the embedded vector of the other prefix information with a lower cache priority can be removed, and the embedded vector of the prefix information is stored. Exemplarily, if the cache priority of all prefix information already in the vector database is higher than the cache priority of the prefix information, the embedded vector of the prefix information is not stored.

[0047] In some examples, a first response to a cache request may be sent to the client. The first response may indicate whether an embedded vector of the prefix information exists in the vector database. Optionally, the first response may include a cache status code to indicate whether the embedded vector of the prefix information has been stored in the vector database. For example, the cache status code may be a Boolean value, such as 0 or 1, where 1 indicates that it has been stored and 0 indicates that it has not been stored. Optionally, the first response may also include a status code to indicate whether the cache request is successful.

[0048] In the second stage, a processing request from the client may be received, and the processing request may include prompt information. Optionally, the processing request may include a second execution code to indicate the type of the processing request. Optionally, the second execution code may indicate that the request is a request of the second stage. For example, the second execution code may be 1, indicating that the request is a processing request of the second stage. It is understandable that the second execution code is different from the first execution code.

[0049] Exemplarily, after receiving the aforementioned first response, the client 110 may determine whether to send a processing request. For example, the client 110 may determine whether to execute the second phase based on the user's cost control strategy.

[0050] Exemplarily, after receiving the processing request, the prefix information in the prompt information can be extracted or determined. Then, the embedded vector of the prefix information can be obtained or extracted from the vector database. In this way, the prefix information in the prompt information does not need to be processed by the large model, and the stored embedded vector is directly obtained from the vector database, which saves computing resources and improves the processing effect.

[0051] Furthermore, a model output corresponding to the prompt information can be obtained through the large model, and the model output can be sent to the client 110. For example, a second response to the processing request can be sent to the client 110, and the second response can include the model output. For example, the model output includes a reply or answer to the input text.

[0052] In addition, optionally, the cache request or processing request in the embodiment of the present disclosure may have a corresponding request priority. For example, a request with a higher request priority may be processed faster to avoid excessive delays. It is also understandable that the server 120 in the embodiment of the present disclosure may perform reasonable and effective resource management, for example, in the process of receiving and processing requests, it is determined that the resources of the server are effectively managed to meet the quality of service requirements corresponding to the request, wherein the resources of the server may include a central processing unit (CPU), a graphics processing unit (GPU), memory, etc.

[0053] Figure 3 A schematic diagram of a communication process 300 according to some embodiments of the present disclosure is shown. The process 300 may involve a client 110 and a server 120, wherein the client 110 may be operated or used by a user.

[0054] At 302, the client 110 may send an API request to request the establishment of the API endpoint 115. For example, the user may input the API request through a browser or an application. For example, the API request may include an API key, an HTTP method, etc. For example, the API component may be integrated into the application of the client 110 through the API request, so as to facilitate the user to use the large model through the application.

[0055] The API component may include an API endpoint, a request, and a response. The API endpoint may be one or more and may be used to receive a request from the client 110 and return a response.

[0056] The client 110 may send a request to the API endpoint, such as the aforementioned cache request, processing request, etc. The request may include an API key for identity authentication. The request may include an HTTP method to indicate that the request is a POST request, that is, the data in the request (cache request, processing request) needs to be transmitted to the server. The request may include a request body to indicate data in a predetermined format. The request body may include indication information of the large model, cache control parameters, prompt information, first / second execution codes, etc. For example, the indication information of the large model may include model parameters, such as the name of the model, version number, etc. In this way, when the server 120 is associated with multiple large models, it can quickly locate a specific large model 122 based on the model parameters.

[0057] The server 120 may return a response to the client 110, such as the aforementioned first response, second response, etc. The response may include a status code to indicate whether the request is successful. The response may include a cache status code to indicate whether the embedded vector of the prefix information is stored in the vector database. The response may include a response body to carry the data to be returned. For example, the response body of the second response includes the model output.

[0058] At 304, the client 110 may perform a first phase of communication. The client 110 may send a first phase request (such as the aforementioned cache request), the request body of which may include cache control parameters and an execution code. The cache control parameters may include a common prefix, such as in the form of a string. The cache control parameters may also include cache priority information, such as a positive integer from 1 to 5. The execution code may be a first execution code, such as 0, indicating that the request is a first phase of communication. Optionally, the common prefix in the cache request may be user-defined.

[0059] At 306, the API endpoint 115 may send the first-stage request to the server 120. The server 120 may call the large model 122 at 308, and use the key-value cache to accelerate the calculation, thereby obtaining the embedding vector of the common prefix at 310. Subsequently, at 312, the embedding vector of the common prefix is ​​stored in the vector database 124.

[0060] At 314, the server 120 may return a first-stage response, such as the aforementioned first response. The response may include a status code indicating that the request is successful. The response also includes a cache status code, such as a Boolean value of 1, indicating that the embedding vector of the common prefix has been stored in the vector database 124. Optionally, the first response may also indicate that the communication between the client 110 and the server 120 has been established or connected.

[0061] Exemplarily, steps 304 to 314 may be repeatedly performed to store more embedding vectors of common prefixes in the vector database 124 .

[0062] At 320, the client 110 may perform the second phase of communication. The client 110 may send a second phase request (such as the aforementioned processing request), and the request body of the request may include prompt information and an execution code. The execution code may be a second execution code, such as 1, indicating that the request is the second phase of communication. For example, the client 110 may switch the execution code from 0 to 1 to initiate the second phase of communication.

[0063] At 322, the API endpoint 115 may send a second-stage request to the server 120. The server 120 may then obtain the embedding vectors of the common prefixes in the vector database 124, and further utilize the large model 122 to obtain a model output corresponding to the prompt information.

[0064] At 324, the server 120 may return a second-stage response, such as the second response described above. The response may include a status code indicating that the request is successful. The response may also include a response body, such as a model output.

[0065] In addition, in 330, the cache in the vector database 124 can be managed. Specifically, based on the storage time threshold associated with the cache priority, the embedded vector of the prefix information whose storage time reaches the corresponding threshold can be removed. Alternatively, when there is a storage demand for high-priority prefix information, the embedded vector of the low-priority prefix information can be removed to store the embedded vector of the high-priority prefix information. It should be understood that the embodiments of the present disclosure do not limit the specific management method of the vector database, for example, it can be implemented using a hash table, an ordered set, etc.

[0066] In this way, in the embodiment of the present disclosure, the embedded vector of the prefix information of the prompt information is pre-stored in the vector database, and when the prompt information is processed, the prefix information can no longer be repeatedly processed. This can reduce the number of processed word units, achieve the length control of the processed text, effectively control the use cost of the large model, and achieve more efficient use of the large model.

[0067] It should be noted that the above describes the main embodiments of the present disclosure, but the present disclosure does not limit the specific implementation details. For example, during the communication process, the requested data packet can be sent to the Internet through the Internet service provider's network. For example, in order to speed up the content transmission, the request can be distributed through the content distribution network. For example, the server 120 shown in the figure can be implemented as a model server, Figure 3 The API endpoint 115 in can be implemented as an API server. Exemplarily, a load balancer can be used to evenly distribute requests to different servers. For example, if there are multiple model servers, requests can be distributed based on the current load conditions, health conditions, and configured policies of the multiple model servers. For example, the number of servers can be automatically increased or decreased based on the load conditions. For example, the available instances of multiple model servers can be determined through service discovery. For example, requests can be routed to the available instances of multiple model servers as needed.

[0068] It should be understood that in the embodiments of the present disclosure, "first", "second", "third", etc. are only used to indicate that multiple objects may be different, but at the same time do not exclude that two objects are the same, and should not be interpreted as any limitation on the embodiments of the present disclosure.

[0069] It should also be understood that the methods, situations, categories and divisions of the embodiments in the present disclosure are only for the convenience of description and should not constitute special limitations. The features in various methods, categories, situations and embodiments can be combined with each other when it is logical.

[0070] It should also be understood that the above content is only to help those skilled in the art better understand the embodiments of the present disclosure, rather than to limit the scope of the embodiments of the present disclosure. Those skilled in the art may make various modifications, changes or combinations based on the above content. Such modifications, changes or combinations are also within the scope of the embodiments of the present disclosure.

[0071] It should also be understood that the description of the above content focuses on emphasizing the differences between the various embodiments, and the same or similar points can be referenced or borrowed from each other. For the sake of brevity, they will not be repeated here.

[0072] Figure 44 shows a schematic block diagram of an example device 400 according to some embodiments of the present disclosure. The device 400 may be implemented in software, hardware, or a combination of both. Figure 4 As shown, the apparatus 400 includes a receiving unit 410 , a providing unit 420 , a vector acquiring unit 430 and an output unit 440 .

[0073] The receiving unit 410 is configured to receive a cache request from a client, the cache request including prefix information. The providing unit 420 is configured to provide a first response to the client, the first response indicating that the embedding vector of the prefix information has been stored in the vector database. The receiving unit 410 is also configured to receive a processing request from the client, the processing request including prompt information, and the prompt information has prefix information. The vector acquisition unit 430 is configured to acquire the embedding vector of the prefix information pre-stored in the vector database. The output unit 440 is configured to obtain a model output based on the embedding vector of the prefix information and other information in the prompt information except the prefix information, using a large model.

[0074] Exemplarily, the apparatus 400 may include a determining unit configured to determine, based on a cache request, whether an embedded vector of prefix information is stored in a vector database; and generate a first response if it is determined that the embedded vector exists in the vector database.

[0075] Exemplarily, the apparatus 400 may include a determining unit configured to determine whether an embedding vector of the prefix information is stored in the vector database based on a cache request; and if it is determined that the embedding vector does not exist in the vector database, determine the embedding vector of the prefix information using a large model.

[0076] Optionally, the cache request also includes cache priority information corresponding to the prefix information.

[0077] Exemplarily, the apparatus 400 may include a clearing unit. The determining unit is configured to: determine a storage duration threshold corresponding to the cache priority information. The clearing unit is configured to: clear the embedded vector of the prefix information from the vector database if the storage time length of the embedded vector of the prefix information in the vector database reaches the storage duration threshold.

[0078] Optionally, the size of the vector database is associated with the user authority of the client. Optionally, the cache request includes a first execution code for indicating the type of the cache request. Optionally, the processing request includes a second execution code for indicating the type of the processing request.

[0079] Exemplarily, the device 400 may include a clearing unit configured to: clear the stored embedded vector of the prefix information based on at least one of the following items: available storage space in the vector database, cache priority information of the prefix information, the length of time the embedded vector has been stored, or the frequency with which the embedded vector has been called within a historical predetermined period of time.

[0080] Figure 4 The device 400 can be used to achieve the above combination Figures 1 to 3 For the sake of brevity, the process will not be described in detail here.

[0081] The division of modules or units in the embodiments of the present disclosure is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in the disclosed embodiments may be integrated into one unit, or may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0082] Figure 5 1 shows a block diagram of an example device 500 that can be used to implement embodiments of the present disclosure. It should be understood that Figure 5 The device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. Figures 1 to 3 The process described.

[0083] like Figure 5 As shown, device 500 is in the form of a general-purpose computing device. Components of computing device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be an actual or virtual processor and is capable of performing various processes according to a program stored in memory 520. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to increase the parallel processing capabilities of computing device 500.

[0084] The computing device 500 typically includes a plurality of computer storage media. Such media may be any available media accessible to the computing device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 may be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 may be a removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which may be capable of being used to store information and / or data (e.g., training data for training) and may be accessed within the computing device 500.

[0085] The computing device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 5 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules that are configured to perform various methods or actions of various implementations of the present disclosure.

[0086] The communication unit 540 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Therefore, the computing device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0087] Input device 550 may be one or more input devices, such as a mouse, keyboard, tracking ball, etc. Output device 560 may be one or more output devices, such as a display, a speaker, a printer, etc. Computing device 500 may also communicate with one or more external devices (not shown) through communication unit 540 as needed, such as storage devices, display devices, etc., communicate with one or more devices that allow users to interact with computing device 500, or communicate with any device that allows computing device 500 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0088] According to an exemplary implementation of the present disclosure, a non-transitory computer-readable storage medium is provided, on which computer executable instructions are stored, wherein the computer executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product is tangibly stored on a non-transitory computer-readable medium and includes computer executable instructions, and the computer executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0089] According to an exemplary implementation of the present disclosure, a chip or a chip system is provided, on which computer executable instructions are carried. When the computer executable instructions are executed by a device or an apparatus or a processor, the method described above can be implemented.

[0090] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0091] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0092] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0093] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0094] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A data processing method, comprising: Receiving a cache request from a client, wherein the cache request includes prefix information; providing a first response to the client, the first response indicating that the embedding vector of the prefix information has been stored in a vector database; receiving a processing request from the client, wherein the processing request includes prompt information, and the prompt information has the prefix information; Obtaining the embedded vector of the prefix information pre-stored in the vector database; as well as Based on the embedding vector of the prefix information and other information in the prompt information except the prefix information, a large model is used to obtain a model output.

2. The method according to claim 1, further comprising: Based on the cache request, determining whether the embedding vector of the prefix information is stored in the vector database; as well as If it is determined that the embedded vector exists in the vector database, the first response is generated.

3. The method according to claim 1, further comprising: Based on the cache request, determining whether the embedding vector of the prefix information is stored in the vector database; as well as If it is determined that the embedded vector does not exist in the vector database, the embedded vector of the prefix information is determined using the large model.

4. The method according to claim 1, wherein the cache request further includes cache priority information corresponding to the prefix information.

5. The method according to claim 4, further comprising: Determine a storage duration threshold corresponding to the cache priority information; as well as If the storage time length of the embedded vector of the prefix information in the vector database reaches the storage time length threshold, the embedded vector of the prefix information is cleared from the vector database. The method according to claim 1 , wherein the size of the vector database is associated with the user rights of the client.

7. The method according to claim 1, wherein the cache request comprises a first execution code for indicating a type of the cache request.

8. The method according to claim 1, wherein the processing request comprises a second execution code for indicating a type of the processing request.

9. The method according to claim 1, further comprising: The stored embedding vector of the prefix information is cleared based on at least one of the following: The available storage space for the vector database, cache priority information of the prefix information, the length of time the embedding vector has been stored, or The frequency with which the embedding vector is called within a historical predetermined period of time.

10. An electronic device, comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 9.

11. A data processing device, comprising: A receiving unit, configured to receive a cache request from a client, wherein the cache request includes prefix information; a providing unit, configured to provide a first response to the client, wherein the first response indicates that the embedding vector of the prefix information has been stored in a vector database; The receiving unit is further configured to receive a processing request from the client, wherein the processing request includes prompt information, and the prompt information has the prefix information; a vector acquisition unit, configured to acquire the embedded vector of the prefix information pre-stored in the vector database; as well as The output unit is configured to obtain a model output using a large model based on the embedding vector of the prefix information and other information in the prompt information except the prefix information.

12. A computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Dialogue method and system based on document retrieval enhanced machine language model

    CN117807199A

  • Question and answer model reasoning optimization and acceleration method and device based on chat robot

    CN118503383A

  • Voice control implementation method and system based on large model

    CN119517026A

  • Large model reasoning method for user request and server

    CN119721261A

  • Request processing method, device, equipment and storage medium

    CN119761374A