Data processing method, device and system and server equipment

By using the relevant requests and response results stored in the database in the AI ​​model for imitation, the problem of improving the QPS of the AI ​​model under limited computing resources is solved, and the effect of efficient utilization of computing resources and improving the accuracy of the response results is achieved.

CN119940531APending Publication Date: 2025-05-06LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411909107.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Under the conditions of limited computing resources, how to improve the number of queries per second (QPS) of the AI ​​model, thereby improving the efficiency of computing resources utilization.

Method used

By determining other requests related to user requests and their response results from requests stored in the database, and imitating these related information using a model of a smaller parameter quantity, the response results of the user request are generated.

Benefits of technology

The response result accuracy of the small parameter quantity model is improved, and the computing resource consumption of the model when processing requests is reduced, thereby improving the utilization efficiency of the overall computing resource.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940531A_ABST
    Figure CN119940531A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, device and system and server equipment, and the method comprises the steps: determining a second request related to a first request from at least one request stored in a database based on the first request; each request in the database has a corresponding response result; inputting the first request, the second request, a response result of the second request and instruction information into a first model to obtain a response result of the first request generated by the first model; the instruction information is used for instructing the first model to generate a response result of the first request by referring to the second request and the response result of the second request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to but is not limited to the field of computer technology, and in particular to a data processing method, device, system and server-side equipment. Background Art

[0002] With the development of artificial intelligence (AI) technology, various AI models for realizing AI functions have appeared in users' daily lives and work. For example, the Large Language Model (LLM) is a natural language processing model based on deep learning, which can realize functions such as text creation, dialogue generation, summary generation, machine translation, text classification and sentiment analysis.

[0003] However, the operation of AI models requires a large amount of computing resources, and the larger the number of parameters of an AI model, the more computing resources it consumes. Therefore, how to increase the number of queries per second (QPS) of the model under limited computing resources, thereby improving the utilization efficiency of computing resources, has become an urgent problem to be solved. Summary of the invention

[0004] In view of this, the present application at least provides a data processing method, apparatus, system and server-side device.

[0005] The technical solution of this application is implemented as follows:

[0006] In one aspect, the present application provides a data processing method, the method comprising:

[0007] Based on the first request, determining a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result;

[0008] The first request, the second request, the response result of the second request and the instruction information are input into the first model to obtain the response result of the first request generated by the first model; the instruction information is used to instruct the first model to generate the response result of the first request with reference to the second request and the response result of the second request.

[0009] In some embodiments, before determining, based on the first request, from at least one request stored in a database, a second request related to the first request, the method further comprises:

[0010] Based on the second request, generating a response result of the second request using the second model; wherein the model parameter amount of the second model is greater than the model parameter amount of the first model;

[0011] The second request and the response result of the second request are stored in the database.

[0012] In some embodiments, before determining, based on the first request, from at least one request stored in a database, a second request related to the first request, the method further comprises:

[0013] Based on the second request, generate a response result of the second request using the first model;

[0014] In response to the user's marking operation on the response result of the second request, the second request and the response result of the second request are stored in the database.

[0015] In some embodiments, based on the first request, determining a second request related to the first request from at least one request stored in a database includes:

[0016] determining a similarity between the first request and each request stored in a database;

[0017] Based on the similarity, a second request is determined from the at least one request.

[0018] In some embodiments, the instruction information includes a first instruction, and the first instruction is used to instruct the first model to generate a response result of the first request with reference to the second request and content information of the response result of the second request.

[0019] In some embodiments, the instruction information includes a second instruction; the second instruction is used to instruct the first model to generate a response result of the first request by referring to the second request and the text structure of the response result of the second request.

[0020] In some embodiments, the method further comprises:

[0021] When the response result of the first request meets the target condition, the first request is sent to the second model to obtain an updated response result of the first request generated by the second model.

[0022] In another aspect, the present application provides a data processing device, comprising:

[0023] A determination module, based on the first request, determines a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result;

[0024] The first processing module inputs the first request, the second request, the response result of the second request and the instruction information into the first model to obtain the response result of the first request generated by the first model; the instruction information is used to instruct the first model to generate the response result of the first request with reference to the second request and the response result of the second request.

[0025] On the other hand, the present application provides a server device, including a communication unit and a processing unit; wherein:

[0026] A communication unit receives a first request from a client device;

[0027] The processing unit determines, based on the first request, a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result; the first request, the second request, the response result of the second request and the instruction information are input into the first model to obtain the response result of the first request generated by the first model; the instruction information is used to instruct the first model to generate the response result of the first request with reference to the second request and the response result of the second request;

[0028] The communication unit also sends a response result of the first request to the client device.

[0029] In another aspect, the present application provides a data processing system, including a client and a server; wherein:

[0030] The client is used to determine the first request and send the first request to the server;

[0031] A server is used to receive a first request from a client; determine a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result; input the first request, the second request, the response result of the second request and the instruction information into a first model to obtain the response result of the first request generated by the first model; the instruction information is used to instruct the first model to generate a response result of the first request with reference to the second request and the response result of the second request.

[0032] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.

[0034] Figure 1 A schematic diagram of the implementation flow of a data processing method provided in this application;

[0035] Figure 2 A schematic diagram of the structure of a data processing device provided in this application;

[0036] Figure 3 A hardware entity diagram of a server device provided for this application;

[0037] Figure 4 A system block diagram of a data processing system provided for this application;

[0038] Figure 5 The present invention is a schematic diagram of an embodiment of executing request processing using the data processing system provided by the present application. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0040] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0041] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.

[0043] The present application provides a data processing method, which can be executed by a processor of an electronic device. The electronic device may refer to a server, a laptop, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant), or other device with data processing capabilities.

[0044] Figure 1 A schematic diagram of the implementation flow of a data processing method provided in this application, such as Figure 1 As shown, the method includes the following steps S101 to S102:

[0045] Step S101: based on a first request, determine a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result.

[0046] Here, the first request refers to model processing request information determined based on user operations.

[0047] In some embodiments, the first request may be determined based on a request input operation of a user, for example, a text image request input by the user.

[0048] In some embodiments, the first request may be actively determined by a designated application based on a user interaction operation, for example, a summary generation request actively determined by an intelligent assistant based on a user's selection operation on content in a document.

[0049] A database refers to a database used to store at least one request and its response result.

[0050] In some embodiments, the requests stored in the database may be requests of multiple types. For example, the requests stored in the database may be text-generated-image requests, text-generated-text requests, question-and-answer requests, and the like. In the case where the request is a text-generated-image request, the database stores text information and generated image information corresponding to the request; in the case where the request is a text-generated-text request, the database stores original text information and generated text information corresponding to the request; in the case where the request is a question-and-answer request, the database stores question information and corresponding answer information corresponding to the request; and the like.

[0051] In some embodiments, the database may be a personal knowledge base of a user and / or enterprise. In this way, the requests and responses stored in the database are historical requests and responses of the user and / or enterprise.

[0052] In some embodiments, the database may be a public database provided by a service provider. In this way, the requests and responses stored in the database may be historical requests and responses of multiple people and / or multiple companies.

[0053] In some embodiments, the database may be a vector database. For example, the database may be a database used for retrieval-augmented generation (RAG) technology, in which the stored requests and their response results are in the form of vectors.

[0054] Thus, in some embodiments, determining a second request related to a first request from at least one request stored in a database may be performed by comparing the semantic similarity between the first request and at least one request stored in the database to determine the second request; or the second request may be determined by calculating the distance between a vector representation of the first request and a vector representation of at least one request stored in the database.

[0055] Step S102: input the first request, the second request, the response result of the second request and the instruction information into the first model to obtain the response result of the first request generated by the first model; the instruction information is used to instruct the first model to generate the response result of the first request with reference to the second request and the response result of the second request.

[0056] Here, the instruction information is used to instruct the first model to generate a response result of the first request with reference to the second request and the response result of the second request, that is, the instruction information is used to instruct the first model to imitate the second request and the response result of the second request to generate a response result of the first request.

[0057] In some embodiments, the instruction information is at least part of the prompt word information.

[0058] In some embodiments, the instruction information is information that can guide the first model to perform reasoning.

[0059] In some embodiments, the instruction information is pre-stored text information.

[0060] In some embodiments, the first model is a generative model; the response information of the first request is text information generated by the first model based on its own weight parameters; the response information of the first request is obtained by splicing characteristic information generated one by one by the first model.

[0061] In some embodiments, the first model is a large model, such as a large language model or a large multimodal model. In some scenarios, the first model is also called a small model because the number of parameters is less than that of the model it is matched with; the first model is a model set locally; the number of weight parameters in the first model is greater than 100 million; the first model includes multiple multi-head attention layers.

[0062] In some embodiments, based on the first request, the second request, the response result of the second request and the instruction information, the first model generates indication information, and then the first model generates a response result of the first request based on the indication information.

[0063] In some embodiments, based on the first request, the second request, the response result of the second request and the instruction information, the designated application generates indication information; the indication information is input into the first model so that the first model generates a response result of the first request based on the indication information.

[0064] In the data processing method provided by the present application, first, based on the first request, a second request related to the first request is determined from at least one request stored in a database; then, the first request, the second request, the response result of the second request and the instruction information are input into the first model, and the response result of the first request is generated by using the first model, wherein the instruction information is used to instruct the first model to generate the response result of the first request with reference to the second request and the response result of the second request. It can be seen that when processing the first request, the first model can be imitated with reference to the second request related to the first request and its response result, therefore, the parameter amount requirement for the first model is relatively low, and the response result of the first request generated by the first model has a similar output accuracy to the response result of the second request, thus realizing the generation of a request response result with a higher accuracy by using a first model with a small parameter amount; further, when the parameter amount is small, the amount of computing resources required to run the first model is relatively low, thereby improving the overall utilization efficiency of computing resources.

[0065] In some embodiments, before determining, based on the first request, from at least one request stored in the database, a second request related to the first request, that is, before the above step S101, the data processing method provided by the present application further includes the following steps S103 to S104:

[0066] Step S103: Based on the second request, generate a response result of the second request using a second model; wherein the model parameter amount of the second model is greater than the model parameter amount of the first model.

[0067] Here, in response to obtaining the second request, the second request is input into the second model, so that a response result of the second request is generated by using the second model.

[0068] In some embodiments, the second request is determined based on a request input operation of a user; it may also be actively determined by a designated application based on a user interaction operation.

[0069] Here, the model parameter amount of the second model is greater than the model parameter amount of the first model, that is, the second model has a stronger model reasoning ability than the first model, and the obtained reasoning result has a higher accuracy.

[0070] In some embodiments, the response information of the second request is text information generated by the second model based on its own weight parameters; the response information of the second request is obtained by splicing characteristic information generated one by one by the second model.

[0071] In some embodiments, the second model is a generative model; the second model is a model set up in the cloud; the second model is a large model, such as a large language model or a multimodal large model; the number of weight parameters in the second model is greater than 100 million; the second model includes multiple multi-head attention layers.

[0072] In some embodiments, relative to the first model, the second model is obtained by training the initial model using sample data with a larger amount of data, so that the model parameter quantity of the second model is greater than the model parameter quantity of the first model.

[0073] In some embodiments, the first model is a model with smaller model parameters obtained by performing knowledge distillation on the second model.

[0074] In this way, the response result of the second request generated by using the second model has higher accuracy.

[0075] Step S104: store the second request and the response result of the second request in the database.

[0076] Here, the second request generated by the second model and its response result are stored in the database, so that when the first request is subsequently processed, the first model can be used to refer to the second request and the response result of the second request to generate the response result of the first request.

[0077] In the above embodiment, a response result of the second request is generated by utilizing a second model having a larger amount of model parameters, and the second request and the response result of the second request are stored in a data block, so that a first model having a smaller amount of model parameters can generate a response result of the first request with reference to the second request and the response result of the second request, thereby improving the accuracy of the response result generated by the first model.

[0078] In some embodiments, before determining, based on the first request, from at least one request stored in the database, a second request related to the first request, that is, before the above step S101, the data processing method provided by the present application further includes the following steps S105 to S106:

[0079] Step S105: Based on the second request, use the first model to generate a response result of the second request.

[0080] In some embodiments, in response to obtaining the second request, the second request is input into the first model to generate a response result of the second request using the first model.

[0081] In some embodiments, in response to obtaining a second request, a third request related to the second request and a response result of the third request are determined from a database; the second request, the third request, the response result of the third request and instruction information are input into a first model, and the response result of the second request is generated using the first model.

[0082] Step S106: In response to the user's marking operation on the response result of the second request, the second request and the response result of the second request are stored in the database.

[0083] Here, the marking operation of the response result of the second request by the user refers to the operation of marking the second request and the response result of the second request as the response result to be stored. For example, when the response result of the second request generated by the first model meets the result expected by the user, the user can store the second request and its response result in the database through the marking operation, so as to serve as the simulation object of the first model when the first model processes the relevant request.

[0084] In some embodiments, the user can perform marking operations through a mouse, keyboard, voice and / or gestures.

[0085] In the above embodiment, when the response result generated by the first model for the second request meets the user's expectations, the second request and the response result of the second request can be stored in response to the user's marking operation. In this way, the user can flexibly store the request and the response result that meet the user's expectations in the database in a variety of ways, thereby making the simulation result of the first model more in line with the user's expectations.

[0086] In some embodiments, based on the first request, determining a second request related to the first request from at least one request stored in a database, that is, the above step S101, can be implemented as the following steps S1011 to S1012:

[0087] Step S1011: determine the similarity between the first request and each request stored in the database.

[0088] In some embodiments, first, semantic information corresponding to the first request is determined, and semantic information of at least one request stored in a database is determined; then, based on the semantic information, the similarity between the first request and at least one request stored in the database is determined.

[0089] In some embodiments, when the database is a database for implementing RAG technology, first, vectorization processing is performed on the first request to obtain a vector representation of the first request; then, the distance between the vector representation of the first request and the vector representation of at least one request in the database is calculated, and the similarity between the first request and at least one request in the database is determined based on the calculated distance.

[0090] Step S1012: Determine the second request from the at least one request based on the similarity.

[0091] Here, based on the determined similarity, the second request is determined from at least one request stored in the database. For example, the request with the highest similarity is used as the second request; or a specified number of requests with the highest similarity are used as the second request.

[0092] In some embodiments, the similarity between the first request and at least one request stored in the database may include multiple types of similarities, such as content similarity, request type similarity, and the like.

[0093] Here, content similarity refers to the similarity between the knowledge content requested by the first request and the knowledge content requested by at least one request stored in the database. For example, the first request is "Who is Xiao Ming's mother?" It can be seen that the knowledge content requested by the first request is "Who is Xiao Ming's mother", and the requests related to the first request may be "Who is Xiao Ming's father?", "Who is Xiao Ming's father's wife?", and so on.

[0094] The request type similarity refers to the similarity between the question type of the first request and the question type of at least one request stored in the database. For example, the first request is "Who is Xiaoming's mother?" It can be seen that the type of the first request is to ask who someone's mother is, or who a certain relative of someone is. Then the request with a high similarity to the request type of the first request may be "Who is Xiaodong's mother?", "Who is Xiaodong's father?" or "Who is Xiaoxi's mother?", etc.

[0095] Thus, in some embodiments, the instruction information includes a first instruction, and the first instruction is used to instruct the first model to generate a response result of the first request with reference to the second request and content information of the response result of the second request.

[0096] Here, the first instruction instructs the first model to generate a response result of the first request by referring to the second request and content information of the response result of the second request.

[0097] For example, when the first request is "Who is Xiao Ming's mother?" and the second request is "Who is Xiao Ming's father's wife?", the first model can refer to the response result of "Who is Xiao Ming's father's wife?" to infer the response result of "Who is Xiao Ming's mother?"

[0098] For another example, when the first request is “generate a JAVA program for performing addition operations” and the second request is “generate a JAVA program for performing multiplication operations”, the first model can refer to the generated JAVA program for performing multiplication operations to generate a JAVA program for performing addition operations.

[0099] In some embodiments, the instruction information includes a second instruction; the second instruction is used to instruct the first model to generate a response result of the first request by referring to the text structure of the second request and the response result of the second request.

[0100] Here, when the first request and the second request are of the same request type, the second instruction is used to instruct the first model to generate a response result of the first request with reference to the text structure of the second request and the response result of the second request.

[0101] For example, when the first request is "Who is Xiao Ming's mother?" and the second request is "Who is Xiao Dong's mother?", the response result of the first request can be generated by referring to the text structure of the response result of the second request.

[0102] For example, when the first request is "Generate minutes of this morning's meeting" and the second request is "Generate minutes of this afternoon's meeting", the response result of the first request can be generated with reference to the text structure of the minutes in the response result of the second request, that is, the text structure of the minutes of the morning meeting can be generated according to the text structure of the minutes of the afternoon meeting.

[0103] In the above embodiment, different instruction information is used to instruct the first model to refer to the second request and the different types of information in the response result of the second request to generate a response result of the first request, so that the first model can more accurately extract valuable information from the second request and the response results of the second request, thereby improving the imitation effect of the first model.

[0104] In some embodiments, the data processing method provided by the present application further includes the following steps S107:

[0105] Step S107: When the response result of the first request meets the target condition, the first request is sent to the second model to obtain an updated response result of the first request generated by the second model.

[0106] In some embodiments, the target condition refers to that the response result of the first request does not meet a preset output condition.

[0107] In some embodiments, the preset output condition refers to that the relevance between the response result of the first request and the first request is higher than a preset threshold. For example, after the first model generates the response result of the first request, based on the context-aware mechanism, it is determined that the relevance between the response result of the first request and the first request is lower than a preset threshold, and then it is determined that the response result of the first request meets the target condition.

[0108] In some embodiments, the target condition refers to negative feedback information of the user regarding the response result of the first request. For example, when the user marks the response result of the first request as worthless information, it is determined that the response result of the first request meets the target condition.

[0109] In this way, when the response result of the first request meets the target condition, the first request is sent to the second model, so as to use the second model to perform model reasoning on the first request and obtain an updated response result of the first request.

[0110] In the above embodiment, when the response result of the first request generated by the first model meets the target condition, the second model is used to perform model inference on the first request, thereby obtaining an updated response result of the first request. Since the weight parameter amount of the second model is greater than the weight parameter amount of the first model, the updated response result of the first request has higher precision and accuracy than the response result of the first request generated by the first model.

[0111] As can be seen from the above, in the data processing method provided by the present application, first, the request response result and the corresponding request generated by the second model with a larger weight parameter amount are stored in the database; then, after obtaining the first request, the second request related to the first request is queried from the database, and the second request and the response result of the second request are imitated by the first model with a smaller weight parameter amount to generate the response result of the first request. In this way, since the weight parameter amount of the first model is small, the computing resources consumed by the first model are small, which can achieve the effect of saving computing resources and improving model processing efficiency; at the same time, by imitating the request response result generated by the second model, the accuracy of the response result generated by the first model can be improved. In addition, in response to the user's marking operation on the response result of the first request, the first request and the response result of the first request are stored in the database, which can enable the user to flexibly store the request and its response result that meet the user's expectations in the database in a variety of ways, thereby making the subsequent imitation results of the first model more in line with user expectations. At the same time, using instruction information to instruct the first model to generate a response result of the first request with reference to the second request and the content information and / or text structure of the response result of the second request, can enable the first model to more accurately extract useful information from the second request and the response results of the second request, thereby improving the imitation effect.

[0112] Based on the foregoing embodiments, the present application provides a data processing device, which includes the units included and the modules included in the units, which can be implemented by a processor in a computer device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0113] Figure 2 A schematic diagram of the structure of a data processing device provided in this application is shown in FIG. Figure 2 As shown, the data processing device 200 includes: a determination module 210 and a first processing module 220, wherein:

[0114] The determination module 210 determines, based on the first request, a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result;

[0115] The first processing module 220 inputs the first request, the second request, the response result of the second request and the instruction information into the first model to obtain the response result of the first request generated by the first model; the instruction information is used to instruct the first model to generate the response result of the first request with reference to the second request and the response result of the second request.

[0116] In some embodiments, the apparatus 200 further includes a second processing module 230; the second processing module 230 is configured to:

[0117] Based on the second request, generating a response result of the second request using a second model; wherein the model parameter amount of the second model is greater than the model parameter amount of the first model;

[0118] The second request and a response result of the second request are stored in the database.

[0119] In some embodiments, the first processing module 220 is further configured to:

[0120] Based on the second request, using the first model to generate a response result of the second request;

[0121] In response to a user's marking operation on a response result of the second request, the second request and the response result of the second request are stored in the database.

[0122] In some embodiments, the determining module 210 is used to:

[0123] determining a similarity between the first request and each request stored in the database;

[0124] Based on the similarity, the second request is determined from the at least one request.

[0125] In some embodiments, the instruction information includes a first instruction, and the first instruction is used to instruct the first model to generate a response result of the first request with reference to the second request and content information of the response result of the second request.

[0126] In some embodiments, the instruction information includes a second instruction; the second instruction is used to instruct the first model to generate a response result of the first request by referring to the text structure of the second request and the response result of the second request.

[0127] In some embodiments, the apparatus 200 further includes a determination module 240;

[0128] The judgment module 240 is used to send the first request to the second model to obtain an updated response result of the first request generated by the second model when the response result of the first request meets the target condition.

[0129] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0130] Based on the aforementioned embodiments, the present application provides a server device. Figure 3 A hardware entity diagram of a server device provided for this application, such as Figure 3 As shown, the server device 300 includes: a communication unit 310 and a processing unit 320; wherein,

[0131] The communication unit 310 receives a first request from a client device;

[0132] The processing unit 320 determines, based on the first request, a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result; the first request, the second request, the response result of the second request and the instruction information are input into a first model to obtain the response result of the first request generated by the first model; the instruction information is used to instruct the first model to generate a response result of the first request with reference to the second request and the response result of the second request;

[0133] The communication unit 310 further sends a response result of the first request to the client device.

[0134] like Figure 3 As shown, the server device 300 also includes a bus 330 so that the communication unit 310 and the processing unit 320 can perform data transmission via the bus 330 .

[0135] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or units included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0136] Based on the aforementioned embodiments, the present application provides a data processing system. Figure 4 A system block diagram of a data processing system provided in this application, such as Figure 4 As shown, the data processing system 400 includes: a client 410 and a server 420; wherein,

[0137] The client 410 is used to determine a first request and send the first request to the server 420;

[0138] The server 420 is used to receive the first request from the client 410; determine a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result; input the first request, the second request, the response result of the second request and instruction information into the first model to obtain the response result of the first request generated by the first model; the instruction information is used to instruct the first model to generate a response result of the first request with reference to the second request and the response result of the second request.

[0139] In some embodiments, the client 410 can be configured on any electronic device with data processing capabilities. For example, the client 410 can be configured on a server, a laptop, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant), and other electronic devices.

[0140] In some embodiments, the server 420 can be configured on any electronic device with data processing capabilities. For example, the server 420 can be configured on an electronic device such as a server, a laptop, a tablet computer, a desktop computer, etc.

[0141] In some embodiments, the client 410 and the server 420 may be configured on the same or different electronic devices.

[0142] In some embodiments, the server 420 may be configured on a gateway device.

[0143] In some embodiments, the database, the first model, the second model and the server 420 may be configured on the same device.

[0144] In some embodiments, the database, the first model, and the second model may be configured on different devices from the server 420. For example, the database may be configured on a third-party server; for another example, the first model and the second model may be configured on a third-party server, or on the same or different devices in a device cluster.

[0145] Next, combine Figure 5 An embodiment of executing request processing using the data processing system provided by the present application is described.

[0146] like Figure 5 As shown, the client 510 sends the target reasoning request initiated by the user to the server 520;

[0147] After receiving the target reasoning request sent by the client 510, the server 520 queries the database 530 to determine whether the database 530 stores a target association request related to the target reasoning request;

[0148] In the case where the database 530 stores a target association request related to the target reasoning request, the database 530 sends the target association request and the request response of the target association request to the server 520; the server 520 sends the target reasoning request, the target association request, the response result of the target association request and the instruction information to the first model 540, wherein the instruction information is used to instruct the first model 540 to generate the response result of the target reasoning request with reference to the target association request and the response result of the target association request; after receiving the target reasoning request, the target association request, the response result of the target association request and the instruction information from the server 520, the first model 540 generates the response result of the target reasoning request and returns the response result of the target reasoning request to the server 520; the server 520 sends the response result of the target reasoning request to the client 510;

[0149] When the database 530 does not store a target association request related to the target reasoning request, the database 530 sends the result of not finding the target association request to the server 520; the server 520 sends the target reasoning request to the second model 550; the second model 550 generates a response result of the target reasoning request and returns the response result of the target reasoning request to the server 520; the server 520 sends the response result of the target reasoning request to the client 510.

[0150] It can be seen from the above embodiments that in the data processing system provided by the present application, a second model with a larger number of model parameters and a first model with a smaller number of model parameters are deployed at the same time. When an inference request is received, the server determines whether to use the second model or the first model to perform request processing based on whether a target association request related to the target inference request is stored in the database, thereby achieving the effect of using a large parameter model to improve model processing accuracy, using a small parameter model to improve model processing efficiency, and saving computing resources.

[0151] The description of the above data processing system embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or devices included in the system provided by the embodiment of the present disclosure can be used to execute the method described in the above method embodiment. For technical details not disclosed in the system embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0152] It should be noted that in the embodiment of the present application, if the above-mentioned data processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software, and firmware.

[0153] The embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.

[0154] An embodiment of the present application provides a computer program, including a computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.

[0155] The embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented specifically by hardware, software or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.

[0156] It should be noted here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above device, storage medium, computer program and computer program product embodiments is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of this application, please refer to the description of the method embodiment of this application for understanding.

[0157] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the serial number of each step / process mentioned above does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The serial numbers of the embodiments of the present application mentioned above are for description only and do not represent the advantages and disadvantages of the embodiments.

[0158] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0159] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0160] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0161] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be separately configured as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0162] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0163] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0164] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A data processing method, comprising: Based on the first request, determining a second request related to the first request from at least one request stored in a database; Each request in the database has a corresponding response result; Inputting the first request, the second request, the response result of the second request and the instruction information into the first model to obtain the response result of the first request generated by the first model; The instruction information is used to instruct the first model to generate a response result of the first request by referring to the second request and the response result of the second request.

2. The method according to claim 1, before determining, based on the first request, from at least one request stored in a database, a second request related to the first request, further comprising: Based on the second request, generating a response result of the second request using a second model; wherein the model parameter amount of the second model is greater than the model parameter amount of the first model; The second request and a response result of the second request are stored in the database.

3. The method according to claim 1, before determining, based on the first request, a second request related to the first request from at least one request stored in a database, further comprising: Based on the second request, using the first model to generate a response result of the second request; In response to a user's marking operation on a response result of the second request, the second request and the response result of the second request are stored in the database.

4. The method according to any one of claims 1 to 3, wherein determining, based on the first request, a second request related to the first request from at least one request stored in a database comprises: determining a similarity between the first request and each request stored in the database; Based on the similarity, the second request is determined from the at least one request.

5. According to the method described in any one of claims 1 to 3, the instruction information includes a first instruction, and the first instruction is used to instruct the first model to generate a response result of the first request with reference to the second request and content information of the response result of the second request.

6. According to the method described in any one of claims 1 to 3, the instruction information includes a second instruction; the second instruction is used to instruct the first model to generate a response result of the first request with reference to the text structure of the second request and the response result of the second request.

7. The method according to claim 2, further comprising: When the response result of the first request meets the target condition, the first request is sent to the second model to obtain an updated response result of the first request generated by the second model.

8. A data processing device, comprising: A determination module, based on the first request, determines a second request related to the first request from at least one request stored in a database; Each request in the database has a corresponding response result; A first processing module inputs the first request, the second request, the response result of the second request and the instruction information into a first model to obtain the response result of the first request generated by the first model; The instruction information is used to instruct the first model to generate a response result of the first request by referring to the second request and the response result of the second request.

9. A server device, comprising a communication unit and a processing unit; wherein: The communication unit receives a first request from a client device; The processing unit determines, based on the first request, a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result; Inputting the first request, the second request, the response result of the second request and the instruction information into the first model to obtain the response result of the first request generated by the first model; The instruction information is used to instruct the first model to generate a response result of the first request with reference to the second request and a response result of the second request; The communication unit also sends a response result of the first request to the client device.

10. A data processing system, comprising a client and a server; wherein: The client is used to determine a first request and send the first request to the server; The server is configured to receive the first request from the client; Determine a second request related to the first request from at least one request stored in a database; each request in the database has a corresponding response result; Inputting the first request, the second request, the response result of the second request and the instruction information into the first model to obtain the response result of the first request generated by the first model; The instruction information is used to instruct the first model to generate a response result of the first request by referring to the second request and the response result of the second request.