Query processing method and device in large model service upgrading period

By generating corresponding input information based on pre-upgraded and post-upgraded prompt words during the upgrade of the big model service, and sending them to the target node for processing, the problem of accuracy and consistency of query processing during the upgrade is solved, and high-quality inference results are achieved.

CN119961319APending Publication Date: 2025-05-09BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411864921.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

During the upgrade of large-model services, it is difficult for the existing technology to effectively process queries, resulting in reduced accuracy of inference results and incompatibility may occur during the upgrade process.

Method used

By obtaining the pending query, and generating the first input information and the second input information according to the first prompt word before the upgrade and the second prompt word after the upgrade, it is sent to the target node in the inference cluster, so that the target node can generate the corresponding response of the query based on the target big model service it deploys.

Benefits of technology

This method can ensure the accuracy and consistency of query processing during the upgrade of the big model service, avoid incompatibility problems, and improve the quality of inference results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961319A_ABST
    Figure CN119961319A_ABST
Patent Text Reader

Abstract

The invention provides a query processing method and device during large model service upgrading, relates to the fields of deep learning, large models, distributed storage, natural language processing and the like, and can be applied to application scenes such as large model online reasoning and the like. The method comprises the following steps: acquiring a to-be-processed query, and in response to the determination that a large model service is currently upgraded, generating first input information according to the query and a first cue word before upgrading, and generating second input information according to the query and a second cue word after upgrading; the first input information and the second input information are sent to a reasoning cluster, so that a target node generates a response corresponding to query according to a target large model service deployed by the target node and the target input information, and the target node is a reasoning node which is determined from all reasoning nodes in the reasoning cluster and used for processing the query; the target input information is input information matched with the target large model service in the first input information and the second input information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of deep learning, big models, distributed storage, and natural language processing, and specifically to a query processing method and device during a big model service upgrade. Background Art

[0002] When using the big model service for online reasoning, the query entered by the user needs to be concatenated with the corresponding prompt of the big model service before it can be submitted to the big model service for model reasoning. Accordingly, when the big model service is upgraded (big model version upgrade), due to changes in training data and other reasons, the corresponding prompts also need to be upgraded to ensure good processing results. Summary of the invention

[0003] The present disclosure provides a query processing method and apparatus during a large model service upgrade.

[0004] A query processing method during a large model service upgrade, comprising:

[0005] Acquire a query to be processed, and in response to determining that the large model service is currently in an upgrade period, generate first input information according to the query and a first prompt word before the upgrade, and generate second input information according to the query and a second prompt word after the upgrade, wherein the prompt word is a prompt word corresponding to the large model service;

[0006] The first input information and the second input information are sent to the inference cluster, so that the target node generates a response corresponding to the query according to the target large model service deployed by itself and the target input information. The target node is an inference node determined from the inference nodes in the inference cluster to process the query, and the target input information is the input information of the first input information and the second input information that matches the target large model service.

[0007] A query processing method during a large model service upgrade, comprising:

[0008] Obtaining first input information and second input information, wherein the first input information and the second input information are generated after the splicing cluster obtains a query to be processed and determines that the current period is during the large model service upgrade, the first input information is generated according to the query and a first prompt word before the upgrade, and the second input information is generated according to the query and a second prompt word after the upgrade, and the prompt word is a prompt word corresponding to the large model service;

[0009] A target node for processing the query is determined from each inference node in the inference cluster, and target sending information is sent to the target node. The target sending information includes at least target input information, which is used by the target node to generate a response corresponding to the query based on the target large model service deployed by itself and the target input information. The target input information is the input information of the first input information and the second input information that matches the target large model service.

[0010] A query processing method during a large model service upgrade, comprising:

[0011] Obtain target sending information from the inference control center, wherein the target sending information includes at least target input information, wherein the target input information is input information that matches the target large model service among the first input information and the second input information, wherein the first input information and the second input information are generated after the splicing cluster obtains a query to be processed and determines that the large model service is currently being upgraded, wherein the first input information is generated based on the query and a first prompt word before the upgrade, wherein the second input information is generated based on the query and a second prompt word after the upgrade, and wherein the prompt word is a prompt word corresponding to the large model service;

[0012] A response corresponding to the query is generated according to the target large model service and the target input information.

[0013] A query processing device during a large model service upgrade, comprising: an information generation module and an information sending module;

[0014] The information generation module is used to obtain a query to be processed, and in response to determining that the large model service is currently in the upgrade period, generate first input information according to the query and a first prompt word before the upgrade, and generate second input information according to the query and a second prompt word after the upgrade, wherein the prompt word is a prompt word corresponding to the large model service;

[0015] The information sending module is used to send the first input information and the second input information to the inference cluster, so that the target node generates a response corresponding to the query according to the target large model service deployed by itself and the target input information. The target node is an inference node determined from each inference node in the inference cluster to process the query, and the target input information is the input information of the first input information and the second input information that matches the target large model service.

[0016] A query processing device during a large model service upgrade, comprising: a first information acquisition module and an information processing module;

[0017] The first information acquisition module is used to acquire first input information and second input information, wherein the first input information and the second input information are generated after the splicing cluster acquires a query to be processed and determines that the service is currently in the large model upgrade period, the first input information is generated based on the query and a first prompt word before the upgrade, and the second input information is generated based on the query and a second prompt word after the upgrade, and the prompt word is a prompt word corresponding to the large model service;

[0018] The information processing module is used to determine the target node for processing the query from each inference node in the inference cluster, and send target sending information to the target node. The target sending information at least includes target input information, which is used for the target node to generate a response corresponding to the query based on the target large model service deployed by itself and the target input information. The target input information is the input information of the first input information and the second input information that matches the target large model service.

[0019] A query processing device during a large model service upgrade, comprising: a second information acquisition module and a response generation module;

[0020] The second information acquisition module is used to acquire target sending information from the inference control center, wherein the target sending information includes at least target input information, wherein the target input information is input information that matches the target large model service among the first input information and the second input information, wherein the first input information and the second input information are generated after the splicing cluster acquires the query to be processed and determines that the large model service is currently being upgraded, wherein the first input information is generated based on the query and a first prompt word before the upgrade, wherein the second input information is generated based on the query and a second prompt word after the upgrade, wherein the prompt word is a prompt word corresponding to the large model service;

[0021] The response generation module is used to generate a response corresponding to the query according to the target large model service and the target input information.

[0022] An electronic device, comprising:

[0023] at least one processor; and

[0024] a memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0026] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0027] A computer program product comprises a computer program / instruction, wherein the computer program / instruction implements the method described above when executed by a processor.

[0028] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0030] Figure 1 It is a flowchart of a first embodiment of a query processing method during a large model service upgrade according to the present disclosure;

[0031] Figure 2 It is a flow chart of a second embodiment of the query processing method during the large model service upgrade described in the present disclosure;

[0032] Figure 3 It is a flowchart of a third embodiment of the query processing method during the large model service upgrade described in the present disclosure;

[0033] Figure 4 A schematic diagram of the connection relationship between the splicing cluster, the inference control center in the inference cluster, and each inference node in the inference cluster described in the present disclosure;

[0034] Figure 5 It is a flowchart of a fourth embodiment of the query processing method during the large model service upgrade described in the present disclosure;

[0035] Figure 6 It is a schematic diagram of the composition structure of the first embodiment 600 of the query processing device during the large model service upgrade period described in the present disclosure;

[0036] Figure 7 It is a schematic diagram of the composition structure of the second embodiment 700 of the query processing device during the large model service upgrade period described in the present disclosure;

[0037] Figure 8 It is a schematic diagram of the composition structure of the third embodiment 800 of the query processing device during the large model service upgrade period described in the present disclosure;

[0038] Fig. 9 A schematic block diagram of an electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0039] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0040] In addition, it should be understood that the term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0041] Figure 1 FIG. 1 is a flowchart of the first embodiment of the query processing method during the large model service upgrade described in the present disclosure. Figure 1 As shown, the following specific implementation methods are included.

[0042] In step 101, a query to be processed is obtained. In response to determining that the large model service is currently being upgraded, first input information is generated based on the query and a first prompt word before the upgrade, and second input information is generated based on the query and a second prompt word after the upgrade, where the prompt word is a prompt word corresponding to the large model service.

[0043] In step 102, the first input information and the second input information are sent to the inference cluster, so that the target node generates a response corresponding to the query based on the target large model service deployed by itself and the target input information. The target node is the inference node determined from the inference nodes in the inference cluster to process the query, and the target input information is the input information in the first input information and the second input information that matches the target large model service.

[0044] In order to improve the efficiency of reasoning and reduce the implementation cost, model reasoning clusters can be used to perform online reasoning of large model services. Model reasoning clusters can include splicing clusters and reasoning clusters. The splicing cluster can combine the query input by the user with the prompt word to obtain the splicing result, and can generate input information based on the splicing result. The reasoning cluster can select the reasoning node for processing the query from each reasoning node, and can send the input information to the selected reasoning node, and then the selected reasoning node uses the large model service deployed by itself and the input information to generate the response corresponding to the query.

[0045] For model inference clusters, when large model services need to be upgraded, the nodes of the splicing cluster and the inference cluster are usually upgraded separately. However, this method can easily lead to the following problems: downstream inference nodes may receive incompatible splicing results generated by the upstream splicing cluster, thereby reducing the accuracy of the inference results.

[0046] By adopting the scheme described in the above method embodiment, after obtaining the query input by the user, corresponding input information can be generated for the prompt words before and after the upgrade. In this way, regardless of whether the target large model service in the target reasoning node is upgraded, it can obtain input information that matches itself and generate a response accordingly, thereby avoiding incompatibility issues as much as possible and improving the accuracy of the reasoning results.

[0047] It should be noted that the queries and responses in the embodiments of the present disclosure are not targeted at a specific user and are not used to reflect the personal information of a specific user. In addition, the executor of the method described in the present disclosure can obtain the queries through various public, legal and compliant methods, such as obtaining them from the user with the user's authorization.

[0048] In practical applications, Figure 1 The execution entity of the illustrated embodiment may be any splicing node in a splicing cluster, which may include multiple splicing nodes and a splicing control center. After the splicing control center obtains the query input by the user, it may randomly select a splicing node, or it may select a splicing node according to a predetermined load balancing strategy, and may forward the obtained query to the selected splicing node for processing.

[0049] The splicing node can determine whether it is currently in the large model service upgrade period by reading the first configuration information. The first configuration information can be manually configured locally in the splicing node, or the first configuration information can also be stored in a predetermined storage location. Accordingly, the splicing node can obtain the first configuration information by accessing the predetermined storage location. The specific implementation method is not limited.

[0050] If it is determined that the large model service is currently being upgraded, the first input information may be generated based on the acquired query and the first prompt word before the upgrade, and the second input information may be generated based on the query and the second prompt word after the upgrade.

[0051] In some embodiments of the present disclosure, the query can be spliced ​​with the first prompt word to obtain a first splicing result, and the first splicing result can be converted into a first token sequence, and then the first input information can be determined based on the first token sequence. Similarly, the query can be spliced ​​with the second prompt word to obtain a second splicing result, and the second splicing result can be converted into a second token sequence, and then the second input information can be determined based on the second token sequence.

[0052] There is no restriction on how to concatenate the query and the prompt word, such as using a traditional concatenation method. In addition, after obtaining the first concatenation result and the second concatenation result, the first concatenation result and the second concatenation result can be converted into a token sequence by calling a tokenizer, thereby obtaining a first token sequence and a second token sequence respectively.

[0053] That is, after obtaining the splicing results, the splicing results can also be converted into a token sequence that meets the input requirements of the large model service, thereby further improving the accuracy of the reasoning results.

[0054] In some embodiments of the present disclosure, when the first input information is determined according to the first token sequence, the first input information can be composed of the first token sequence and the first version identifier, and the first version identifier is used to identify the prompt word corresponding to the first input information as the first prompt word. When the second input information is determined according to the second token sequence, the second input information can be composed of the second token sequence and the second version identifier, and the second version identifier is used to identify the prompt word corresponding to the second input information as the second prompt word. The specific forms of the first version identifier and the second version identifier are not limited.

[0055] That is to say, in addition to the first token sequence, the first input information may also include a first version identifier corresponding to the first token sequence; in addition to the second token sequence, the second input information may also include a second version identifier corresponding to the second token sequence. Accordingly, subsequent related devices can use the version identifier to quickly and accurately identify the first input information corresponding to the first prompt word and the second input information corresponding to the second prompt word.

[0056] The inference cluster may include multiple inference nodes and an inference control center. Accordingly, the generated first input information and second input information can be sent to the inference control center. The inference control center can determine the inference node used to process this query from each inference node, that is, the target node. The target node can then generate a response corresponding to this query based on the target large model service deployed by itself and the target input information. The target input information is the input information in the first input information and the second input information that matches the target large model service.

[0057] In addition, in some embodiments of the present disclosure, after obtaining a query input by a user, in response to determining that the big model service upgrade is complete, second input information can be generated based on the query and the second prompt word, and the second input information can be sent to the inference cluster for the target node to generate a response corresponding to the query based on the target big model service and the second input information.

[0058] When the upgrade is completed, the first configuration information can be adjusted in time. Accordingly, if the upgrade is determined to be completed based on the first configuration information, since the first prompt word is no longer used at this time, the second input information can be generated only based on the second prompt word, thereby reducing the consumption of computing resources and network transmission resources.

[0059] The above mainly describes the solution of the present disclosure from the side of the splicing node. The following further describes the solution of the present disclosure from the side of the inference control center and the side of the inference node.

[0060] Figure 2 FIG. 1 is a flow chart of a second embodiment of the query processing method during the large model service upgrade described in the present disclosure. Figure 2 As shown, the following specific implementation methods are included.

[0061] In step 201, first input information and second input information are obtained. The first input information and the second input information are generated after the splicing cluster obtains the query to be processed and determines that it is currently in the large model service upgrade period. The first input information is generated based on the query and the first prompt word before the upgrade, and the second input information is generated based on the query and the second prompt word after the upgrade. The prompt word is the prompt word corresponding to the large model service.

[0062] In step 202, a target node for processing the query is determined from each inference node in the inference cluster, and target sending information is sent to the target node. The target sending information includes at least target input information, which is used by the target node to generate a response corresponding to the query based on the target large model service deployed by itself and the target input information. The target input information is the input information between the first input information and the second input information that matches the target large model service.

[0063] By adopting the scheme described in the above method embodiment, after obtaining the query input by the user, corresponding input information can be generated for the prompt words before and after the upgrade. In this way, regardless of whether the target large model service in the target reasoning node has been upgraded, it can obtain input information that matches itself and generate a response accordingly, thereby improving the accuracy of the reasoning results.

[0064] In practical applications, Figure 2 The execution subject of the illustrated embodiment may be an inference control center.

[0065] After the inference control center obtains the first input information and the second input information, it can determine the target node for processing the current query from each inference node in the inference cluster.

[0066] In some embodiments of the present disclosure, an inference node may be first selected from each inference node, and the selected inference node may be determined as a candidate node, and then the following first processing may be performed: in response to determining that the machine video memory of the candidate node meets predetermined requirements, the candidate node is determined as a target node, in response to determining that the machine video memory of the candidate node does not meet the predetermined requirements, the inference node is reselected, and the reselected inference node is determined as a candidate node, and then the first processing is repeated.

[0067] For example, an inference node can be selected from each inference node according to a predetermined load balancing strategy, and there is no restriction on the specific load balancing strategy. Afterwards, it can be determined whether the machine video memory of the selected inference node (i.e., the candidate node) meets the predetermined requirements. If so, the selected inference node can be determined as the target node. Otherwise, the inference node can be reselected and the judgment can be made again until the required target node is determined.

[0068] In some embodiments of the present disclosure, the first input information may include: a first token sequence, the first token sequence is a token sequence obtained by converting a first splicing result, and the first splicing result is a splicing result obtained by splicing a query with a first prompt word; the second input information may include: a second token sequence, the second token sequence is a token sequence obtained by converting a second splicing result, and the second splicing result is a splicing result obtained by splicing the query with a second prompt word; accordingly, a method for determining whether a machine video memory of a candidate node meets predetermined requirements may include: obtaining a larger value of a sequence length of the first token sequence and a sequence length of the second token sequence; and in response to determining that the remaining amount of the machine video memory of the candidate node is greater than the machine video memory demand corresponding to the larger value, determining that the machine video memory of the candidate node meets the predetermined requirements.

[0069] For example, the sequence length of the first token sequence is A, the sequence length of the second token sequence is B, A and B are both positive integers greater than 1, and B is greater than A, then B is the determined larger value, if it is determined that the remaining machine video memory of the candidate node is greater than the machine video memory requirement corresponding to B, then it can be determined that the machine video memory of the candidate node meets the predetermined requirements, otherwise, it can be determined that the machine video memory of the candidate node does not meet the predetermined requirements.

[0070] Through the above processing, the selected target node can have enough machine video memory to process this query, thereby improving the processing effect.

[0071] Afterwards, the target sending information may be sent to the target node. In some embodiments of the present disclosure, the first input information and the second input information may be determined as the target sending information and sent to the target node, or, in response to determining that the target large model service has not completed the upgrade, the first input information may be determined as the target sending information and sent to the target node, and in response to determining that the target large model service has completed the upgrade, the second input information may be determined as the target sending information and sent to the target node.

[0072] That is, the reasoning control center can directly send both the first input information and the second input information as target sending information to the target node, or the reasoning control center can also first determine the target input information from the first input information and the second input information. A large model service is usually deployed on each reasoning node. Since the upgrade is carried out step by step, at the current moment, the large model services on some reasoning nodes have completed the upgrade, while the large model services on the remaining reasoning nodes have not completed the upgrade. Correspondingly, for the target node, the target large model service deployed thereon may be a large model service that has not completed the upgrade, or it may be a large model service that has completed the upgrade. If it is the former case, the first input information can be determined as the target input information. If it is the latter case, the second input information can be determined as the target input information. Then, the target input information can be sent to the target node as the target sending information. The specific method to be adopted can be determined according to actual needs, which is very flexible and convenient.

[0073] Figure 3 FIG. 1 is a flowchart of a third embodiment of the query processing method during the large model service upgrade described in the present disclosure. Figure 3 As shown, the following specific implementation methods are included.

[0074] In step 301, target sending information is obtained from the inference control center, and the target sending information includes at least target input information. The target input information is input information that matches the target large model service between the first input information and the second input information. The first input information and the second input information are generated after the splicing cluster obtains the query to be processed and determines that it is currently in the large model service upgrade period. The first input information is generated based on the query and the first prompt word before the upgrade, and the second input information is generated based on the query and the second prompt word after the upgrade, and the prompt word is the prompt word corresponding to the large model service.

[0075] In step 302, a response corresponding to the query is generated according to the target large model service and the target input information.

[0076] By adopting the scheme described in the above method embodiment, after obtaining the query input by the user, corresponding input information can be generated for the prompt words before and after the upgrade. In this way, regardless of whether the target large model service in the target reasoning node has been upgraded, it can obtain input information that matches itself and generate a response accordingly, thereby improving the accuracy of the reasoning results.

[0077] In practical applications, Figure 3 The execution subject of the illustrated embodiment may be a target node.

[0078] In some embodiments of the present disclosure, in response to determining that the target sending information only includes the first input information or the second input information, the target sending information can be determined as the target input information, and the target input information can be input into the target large model service to obtain an output response; in response to determining that the target sending information includes both the first input information and the second input information, the target input information can be selected from the target sending information, and the target input information can be input into the target large model service to obtain an output response; or, in response to determining that the target sending information includes both the first input information and the second input information, the target sending information can be input into the target large model service, for the target large model service to generate a response based on the target input information selected from the target sending information.

[0079] If the target input information is determined by the inference control center, then the target sending information will only include the first input information or the second input information. In this case, the first input information or the second input information can be directly determined as the target input information, and the target input information can be input into the target large model service to obtain an output response. If the target sending information includes both the first input information and the second input information, then the target node can select the target input information from the first input information and the second input information. For example, it can be determined whether the target large model service has completed the upgrade according to the splicing version parameter configured by the local environment variable, and then the first input information or the second input information can be selected and input into the target large model service as the target input information to obtain an output response. Alternatively, the first input information and the second input information can also be input into the target large model service, and the target large model service selects the target input information from the target sending information and generates a response. The specific method to be adopted can be determined according to actual needs, which is very flexible and convenient. In addition, no matter which method is adopted, the target large model service can generate a response according to the matching input information, thereby improving the accuracy of the generated response.

[0080] In addition, the response generated by the target large model service is usually in the form of a token sequence, that is, the target large model service infers based on a token sequence and outputs another token sequence. For the convenience of expression, the token sequence generated by the target large model service can be called the third token sequence. For the third token sequence, it is also necessary to convert it into a form such as text content that can be understood by the user. For example, the third token sequence can be converted by calling a detokenizer.

[0081] In actual applications, the operation of converting the third token sequence can be completed by a predetermined device, for example, it can be completed by the target node. After the target node obtains a textual response through the conversion, it can be returned to the splicing control center through the reasoning control center, and the splicing control center returns the textual response to the user. Alternatively, the target node can also return the third token sequence to the splicing control center through the reasoning control center, and the splicing control center converts the third token sequence into a textual response and returns it to the user.

[0082] Based on the above introduction, Figure 4 The figure is a schematic diagram of the connection relationship between the splicing cluster, the reasoning control center in the reasoning cluster and the reasoning nodes in the reasoning cluster described in the present disclosure. In order to simplify the figure, the splicing control center and the splicing nodes in the splicing cluster are not illustrated.

[0083] in addition, Figure 5 This is a flowchart of the fourth embodiment of the query processing method during the large model service upgrade described in the present disclosure. Figure 4 The connection relationship shown, Figure 5 The illustrated embodiment may include the following specific implementations.

[0084] In step 501, the splicing cluster obtains a query input by a user and generates first input information and second input information respectively.

[0085] For example, the query and the first prompt word can be spliced ​​to obtain a first splicing result, and the first splicing result can be converted into a first token sequence, and then the first token sequence and the first version identifier can be used to form the first input information. In addition, the query and the second prompt word can be spliced ​​to obtain a second splicing result, and the second splicing result can be converted into a second token sequence, and then the second token sequence and the second version identifier can be used to form the second input information.

[0086] In step 502, the splicing cluster sends the first input information and the second input information to the inference control center.

[0087] In step 503, the inference control center determines the target node for processing the current query from the inference nodes in the inference cluster.

[0088] For example, one inference node can be selected from each inference node, and the selected inference node can be determined as a candidate node, and then the following first processing can be performed: in response to determining that the machine video memory of the candidate node meets the predetermined requirements, the candidate node is determined as the target node; in response to determining that the machine video memory of the candidate node does not meet the predetermined requirements, the inference node is reselected, and the reselected inference node is determined as the candidate node, and then the first processing is repeated.

[0089] Among them, for the candidate node, the larger value of the sequence length of the first token sequence and the sequence length of the second token sequence can be obtained. In response to determining that the remaining amount of machine video memory of the candidate node is greater than the machine video memory demand corresponding to the larger value, it can be determined that the machine video memory of the candidate node meets the predetermined requirements. Otherwise, it can be determined that the machine video memory of the candidate node does not meet the predetermined requirements.

[0090] In step 504, the inference control center sends the first input information and the second input information to the target node as target sending information.

[0091] In step 505, the target node selects target input information that matches the target large model service deployed by itself from the target sent information, and inputs the target input information into the target large model service to obtain a response in the form of an output token sequence.

[0092] For example, if it is determined that the target large model service has not completed the upgrade, the first input information can be determined as the target input information; if it is determined that the target large model service has completed the upgrade, the second input information can be determined as the target input information.

[0093] In step 506, the target node converts the response in the form of a token sequence into a response in the form of text, and returns it to the user by concatenating the clusters.

[0094] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the described order of actions, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure. In addition, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0095] The above is an introduction to the method embodiment. The following is a further explanation of the scheme disclosed in the present invention through an apparatus embodiment.

[0096] Figure 6FIG. 6 is a schematic diagram of the composition structure of the first embodiment 600 of the query processing device during the large model service upgrade period described in the present disclosure. Figure 6 As shown, it includes: an information generating module 601 and an information sending module 602.

[0097] The information generation module 601 is used to obtain queries to be processed. In response to determining that the large model service is currently being upgraded, the first input information is generated based on the query and the first prompt word before the upgrade, and the second input information is generated based on the query and the second prompt word after the upgrade, where the prompt word is the prompt word corresponding to the large model service.

[0098] The information sending module 602 is used to send the first input information and the second input information to the inference cluster, so that the target node generates a response corresponding to the query according to the target large model service deployed by itself and the target input information. The target node is the inference node determined from the inference nodes in the inference cluster to process the query, and the target input information is the input information in the first input information and the second input information that matches the target large model service.

[0099] In some embodiments of the present disclosure, the information generation module 601 may splice the query with the first prompt word to obtain a first splicing result, and may convert the first splicing result into a first token sequence, and then determine the first input information based on the first token sequence. Similarly, the query may be spliced ​​with the second prompt word to obtain a second splicing result, and may convert the second splicing result into a second token sequence, and then determine the second input information based on the second token sequence.

[0100] In some embodiments of the present disclosure, when the information generation module 601 determines the first input information according to the first token sequence, the first token sequence and the first version identifier may be used to form the first input information, and the first version identifier is used to identify the prompt word corresponding to the first input information as the first prompt word. When the information generation module 601 determines the second input information according to the second token sequence, the second token sequence and the second version identifier may be used to form the second input information, and the second version identifier is used to identify the prompt word corresponding to the second input information as the second prompt word.

[0101] In addition, in some embodiments of the present disclosure, after obtaining the query input by the user, the information generation module 601 may generate second input information based on the query and the second prompt word in response to determining that the big model service upgrade is complete. Accordingly, the information sending module 602 may send the second input information to the inference cluster, so that the target node can generate a response corresponding to the query based on the target big model service and the second input information.

[0102] Figure 7 FIG. 7 is a schematic diagram of the structure of the second embodiment 700 of the query processing device during the large model service upgrade period described in the present disclosure. Figure 7 As shown, it includes: a first information acquisition module 701 and an information processing module 702.

[0103] The first information acquisition module 701 is used to obtain the first input information and the second input information. The first input information and the second input information are generated after the splicing cluster obtains the query to be processed and determines that it is currently in the large model service upgrade period. The first input information is generated based on the query and the first prompt word before the upgrade, and the second input information is generated based on the query and the second prompt word after the upgrade, and the prompt word is the prompt word corresponding to the large model service.

[0104] The information processing module 702 is used to determine the target node for processing the query from each inference node in the inference cluster, and send the target sending information to the target node. The target sending information includes at least target input information, which is used for the target node to generate a response corresponding to the query based on the target large model service deployed by itself and the target input information. The target input information is the input information between the first input information and the second input information that matches the target large model service.

[0105] In some embodiments of the present disclosure, the information processing module 702 may first select an inference node from each inference node, and may determine the selected inference node as a candidate node, and then may perform the following first processing: in response to determining that the machine video memory of the candidate node meets predetermined requirements, determine the candidate node as a target node, in response to determining that the machine video memory of the candidate node does not meet the predetermined requirements, re-select the inference node, and determine the re-selected inference node as a candidate node, and then repeat the first processing.

[0106] In some embodiments of the present disclosure, the first input information may include: a first token sequence, the first token sequence is a token sequence obtained by converting a first splicing result, and the first splicing result is a splicing result obtained by splicing a query with a first prompt word; the second input information may include: a second token sequence, the second token sequence is a token sequence obtained by converting a second splicing result, and the second splicing result is a splicing result obtained by splicing a query with a second prompt word; accordingly, the information processing module 702 determines that the machine video memory of the candidate node meets the predetermined requirements in a manner that may include: obtaining a larger value of a sequence length of the first token sequence and a sequence length of the second token sequence; in response to determining that the remaining amount of the machine video memory of the candidate node is greater than the machine video memory demand corresponding to the larger value, determining that the machine video memory of the candidate node meets the predetermined requirements.

[0107] Afterwards, the target sending information may be sent to the target node. In some embodiments of the present disclosure, the information processing module 702 may determine the first input information and the second input information as the target sending information and send them to the target node, or, in response to determining that the target large model service has not completed the upgrade, the first input information may be determined as the target sending information and sent to the target node, and in response to determining that the target large model service has completed the upgrade, the second input information may be determined as the target sending information and sent to the target node.

[0108] Figure 8 FIG. 8 is a schematic diagram of the composition structure of the third embodiment 800 of the query processing device during the large model service upgrade period described in the present disclosure. Figure 8 As shown, it includes: a second information acquisition module 801 and a response generation module 802.

[0109] The second information acquisition module 801 is used to obtain target sending information from the inference control center. The target sending information includes at least target input information. The target input information is the input information that matches the target large model service between the first input information and the second input information. The first input information and the second input information are generated after the splicing cluster obtains the query to be processed and determines that it is currently in the large model service upgrade period. The first input information is generated based on the query and the first prompt word before the upgrade, and the second input information is generated based on the query and the second prompt word after the upgrade, and the prompt word is the prompt word corresponding to the large model service.

[0110] The response generation module 802 is used to generate a response corresponding to the query according to the target large model service and the target input information.

[0111] In some embodiments of the present disclosure, in response to determining that the target sending information only includes the first input information or the second input information, the response generation module 802 may determine the target sending information as the target input information, and may input the target input information into the target large model service to obtain an output response; in response to determining that the target sending information includes both the first input information and the second input information, the target input information may be selected from the target sending information, and may be input into the target large model service to obtain an output response; or, in response to determining that the target sending information includes both the first input information and the second input information, the target sending information may be input into the target large model service, for the target large model service to generate a response based on the target input information selected from the target sending information.

[0112] The specific working processes of the above-mentioned device embodiments can refer to the relevant descriptions in the above-mentioned method embodiments and will not be repeated here.

[0113] In summary, the solution disclosed in the present invention can improve the accuracy of the reasoning results. Moreover, the implementation is simple, does not increase excessive resource consumption, and is easy to maintain.

[0114] The scheme disclosed in the present invention can be applied to the field of artificial intelligence, especially to the fields of deep learning, large models, distributed storage, and natural language processing. Artificial intelligence is a discipline that studies how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.

[0115] In addition, in the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0116] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0117] Fig. 9 A schematic block diagram of an electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0118] like Fig. 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0119] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0120] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI, Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP, Digital Signal Processing), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the methods described in the present disclosure may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the method described in the present disclosure in any other appropriate manner (for example, by means of firmware).

[0121] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0123] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, Electronically Programmable Read-Only Memory), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM, Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0125] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0126] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0127] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0128] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A query processing method during a large model service upgrade, comprising: Acquire a query to be processed, and in response to determining that the large model service is currently in an upgrade period, generate first input information according to the query and a first prompt word before the upgrade, and generate second input information according to the query and a second prompt word after the upgrade, wherein the prompt word is a prompt word corresponding to the large model service; The first input information and the second input information are sent to the inference cluster, so that the target node generates a response corresponding to the query according to the target large model service deployed by itself and the target input information. The target node is an inference node determined from the inference nodes in the inference cluster to process the query, and the target input information is the input information of the first input information and the second input information that matches the target large model service.

2. The method according to claim 1, wherein: The generating the first input information according to the query and the first prompt word before the upgrade comprises: splicing the query with the first prompt word to obtain a first splicing result, converting the first splicing result into a first token sequence, and determining the first input information according to the first token sequence; Generating the second input information according to the query and the upgraded second prompt word includes: The query is concatenated with the second prompt word to obtain a second concatenation result, the second concatenation result is converted into a second token sequence, and the second input information is determined according to the second token sequence.

3. The method according to claim 2, wherein: The determining the first input information according to the first token sequence includes: using the first token sequence and a first version identifier to form the first input information, the first version identifier being used to identify that the prompt word corresponding to the first input information is the first prompt word; Determining the second input information according to the second token sequence includes: using the second token sequence and a second version identifier to form the second input information, and the second version identifier is used to identify that the prompt word corresponding to the second input information is the second prompt word.

4. The method according to any one of claims 1 to 3, further comprising: In response to determining that the large model service upgrade is complete, the second input information is generated based on the query and the second prompt word, and the second input information is sent to the inference cluster, so that the target node generates the response based on the target large model service and the second input information.

5. A query processing method during a large model service upgrade, comprising: Obtaining first input information and second input information, wherein the first input information and the second input information are generated after the splicing cluster obtains a query to be processed and determines that the current period is during the large model service upgrade, the first input information is generated according to the query and a first prompt word before the upgrade, and the second input information is generated according to the query and a second prompt word after the upgrade, and the prompt word is a prompt word corresponding to the large model service; A target node for processing the query is determined from each inference node in the inference cluster, and target sending information is sent to the target node. The target sending information includes at least target input information, which is used by the target node to generate a response corresponding to the query based on the target large model service deployed by itself and the target input information. The target input information is the input information of the first input information and the second input information that matches the target large model service.

6. The method according to claim 5, wherein: Determining a target node for processing the query from each inference node in the inference cluster includes: Selecting an inference node from each inference node; The selected inference node is determined as a candidate node, and the following first processing is performed: in response to determining that the machine video memory of the candidate node meets the predetermined requirements, the candidate node is determined as the target node; in response to determining that the machine video memory of the candidate node does not meet the predetermined requirements, the inference node is reselected, and the reselected inference node is determined as the candidate node, and then the first processing is repeated.

7. The method according to claim 6, wherein: The first input information includes: a first token sequence, the first token sequence is a token sequence obtained by converting a first splicing result, the first splicing result is a splicing result obtained by splicing the query and the first prompt word; The second input information includes: a second token sequence, the second token sequence is a token sequence obtained by converting a second splicing result, the second splicing result is a splicing result obtained by splicing the query and the second prompt word; Determining whether the machine video memory of the candidate node meets the predetermined requirements includes: obtaining the larger value of the sequence length of the first token sequence and the sequence length of the second token sequence, and in response to determining that the remaining amount of the machine video memory of the candidate node is greater than the machine video memory demand corresponding to the larger value, determining that the machine video memory of the candidate node meets the predetermined requirements.

8. The method according to claim 5, wherein: The sending of target sending information to the target node comprises: Determine the first input information and the second input information as the target sending information, and send them to the target node; Alternatively, in response to determining that the target large model service has not completed the upgrade, the first input information is determined as the target sending information and sent to the target node; in response to determining that the target large model service has completed the upgrade, the second input information is determined as the target sending information and sent to the target node.

9. A query processing method during a large model service upgrade, comprising: Obtain target sending information from the inference control center, wherein the target sending information includes at least target input information, wherein the target input information is input information that matches the target large model service among the first input information and the second input information, wherein the first input information and the second input information are generated after the splicing cluster obtains a query to be processed and determines that the large model service is currently being upgraded, wherein the first input information is generated based on the query and a first prompt word before the upgrade, wherein the second input information is generated based on the query and a second prompt word after the upgrade, and wherein the prompt word is a prompt word corresponding to the large model service; A response corresponding to the query is generated according to the target large model service and the target input information.

10. The method according to claim 9, wherein: Generating a response corresponding to the query according to the target large model service and the target input information includes: In response to determining that the target sent information only includes the first input information or the second input information, determining the target sent information as the target input information, and inputting the target input information into the target large model service to obtain the outputted response; In response to determining that the target sending information includes both the first input information and the second input information, the target input information is selected from the target sending information, and the target input information is input into the target big model service to obtain the output response; or, in response to determining that the target sending information includes both the first input information and the second input information, the target sending information is input into the target big model service, so that the target big model service generates the response according to the target input information selected from the target sending information.

11. A query processing device during a large model service upgrade, comprising: Information generating module and information sending module; The information generation module is used to obtain a query to be processed, and in response to determining that the large model service is currently in the upgrade period, generate first input information according to the query and a first prompt word before the upgrade, and generate second input information according to the query and a second prompt word after the upgrade, wherein the prompt word is a prompt word corresponding to the large model service; The information sending module is used to send the first input information and the second input information to the inference cluster, so that the target node generates a response corresponding to the query according to the target large model service deployed by itself and the target input information. The target node is an inference node determined from each inference node in the inference cluster to process the query, and the target input information is the input information of the first input information and the second input information that matches the target large model service.

12. The device according to claim 11, wherein The information generation module splices the query with the first prompt word to obtain a first splicing result, converts the first splicing result into a first token sequence, determines the first input information according to the first token sequence, and splices the query with the second prompt word to obtain a second splicing result, converts the second splicing result into a second token sequence, and determines the second input information according to the second token sequence.

13. The device according to claim 12, wherein: The information generation module uses the first token sequence and the first version identifier to form the first input information, and the first version identifier is used to identify that the prompt word corresponding to the first input information is the first prompt word. The information generation module uses the second token sequence and the second version identifier to form the second input information, and the second version identifier is used to identify that the prompt word corresponding to the second input information is the second prompt word.

14. The device according to any one of claims 11 to 13, wherein: The information generation module is further used to, in response to determining that the large model service upgrade is completed, generate the second input information according to the query and the second prompt word; The information sending module is further used to send the second input information to the reasoning cluster, so that the target node generates the response according to the target large model service and the second input information.

15. A query processing device during a large model service upgrade, comprising: A first information acquisition module and an information processing module; The first information acquisition module is used to acquire first input information and second input information, wherein the first input information and the second input information are generated after the splicing cluster acquires a query to be processed and determines that the service is currently in the large model upgrade period, the first input information is generated based on the query and a first prompt word before the upgrade, and the second input information is generated based on the query and a second prompt word after the upgrade, and the prompt word is a prompt word corresponding to the large model service; The information processing module is used to determine the target node for processing the query from each inference node in the inference cluster, and send target sending information to the target node. The target sending information at least includes target input information, which is used for the target node to generate a response corresponding to the query based on the target large model service deployed by itself and the target input information. The target input information is the input information of the first input information and the second input information that matches the target large model service.

16. The device according to claim 15, wherein: The information processing module selects an inference node from each inference node, determines the selected inference node as a candidate node, and performs the following first processing: in response to determining that the machine video memory of the candidate node meets the predetermined requirements, the candidate node is determined as the target node; in response to determining that the machine video memory of the candidate node does not meet the predetermined requirements, reselects the inference node, determines the reselected inference node as the candidate node, and then repeats the first processing.

17. The device according to claim 16, wherein: The first input information includes: a first token sequence, the first token sequence is a token sequence obtained by converting a first splicing result, the first splicing result is a splicing result obtained by splicing the query and the first prompt word; The second input information includes: a second token sequence, the second token sequence is a token sequence obtained by converting a second splicing result, the second splicing result is a splicing result obtained by splicing the query and the second prompt word; The information processing module obtains the larger value of the sequence length of the first token sequence and the sequence length of the second token sequence, and in response to determining that the remaining amount of machine video memory of the candidate node is greater than the machine video memory demand corresponding to the larger value, determines that the machine video memory of the candidate node meets the predetermined requirements.

18. The device according to claim 15, wherein: The information processing module determines the first input information and the second input information as the target sending information, and sends them to the target node; Alternatively, in response to determining that the target large model service has not completed the upgrade, the information processing module determines the first input information as the target sending information and sends it to the target node; in response to determining that the target large model service has completed the upgrade, the information processing module determines the second input information as the target sending information and sends it to the target node.

19. A query processing device during a large model service upgrade, comprising: A second information acquisition module and a response generation module; The second information acquisition module is used to acquire target sending information from the inference control center, wherein the target sending information includes at least target input information, wherein the target input information is input information that matches the target large model service among the first input information and the second input information, wherein the first input information and the second input information are generated after the splicing cluster acquires the query to be processed and determines that the large model service is currently being upgraded, wherein the first input information is generated based on the query and a first prompt word before the upgrade, wherein the second input information is generated based on the query and a second prompt word after the upgrade, wherein the prompt word is a prompt word corresponding to the large model service; The response generation module is used to generate a response corresponding to the query according to the target large model service and the target input information.

20. The device according to claim 19, wherein In response to determining that the target sending information only includes the first input information or the second input information, the response generation module determines the target sending information as the target input information, and inputs the target input information into the target large model service to obtain the output response; in response to determining that the target sending information includes both the first input information and the second input information, the target input information is selected from the target sending information, and the target input information is input into the target large model service to obtain the output response; or, in response to determining that the target sending information includes both the first input information and the second input information, the target sending information is input into the target large model service, so that the target large model service generates the response according to the target input information selected from the target sending information.

21. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 10.

23. A computer program product, comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 10 is implemented.