Multi-source artificial intelligence service adaptation method and device, medium, equipment and product
By creating target program instances based on low-bytecode files in the gateway, stability and security issues caused by differences in various artificial intelligence services are resolved, achieving unified access and adaptation, and improving the stability and security of the gateway.
Patent Information
- Application Number
- CN202511892750.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-27
AI Technical Summary
The differences between various artificial intelligence services pose challenges to the stability and security of gateways, making it difficult to achieve unified access.
By acquiring the target program module that matches the model service request, a target program instance based on a low bytecode file is created, runs in an independent environment, and performs processing to comply with the standardized provisions of artificial intelligence services and gateways, thereby achieving the adaptation of model service requests and responses.
It enables unified access to different artificial intelligence services through the gateway, improving stability and security and avoiding the impact of malicious bytecode files.
Smart Images

Figure CN121743075A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular, to a multi-source artificial intelligence service adaptation method, device, medium, equipment and product. BACKGROUND
[0002] With the rapid development of AI (Artificial Intelligence) services, different providers have launched different artificial intelligence services, such as large model services.
[0003] In the related art, as an intermediate party between the business party and the provider, the gateway needs to break the differences between different artificial intelligence services provided by the provider to ensure that the business party can use multi-source artificial intelligence services stably and safely. SUMMARY
[0004] This summary is provided to introduce a selection of concepts, which will be described with greater specificity below in the detailed description section. This summary does not intend to identify key or essential features of the claimed technology nor does it intend to limit the scope of the claimed technology.
[0005] In a first aspect, the present disclosure provides a multi-source artificial intelligence service adaptation method, comprising: obtaining a target program module matched with a model service request; creating a target program instance of the target program module, the target program module being obtained based on a low bytecode file, and the target program instance supporting running in an independent environment; performing first processing on the model service request by calling the target program instance running in the independent environment to obtain a target model service request, the target model service request conforming to a standardized provision of an artificial intelligence service corresponding to the target program module; performing second processing on an obtained model service response by calling the target program instance to obtain a target model service response, the model service response being obtained by calling the artificial intelligence service to respond to the target model service request, and the target model service response conforming to a standardized provision of the gateway.
[0006] In a second aspect, the present disclosure provides a multi-source artificial intelligence service adaptation device, comprising: an obtaining module configured to obtain a target program module matched with a model service request; a creating module configured to create a target program instance of the target program module, the target program module being obtained based on a low bytecode file, and the target program instance supporting running in an independent environment; The first processing module is configured to perform first processing on the model service request by invoking the target program instance running in the independent environment, to obtain a target model service request, wherein the target model service request conforms to a standardized regulation of an artificial intelligence service corresponding to the target program module. The second processing module is configured to perform second processing on the obtained model service response by invoking the target program instance, to obtain a target model service response, wherein the model service response is obtained by invoking the artificial intelligence service to respond to the target model service request, and the target model service response conforms to the standardized regulation of the gateway.
[0007] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, wherein the computer program is executed by a processing device to implement the steps of the method in the first aspect.
[0008] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the method in the first aspect.
[0009] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the steps of the method in the first aspect.
[0010] According to the above technical solution, since the target program module is loaded based on the low bytecode file, the target program instance can support running in an independent environment, thereby avoiding the influence of malicious low bytecode files on the stability and security of the gateway. In addition, the target program instance is used to perform first processing on the model service request, to obtain a target model service request conforming to the standardized regulation of the artificial intelligence service corresponding to the target program module, so that the artificial intelligence service can be used to respond to the target model service request. The second processing is performed on the obtained model service response by invoking the target program instance, to obtain a target model service response conforming to the standardized regulation of the gateway, thereby solving the differences between different artificial intelligence services and realizing the unified access of the gateway to different artificial intelligence services.
[0011] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings. The same or similar components have the same or similar reference numbers regardless of the drawing in which they are depicted. It should be noted that the drawings are schematic and elements in the drawings have not necessarily been drawn to scale. In the drawings: Figure 1 is a schematic diagram of an application scenario for implementing a multi-source artificial intelligence service adaptation method according to an embodiment of the present disclosure; Figure 2 is a flowchart of a multi-source artificial intelligence service adaptation method according to an embodiment of the present disclosure; Figure 3 is another flowchart of a multi-source artificial intelligence service adaptation method according to an embodiment of the present disclosure; Figure 4 is a block diagram of a multi-source artificial intelligence service adaptation apparatus according to an embodiment of the present disclosure; Figure 5 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0013] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.
[0014] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0015] The term “comprising” and variations thereof as used herein are used inclusively, i.e., “comprising but not limited to.” The term “based on” is “based at least in part on.” The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments.” Related definitions will be given in the description below.
[0016] It should be noted that the concepts of “first”, “second”, and the like mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0017] It should be noted that the modification of "one" and "multiple" mentioned in the present disclosure is illustrative rather than limiting, and those skilled in the art should understand that "one or more" should be understood unless the context clearly indicates otherwise.
[0018] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0019] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the permission of the user should be obtained through appropriate means according to relevant laws and regulations.
[0020] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0021] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0022] It can be understood that the above notification and user permission obtaining process is only illustrative and does not limit the implementation manners of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0023] At the same time, it can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the present technical solutions should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0024] Figure 1 is a schematic diagram of an application scenario for implementing a multi-source artificial intelligence service adaptation method according to an embodiment of the present disclosure. The application scenario includes a client, a gateway and artificial intelligence services (for example Figure 1 artificial intelligence service 1, artificial intelligence service 2, artificial intelligence service N) provided by each provider, where N is an integer greater than 2.
[0025] Among them, the client is a carrier for user interaction with the artificial intelligence service, responsible for transmitting the model service request initiated by the user to the gateway, and presenting the processing result returned by the gateway in a visual manner.
[0026] A gateway is a middleware that connects clients and artificial intelligence services. Its core functions include request routing, traffic control, capability adaptation, and security protection. It handles diverse model service requests from clients, converts these requests into standardized requests that conform to the requested artificial intelligence service, receives model service responses from artificial intelligence services, converts these responses into standardized responses that conform to the gateway's specifications, and then returns them to the client.
[0027] Artificial intelligence services are the underlying services that provide core AI capabilities. Examples of artificial intelligence services include large language model services and embedding (vector) models.
[0028] Based on the above application scenario, the user initiates a model service request through the client. The gateway constructs a target model service request through the target program instance of the target program module, completes the adaptation of the artificial intelligence service requested by the model service request, receives the model service response of the artificial intelligence service, processes the model service response through the target program instance, obtains a standardized target model service response, and finally returns the standardized target model service response to the user through the client, thus completing the processing of the model service request.
[0029] Figure 2 This is a flowchart illustrating a multi-source artificial intelligence service adaptation method according to an embodiment of the present disclosure. This multi-source artificial intelligence service adaptation method can be... Figure 1 The gateway shown is executed, refer to Figure 2 The multi-source AI service adaptation method may include the following steps: In step 210, the target program module that matches the model service request is obtained.
[0030] Among them, model service requests can be made by Figure 1 The client shown sends a model service request to request an artificial intelligence service, such as a large model service. This model service request can be, for example, an HTTP (Hypertext Transfer Protocol) request. More specifically, it can be a knowledge-based question-and-answer request. This request is used to request an AI service that provides knowledge-based question-and-answer services. The AI service responds to the knowledge-based question-and-answer request by generating a question-and-answer result (i.e., the model service response). The following explanation and illustration of this embodiment uses an HTTP request as an example of a model service request.
[0031] In some embodiments, step 210 can be implemented by determining, based on the model identifier carried in the obtained model service request, a target program module matching the model identifier from the candidate program modules according to priorities of a plurality of preset matching strategies, wherein the matching accuracy of a preset matching strategy with a higher priority is higher than that of a preset matching strategy with a lower priority.
[0032] As an example, the model identifier can include a first model identifier, a second model identifier, and a third model identifier, wherein the first model identifier is used to represent the provider of the artificial intelligence service, the second model identifier is used to represent the type of the artificial intelligence service, and the third model identifier is used to represent the name of the artificial intelligence service. Here, the type is, for example, the large language model service and the Embedding (vector) model described above.
[0033] Here, the preset matching strategies with different priorities can be set according to the number of types of model identifiers required to be matched by the strategy, and the types here are, for example, the provider, the type, and the name described above. Continuing the above example, the preset matching strategy with the highest priority (hereinafter referred to as the first preset matching strategy) can be to match according to the first model identifier, the second model identifier, and the third model identifier; the preset matching strategy with an intermediate priority (hereinafter referred to as the second preset matching strategy) can be to match according to the first model identifier and the second model identifier; and the preset matching strategy with the lowest priority (hereinafter referred to as the third preset matching strategy) can be to match according to the first model identifier.
[0034] Thus, when determining the target program module matching the model identifier carried in the model service request, the first preset matching strategy can be used first, i.e., the target program module matching the first model identifier, the second model identifier, and the third model identifier carried in the model service request is determined from the candidate program modules according to the first model identifier, the second model identifier, and the third model identifier. If the target program module cannot be matched by using the first preset matching strategy, the second preset matching strategy is used, i.e., the target program module matching the first model identifier and the second model identifier carried in the model service request is determined from the candidate program modules according to the first model identifier and the second model identifier. If the target program module still cannot be matched, the third preset matching strategy is used, i.e., the target program module matching the first model identifier carried in the model service request is determined from the candidate program modules according to the first model identifier. It should be noted that the candidate program modules are bound with corresponding model identifiers, and the candidate program modules can be referred to the related embodiments described below, which will not be repeated here.
[0035] In this way, the precise matching strategy (i.e., the first preset matching strategy) requiring matching multiple model identifiers is preferentially adopted to match the target program module for the model service request, so as to realize precise matching of the target program module and the model service request. In addition, the bottom-up fuzzy matching strategy (i.e., the second preset matching strategy and the third preset matching strategy) is provided to provide a data basis for responding to the model service request.
[0036] In step 220, a target program instance of the target program module is created, and the target program module is loaded based on the low bytecode file. The target program instance supports running in an independent environment.
[0037] As an example, the low bytecode file of the target program module can be a WebAssembly (WASM) file. The WASM file represents a customized code logic, which can implement standardized processing of HTTP requests and HTTP responses. Taking the standardized processing of the HTTP request as an example, the standardized processing of the HTTP request is used to convert the HTTP request into a standardized regulation of an artificial intelligence service. The WASM file is close to the execution performance of the native code and can support multiple programming languages.
[0038] Here, the standardized regulation of the HTTP request includes the standardized regulation of the URL (Uniform Resource Locator), the standardized regulation of the request header, and the standardized regulation of the request body. The standardized regulation of the HTTP response includes the standardized regulation of the response header and the standardized regulation of the response body. The implementation of these standardized regulations can refer to the related embodiments described below, and this embodiment will not be described here.
[0039] Here, the running environment can be understood as a sandbox environment, and the running environment can be provided by the gateway to the target program instance, so that the target program instance cannot directly access the resources of the gateway. Not only can the unlimited use of the gateway resources by the target program instance be avoided, but also the impact of malicious or erroneous target program instances on the gateway and other running target program instances can be avoided, thereby improving the stability of the gateway.
[0040] In some embodiments, the target program instances of different model service requests are independent of each other, that is, each model service request corresponds to an independent target program instance. In this way, fault isolation of a single target program instance can be realized, and the impact of a fault of a target program instance on a large number of model service requests when the target program instance is associated with multiple model service requests can be avoided.
[0041] In step 230, a first processing is performed on the model service request by calling the target program instance running in the independent environment, and a target model service request is obtained, which conforms to the standardized regulation of the artificial intelligence service corresponding to the target program module.
[0042] Here, the target model service request specifically refers to the standardized regulation of the API (Application Programming Interface) of the artificial intelligence service.
[0043] The first processing can be referred to the related embodiments described below, and the embodiments will not be repeated here.
[0044] In step 240, a second processing is performed on the obtained model service response by calling the target program instance, and a target model service response is obtained, which conforms to the standardized regulation of the gateway, and the model service response is obtained by calling the artificial intelligence service to respond to the target model service request.
[0045] The second processing can be referred to the related embodiments described below, and the embodiments will not be repeated here.
[0046] The first processing is performed on the model service request by using the target program instance, and the target model service request conforming to the standardized regulation of the artificial intelligence service corresponding to the target program module is obtained, so that the artificial intelligence service can respond to the target model service request. The second processing is performed on the obtained model service response by calling the target program instance, and the target model service response conforming to the standardized regulation of the gateway is obtained, which solves the difference between different artificial intelligence services and realizes the unified access of the gateway to different artificial intelligence services.
[0047] In some embodiments, the above-mentioned standardized regulation includes the regulation of data protocol and the security regulation, and the regulation of data protocol and the security regulation can be referred to the related embodiments described below.
[0048] In some embodiments, the above-mentioned method can further include the following steps: in the running process of the gateway, in response to an update operation for the target directory, loading a low bytecode file associated with the update operation to obtain a program module corresponding to the low bytecode file, and the low bytecode file corresponds to a unique model identifier; constructing a mapping relationship between the program module and the model identifier, and updating the target mapping table based on the mapping relationship.
[0049] As an example, the update operation can be an addition operation of adding a WASM file, and the low bytecode file associated with the update operation is understood as the added WASM file. The corresponding constructed mapping relationship is added to the target mapping table to achieve the purpose of updating the target mapping table.
[0050] As another example, the update operation can be a modification operation of modifying an existing WASM file in the target directory, and the low bytecode file associated with the update operation is understood as the modified WASM file, and the mapping relationship corresponding to the build is used to replace the target mapping relationship in the target mapping table, so as to achieve the purpose of updating the target mapping table, and the model identifier of the target mapping relationship and the mapping relationship corresponding to the build is the same.
[0051] From the above, it can be seen that the unique model identifier corresponding to the low bytecode file in the embodiment can include the first model identifier, as another example, the unique model identifier corresponding to the low bytecode file can include the first model identifier and the second model identifier, as another example, the unique model identifier corresponding to the low bytecode file can include the first model identifier, the second model identifier and the third model identifier, and the configuration of the low bytecode file is set by the developer.
[0052] The program module corresponding to the mapping relationship in the target mapping table can be the above-mentioned candidate program module, which is used to support the acquisition of the above-mentioned target program module.
[0053] In the above manner, the update operation for the target directory is detected during the running of the gateway, and the dynamic loading of the low bytecode file associated with the update operation is performed, and then the hot update of the program module and the target mapping table is realized. In this way, the hot update of the program module and the target mapping table is realized without interrupting the running of the gateway.
[0054] In some embodiments, the steps of building the mapping relationship between the program module and the model identifier, and updating the target mapping table based on the mapping relationship can be implemented in the following manner: detecting the program module to obtain a detection result; building the mapping relationship between the program module and the model identifier if the detection result represents a pass; and updating the target mapping table based on the mapping relationship.
[0055] It should be noted that the gateway can define the core functions that the WASM module (i.e. the program module) must implement in advance, and on this basis, the gateway parses the program module through the reflection mechanism, exports the function functions in the program module, and detects whether the exported function functions contain the core functions. If any core function is missing, the detection result is determined to be a result representing a pass, and if any core function is missing, the detection result is determined to be a result representing a pass, and it can be understood that in the case where the detection result represents a pass, the mapping relationship between the program module and the model identifier is prohibited.
[0056] In some embodiments, as can be known from the above, the updating operation can include a modification operation on the existing low bytecode file in the target directory, in which case the above method can further include the following steps: in response to the updating operation on the target directory, locking the mapping relationship corresponding to the target low bytecode file in the target mapping table, the program module corresponding to the locked mapping relationship is prohibited from being instantiated, and the target low bytecode file is the low bytecode file obtained through the updating operation.
[0057] In order to enable the model service request to be implemented based on the instance of the WASM module corresponding to the latest WASM file, the first processing or the second processing, it is necessary to lock the program module corresponding to the modified WASM file, so as to avoid taking the program module that needs to be modified as the target program module, and to protect the concurrent access of the target mapping table, that is, to avoid the influence caused by simultaneously implementing the hot update of the program module based on the target mapping table and the acquisition of the target program model.
[0058] In some embodiments, in response to completing the updating of the target mapping table based on the mapping relationship, the lock of the mapping relationship corresponding to the target low bytecode file in the target mapping table is released.
[0059] In some embodiments, the target data is processed by: passing the location parameter of the target data in the memory to the running target program instance, the target data being written into the memory through the memory buffer of the gateway; and calling the function function corresponding to the target data in the target program instance to perform processing on the target data accessed based on the location parameter through the function function.
[0060] In the present embodiment, when the processing is the first processing, the target data is a model service request, and the function function is a function that converts the target model service request into a standardized provision of an application programming interface of an artificial intelligence service. As an example, taking an HTTP request as an example, the target data can be a URL, a request body and a request header in the HTTP request. Further, the function function includes a function function for rewriting the URL, a function function for rewriting the request header and a function function for rewriting the request body.
[0061] The function function for rewriting the URL is used to implement directional routing and parameter normalization. The directional routing is used to forward the model service request of the gateway to the real interface address of the corresponding artificial intelligence service, and is the basis of request adaptation. The parameter normalization is used to convert the general parameters of the gateway into the URL parameter format required by the artificial intelligence service. Specifically, the function function can be used to rewrite the API of the artificial intelligence service in the domain name, the interface path, the request method and the URL parameter to implement directional routing and parameter normalization.
[0062] The function function of rewriting the request header, the request header is used to carry the meta information of the HTTP request, the function function is used to realize the authentication of the provider of the artificial intelligence service, the format adaptation of the request header information and the compliance adaptation, the function function adds the authentication credential required by the provider of the artificial intelligence service to the request header, so as to ensure that the HTTP request passes the authentication of the requested artificial intelligence service, the function function rewrites the format of the request header information to meet the format requirements of the requested artificial intelligence service, the function function rewrites the data of the request header information to meet the compliance requirements of the requested artificial intelligence service, for example, some artificial intelligence services require that the received request header must remove redundant header information, such as non-standard header information inside the gateway.
[0063] The function function of rewriting the request body, the request body is used to carry the core data of the HTTP request (for example, the user's question and the configuration parameter of the model), the function function is used to realize the format adaptation and the compliance adaptation of the request body, the function function rewrites the format of the request body to convert to the format required by the requested artificial intelligence service, the function function rewrites the data of the request body to meet the compliance requirements of the artificial intelligence service, for example, supplementing the required mandatory parameters of the requested artificial intelligence service and adjusting the parameter value range.
[0064] In the embodiment, when the processing is the second processing, the target data is the model service response, and the function function is a function of converting the model service response into a function that meets the standardized provisions of the gateway. As an example, taking the HTTP response as an example, the target data can be the response header and the response body in the HTTP response, further, the function function of rewriting the response header and the function function of rewriting the response body. Among them, the function function of rewriting the response header and the function function of rewriting the request header realize similar functions, and the function function of rewriting the response body and the function function of rewriting the request body realize similar functions, which will not be repeated here.
[0065] The memory in the embodiment can be a linear memory allocated by the gateway to the target program instance for use, and the linear memories of different target program instances are different.
[0066] In the embodiment, the location parameter of the target data in the memory includes a pointer and a length, the pointer is used to indicate the starting address of the target data in the memory, and the length is used to represent the length of the target data.
[0067] In the embodiment, before the position parameter of the target data in the memory is passed to the running target program instance, the gateway needs to write the obtained target data into the corresponding linear memory, and the target data is written into the memory through the memory buffer of the gateway. Since the target data is directly from the memory buffer to the linear memory of the WASM module, no intermediate copying link is needed, that is, the access of the target program instance to the target data is realized in a zero-copy manner.
[0068] In some embodiments, the target data includes first target data, the first target data is data in a request body in a model service request or data in a response body in a model service response, and the first target data is streaming data. The step of passing the position parameter of the target data in the memory to the running target program instance can include passing the position parameter of the read chunk data of the first target data in the memory to the running target program instance.
[0069] In the above manner, the chunk reading and chunk rewriting of the target data are realized, and the situation that the request body or the response body is too large to be loaded into the memory at one time and causes the memory to be exhausted is avoided.
[0070] In the target program instance, in addition to the conversion of the first target data into the API of the artificial intelligence service or the standardized provision of the gateway, the second target data also needs to be converted into the API of the artificial intelligence service or the standardized provision of the gateway. It should be noted that the second target data is other data in the target data except the first target data. Taking the HTTP request as an example, the second target data includes the request header and the URL. Taking the HTTP response as an example, the second target data includes the response header. The obtaining process of the target model service request is exemplarily described below taking the HTTP request as an example.
[0071] When the gateway receives the HTTP request, a memory allocation function (i.e. a kind of functional function) in the target program instance allocates linear memory of the target program instance for the URL, the request body and the request header in the HTTP request; The gateway passes the position parameters of the linear memories of the URL, the request body and the request header to the target program instance, and calls a stream initialization function (i.e. a kind of functional function) provided by the target program instance to realize the initialization operation, so as to provide a basis for other processing, such as allocating a stream resource, returning a stream identifier to the gateway, etc. The gateway writes the obtained URL, request body and request header into the corresponding linear memory based on the position parameters of the allocated linear memory, wherein the reading and writing of the request body are realized based on the chunk data; The gateway passes the URL, the request header, and the location parameter of the chunk data in the linear memory to the target program instance based on the write result, so that the data processing function in the target program instance accesses the corresponding data from the linear memory based on the location parameter, and performs processing on the accessed data based on the data processing function (i.e., a functional function). When the data is the chunk data of the request body, the data processing function implements rewriting processing of the request body; when the data is the URL, the data processing function implements rewriting processing of the URL; and when the data is the request header, the data processing function implements rewriting processing of the request header.
[0072] In this way, the processing results of the first target data and the second target data can be bound to the same flow identifier, so that the gateway can call the flow result acquisition function (i.e., a functional function) based on the flow identifier, and simultaneously acquire the processing results of the first target data and the second target data. The gateway can also call the flow cleaning function (i.e., a functional function) based on the flow identifier, and synchronously clean up resources related to the processing results of the first target data and the second target data, such as the allocated flow resource, the flow identifier, and the like.
[0073] In some embodiments, the pointer and the length in the location parameter can be packed into one parameter, and the parameter can be passed to the target program instance, so as to reduce the communication cost of function calling and improve the data transmission efficiency.
[0074] In some embodiments, a boundary detection operation is performed before the target program instance accesses the data from the memory based on the location parameter, so as to ensure that the data read by the target program instance does not exceed the pre-allocated memory range, and to avoid data leakage caused by out-of-bound access.
[0075] In some embodiments, the above method can further include the following steps: determining the length of the data to be written, the data to be written being data that needs to be written into the memory; calling a memory allocation function in the target program instance to allocate memory for the data to be written according to the length and the memory instance in the target program instance; writing the data to be written into the memory based on the pointer of the allocated memory; and in response to a failure of the writing, calling a memory release function in the target program instance to release the memory.
[0076] In this embodiment, the data to be written can include at least one of the following: a model service request read by the gateway; rewritten data of the model service request, the rewritten data being data obtained by performing first processing on the model service request; a model service response; rewritten data of the model service response, the rewritten data being data obtained by performing second processing on the model service response.
[0077] In this embodiment, based on the memory instance in the target program instance, it can be determined which memory region needs to be allocated, thereby determining the pointer to the memory to store the data to be written.
[0078] Taking an HTTP request read by the gateway as an example, as described above, the gateway needs to write the read HTTP request (e.g., request headers) into the linear memory of the target program instance. Furthermore, when the gateway writes target data into this linear memory, it needs to request the target program instance so that the target program instance can allocate memory for the HTTP request. The target program instance can return a pointer to this linear memory, allowing the gateway to write the target data based on the starting address indicated by the pointer.
[0079] By using the above method, if a write operation fails, the memory release function in the target program instance is called to release the memory and avoid memory leaks.
[0080] In some embodiments, memory alignment can be performed before writing the data to be written. This memory alignment process checks whether the starting address of the data to be written meets the alignment requirements of the data; for example, if the data to be written is 64 bits, the starting address must be a multiple of 8.
[0081] In some embodiments, the above method may further include the following steps: in response to failure to call the target program instance, obtaining default adaptation logic; and performing corresponding processing on the model service request or model service response through the default adaptation logic.
[0082] The corresponding processing here refers to converting model service requests into standardized provisions that conform to the AI service corresponding to the target program module, or converting model service responses into standardized provisions that conform to the gateway.
[0083] The default adaptation logic can be pre-configured in the gateway. For example, the default adaptation logic can be either an adaptation logic coupled to the gateway or an adaptation logic based on a configuration file.
[0084] It should be noted that the adaptation logic coupled with the gateway is implemented through hard coding, making it impossible to hot update the target mapping table during gateway operation. The configuration file-based adaptation logic can only implement simple field mappings and cannot handle complex transformations, thus limiting the adaptation support capabilities of the model service.
[0085] By using the above method, when the target model service request or response cannot be successfully obtained using the target program instance, the default adaptation logic configured in the gateway is used to perform corresponding fallback processing on the model service request or response, thereby improving the user experience.
[0086] In some embodiments, the interfaces of different program modules are unified, which facilitates the development of WASM files from different providers. These interfaces expose the aforementioned functionalities, enabling the gateway to call these functionalities. Based on the above, the interfaces include the following: The memory management interface is used to allocate memory for writing data and to release memory allocated by the gateway. The model identifier interface is used to return the model identifier corresponding to the program module; The request header modification interface is used to rewrite request headers. The URL modification interface is used to implement URL rewriting processing; The response header modification interface is used to implement response header rewriting. The stream processing interface is used to rewrite and process streaming data.
[0087] In some embodiments, the gateway is configured with a logging interface that allows program modules to output logs of the target program instance during runtime to the gateway. Specifically, the target program instance passes a pointer to the memory location and length of the log to the logging interface, enabling the gateway to retrieve the corresponding log based on the pointer to the memory location and length of the log, where the memory refers to the aforementioned linear memory.
[0088] Figure 3 This is another flowchart illustrating a multi-source artificial intelligence service adaptation method according to an embodiment of this disclosure, which is described below in conjunction with... Figure 3 The embodiments disclosed herein will be further explained and described.
[0089] The monitoring module in the gateway monitors update operations on the target directory. When an update operation is detected, a dynamic hot update is performed on the target mapping table. HTTP requests initiated by clients enter the gateway through the gateway entry point. The matching module in the gateway retrieves the target WASM module that matches the HTTP request. The WASM manager in the gateway manages the target WASM instance, such as creating and destroying the target program instance. The target WASM instance is called to process the HTTP request, obtaining a target request that conforms to the standardized specifications of the target large model service corresponding to the target program module. The target large model service responds to the target request, obtaining an HTTP response. The target WASM instance is called to process the HTTP response, obtaining a target response that conforms to the standardized specifications of the gateway, and the target response is returned to the client.
[0090] Based on the same concept, embodiments of this disclosure provide a multi-source artificial intelligence service adaptation device. Figure 4 This is a block diagram illustrating a multi-source artificial intelligence service adaptation device according to an embodiment of this disclosure, with reference to... Figure 4The multi-source artificial intelligence service adapter 400 includes: Module 401 is used to obtain the target program module that matches the model service request; Module 402 is used to create a target program instance of the target program module, wherein the target program module is loaded based on a low bytecode file and the target program instance supports running in an independent environment. The first processing module 403 is used to perform first processing on the model service request by calling the target program instance running in an independent environment to obtain a target model service request, wherein the target model service request conforms to the standardization provisions of the artificial intelligence service corresponding to the target program module; The second processing module 404 is used to perform a second processing on the obtained model service response by calling the target program instance to obtain a target model service response. The model service response is obtained by calling the artificial intelligence service to respond to the target model service request. The target model service response conforms to the standardization provisions of the gateway.
[0091] Optionally, the multi-source artificial intelligence service adaptation device 400 further includes: The loading module is used during the operation of the gateway to load the low bytecode file associated with the update operation in response to the update operation of the target directory, so as to obtain the program module corresponding to the low bytecode file, wherein the low bytecode file corresponds to a unique model identifier. A construction module is used to construct a mapping relationship between the program module and the model identifier, and to update the target mapping table based on the mapping relationship. The target mapping table is used to support the acquisition of the target program module.
[0092] Optionally, the building module is further configured to: The program module is tested to obtain a test result; if the test result indicates that the test is passed, a mapping relationship between the program module and the model identifier is constructed; the target mapping table is updated based on the mapping relationship.
[0093] Optionally, the update operation includes modification operations on existing low-bytecode files in the target directory, and the multi-source artificial intelligence service adaptation device 400 further includes: The locking module is used to lock the mapping relationship corresponding to the target low bytecode file in the target mapping table in response to the update operation for the target directory. The program module corresponding to the locked mapping relationship is prohibited from being instantiated. The target low bytecode file is the low bytecode file obtained through the update operation.
[0094] Optionally, the target data can be processed in the following ways: The location parameter of the target data in memory is passed to the running target program instance, and the target data is written into the memory through the gateway's memory buffer; Call the function corresponding to the target data in the target program instance, so as to perform processing on the target data accessed from the memory based on the position parameter through the function; When the target data is the model service request, the processing is the first processing; when the target data is the model service response, the processing is the second processing.
[0095] Optionally, the target data includes first target data, which is the request body in the model service request or the response body in the model service response. The first target data is streaming data, and the target data is processed in the following way: the location parameters of the read first target data chunks in memory are passed to the running target program instance; the function corresponding to the chunks in the target program instance is called to process the chunks accessed from memory based on the location parameters.
[0096] Optionally, the multi-source artificial intelligence service adaptation device 400 further includes: The first calling module is used to call the memory allocation function in the target program instance, and allocate target memory for the data to be written according to the length of the data to be written and the memory instance in the target program instance, wherein the data to be written is the data that needs to be written to the memory; A writing module is used to write the data to be written into the target memory based on a pointer to the target memory. The second calling module is used to call the memory release function in the target program instance in response to the write failure, and release the target memory.
[0097] Optionally, the acquisition module 401 is further configured to: based on the model identifier carried in the acquired model service request, determine the target program module that matches the model identifier from the candidate program modules according to the priority of multiple preset matching strategies, wherein the matching accuracy of the preset matching strategy with higher priority is higher than the matching accuracy of the preset matching strategy with lower priority.
[0098] Optionally, the interfaces of different program modules are unified.
[0099] The implementation methods of each module in the multi-source artificial intelligence service adaptation device 400 can be found in the embodiments of the multi-source artificial intelligence service adaptation method, and will not be described in detail here.
[0100] Based on the same concept, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described multi-source artificial intelligence service adaptation method.
[0101] Based on the same concept, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described multi-source artificial intelligence service adaptation method.
[0102] Based on the same concept, embodiments of this disclosure provide an electronic device, including: A storage device on which computer programs are stored; A processing device is configured to execute the computer program in the storage device to implement the steps of the above-described multi-source artificial intelligence service adaptation method.
[0103] The following is for reference. Figure 5 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 1 The diagram shows the structure of the gateway 500. The terminal devices in this embodiment may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0104] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. RAM 503 also stores various programs and data required for the operation of electronic device 500. Processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0105] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0106] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0107] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0108] In some implementations, electronic devices can communicate using any currently known or future-developed network protocol, such as HyperText Transfer Protocol, and can interconnect with digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0109] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0110] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following to occur: acquire a target program module matching a model service request; create a target program instance of the target program module, wherein the target program module is loaded based on a low-bytecode file, and the target program instance supports running in an independent environment; perform a first processing on the model service request by calling the target program instance running in the independent environment to obtain a target model service request, wherein the target model service request conforms to the standardization provisions of the artificial intelligence service corresponding to the target program module; and perform a second processing on the acquired model service response by calling the target program instance to obtain a target model service response, wherein the model service response is obtained by calling the artificial intelligence service to respond to the target model service request, and the target model service response conforms to the standardization provisions of the gateway.
[0111] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0113] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a module does not necessarily limit the module itself; for example, an acquisition module can also be described as "a module that acquires a target program module that matches a model service request".
[0114] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0115] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0117] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
Claims
1. A multi-source artificial intelligence service adaptation method, characterized in that, The method comprises: obtaining a target program module matched with a model service request; creating a target program instance of the target program module, the target program module being loaded based on a low bytecode file, the target program instance supporting running in an independent environment; performing first processing on the model service request by calling the target program instance running in the independent environment to obtain a target model service request, the target model service request conforming to a standardized regulation of an artificial intelligence service corresponding to the target program module; performing second processing on an obtained model service response by calling the target program instance to obtain a target model service response, the model service response being obtained by calling the artificial intelligence service to respond to the target model service request, the target model service response conforming to a standardized regulation of the gateway.
2. The method of claim 1, wherein, The method further comprises: during running of the gateway, in response to an update operation on a target directory, loading a low bytecode file associated with the update operation to obtain a program module corresponding to the low bytecode file, the low bytecode file corresponding to a unique model identifier; constructing a mapping relationship between the program module and the model identifier, and updating a target mapping table based on the mapping relationship, the target mapping table being used to support obtaining of the target program module.
3. The method of claim 2, wherein, The constructing of the mapping relationship between the program module and the model identifier, and the updating of the target mapping table based on the mapping relationship, comprise: detecting the program module to obtain a detection result; in a case where the detection result represents that the detection is passed, constructing the mapping relationship between the program module and the model identifier; updating the target mapping table based on the mapping relationship.
4. The method of claim 2, wherein, The update operation comprises a modification operation on an existing low bytecode file in the target directory, and the method further comprises: in response to the update operation on the target directory, locking a mapping relationship corresponding to a target low bytecode file in the target mapping table, the program module corresponding to the locked mapping relationship being prohibited from being instantiated, the target low bytecode file being a low bytecode file obtained through the update operation.
5. The method of claim 1, wherein, The target data is processed in the following manner: passing a location parameter of the target data in the memory to the running target program instance, the target data being written into the memory through a memory buffer of the gateway; calling a function function corresponding to the target data in the target program instance to perform processing on the target data accessed from the memory based on the location parameter through the function function; in a case where the target data is the model service request, the processing is the first processing, and in a case where the target data is the model service response, the processing is the second processing.
6. The method of claim 5, wherein, The target data comprises first target data, the first target data being a request body in the model service request or a response body in the model service response, and the first target data being streaming data, and the passing of the location parameter of the target data in the memory to the running target program instance comprises: The position parameter of the read first target data block data in the memory is passed to the running target program instance.
7. The method of claim 5, wherein, The method further comprises: A memory allocation function in the target program instance is called to allocate target memory for the to-be-written data according to the length of the to-be-written data and the memory instance in the target program instance, the to-be-written data being data that needs to be written into the memory; The to-be-written data is written into the target memory based on a pointer of the target memory; In response to a failure of the writing, a memory release function in the target program instance is called to release the target memory.
8. The method of claim 1, wherein, The obtaining of the target program module matched with the model service request comprises: Based on the model identifier carried in the obtained model service request, a target program module matched with the model identifier is determined from the candidate program modules according to the priority of the plurality of preset matching strategies, wherein the matching accuracy of the preset matching strategy with higher priority is higher than the matching accuracy of the preset matching strategy with lower priority.
9. The method of any one of claims 1-8, wherein, The interfaces of different program modules are unified. 10.A multi-source artificial intelligence service adaptation apparatus characterized by comprising: Comprise: An obtaining module is configured to obtain a target program module matched with a model service request; A creating module is configured to create a target program instance of the target program module, the target program module being loaded based on a low bytecode file, and the target program instance supporting running in an independent environment; A first processing module is configured to perform first processing on the model service request by calling the target program instance running in the independent environment to obtain a target model service request, the target model service request conforming to a standardized regulation of an artificial intelligence service corresponding to the target program module; A second processing module is configured to perform second processing on an obtained model service response by calling the target program instance to obtain a target model service response, the model service response being obtained by calling the artificial intelligence service to respond to the target model service request, and the target model service response conforming to a standardized regulation of the gateway.
11. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by the processing device to implement the steps of the method of any one of claims 1-9.
12. An electronic device, comprising: Comprise: A storage device having a computer program stored thereon; A processing device configured to execute the computer program in the storage device to implement the steps of the method of any one of claims 1-9.
13. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1-9.