Model service cross-framework request conversion method and device, equipment and medium

Through the cross-frame request conversion method, the protocol incompatibility problem of model inference services in GPU to NPU migration is solved, cross-platform deployment and call compatibility is achieved, transformation costs are reduced, migration efficiency and stability are improved, and it is suitable for fields such as financial technology and medical health.

CN120358269APending Publication Date: 2025-07-22CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510589876.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When the model inference service is migrating from the GPU system to the NPU system, the existing technology faces problems such as incompatibility of service protocols, high cost of call-end transformation, and large differences in request formats. It lacks an efficient, stable and automated request protocol conversion mechanism, resulting in low system stability and online efficiency.

Method used

It provides a cross-frame request conversion method for model service. By receiving client requests, analyzing and extracting input parameters and model identification information, determining the target request path and format based on the identification information, generating target input data, sending it to the target inference service, receiving and converting output parameters, and returning to the client, realizing automatic conversion of cross-frame request protocol and input and output formats.

Benefits of technology

It realizes cross-platform deployment and call compatibility without modifying the call side structure, reduces transformation costs, improves migration efficiency and stability, and is suitable for business scenarios such as financial technology and medical health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358269A_ABST
    Figure CN120358269A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of model deployment, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a model service cross-framework request conversion method, device, equipment and medium, and the method comprises the steps: receiving a reasoning request based on a first model service framework, analyzing input data and a request path, and extracting input parameters and model identification information; determining a target request path of the second model service framework according to the model identification information, and converting the input parameters into a target input format; sending the converted input data to a target reasoning service, receiving a reasoning result and extracting an output parameter; and converting the output parameter into an output format of the first model service framework and returning the output format to the client. According to the method, an adaptation mechanism is introduced between the client and the target model service, so that calling logic which originally depends on a specific reasoning framework can be accessed to heterogeneous reasoning service without changing, and structural compatibility and communication unification of cross-platform deployment and calling of the model service are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model deployment, and in particular, to a method, apparatus, device, and storage medium for cross-frame request conversion of model services. Background Art

[0002] In the process of deep learning model deployment, the model inference service framework plays a key supporting role. Inference service frameworks represented by TFServing have been widely used in various business systems. Especially when deploying models trained by TensorFlow, they can provide relatively complete service capabilities. However, TFServing is essentially designed around the GPU system, and model loading, request format, calculation acceleration strategies, etc. are highly bound to GPU devices. In the context of the increasing richness of current heterogeneous computing devices, its adaptation ability gradually shows limitations.

[0003] With the development of domestic AI chips, in many business scenarios with intensive computing power requirements, more and more deployment requirements are migrating from the GPU system to the NPU system. For example, in the field of medical and health services, in order to accelerate the real-time response capabilities of models such as medical image recognition and auxiliary diagnosis, medical institutions are gradually introducing NPU devices to improve inference efficiency; in the field of fintech, a large number of inference services involving tasks such as risk assessment and credit approval are also being deployed to heterogeneous chip environments to reduce computing power costs and improve system stability. However, since TFServing is not compatible with the inference engines of new chips such as Ascend NPU, the model service system originally built on TFServing cannot directly run in the NPU environment.

[0004] To adapt to NPU devices, it is often necessary to introduce a new generation of inference frameworks such as MindSpore Serving. However, such frameworks have significant differences from TFServing in terms of service interfaces, input / output structures, protocol specifications, etc. Especially when using HTTP or GRPC protocols for remote calls, the calling end often needs to reconstruct the request message, recompile the client code, and even make large-scale modifications to the IDL definition and protocol buffer structure in the GRPC scenario. This poses a great impact on the existing deployment system. Especially in the production environment that has been put into operation, the modification cost of the calling ends distributed in different subsystems is extremely high, and collaborative testing and regression verification are also extremely complex.

[0005] In addition, there is currently a lack of a general adaptation mechanism to achieve protocol conversion and input / output mapping between different inference service frameworks without changing the structure of the calling end. Due to the lack of this ability, during the process of cross-chip and cross-platform migration of model inference services, a large amount of underlying protocol adaptation work still needs to be manually completed by developers, which is extremely likely to lead to problems such as interface inconsistency and request failure, affecting system stability and online efficiency.

[0006] In summary, when the existing technology migrates the model inference service from the GPU system to the NPU system, it faces problems such as service protocol incompatibility, high calling end transformation cost, and large differences in request formats. There is a lack of an efficient, stable, and automated request protocol conversion mechanism, which has become an important obstacle restricting the diversification of inference service deployment and the heterogeneity of computing power platforms. Summary of the Invention

[0007] The main purpose of the present invention is to provide a method, device, equipment, and storage medium for cross-framework request conversion of model services, aiming to solve the technical problem in the existing technology that there is a lack of a mechanism capable of automatically converting request protocols and input / output formats between different model service frameworks without modifying the structure of the calling end.

[0008] To achieve the above object, the present invention provides a method for cross-framework request conversion of model services, including:

[0009] Receiving an initial inference request defined based on a first model service framework sent by a client;

[0010] Parsing the initial input data and the initial request path in the initial inference request, and extracting the input parameters in the initial input data and the model identification information in the initial request path;

[0011] Determining a target request path of a second model service framework according to the model identification information, and mapping the input parameters to the input format defined by the second model service framework to generate converted target input data;

[0012] Sending the target input data and the target request path to a target inference service corresponding to the second model service framework to generate a target inference result;

[0013] Receiving the target inference result returned by the second model service framework, and extracting the output parameters in the target inference result;

[0014] Mapping the output parameters to the output format defined by the first model service framework to generate converted target output data;

[0015] Returning the target output data to the client.

[0016] Furthermore, to achieve the above object, the present invention provides a model service cross-frame request conversion device, including:

[0017] A request receiving module, configured to receive an initial inference request defined based on a first model service framework sent by a client;

[0018] A request parsing module, configured to parse initial input data and an initial request path in the initial inference request, and extract input parameters in the initial input data and model identification information in the initial request path;

[0019] An input mapping module, configured to determine a target request path of a corresponding second model service framework according to the model identification information, and map the input parameters into an input format defined by the second model service framework to generate converted target input data;

[0020] An inference forwarding module, configured to send the target input data and the target request path to a target inference service corresponding to the second model service framework to generate a target inference result;

[0021] A result receiving module, configured to receive the target inference result returned by the second model service framework, and extract output parameters in the target inference result;

[0022] An output mapping module, configured to map the output parameters into an output format defined by the first model service framework to generate converted target output data;

[0023] A response returning module, configured to return the target output data to the client.

[0024] Furthermore, to achieve the above object, the present invention further provides a computer device, where the computer device includes a memory, a processor, and a model service cross-frame request conversion program stored on the memory and executable on the processor. When the model service cross-frame request conversion program is executed by the processor, the steps of the model service cross-frame request conversion method as described above are implemented.

[0025] Furthermore, to achieve the above object, the present invention further provides a computer-readable storage medium, where a model service cross-frame request conversion program is stored on the storage medium. When the model service cross-frame request conversion program is executed by a processor, the steps of the model service cross-frame request conversion method as described above are implemented.

[0026] Beneficial effects: The present invention relates to the technical field of model deployment and can be applied to business scenarios such as fintech and healthcare. It discloses a method for cross-frame request conversion of model services, including: receiving an initial inference request defined based on a first model service framework sent by a client, parsing the initial input data and the initial request path, and extracting input parameters and model identification information; determining the target request path of the second model service framework according to the model identification information, and converting the input parameters into target input data; sending the target input data and the target request path to the target inference service corresponding to the second model service framework, receiving the target inference result and extracting output parameters; converting the output parameters into the output format defined by the first model service framework, generating target output data and returning it to the client. By introducing an adaptation mechanism between the client and the target model service, the present invention completes the whole process of model identification extraction, input parameter structure conversion, request protocol reconstruction, and response format restoration, enabling the call logic originally dependent on a specific inference framework to access heterogeneous inference services without modification, realizing the structural compatibility and communication unity of cross-platform deployment and call of model services, thereby effectively reducing the transformation cost and improving the migration efficiency and stability of the inference service among heterogeneous computing power platforms. Brief Description of the Drawings

[0027] The following will further illustrate the present invention in conjunction with the drawings. In the drawings:

[0028] Figure 1 It is a schematic diagram of an application environment of the method for cross-frame request conversion of model services in an embodiment of the present invention;

[0029] Figure 2 It is a schematic flowchart of an embodiment of the method for cross-frame request conversion of model services of the present invention;

[0030] Figure 3 It is a schematic diagram of functional modules of a preferred embodiment of the device for cross-frame request conversion of model services of the present invention;

[0031] Figure 4 It is a schematic diagram of the structure of a computer device in an embodiment of the present invention;

[0032] Figure 5 It is another schematic diagram of the structure of a computer device in an embodiment of the present invention. Detailed Embodiments

[0033] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] The method for cross-frame request conversion of model services provided by the embodiments of the present invention can be applied in, for example Figure 1In the application environment, the client communicates with the server through the network. The server can receive the initial inference request defined based on the first model service framework sent by the client through the user terminal, parse the initial input data and the initial request path, and extract the input parameters and the model identification information; determine the target request path of the second model service framework according to the model identification information, and convert the input parameters into the target input data; send the target input data and the target request path to the target inference service corresponding to the second model service framework, receive the target inference result and extract the output parameters; convert the output parameters into the output format defined by the first model service framework, generate the target output data and return it to the client. By introducing an adaptation mechanism between the client and the target model service, the present invention completes the whole process of model identification extraction, input parameter structure conversion, request protocol reconstruction and response format restoration, so that the call logic originally dependent on a specific inference framework can be accessed to heterogeneous inference services without modification, realizing the structural compatibility and communication unity of cross-platform deployment and call of model services, thereby effectively reducing the transformation cost and improving the migration efficiency and stability of the inference service between heterogeneous computing power platforms. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0035] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the method for converting cross-frame requests of model services provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.

[0036] As Figure 2 shown, the method for converting cross-frame requests of model services proposed by the present invention includes the following steps:

[0037] S10, receiving an initial inference request sent by a client and defined based on a first model service framework;

[0038] In this embodiment, receiving the initial inference request defined based on the first model service framework sent by the client is the starting link of the entire model service structure conversion process. Its purpose is to accurately obtain and carry the original model inference request without interrupting the client's existing call method, and establish a unified entry for subsequent format conversion and cross-frame adaptation. The client here refers to the model call initiator in the business system, which may be deployed at positions such as microservice units, front-end transfer interfaces, or data processing flow nodes. Its requests usually carry structured input data and strictly follow the service path specifications and parameter format requirements defined by the first model service framework. The first model service framework here refers to the model service deployment specification closely associated with the request structure, such as TensorFlow Serving. Its input path generally follows a fixed REST style (such as / v1 / models / model_name:predict), and the request body is a JSON object containing one or more inference instances, and each instance is a set of field collections consistent with the model input definition.

[0039] To accurately receive the above inference request, it is necessary to construct a service listening mechanism for the access layer. The service listening mechanism first establishes a listening channel at the specified port through a network socket, and the protocol type is usually one of the two mainstream inference service call protocols, HTTP or GRPC. In the HTTP protocol scenario, the listening logic can be constructed through a general HTTP server component (such as FastAPI, Flask, or NGINX, etc.), while in the GRPC protocol scenario, it is necessary to generate a service interface definition based on Protocol Buffers and use the GRPC framework to bind the listening port and the service implementation class. On this basis, prepare corresponding message decoders for different protocols. HTTP request decoding usually identifies JSON or form structures based on the Content-Type type, while GRPC decoding needs to restore the byte stream to a service call message object according to the Protocol Buffers defined structure.

[0040] The key contents of the initial inference request include the request path, input data, request method, etc. Among them, the request path is used to specify the target model and service version and is the basis for subsequent path mapping; the input data contains all the input fields required for model inference and is the core source of format conversion. To ensure that the request can be received and recognized by the system, the receiving end also needs to implement mechanisms such as interface signature verification and data integrity check, and filter abnormal request structures.

[0041] In the scenario of accessing model services based on HTTP, the Flask framework can be used as a service listener. When the service starts, it binds to a specified port and registers API routes for receiving inference requests. After receiving an inference request, use request.get_json() to extract the JSON-formatted data body, and at the same time parse the URL path to obtain the model name and service version information. To improve concurrency performance, a WSGI server such as Gunicorn can be used as the concurrency scheduler for the Flask service to support multi-process processing.

[0042] In the GRPC-based scenario, service-side interface code can be generated using Protocol Buffers based on the service interface definition file, and the service listening can be started through the GRPC framework. After receiving the client byte stream request, it is automatically decoded into a predefined data object, and information such as the model name, version, and call parameter fields are extracted. For the unfixed input structure of GRPC, a path-level verification mechanism can be added to ensure that the model identifier in the path is consistent with the system support scope.

[0043] To adapt to the diversity of inference request structures in different business scenarios, the access layer can be designed as a plug-in structure. It judges the protocol type according to the Header parameters sent by the client and dynamically switches the receiving and parsing strategies to achieve multi-protocol hybrid access. At the same time, dynamic thresholds are set for request frequency, data size, etc. to prevent abnormal call behaviors from affecting service stability.

[0044] Example illustration: In the field of medical and health, a certain hospital has deployed a medical image recognition model based on TFServing to receive the image feature data uploaded by the front-end inspection equipment and return the prediction result of the malignant probability of pulmonary nodules. This model service provides an inference interface using the HTTP protocol, and the request path format is / v1 / models / nodule_model:predict. The input data is a JSON structure containing image encoding, and the format is as follows:

[0045]

[0046] When the hospital decides to migrate model inference to Ascend NPU devices to reduce costs and improve throughput, the original TFServing service cannot run in the new hardware environment. To maintain the existing client request format unchanged, the intermediate service will first listen for client requests on the HTTP port. After receiving an inference request with the above structure, it will receive it as an initial inference request. After parsing the request, the service extracts the model name "nodule_model", the service interface "predict", and the model input "image_tensor". By encapsulating these contents into a structured request object, the intermediate service can complete the adaptation process with the subsequent target model service without modifying any code in the hospital's front-end system.

[0047] In the financial field, a risk control platform calls a deployed credit scoring model through the GRPC protocol. The original service is built based on the TFServing framework, with the request path being / v1 / models / credit_model:predict, and the request body including multiple numerical behavioral characteristics such as transaction frequency, number of device changes, and login location distribution. These characteristics are encapsulated into a Protocol Buffers structure and sent to the inference service in GRPC form. When the financial institution decides to adopt MindSpore Serving optimized and deployed based on Ascend chips to improve the scoring efficiency, the GRPC request structure defined by the original TFServing cannot be directly transmitted to the new service. Under the action of this mechanism, when the system receives a GRPC request, it automatically parses the original byte stream according to the service interface generated by the Proto file, obtains the content of the input fields, and extracts the model name and interface call method. At the same time, through preset access rules, format recognition and structure restoration are completed to ensure that the inference request is accurately received, laying a structural foundation for subsequent input parameter conversion and cross-service framework adaptation. This method can keep the call interface between the original business system and the scoring engine unchanged in the risk approval scenario, significantly shortening the cycle from model optimization to online, and reducing testing and collaboration costs.

[0048] By establishing a structured and configurable service reception mechanism, the initial inference requests defined by the first model service framework can be received stably and accurately, realizing the unified bearing and recognition of the request path, input data, and communication protocol, providing a standardized entry for subsequent structure conversion, thus avoiding changes to the structure of the calling end and reducing the complexity of system transformation.

[0049] S20, parse the initial input data and the initial request path in the initial inference request, and extract the input parameters in the initial input data and the model identification information in the initial request path;

[0050] In this embodiment, during the inference request adaptation process, it is necessary to first completely parse the original data structure included in the initial inference request. This process mainly includes the structured extraction of the initial input data and the disassembling and identification of the model information included in the initial request path. The initial inference request is generally a standard call request generated by the client based on the first model service framework, and this request includes two key parts: one is the data body used to represent the model input, and the other is the request path used to identify the service call target.

[0051] The initial input data usually has a nested structure, such as multi-level dictionaries and array nesting in the JSON format, or Protocol Buffers objects in the GRPC scenario. In actual parsing, recursive parsing or structure expansion techniques are required to convert the nested structure into a sequence of key-value pairs. The keys in the key-value pairs are usually input field names, such as "input_ids", "image_tensor", etc., and the corresponding values are specific tensors or feature arrays. These fields must be consistent with the input key names defined by the first model service framework to be regarded as valid input parameters. The extraction logic of the input parameters depends on the preset input field rules or field template mapping relationships, and supports filtering target fields in complex nested structures and extracting them into a set of structured input parameters for subsequent format conversion.

[0052] The initial request path is used to identify the model requested by the client, its version, interface actions, and other key information. Under the HTTP protocol, it usually appears in a path format similar to " / v1 / models / model_name / versions / version_id:method", while in the GRPC protocol, it is implicitly expressed through the interface service name and method name. To support the unified extraction of cross-protocol structures, it is necessary to normalize the path string, such as splitting and extracting the model name, model version number, interface identifier, and other contents in the path. The extracted information is uniformly encapsulated as model identification information and is used as the core identification field throughout the protocol adaptation process. The generation of the model identification information does not depend on the specific field order and has a certain fault tolerance ability to adapt to variant formats of different client request paths.

[0053] To enhance the adaptation stability and request generalization ability, this process can also introduce mechanisms such as field regular matching and path template reverse parsing to achieve compatible recognition of non-standard input structures. In a multi-model deployment environment, the model identification information can also be associated with the model mapping table maintained in the service registry or configuration file to achieve dynamic routing.

[0054] In one specific implementation, for an inference request using the HTTP protocol, the reverse proxy service can be first called to intercept the original request packet, extract the Body part as the initial input data, and at the same time obtain the URI path as the initial request path. Then, the JSON parser is used to expand the structure hierarchically, extract all leaf node fields, and compare them with the preset input key name table to screen and extract them as the input parameter set. In the URI path parsing, string slicing or regular parsing methods are used to identify key identifiers such as "models", "versions", "predict", etc., and the model name and version number are extracted from them and combined into the standardized model identification information.

[0055] In another specific implementation, for an inference request using the GRPC protocol, the Proto definition file of the first model service framework can be pre-loaded to generate the corresponding service interface parsing class. After receiving the GRPC byte stream request, these classes are used to deserialize the request into a structured object, and information such as model input fields, interface call methods, service names, etc. are extracted from the object structure to construct the input parameter set and model identification information. In a high-concurrency environment, this process can be processed concurrently through a thread pool or an asynchronous task management mechanism to reduce blocking latency.

[0056] An abstract parsing module can also be built based on the multi-protocol unified processing engine, and the corresponding parsing method can be dynamically selected according to the protocol header fields when the request is accessed, improving the versatility and scalability of the system.

[0057] Example illustration: In the medical and health scenario, the front-end imaging device sends a lung CT image inference request based on the TFServing standard format, the request path is / v1 / models / lung_ct:predict, and the input data is a JSON structure containing a three-dimensional image matrix. After the intermediate service receives this request, it parses the URI path to extract the model name "lung_ct" and the interface identifier "predict", and at the same time traverses all the image fields under the "instances" field in the JSON structure, identifies the "image_tensor" field and extracts the image content as the input parameter set.

[0058] In the field of financial risk control, a certain credit approval system calls the credit assessment model service through GRPC. The request message structure contains fields such as user behavior sequence, historical overdue records, and device binding status. The adaptation service extracts the input field mapping relationship through the GRPC interface parser, identifies the information containing key-value pairs such as "device_change_count" and "overdue_days" as the input parameter set, and at the same time extracts the model name "credit_score" and the interface action "Predict" through the service path and method name, and assembles them into model identification information. This operation provides structural support for the input-output structure adaptation and call link conversion between different inference engines, ensuring the smooth migration and stable operation of the inference service.

[0059] By performing structured parsing and field extraction on the initial input data and the initial request path, the accuracy of the model input and the unique identification of the target model can be ensured, and a standardized input parameter set and model identification information can be constructed, laying a structural foundation for subsequent format conversion and protocol adaptation. This processing method does not rely on client code adjustment and can access the original request in a non-intrusive manner, improving the adaptation versatility and system deployment flexibility.

[0060] S30. Determine the target request path of the corresponding second model service framework according to the model identification information, and map the input parameters to the input format defined by the second model service framework to generate the converted target input data;

[0061] In this embodiment, determining the target request path of the second model service framework according to the model identification information means reconstructing the path information (such as model name, version number, interface identifier, etc.) originally in the first model service framework and adapting it to the path specification defined in the second model service framework. Since there are differences in the composition rules of the request paths of different frameworks, for example, TFServing uses the RESTful style path / v1 / models / {model} / versions / {version}:{method}, while other frameworks such as MindSpore Serving may use the combination mode of / models / {model} / instances / {version} / {method}, so the path conversion needs to be based on the mapping template mechanism.

[0062] Path conversion usually requires the help of a predefined path template mapping table, which defines the mapping relationship of path patterns between different model service frameworks. Through the interface identifier in the model identification information, a matching template is selected in the mapping table, and placeholders are replaced with fields such as model name and version number to generate the target request path that conforms to the target framework. This method supports dynamically expanding the support capabilities of new frameworks and avoids the maintenance costs brought by hard-coded path logic.

[0063] On the other hand, mapping the input parameters to the input format defined by the second model service framework mainly involves two aspects: renaming of input fields and restructuring of input structures. There are inconsistencies in the naming of input fields among different model service frameworks. For example, the field named "input_ids" in TFServing may be named "input_tensor" in the target framework. Therefore, through an input key name mapping table, the key names in the original input parameter set are replaced with the corresponding field names in the target framework.

[0064] In addition, the organizational form of the input structure may also vary. TFServing often uses a flat or list structure nested under the "instances" field, while frameworks such as MindSpore Serving may require fields to appear nested at different levels. Therefore, after completing the key name mapping, according to the input format rules of the second model service framework, the structure needs to be re-encapsulated, such as constructing nested arrays, merging fields into a unified structure, or adjusting the field order. This process can be regarded as structure reorganization, and the goal is to generate a data object that fully conforms to the input format requirements of the target inference service, facilitating the correct identification and processing by the subsequent inference engine.

[0065] In a specific implementation, a path template mapping table and an input key name mapping table can be pre-maintained. The former binds path templates of different frameworks to interface actions, and the latter establishes a mapping relationship of input field names between the first and second model service frameworks. First, the system selects the path template of the second framework (such as / models / {model} / versions / {version}:{method}) in the path template mapping table according to the interface identifier (such as "predict") in the model identification information, and replaces it with the model name and version number to form the target request path. During the mapping of input parameters, first traverse the original input parameter set. For each key-value pair, query the input key name mapping table, convert its key name to the target field name, and keep the value structure unchanged. After the mapping is completed, according to the input format specification of the second model service framework, combine the mapped fields into the target input data. For example, if the target framework requires all fields to be encapsulated in a nested object named "input_data", insert all fields into this structure in the order of field names and generate standard format data.

[0066] In the scenario of multi-framework parallel support, it can also be designed as a plug-in structure, that is, define independent path conversion plug-ins and input format constructors for each model service framework, and dynamically load them at runtime according to the framework type in the model identification information to improve generality and maintainability.

[0067] By generating the target request path based on the model identification information and converting the input parameters into the target input data format, the call request can be seamlessly adapted from the first model service framework structure to the second model service framework structure without modifying the call-side logic or the model ontology, ensuring the portability and platform independence of the model service. This operation establishes a unified communication bridge between different inference frameworks, improves the flexibility and migration efficiency of model inference deployment, and reduces the cross-platform adaptation cost.

[0068] S40. Send the target input data and the target request path to the target inference service corresponding to the second model service framework to generate a target inference result;

[0069] In this embodiment, sending the target input data and the target request path to the target inference service corresponding to the second model service framework is a key link in the implementation of the inference execution process, which involves the selection of communication protocols, the construction of request bodies, the invocation of remote services, and the reception of response data. The target input data is an adapted format data object generated based on the input parameter structure conversion in the previous step, and the target request path is the service entry address constructed according to the target framework path specification. The core is to combine the two into one, construct an inference call request that conforms to the communication specification of the second model service framework, and achieve the secure and reliable reception of the results.

[0070] Different model service frameworks may support multiple communication protocol types, commonly including HTTP, HTTPS, gRPC, etc. Therefore, before sending, it is first necessary to determine the applicable transport protocol type according to the protocol identifier included in the target request path or external configuration. Taking gRPC as an example, a request object that conforms to the Protocol Buffers definition needs to be generated and sent through the established channel; while in the HTTP scenario, a standard JSON request body needs to be constructed and a POST call is made through the REST interface.

[0071] During the process of constructing the request body, the request format defined by the second model service framework needs to be followed. For example, in MindSpore Serving, it may be required to nest the input data under fields named "instances" or "inputs", and at the same time, request header fields such as "X-Model-Name", "X-Version-ID", etc. may also need to be set to assist the service side in identifying the specific model version. When serializing the request content, a JSON serializer or ProtoBuf serializer should be selected according to the protocol, and the request body should be compressed, encrypted, or signed (if the target framework has requirements for transmission security).

[0072] After the request body is sent, the system needs to receive the response data returned by the target inference service. The return result usually includes two parts: the status code and the inference output. The status code is used to judge whether the service has processed the request normally. Commonly, 200 in HTTP indicates success, 400 or 500 indicates an exception, and in gRPC, internal status enums are used for identification. The inference output is a data field containing the model calculation results, usually encapsulated in a structured format (JSON object or ProtoBuf structure) for the data processing module in the next stage to parse and extract.

[0073] In one implementation, the adaptation module makes a judgment based on the protocol prefix in the target request path. For example, if it starts with "grpc: / / ", the gRPC protocol is selected, otherwise, HTTP is used by default. In the gRPC mode, a request object is generated through the IDL definition corresponding to the target service, the target input data is filled into the message body, and the preset service interface method is called to initiate the request, and the blocking mode or asynchronous callback mechanism is used to wait for the response to return.

[0074] In the HTTP scenario, the adaptation module first constructs a JSON request body containing the target input data and concatenates the complete target request path through the URL path. For example, it combines the target address http: / / npu-server:8500 with the path / models / model_x / versions / 3:predict to form the final URL. Subsequently, it calls an HTTP client tool such as the requests library (Python) or the curl command (Shell) to make a POST request and sets necessary request header parameters such as "Content-Type: application / json", "Authorization", etc. After the request is completed, it parses the response status code and return body, records or reports exception information (such as timeouts, network errors, service denials) and encapsulates them into an internal unified error format.

[0075] When the target inference service returns response data, the system extracts the original inference result field from it and simultaneously reads the status information. If the status code indicates failure, the system can automatically re-initiate the request according to the preset retry policy, supporting retry count control, exponential backoff waiting, integration of failure reporting and circuit breaker policies, and enhancing stability and robustness in the production environment.

[0076] By sending the target input data and the target request path to the target inference service and combining the processes of protocol recognition, request body encapsulation, communication interface invocation, and result reception, a unified and transparent inference invocation intermediate layer can be constructed, avoiding direct coupling between the calling end and various model service frameworks. This mechanism improves the portability of requests and the cross-framework execution ability, maintains the consistency and stability of the communication process in a heterogeneous deployment environment, reduces the impact of communication exceptions on the overall inference process, and enhances the system's adaptability to the diversity of model deployment platforms.

[0077] S50, receive the target inference result returned by the second model service framework and extract the output parameters in the target inference result;

[0078] In this embodiment, receiving the target inference result returned by the second model service framework and extracting the output parameters is the key bridge between model execution and result standardization. This process can consist of two closely related stages. One is the reception and format parsing of the response data, and the other is to extract the output parameters with semantic meaning from the response structure and construct them into a unified result expression. The reception of the response data usually depends on the return mechanism defined by the transport protocol (such as HTTP or gRPC). In this process, the issue of inconsistent response formats between heterogeneous frameworks should be considered, such as differences in field naming, different depths of nested structures, and redundancy of unstructured additional fields.

[0079] After successfully receiving the response data, it is necessary to perform structured parsing on it according to the response format defined by the second model service framework. Specifically, the system needs to extract the status code field and the inference result body from the response data. The status code is used to determine whether the request has been successfully completed, and the inference result body may be a JSON object, a Protocol Buffers sequence structure, or a more complex nested data collection. The inference result structure may contain multiple fields, such as prediction results, confidence levels, label serial numbers, etc. To adapt to the requirements of constructing a unified output format in the future, it is necessary to filter out the truly valuable output parameter fields from the response results and ensure that their data types and semantic consistency meet the expectations of the target system.

[0080] In the stage of extracting output parameters, the system needs to introduce an output field screening mechanism. This mechanism usually relies on a predefined output key name mapping table or a field whitelist provided by the second model service framework. The mapping table is used to identify which fields in the target inference service response need to be retained and provides a basis for field name conversion. The extraction operation not only needs to judge whether the field exists in the mapping table but also needs to combine the required field list to judge the retention priority of the field. For fields declared in the mapping table but missing in the response, a default value mechanism or an exception path can be introduced. For fields that are not declared and are not required output fields, they should be explicitly excluded to prevent format pollution or field misuse.

[0081] The extracted output parameters are usually encapsulated into a structured set of key-value pairs for subsequent standard format conversion steps to call. In actual implementation, it is also necessary to perform type verification and legality judgment on the values of the fields to avoid parsing errors or type inconsistency problems. For example, the output fields of a classification task should be a numeric array or classification labels. If the return is a nested dictionary or an inappropriate type, it is necessary to repair or mark the error according to the preset logic.

[0082] In an implementation method under a gRPC protocol, the response body returned by the target inference service is a message structure defined by ProtoBuf. The system first decodes the response body using the automatically generated deserialization interface of gRPC. After parsing, the system obtains a message object containing several fields, extracts fields such as outputs, scores, labels, etc. according to the interface method return definition, and determines the fields to be retained and rename them according to the configured field mapping table.

[0083] In the HTTP protocol scenario, the results returned by the inference service are usually in JSON format. The adaptation module uses a standard JSON parser to parse the response body and traverses the fields item by item. During the traversal, candidate fields such as result, output, predictions, etc. are first read, and it is determined whether to retain and how to rename and convert them according to the pre-stored mapping rules. For example, result is mapped to prediction_result, and it is verified whether its value is a floating-point array or a scalar value. If it is not the expected type, type conversion is performed according to the configuration or it is marked as an exception.

[0084] In the multi-field output scenario, a strategic extraction mechanism can be adopted. For example, the top k items are retained by sorting according to confidence, or the output structures of multiple models are grouped and extracted according to the main model priority strategy. For image classification tasks, the label_id and confidence fields can be extracted and combined into a composite output parameter; for financial scoring tasks, the risk_score field can be extracted as the key output.

[0085] By receiving the target inference results and extracting output parameters based on field mapping and format parsing strategies to construct a set of output parameters with consistent structures, the format unification and data standardization of cross-frame results can be achieved without changing the logic of the downstream calling end. This mechanism enhances the interface stability and maintainability of the system in heterogeneous deployment environments, avoids problems such as chaotic result fields or incorrect use of inference results caused by differences in model service frameworks, and improves the robustness and scalability of the overall system.

[0086] S60, map the output parameters to the output format defined by the first model service framework to generate the converted target output data;

[0087] In this embodiment, mapping the output parameters to the output format defined by the first model service framework and generating the converted target output data is the final key step in the process of realizing the result landing in the model inference adaptation. The essence of this process is to perform format conversion, structure reorganization, data type standardization, and meta-information supplementation on the set of structured output parameters extracted in the previous stage according to the format specifications required by the target framework without changing the semantic expression, so as to construct a data output result that meets the recognition requirements of the target caller.

[0088] First, this mapping behavior involves the standardization of data types. That is, the values of output parameters may have different type precisions or representation structures under different inference frameworks, such as integer and floating-point, boolean and string, array and nested object, etc. To ensure that the target system can correctly receive and parse these parameters, it is necessary to perform type conversion on the value of each output parameter according to the type specifications of the first model service framework. For example, when the first model service framework requires the classification label to be represented as a string, and the current output parameter is an integer label number, a classification dictionary must be introduced to complete the conversion from the number to the label and convert the result to a string type.

[0089] Next, the nested restructuring of the output structure is required. Since different frameworks organize the hierarchical structure of output fields in different ways. For example, TensorFlow Serving may expect the data to be organized as {"outputs":{"label":"cat","score":0.98}}, while some frameworks may use a flat structure like {"label":"cat","score":0.98}. The adaptation system needs to re-nest the parameters according to the nested rules of the target framework, and may even need to encapsulate them under a unified namespace field or use specific hierarchical tags for wrapping (such as "results", "data").

[0090] After completing the data structure restructuring, the adaptation module also needs to insert necessary metadata fields, such as inference timestamps, model version numbers, request tracking IDs, etc. These fields are generally defined uniformly by the first model service framework and are commonly used to trace the call chain, trace the source of results, or for subsequent log analysis and system monitoring. During the process of inserting metadata fields, it is necessary to ensure that they are clearly separated from business fields and follow a unified format specification. For example, the timestamp is in ISO standard format or Unix time format.

[0091] After the final output structure is constructed, it also needs to be converted into the corresponding byte stream form through a serialization protocol for subsequent network transmission or data writing. Different first model service frameworks may use different serialization methods. For example, TensorFlow Serving usually uses the JSON format, while some systems using gRPC require serialization in the ProtocolBuffers format. The adaptation logic needs to dynamically select a serializer according to the transmission protocol identifier of the request channel to complete the data conversion.

[0092] By converting the output parameters to the target format, completing the metadata, and serializing the output as required, the problem of different response result formats between different model service frameworks is effectively solved. The unified output structure generation across the inference framework is achieved, avoiding the need for the caller to rewrite the parsing logic for different frameworks, reducing the system coupling and maintenance complexity. Further, through protocol adaptation and dynamic selection of serialization formats, the system's compatibility and elastic expansion capabilities for heterogeneous deployment environments are enhanced.

[0093] S70: Return the target output data to the client.

[0094] In this embodiment, returning the target output data to the client is the final stage in the model reasoning adaptation process, and its main function is to complete the delivery of results and the communication loop. This process is not only an act of data transmission, but also involves multiple key sub-links such as channel selection, encoding processing, return format configuration, and response status generation for the target output data. The target output data is generally a result byte stream after format mapping, structural reorganization, and serialization processing. The byte stream needs to be returned through a protocol stack acceptable to the client and comply with the client's communication specification requirements for multiple dimensions such as data packet structure, status code, and content type.

[0095] First, you need to clarify the type of the return channel. Usually, when the model service receives the initial inference request from the client, the adaptation system has established a communication channel based on the transport protocol identifier (such as HTTP or gRPC). The context structure of the channel contains the client request object, connection information, and the expected response format. The return process needs to maintain the consistency of the channel state and correctly associate the response context. For example, in the HTTP scenario, data needs to be returned through the HTTP response body with the correct Content-Type field (such as application / json); in the gRPC scenario, the results need to be written to the return stream and encoded and encapsulated according to the proto definition format.

[0096] Secondly, before writing the target output data in the form of a byte stream to the return channel, the system usually needs to construct the response header information, including auxiliary fields such as the response status code, request tracking ID, and model version information. The choice of status code is usually determined by the response verification result in the previous reasoning process. If the target output data verification passes, 200 (HTTP) or OK (gRPC) is returned; if the output field is missing or the data is damaged, the corresponding error code is returned with error description information. The tracking ID is usually consistent with the request identifier recorded in the initial reasoning request, and is used for link log positioning or multi-system request context tracking.

[0097] To support multiple client types and system environments, the return process may also include extended capabilities such as output compression, streaming output support, and sharded return. For example, when the returned data is large or contains multiple segments of inference results, the system can enable GZIP compression or transmit the byte stream in segments to improve transmission efficiency and client compatibility.

[0098] One implementation is to implement client response transmission through a context-based general return module. This module reads the transport protocol identifier recorded in the initial request from the current thread context and dynamically selects the corresponding response processor. If it is the HTTP protocol, the system uses the response object built into the standard Web framework, writes the header information with Content-Type as application / json, then writes the target output data byte stream into the response body, and finally triggers the commit or flush mechanism to push the response to the client; if it is the gRPC protocol, the method call mode of the stub object is used, a Response object is constructed according to the return format structure of protobuf, and the response stream write method is called for data writing and the current RPC call context is terminated.

[0099] Another implementation is applicable to edge inference service nodes deployed in a cloud-native environment. The adaptation module runs as middleware in the service gateway. It encapsulates the target output data into a standard JSON response body through reverse proxy and uniformly encapsulates additional fields such as status codes and metadata into the top-level fields of the response body to return to the client, adapting to front-end microservice call components with different versions or protocol stacks.

[0100] It is also possible to set response templates for different client types through configuration files. For example, mobile, Web, or third-party API docking systems use different data formats and field naming conventions. The adaptation module completes compatibility adaptation for multiple callers through the structure conversion of the target output data.

[0101] Through the unified data return mechanism, not only is the complete delivery of the output data by the inference service adaptation system achieved, but also the client communication compatibility and result structure consistency are effectively guaranteed. This mechanism avoids the overhead of clients repeatedly parsing or reconstructing the response structure, supports automatic channel selection and response automatic formatting in multiple protocol scenarios, and further enhances the cross-platform deployment ability and system stability of the inference service.

[0102] The present invention relates to the technical field of model deployment and can be applied to business scenarios such as fintech and healthcare. It discloses a method for cross-frame request conversion of model services, including: receiving an initial inference request defined based on a first model service framework sent by a client, parsing the initial input data and the initial request path, and extracting input parameters and model identification information; determining the target request path of a second model service framework according to the model identification information, and converting the input parameters into target input data; sending the target input data and the target request path to a target inference service corresponding to the second model service framework, receiving the target inference result and extracting output parameters; converting the output parameters into an output format defined by the first model service framework, generating target output data and returning it to the client. By introducing an adaptation mechanism between the client and the target model service, the present invention completes the whole process of model identification extraction, input parameter structure conversion, request protocol reconstruction, and response format restoration, enabling the call logic originally dependent on a specific inference framework to access heterogeneous inference services without modification, realizing the structural compatibility and communication unity of cross-platform deployment and call of model services, thereby effectively reducing the transformation cost and improving the migration efficiency and stability of inference services among heterogeneous computing power platforms.

[0103] In one embodiment, step S10 above includes:

[0104] S101, starting a service listening program to establish a receiving channel facing the client and setting a transmission protocol identifier;

[0105] S102, initializing a corresponding protocol parser based on the transmission protocol identifier;

[0106] S103, listening to the receiving channel to wait for a connection request initiated by the client;

[0107] S104, accepting the connection request of the client and establishing a communication link;

[0108] S105, receiving the original data stream of the initial inference request through the communication link;

[0109] S106, using the protocol parser to decode the original data stream to generate a structured request object;

[0110] S107, performing data integrity verification on the structured request object;

[0111] S108, after the data integrity verification passes, verifying whether the interface fields of the structured request object match the format policy defined by the first model service framework;

[0112] S109. If the interface field of the structured request object matches the format policy defined by the first model service framework, mark the structured request object as a valid request and generate a unique request identification value for each valid request.

[0113] S110. Write the valid request and the corresponding request identification value into the request buffer.

[0114] In this embodiment, receiving the initial inference request sent by the client based on the first model service framework is the starting link of the model service request adaptation process. Its core role is to build a communication basic environment, identify the request structure, ensure data validity, and complete the preliminary enqueue processing.

[0115] First of all, starting the service listener program is a prerequisite for establishing the client communication ability. Usually, the service process is bound to a predefined port to listen for external network connections, and requests can arrive in different protocol forms such as HTTP, HTTPS, and gRPC. The listener program is usually integrated in the server framework, such as Flask in Python, Spring Boot in Java, or the gRPC server library in C++. The service listener will also initialize the transport protocol identifier, which not only reflects the underlying transport protocol type but also can identify the request semantic model, such as the RESTful path, gRPC service signature, etc.

[0116] The initialization of the protocol parser depends on the transport protocol identifier, and this parser determines the reading strategy and decoding logic for the original data stream. For example, HTTP requests use a structure parsing method based on Header + Body, while gRPC uses Protocol Buffers for structured parsing. The protocol parser must be able to automatically identify the data boundary, content type, and generate a unified structure for subsequent processing according to the request type.

[0117] The process of listening and establishing a connection for the receiving channel needs to support two processing models: asynchronous or blocking, and can be flexibly selected according to service pressure and communication delay requirements. When the client initiates a connection, the listener program needs to complete operations such as handshake, protocol version negotiation, and initialization of the connection state, and finally establish a readable and writable data transmission channel. Some frameworks support encryption methods such as TLS to protect data security.

[0118] The original data stream of the initial inference request is usually an undecoded byte stream. After being received, it is converted into a structured request object by the protocol parser. The structure should at least contain key contents such as the request path, input data, and header parameters. The generated structured request object needs to be subjected to data integrity verification, which usually includes whether the request fields are complete, whether the data types match, whether the numerical range is reasonable, whether there are illegal characters or empty fields, etc. Verification is an important guarantee to prevent illegal or abnormal requests from entering the subsequent inference process.

[0119] The matching of the interface field and the format strategy is a consistency check operation to determine whether the structured request is consistent with the definition of the first model service framework. The format strategy is a parameter template defined by the server according to the model service protocol, such as the input field name, number of fields, hierarchical structure, etc., which is aligned with the request object at the field level. If it matches, the request is marked as a valid request.

[0120] For each valid request, the system needs to generate a unique request identification value, which is generally generated by combining a timestamp, a service instance identifier, and an auto-incrementing serial number, such as a UUID or an ID generated by the Snowflake algorithm, for subsequent request link tracking, exception backtracking, performance monitoring, and other scenarios. The generated valid request object and identification value need to be written into the request cache, which is generally a thread-safe data structure, such as a blocking queue or memory hash map, for subsequent scheduling threads to obtain and execute.

[0121] In the service deployment based on HTTP protocol, Nginx or Gunicorn can be used as a listener, and the listening port can be configured and forwarded to the application service. When the service application starts, the configuration file is read to initialize the listening strategy and protocol identifier, such as the REST interface path and supported data format types. The protocol parser can use the FastAPI Request object parsing method to convert the original data into a Pydantic model.

[0122] When the client initiates an HTTP request through a browser or Postman, the listener accepts the connection, automatically initializes the connection context, reads the HTTP request header and body, and passes them to the parsing module. The parsing module decodes the data stream using a predefined field mapping table to generate a request structure in a unified format. The model input fields are matched through the configured format strategy, such as checking whether the "input_ids" and "attention_mask" fields are included, and checking whether the data format complies with the JSON object specification. If it complies with the specification, a UUID is generated as the request ID, and the queue module is called to write the request structure and ID into the request cache.

[0123] In the gRPC scenario, the server uses Protobuf to define the service interface and generate a server stub, which listens to the port and waits for the client to connect. The client request is sent through Protocol Buffers encoding, and the server stub method automatically parses the request into a Request structure and calls the internal verification function to check the integrity of the field. If the verification passes, the timestamp, request source information, and call chain ID are recorded, and a structured valid request is generated and passed to the service processing queue.

[0124] In this embodiment, by constructing a standardized receiving process, not only the adaptation ability of the model service in a multi-protocol environment is improved, but also the consistency and security of the request data are ensured. The structured processing and the generation mechanism of the request identifier provide stable input for subsequent inference scheduling, and also provide an accurate link tracing basis for multi-service collaboration, significantly reducing the system abnormal response rate and the cost of troubleshooting.

[0125] In one embodiment, the above step S20 includes:

[0126] S201, separating the initial input data and the initial request path from the initial inference request;

[0127] S202, parsing the nested data structure of the initial input data to generate a hierarchical key-value pair list;

[0128] S203, traversing the hierarchical key-value pair list, screening matching key-value pairs from the hierarchical key-value pairs according to the input parameter key name list defined by the first model service framework, and encapsulating the screened key-value pairs into an input parameter set;

[0129] S204, splitting the string of the initial request path to obtain the model name, version number, and interface identifier in the initial request path;

[0130] S205, combining the model name, version number, and interface identifier into complete model identification information.

[0131] In this embodiment, the core purpose of parsing the initial input data and the initial request path in the initial inference request is to deconstruct the original unstructured request content into two types of core information: one is the input parameters that can be used in the subsequent inference process, and the other is the model identification information that represents the current inference service target. This processing flow belongs to the intermediate conversion link from the general request format to the internal call interface of the model service, and has the functions of decoupling the protocol structure, eliminating request redundancy, and extracting the core semantics.

[0132] First, separating the initial input data and the initial request path from the initial inference request is a logical splitting behavior, and the purpose is to clarify the boundary between the data entity and the path description. The initial inference request is usually represented in formats such as JSON or Protocol Buffers, which contain multiple fields, such as "input", "signature_name", "model_path", etc. By locating and extracting through the field names, the input data part representing the request data body and the request path part representing the routing or resource identifier are obtained.

[0133] After separating the initial input data, it is necessary to deconstruct its nested structure. In practical applications, the input data is usually a deeply nested structure, such as containing nested arrays, nested dictionaries, or even structured nested objects (e.g., {"instances": [{"input_ids": [1, 2], "attention_mask": [1, 1]}]}). To achieve flexible extraction, a recursive parsing strategy is required to expand the nested structure into a list of hierarchical key-value pairs, that is, to generate key name paths (e.g., "instances.0.input_ids") and corresponding value pairs through depth traversal, thus forming an intermediate abstraction layer that adapts to the input structures of different frameworks.

[0134] After obtaining the list of hierarchical key-value pairs, the system filters these key-value pairs according to the list of input parameter key names defined by the first model service framework. This process can be understood as a pattern matching process. Based on the set of parameter names defined in the configuration or model meta-information, such as ["input_ids", "attention_mask"], relevant fields are identified from the key name paths, and their corresponding values are retained to form a filtered set of key-value pairs.

[0135] The filtered key-value pairs need to be further encapsulated into a set of input parameters. This operation is not only a simple aggregation but also includes semantic layer normalization. That is, mapping flat key-value pairs to input format objects required by the framework, such as {"instances": [{"input_ids": [...]}]} required by TensorFlow Serving, or the input tensor dictionary structure required by ONNX Runtime. This encapsulation operation provides a structural alignment basis for subsequent conversion to the input format of the target model service.

[0136] At the same time, it is also necessary to parse the initial request path and extract the model identification information contained in the path. The request path usually includes information such as the model name, model version, interface type, or call signature. For example, in the path " / models / bert / versions / 1:predict", the model name "bert" and version number "1" and interface identifier "predict" can be extracted through string splitting and regular matching methods.

[0137] Finally, the model name, version number, and interface identifier are combined into a complete model identification information object, which is an important basis for generating the inference path and call parameters of the subsequent target model service. This combination operation usually manifests as the construction of a standardized structure body, with unified field naming and normalized value types, and serves as the key input for subsequent path template filling and interface protocol conversion.

[0138] In this embodiment, by deconstructing the request data into an input parameter set and model identification information, the standardization and abstraction of the model service call parameters are achieved, effectively isolating the coupling relationship between the client request structure and the internal service call structure. Based on this, subsequent protocol adaptation and input format conversion can be carried out on clear boundaries and structured data, avoiding problems such as manual parsing errors and interface field mismatches, and significantly improving the robustness of service calls and the automatic conversion ability.

[0139] In one embodiment, the above step S30 includes:

[0140] S301, according to the interface identifier in the model identification information, match the target path template of the second model service framework from the pre-stored path template mapping table;

[0141] S302, replace the model name placeholder and version number placeholder in the target path template with the model name and version number in the model identification information to generate a complete target request path;

[0142] S303, according to the interface identifier, obtain the input key name list of the second model service framework from the pre-stored input key name mapping table;

[0143] S304, traverse the key-value pairs in the input parameter set, and replace the original key name in the key-value pair with the corresponding target key name according to the input key name mapping table;

[0144] S305, according to the input format defined by the second model service framework, re-encapsulate the replaced key-value pairs in a hierarchical nested structure to generate the converted target input data.

[0145] In this embodiment, determining the target request path and generating the converted target input data according to the model identification information belongs to the key conversion link in the model service request adaptation process. The main task of this process is to map the model call intention and input parameters in the source framework to the request path and input data structure supported by the target model service framework, so as to complete the structural bridging between different frameworks. This process involves multiple sub-operations such as path reconstruction, parameter name mapping, and input format reorganization, aiming to achieve transparent call migration at the framework level.

[0146] First, it is necessary to match the target path template according to the interface identifier in the model identification information. The interface identifier is usually used to distinguish specific model service capabilities, such as "predict", "classify", "generate", etc. The system can use the interface identifier as a keyword to search for the matching target path structure in the path template mapping table. The path template is a pre-configured rule path format, often containing wildcard fields in the form of placeholders, such as / model / {model} / version / {version}:{signature}, which is used to define the actual request routing in the target framework. This template mapping table is maintained in the form of a configuration file or a database table, and corresponding path structures can be configured for different service interfaces and different model systems.

[0147] Subsequently, it is necessary to replace the placeholder fields in the path template. Identifiers such as {model} and {version} in the template are replaced by the model name and version number in the model identification information respectively, so as to generate a complete target request path that can be specifically used for request sending. This string filling operation can be implemented through a placeholder replacement function, or a dynamic path splicing logic based on routing mapping rules can be constructed to ensure the accuracy and callability of the path.

[0148] Next, the system needs to obtain the list of key names of the target input parameters from the input key name mapping table according to the interface identifier. The input key name mapping table is a predefined key name conversion dictionary, which is used to solve the problem of inconsistent input field naming between different frameworks. For example, TFServing uses "input_ids", while MindSpore Serving may use "input_tensor". By accurately positioning the parameter name mapping rule required by the currently called interface through the interface identifier, the context consistency of key name replacement can be guaranteed.

[0149] After obtaining the key name mapping relationship, the system starts to traverse the input parameter set. Each key-value pair checks whether its original key name exists in the mapping table. If it exists, it is replaced with the key name supported by the target framework. The replacement process is not limited to simple mapping, but also needs to consider the semantic consistency issues in the key name hierarchy and nested paths. When necessary, path merging or structure normalization processing needs to be carried out.

[0150] After completing the key name replacement, all parameters need to be repackaged into a structured input object according to the input format defined by the second model service framework. The target input format may require nested hierarchical structures, batch dimension encapsulation, or type annotations, such as being encapsulated into a format structure like {"instances": [{"input_tensor": [...]}]}. This process not only includes dictionary reconstruction and list encapsulation but also involves operations such as data type coercion and structure compatibility checks to ensure that the final input data meets the target service reception constraints.

[0151] In this embodiment, by constructing a mapping mechanism between path templates and parameter key names, the dynamic adaptation and structural unification of model identifiers and input parameters among different inference service frameworks are achieved. This process can automatically map the request content into a path structure and input format recognizable by the target inference framework without modifying the original calling end, reducing the adaptation threshold for service migration and avoiding the risk of manual code reconstruction, thereby enhancing the deployment flexibility and automation of inference services in multi-chip and multi-platform environments.

[0152] In one embodiment, the above step S40 includes:

[0153] S401, select the corresponding transport protocol type according to the protocol type identifier of the target request path;

[0154] S402, encapsulate the target input data into the request body format defined by the second model service framework according to the requirements of the transport protocol type to generate a target framework request body;

[0155] S403, send the target framework request body and the target request path to the target inference service corresponding to the second model service framework through the communication interface corresponding to the transport protocol type, and receive the original response data returning the target inference result;

[0156] S404, extract the status code and the inference result data from the original response data according to the response format defined by the second model service framework;

[0157] S405, if the status code indicates that the request is successful, mark the inference result data as a valid target inference result.

[0158] In this embodiment, sending the target input data and the target request path to the inference service corresponding to the second model service framework is the core execution process in the entire model inference adaptation link, undertaking the bridging task of request distribution and result response between different service protocols. It not only covers the selection of transport protocol types and data encapsulation but also involves communication interface adaptation, response format parsing, and exception handling logic design to ensure the accuracy, robustness, and recoverability of the communication process between heterogeneous inference frameworks.

[0159] First, select the transport protocol type according to the protocol type identifier carried in the target request path. The protocol type identifier is an implicit or explicit field in the path structure, usually determined by specific marker fields (such as "grpc", "http") in the request path or configuration meta-information. This identifier is used to dynamically determine whether to perform JSON body transmission via an HTTP POST request or communicate in Protocol Buffers format via a gRPC channel, thereby triggering different communication module loading logics.

[0160] After identifying the transport protocol type, it is necessary to encapsulate the request body data according to the request format defined by the target framework. The target input data may be a collection of key-value pairs or a nested structure in the original structure, but different inference service frameworks have specific agreed-upon formats for the request body. For example, HTTP interfaces usually require the input to be encapsulated into the "inputs" or "instances" fields in JSON, while gRPC interfaces need to strictly follow the field structure and types defined in the IDL. The encapsulation process needs to map the data to the corresponding structured request body according to the field naming specifications, structure templates, or interface protocol documents configured for the target service framework. This process is usually completed by a serializer, such as a JSON serialization tool or a Protocol Buffers codec.

[0161] After encapsulation, send the request through the communication interface corresponding to the transport protocol type. The communication interface can be a RESTful client library, such as the requests module, or a gRPC client tool, such as a call chain built based on grpc.Channel and Stub. The system constructs the target request path and binds the encapsulated request body to establish a communication link to send the inference request. To improve the response robustness, this stage usually cooperates with connection timeout, reconnection strategy, and injection mechanism of request tracking identifiers.

[0162] After sending the request, the system receives the raw response data returned by the target inference service. The response data format depends on the response protocol definition of the target service framework and usually includes a status code, a data body, exception prompts, or diagnostic information, etc. It is necessary to call the corresponding response parser to extract key fields from the service return format. The status code is used to determine whether the request has been successfully completed. If the status code falls within the successful response range (such as 200 or OK), the inference result data is extracted from the response body; if the status code indicates that the request has failed or an execution error has occurred, the system needs to trigger an exception recovery mechanism.

[0163] Finally, in the scenario where the status code indicates a successful request, mark the parsed inference result data as the valid target inference result and enter the next format adaptation stage. If the request fails, perform an automatic retry operation according to the configured retry policy. The retry policy can include the fixed number of retries, exponential backoff retry intervals, error code filtering trigger rules, etc., to ensure the request recovery ability under non-fatal exceptions such as network fluctuations and service jitters.

[0164] In this embodiment, a cross-protocol inference request distribution mechanism with adaptive capabilities is constructed through dynamic protocol type recognition, format encapsulation, and communication interface abstraction. The system can automatically complete request protocol adaptation and response extraction between different frameworks according to the model configuration and target platform protocol requirements, avoiding the need to manually write the call logic for multiple protocol branches and improving the consistency and flexibility of system calls. The automatic retry mechanism introduced in case of communication exceptions further enhances the fault tolerance during the request sending stage, effectively improving the stability and recoverability of the model inference service in complex network environments.

[0165] In one embodiment, the above step S50 includes:

[0166] S501, receive the target inference result returned by the second model service framework;

[0167] S502, perform structured parsing on the target inference result according to the output format specification defined by the second model service framework to generate a standardized output data object containing the original field key-value pairs;

[0168] S503, traverse each field in the standardized output data object;

[0169] S504, if the original key name of the field exists in the pre-stored output key name mapping table, replace the key name of the field with the corresponding target key name in the output key name mapping table and retain the key value;

[0170] S505, if the original key name of the field is not in the output key name mapping table but exists in the required output field list of the second model service framework, retain the original key name and key value of the field; if the original key name of the field is not declared in the output key name mapping table and is not declared in the required output field list, delete the field from the standardized output data object;

[0171] S506, summarize the key-value pairs in the processed standardized output data object into an output parameter key-value pair set corresponding to the first model service framework.

[0172] In this embodiment, receiving the target inference result returned by the second model service framework and extracting the output parameters in the target inference result belong to the processing stage of obtaining the effective inference result from the response data in the model service adaptation process. In this stage, the response data needs to be deconstructed, parsed, field-filtered, and key-name mapped to extract the set of output parameters that meet the format requirements of the first model service framework. Its core purpose is to convert the response format returned by heterogeneous inference services into a unified, standardized output structure that can be used for upper-layer service calls through a series of parsing and structural transformations.

[0173] The behavior of receiving the inference result depends on the request-sending logic of the previous stage. The system obtains the response data returned by the second model service framework through asynchronous callbacks or blocking waits. The response data is usually presented in the form of a serialized byte stream, and its format needs to be parsed in combination with the response protocol standard of the target inference framework.

[0174] The parsing of the response content needs to be structured according to the output format specification defined by the target service. For example, in a JSON-based HTTP service, the inference result is usually located in fields such as "outputs", "predictions", or "result"; while in a GRPC-based Protocol Buffers service, the returned object may contain nested fields, and the structure defined by the IDL needs to be parsed step by step to restore. The system needs to construct an intermediate standardized output data object during the parsing process, which encapsulates the parsed fields and content in the form of key-value pairs for subsequent operations.

[0175] Performing a traversal process on each field in the standardized output data object is a key step in field filtering and key-name conversion. First, check whether the original key name of the current field exists in the pre-stored output key-name mapping table. This mapping table records the correspondence between the output field names of the second model service framework and the output fields of the first model service framework. For example, "result" is mapped to "predictions", and "prob" is mapped to "confidence". If the match is successful, the original field key name needs to be replaced with the target key name, and the key-value content of the original field is retained to achieve consistent connection of cross-frame field semantics.

[0176] If a field does not appear in the mapping table but is included in the list of required output fields of the target inference framework, it means that although there is no mapping relationship for this field, it has an indispensable attribute for the result integrity. In this case, the original key name and key value should be retained to avoid accidentally deleting key inference content. The list of required fields is usually determined through model registration information, framework configuration files, or service interface documents, and can be used as an important supplementary basis for field filtering.

[0177] If a field is neither in the mapping table nor in the list of required output fields, it is regarded as content such as debugging information or statistical fields that is irrelevant to the current business or added by the framework, and should be removed from the standardized output data object to reduce the redundancy of the result body and improve the subsequent format encapsulation and transmission efficiency. This deletion operation constitutes an output cleaning operation, which helps to improve the consistency and lightweight of the response structure.

[0178] Finally, all the field key-value pairs that have undergone key name replacement or retention operations are aggregated into a set of output parameter key-value pairs, which is used as the effective model inference result extracted in the current adaptation process and handed over to the subsequent format conversion module for processing.

[0179] This embodiment realizes the automatic conversion and cleaning of result data with inconsistent structures extracted between different model inference services by constructing an output field mapping mechanism and a keyword field retention logic, and solves the problem that the calling end cannot parse due to the differences in the response structures of heterogeneous model services. This mechanism not only improves the format uniformity of the model service response, but also avoids unnecessary field redundancy and the risk of incorrect parsing, providing a clear and semantically consistent intermediate result for subsequent data standardization and service output.

[0180] In one embodiment, the above step S60 includes:

[0181] S601, convert the data type of each key-value pair in the output parameters according to the data type specification of the first model service framework to generate a standardized data set;

[0182] S602, combine the key-value pairs in the standardized data set into a multi-level nested structure according to the hierarchical nesting strategy of the first model service framework;

[0183] S603, add the metadata fields defined by the first model service framework to the multi-level nested structure to generate a complete output structure;

[0184] S604, convert the complete output structure into a byte stream in the target serialization format according to the serialization protocol of the first model service framework;

[0185] S605, verify whether the byte stream contains the required output fields defined by the first model service framework;

[0186] S606, if the verification passes, mark the byte stream as the converted target output data.

[0187] In this embodiment, mapping the output parameters to the output format defined by the first model service framework is a key step in achieving the unified presentation of the model inference response results in a heterogeneous inference system, involving multiple sub - processes such as data type conversion, structure reorganization, meta - information injection, serialization, and legality verification. This processing not only undertakes the format conversion task but also ensures that the data can be correctly parsed and used by the calling end.

[0188] The data type conversion is based on the data type specification of the first model service framework, and static or dynamic type matching needs to be performed on each key - value pair of the output parameters. For example, when the first model service framework defines that a certain field should be int64, the system needs to explicitly convert the original value from types such as float32 or string to the int64 format. This process may adopt strategies such as forced type conversion, precision truncation, or format encoding to ensure that the generated result strictly meets the input specifications of the upper - layer system.

[0189] The structure reorganization is organized according to the nesting strategy of the first model service framework. It is common in HTTP interfaces where multiple result fields need to be embedded into a unified result node (such as "outputs", "predictions", etc.), or in GRPC interfaces where they need to be nested as IDL structure fields. The reorganization process combines the standardized key - value pairs into a multi - level nested object that conforms to the target structure according to a preset template or mapping relationship. If the structure hierarchy is not encapsulated according to the preset specification, it may cause the calling end to be unable to correctly parse the field path.

[0190] The injection of metadata fields enables the output structure to have cross - module tracking capabilities and request context awareness capabilities. It usually includes a timestamp field for the current response time, which is used for service performance analysis; an interface version number field, which is used to identify the model service version on which the current return structure depends; and a request tracking identifier field, which is used to mark the full - link context information of this request in the call path in the microservice link. The addition of these fields is usually embedded in the output structure in the form of reserved fields or system - reserved key - value pairs, such as under the metadata field at the top - level node.

[0191] Serialization is the process of converting the structured output structure into a byte - stream form, and the serialization method is determined according to the transport protocol identifier of the first model service framework. If the transport protocol is HTTP, JSON format is used for serialization; if it is GRPC, Protocol Buffers is used for binary encoding. This mapping relationship ensures that the output format is consistent with the transport protocol, avoiding communication failures due to inconsistent formats.

[0192] After completing the structure generation and serialization, the byte stream needs to be checked for structural legitimacy, that is, to verify whether the generated byte stream contains the required output fields defined by the first model service framework. Required fields include key content such as prediction results and return status. If missing, the calling end behavior will be uncontrollable. If the verification is successful, the current byte stream is marked as the converted target output data, indicating that the response is ready to be returned to the client; otherwise, the exception process is entered, which may include retry, error response generation or reporting mechanism.

[0193] Example description: In the financial technology business, credit approval models often build inference models based on large-scale behavioral data, financial indicators and credit information. Traditional deployments rely on TFServing and GPU devices to complete model services. At this stage, in order to reduce deployment costs and improve the adaptability of model deployment, the system hopes to migrate the original TFServing service to MindSpore Serving to adapt to NPU devices. However, the existing client call logic depends on the input format, service path structure and serialization protocol of TFServing, which is difficult to be directly compatible with the new NPU inference engine. When the client initiates an inference request in TFServing format through the HTTP protocol, the system first establishes a listening channel to receive the initial inference request, decodes it into a structured request object through the protocol parser, and completes data integrity verification. After parsing the input data and request path, the system extracts parameter fields such as user behavior characteristics and credit scores, and obtains the name, version and interface type of the current inference model from the path. The system queries the service path template and input format rules corresponding to MindSpore Serving through the preset mapping table, converts the extracted input parameter fields into key names, and repackages the data according to the nested format of the target service framework to construct the target input data suitable for the MindSpore reasoning engine. At the same time, the HTTP target request path is determined according to the characteristics of the reasoning interface, and the call is completed. After receiving the reasoning response from the NPU backend, the system extracts output fields such as risk score and score explanation from the original response data according to the mapping table, removes non-TFServing defined fields such as debug log and NPU execution time, and unifies the mapping field names and types. Finally, the system combines the output parameters into a standard structure, supplements metadata such as version identification and request tracking ID, and serializes them into the JSON format response data expected by the TFServing client and returns them. During the entire process, the client can complete the seamless reasoning service migration from GPU to NPU without making any changes to the docking logic.

[0194] In the scenario of medical image assisted diagnosis, a hospital has deployed a pulmonary nodule recognition model service based on TFServing. The model accepts tensor data after processing DICOM format images as input and outputs information such as nodule type and malignancy probability. To meet the requirement of improving processing performance, the hospital introduced NPU to deploy the inference service and switched to use MindSpore Serving to support the new computing power platform. However, due to the inconsistencies in input and output field naming, hierarchical structure, and protocol encoding between TFServing and MindSpore Serving, directly replacing the backend service will cause a large number of upper-layer service interface incompatibility problems. After the system receives an inference request initiated by the CT scan image preprocessing module, it first completes the decoding process and structure verification of the HTTP data stream, extracts the image tensor field (such as image_tensor) and the model information in the request path (such as the model nodule_detector, version v3.4, and the interface type is predict). Through the configured path template and input mapping rules, the system automatically maps image_tensor to the field name expected by MindSpore Serving (such as input_tensor) and nests it under the instances field according to its definition. After the inference service returns the result, the system parses the response structure returned by MindSpore, extracts the key output fields such as classification_score and type_label, removes unnecessary fields such as npu_inference_time and debugging data, and then uniformly maps them to the predictions field in the TFServing format, and inserts meta-information such as the request timestamp and model version according to the current service configuration. After completing the format reconstruction and type conversion, the system serializes the result into a JSON byte stream and returns it to the calling system through the original TFServing interface, realizing a cross-platform inference service upgrade that is imperceptible to the image analysis module.

[0195] In this embodiment, through the type specification conversion, structure encapsulation, and protocol serialization mechanisms, the format consistency processing of the inference results among heterogeneous model services is realized, avoiding the client parsing failure problems caused by inconsistent data types, mismatched structure levels, or protocol differences. Especially in systems that need to seamlessly access different inference backends, this mechanism significantly reduces the client transformation cost and improves the compatibility and engineering controllability of the model inference service.

[0196] In one embodiment, a model service cross-framework request conversion device is provided. The model service cross-framework request conversion device corresponds one-to-one to the model service cross-framework request conversion method in the above embodiment. Refer to Figure 3 , Figure 3It is a schematic diagram of functional modules of a preferred embodiment of a cross-frame request conversion device for a model service of the present invention. Request receiving module 10, request parsing module 20, input mapping module 30, inference forwarding module 40, result receiving module 50, output mapping module 60, and response returning module 70. The detailed description of each functional module is as follows:

[0197] The request receiving module 10 is used to receive an initial inference request defined based on a first model service framework sent by a client;

[0198] The request parsing module 20 is used to parse the initial input data and the initial request path in the initial inference request, and extract the input parameters in the initial input data and the model identification information in the initial request path;

[0199] The input mapping module 30 is used to determine a target request path of a second model service framework according to the model identification information, and map the input parameters into an input format defined by the second model service framework to generate converted target input data;

[0200] The inference forwarding module 40 is used to send the target input data and the target request path to a target inference service corresponding to the second model service framework to generate a target inference result;

[0201] The result receiving module 50 is used to receive the target inference result returned by the second model service framework, and extract the output parameters in the target inference result;

[0202] The output mapping module 60 is used to map the output parameters into an output format defined by the first model service framework to generate converted target output data;

[0203] The response returning module 70 is used to return the target output data to the client.

[0204] In an embodiment, the request receiving module 10 is specifically used for:

[0205] Start a service listening program to establish a receiving channel facing the client, and set a transport protocol identifier;

[0206] Initialize a corresponding protocol parser based on the transport protocol identifier;

[0207] Listen to the receiving channel to wait for a connection request initiated by the client;

[0208] Accept the connection request of the client and establish a communication link;

[0209] Receive the original data stream of the initial inference request through the communication link;

[0210] Decode the original data stream using the protocol parser to generate a structured request object;

[0211] Perform data integrity verification on the structured request object;

[0212] After the data integrity verification passes, verify whether the interface fields of the structured request object match the format policy defined by the first model service framework;

[0213] If the interface fields of the structured request object match the format policy defined by the first model service framework, mark the structured request object as a valid request and generate a unique request identifier value for each valid request;

[0214] Write the valid request and the corresponding request identifier value to the request buffer.

[0215] In one embodiment, the request parsing module 20 is specifically configured to:

[0216] Separate the initial input data and the initial request path from the initial inference request;

[0217] Parse the nested data structure of the initial input data to generate a hierarchical key-value pair list;

[0218] Traverse the hierarchical key-value pair list, and according to the input parameter key name list defined by the first model service framework, filter the matching key-value pairs from the hierarchical key-value pairs and encapsulate the filtered key-value pairs into an input parameter set;

[0219] Split the string of the initial request path to obtain the model name, version number, and interface identifier in the initial request path;

[0220] Combine the model name, version number, and interface identifier into complete model identification information.

[0221] In one embodiment, the input mapping module 30 is specifically configured to:

[0222] According to the interface identifier in the model identification information, match the target path template of the second model service framework from the pre-stored path template mapping table;

[0223] Replace the model name placeholder and version number placeholder in the target path template with the model name and version number in the model identification information to generate a complete target request path;

[0224] According to the interface identifier, obtain the input key name list of the second model service framework from the pre-stored input key name mapping table;

[0225] Traverse the key-value pairs in the set of input parameters, and replace the original key names in the key-value pairs with the corresponding target key names according to the input key name mapping table;

[0226] According to the input format defined by the second model service framework, repackage the replaced key-value pairs in a hierarchical nested structure to generate the converted target input data.

[0227] In one embodiment, the inference forwarding module 40 is specifically configured to:

[0228] Select the corresponding transmission protocol type according to the protocol type identifier of the target request path;

[0229] Package the target input data into the request body format defined by the second model service framework according to the requirements of the transmission protocol type to generate a target framework request body;

[0230] Send the target framework request body and the target request path to the target inference service corresponding to the second model service framework through the communication interface corresponding to the transmission protocol type, and receive the original response data returning the target inference result;

[0231] Extract the status code and inference result data from the original response data according to the response format defined by the second model service framework;

[0232] If the status code indicates that the request is successful, mark the inference result data as a valid target inference result.

[0233] In one embodiment, the result receiving module 50 is specifically configured to:

[0234] Receive the target inference result returned by the second model service framework;

[0235] Perform a structured parsing on the target inference result according to the output format specification defined by the second model service framework to generate a standardized output data object containing the original field key-value pairs;

[0236] Traverse each field in the standardized output data object;

[0237] If the original key name of the field exists in the pre-stored output key name mapping table, replace the key name of the field with the corresponding target key name in the output key name mapping table and retain the key value;

[0238] If the original key name of a field is not in the output key name mapping table but exists in the list of required output fields of the second model service framework, the original key name and key value of the field are retained; if the original key name of a field is not declared in the output key name mapping table and is not declared in the list of required output fields, the field is deleted from the standardized output data object;

[0239] Summarize the key-value pairs in the processed standardized output data object into an output parameter key-value pair set corresponding to the first model service framework.

[0240] In one embodiment, the output mapping module 60 is specifically configured to:

[0241] According to the data type specification of the first model service framework, convert the data type of each key-value pair in the output parameters to generate a standardized data set;

[0242] According to the hierarchical nesting strategy of the first model service framework, combine the key-value pairs in the standardized data set into a multi-level nested structure;

[0243] Add the metadata fields defined by the first model service framework to the multi-level nested structure to generate a complete output structure;

[0244] According to the serialization protocol of the first model service framework, convert the complete output structure into a byte stream in the target serialization format;

[0245] Verify whether the byte stream contains the required output fields defined by the first model service framework;

[0246] If the verification passes, mark the byte stream as the converted target output data.

[0247] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the server side of a model service cross-frame request conversion method.

[0248] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as shown in Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a method for converting cross-frame requests of a model service

[0249] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:

[0250] Receiving an initial inference request sent by a client based on a first model service framework definition;

[0251] Parsing the initial input data and the initial request path in the initial inference request, and extracting the input parameters in the initial input data and the model identification information in the initial request path;

[0252] Determining a target request path of a corresponding second model service framework according to the model identification information, and mapping the input parameters to an input format defined by the second model service framework to generate converted target input data;

[0253] Sending the target input data and the target request path to a target inference service corresponding to the second model service framework to generate a target inference result;

[0254] Receiving the target inference result returned by the second model service framework, and extracting the output parameters in the target inference result;

[0255] Mapping the output parameters to an output format defined by the first model service framework to generate converted target output data;

[0256] Returning the target output data to the client.

[0257] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are realized:

[0258] Receiving an initial inference request sent by a client based on a first model service framework definition;

[0259] Analyze the initial input data and the initial request path in the initial inference request, and extract the input parameters in the initial input data and the model identification information in the initial request path;

[0260] Determine the target request path of the corresponding second model service framework according to the model identification information, map the input parameters to the input format defined by the second model service framework, and generate the converted target input data;

[0261] Send the target input data and the target request path to the target inference service corresponding to the second model service framework to generate a target inference result;

[0262] Receive the target inference result returned by the second model service framework, and extract the output parameters in the target inference result;

[0263] Map the output parameters to the output format defined by the first model service framework, and generate the converted target output data;

[0264] Return the target output data to the client.

[0265] It should be noted that the functions or steps that can be realized by the above computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0266] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0267] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0268] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for example introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A method for cross-frame request conversion of model services, characterized in that, Including the following steps: Receiving an initial inference request sent by a client and defined based on a first model service framework; Parsing the initial input data and the initial request path in the initial inference request, and extracting the input parameters in the initial input data and the model identification information in the initial request path; Determining a target request path of a corresponding second model service framework according to the model identification information, mapping the input parameters into an input format defined by the second model service framework, and generating converted target input data; Sending the target input data and the target request path to a target inference service corresponding to the second model service framework to generate a target inference result; Receiving the target inference result returned by the second model service framework, and extracting the output parameters in the target inference result; Mapping the output parameters into an output format defined by the first model service framework, and generating converted target output data; Returning the target output data to the client.

2. The model service cross-frame request conversion method according to claim 1, wherein Receiving an initial inference request sent by a client and defined based on a first model service framework, including: Starting a service listening program to establish a receiving channel facing the client, and setting a transmission protocol identifier; Initializing a corresponding protocol parser based on the transmission protocol identifier; Listening to the receiving channel to wait for a connection request initiated by the client; Accepting the connection request of the client and establishing a communication link; Receiving a raw data stream of the initial inference request through the communication link; Decoding the raw data stream using the protocol parser to generate a structured request object; Performing data integrity verification on the structured request object; After the data integrity verification passes, verifying whether the interface fields of the structured request object match the format policy defined by the first model service framework; If the interface fields of the structured request object match the format policy defined by the first model service framework, marking the structured request object as a valid request, and generating a unique request identifier value for each valid request; Writing the valid request and the corresponding request identifier value into a request buffer.

3. The model service cross-frame request conversion method according to claim 1, wherein Parsing the initial input data and the initial request path in the initial inference request, and extracting the input parameters in the initial input data and the model identification information in the initial request path, including: Separating the initial input data and the initial request path from the initial inference request; Parsing the nested data structure of the initial input data to generate a hierarchical key-value pair list; Traversing the hierarchical key-value pair list, screening out matching key-value pairs from the hierarchical key-value pairs according to the input parameter key name list defined by the first model service framework, and encapsulating the screened key-value pairs into an input parameter set; Splitting the string of the initial request path to obtain the model name, version number, and interface identifier in the initial request path; Combining the model name, version number, and interface identifier into complete model identification information.

4. The model service cross-frame request conversion method according to claim 1, characterized in that Determine the target request path of the corresponding second model service framework according to the model identification information, and map the input parameters to the input format defined by the second model service framework to generate the converted target input data, including: Match the target path template of the second model service framework from the pre-stored path template mapping table according to the interface identifier in the model identification information; Replace the model name placeholder and version number placeholder in the target path template with the model name and version number in the model identification information to generate a complete target request path; Obtain the input key name list of the second model service framework from the pre-stored input key name mapping table according to the interface identifier; Traverse the key-value pairs in the input parameter set, and replace the original key name in the key-value pair with the corresponding target key name according to the input key name mapping table; According to the input format defined by the second model service framework, repackage the replaced key-value pairs in a hierarchical nested structure to generate the converted target input data.

5. The model service cross-frame request conversion method according to claim 1, wherein, Send the target input data and the target request path to the target inference service corresponding to the second model service framework to generate a target inference result, including: Select the corresponding transmission protocol type according to the protocol type identifier of the target request path; Package the target input data into the request body format defined by the second model service framework according to the requirements of the transmission protocol type to generate a target framework request body; Through the communication interface corresponding to the transmission protocol type, send the target framework request body and the target request path to the target inference service corresponding to the second model service framework, and receive the original response data returning the target inference result; Extract the status code and inference result data from the original response data according to the response format defined by the second model service framework; If the status code indicates a successful request, mark the inference result data as a valid target inference result.

6. The method for converting cross-frame requests of model services according to claim 1, characterized in that Receive the target inference result returned by the second model service framework and extract the output parameters in the target inference result, including: Receive the target inference result returned by the second model service framework; Perform a structured parsing on the target inference result according to the output format specification defined by the second model service framework to generate a standardized output data object containing the original field key-value pairs; Traverse each field in the standardized output data object; If the original key name of the field exists in the pre-stored output key name mapping table, replace the key name of the field with the corresponding target key name in the output key name mapping table and retain the key value; if the original key name of the field is not in the output key name mapping table but exists in the required output field list of the second model service framework, retain the original key name and key value of the field; if the original key name of the field is not declared in the output key name mapping table and is not declared in the required output field list, delete the field from the standardized output data object; Summarize the key-value pairs in the processed standardized output data object into an output parameter key-value pair set corresponding to the first model service framework.

7. The model service cross-frame request conversion method according to claim 1, wherein Map the output parameters to the output format defined by the first model service framework to generate the converted target output data, including: Convert the data types of each key-value pair in the output parameters according to the data type specification of the first model service framework to generate a standardized data set; Combine the key-value pairs in the standardized data set into a multi-level nested structure according to the hierarchical nesting strategy of the first model service framework; Add the metadata fields defined by the first model service framework to the multi-level nested structure to generate a complete output structure; Convert the complete output structure into a byte stream in the target serialization format according to the serialization protocol of the first model service framework; Verify whether the byte stream contains the required output fields defined by the first model service framework; If the verification passes, mark the byte stream as the converted target output data.

8. A cross-frame request conversion device for model services, characterized in that, The model service cross-frame request conversion device includes: A request receiving module for receiving an initial inference request sent by a client based on the definition of the first model service framework; A request parsing module for parsing the initial input data and the initial request path in the initial inference request, and extracting the input parameters in the initial input data and the model identification information in the initial request path; An input mapping module for determining the target request path of the corresponding second model service framework according to the model identification information, and mapping the input parameters to the input format defined by the second model service framework to generate the converted target input data; An inference forwarding module for sending the target input data and the target request path to the target inference service corresponding to the second model service framework to generate a target inference result; A result receiving module for receiving the target inference result returned by the second model service framework and extracting the output parameters in the target inference result; An output mapping module for mapping the output parameters to the output format defined by the first model service framework to generate the converted target output data; A response returning module for returning the target output data to the client.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a model service cross-frame request conversion program stored on the memory and executable on the processor. When the model service cross-frame request conversion program is executed by the processor, it implements the steps of the model service cross-frame request conversion method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A model service cross-frame request conversion program is stored on the storage medium. When the model service cross-frame request conversion program is executed by a processor, it implements the steps of the model service cross-frame request conversion method according to any one of claims 1-7.

Citation Information

Cited By

  • Method, device and equipment for realizing unified API (Application Program Interface) of call center and storage medium

    CN120897002A

  • Call center unified api implementation method, apparatus, device, and storage medium

    CN120897002B

  • Business processing method, system and terminal based on large model middleware engine

    CN120956790A

  • Data processing method and device, data encoding and decoding method and device, electronic equipment, storage medium and program product

    CN121636448A

  • Model service-oriented content conversion method and device, medium, equipment and product

    CN122240172A