Model operation method and device
By decoupling model files and services, and assembling the model combination call interface according to request parameters, the problem of low utilization of model service resources is solved, and the model is efficient, flexible deployment and rapid response is achieved.
Patent Information
- Application Number
- CN202510368176.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, model services have the problem of low resource utilization, resulting in high cost of model deployment and waste of resources.
By assembling the model combination call interface based on the request parameters, determining the target model files and services, decoupling the model files and services, and achieving flexible calling and combination of models, including storage of model files, preloading and updating configuration information.
It improves resource utilization, shortens the time from development to completion of the model, improves the response speed and flexibility of model services, and reduces operational costs.
Smart Images

Figure CN120371377A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a method and device for model operation. Background Art
[0002] In a service platform, a user may need to perform different business operations to call different models provided by a service provider. Therefore, the service provider needs to prepare multiple models to cope with the needs.
[0003] In the existing model deployment process, the model file and the model service are coupled, that is, a model service can only run a specific model file. Each time a model is put into production, a complete set of services must be deployed, which requires huge costs and causes waste of resources. That is, the model service in the related art has the problem of low resource utilization rate.
[0004] For the above problems in the related art, no solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present invention provide a method and device for model operation to at least solve the problem of low resource utilization rate of the model service in the related art.
[0006] According to an embodiment of the present invention, a method for model operation is provided, including: assembling a model combination call interface according to a target service request in request parameters, where the target service request is used to indicate a target model; determining configuration information of a target model combination according to the target service request; determining a target model file according to the configuration information, and obtaining a first model service and a second model service corresponding to the target model, where the first model service is used to indicate a main model service, and the second model service is used to indicate a slave model service; executing the first model service, obtaining an output result, and recording the output result and the request parameters.
[0007] Optionally, before assembling the model combination call interface according to the target service request in the request parameters, it includes: storing at least one received model file in an object storage, where the model file is used to encapsulate a calculation logic; preloading the model file into a memory, and calling an execution engine corresponding to the model file; updating configuration information corresponding to the model file.
[0008] Optionally, preloading the model file into the memory and calling the execution engine corresponding to the model file includes: determining model identification information of the model file, and determining path information of the model file in the object storage according to the model identification information; loading information of the model file from the object storage into the corresponding execution engine according to the path information of the model file in the object storage.
[0009] Optionally, update the configuration information corresponding to the model file, including: determining the configuration information of the model file, and updating the model encoding, model file name, and version number in the configuration file, where the model encoding is used to indicate the corresponding model service.
[0010] Optionally, determine the configuration information of the model file, and update the model encoding, model file name, and version number in the configuration file, including: storing the configuration information of the model file in a configuration cache variable, and reading the changed configuration information in case of configuration information change; comparing the configuration cache variable with the changed configuration information, and in case the configuration cache variable is different from the changed configuration information, storing the changed configuration information in the configuration cache variable and synchronizing the update to the corresponding node.
[0011] Optionally, determine the configuration information of the target model combination according to the target service request, including: determining the node information corresponding to the target model combination according to the node encoding in the target service request, where the node is used to indicate a model service in the model combination; determining the configuration information according to the model encoding in the node information.
[0012] Optionally, after executing the first model service to obtain an output result and recording the output result and request parameters, it further includes: executing the second model service to obtain a reference result and recording the reference result and request parameters; updating the parameters of the first model service according to the reference result and the output result.
[0013] According to another embodiment of the present invention, there is also provided a model running device, including:
[0014] A model combination module, configured to assemble a model combination call interface according to the target service request in the request parameters, where the target service request is used to indicate the target model;
[0015] A configuration information determination module, configured to determine the configuration information of the target model combination according to the target service request;
[0016] A service determination module, configured to determine a target model file according to the configuration information, and obtain a first model service and a second model service corresponding to the target model, where the first model service is used to indicate the main model service, and the second model service is used to indicate the slave model service;
[0017] A result output module, configured to execute the first model service to obtain an output result and record the output result and request parameters.
[0018] Optionally, the model running device further includes: a storage module for storing at least one received model file into an object storage, where the model file is used to encapsulate computing logic; a preloading module for preloading the model file into the memory and invoking an execution engine corresponding to the model file; a configuration update module for updating the configuration information corresponding to the model file.
[0019] Optionally, the preloading module includes: a path determination unit for determining the model identification information of the model file and determining the path information of the model file in the object storage according to the model identification information; a loading unit for loading the information of the model file from the object storage into the corresponding execution engine according to the path information of the model file in the object storage.
[0020] Optionally, the configuration update module is further configured to: determine the configuration information of the model file and update the model encoding, model file name, and version number in the configuration file, where the model encoding is used to indicate the corresponding model service.
[0021] Optionally, the configuration update module is further configured to: store the configuration information of the model file into a configuration cache variable, and read the changed configuration information in the case of a configuration information change; compare the configuration cache variable with the changed configuration information, and in the case where the configuration cache variable is different from the changed configuration information, store the changed configuration information into the configuration cache variable and synchronize the update to the corresponding node.
[0022] Optionally, the configuration information determination module includes: a node determination unit for determining the node information corresponding to the target model combination according to the node encoding in the target service request, where the node is used to indicate a model service in the model combination; a configuration information determination unit for determining the configuration information according to the model encoding in the node information.
[0023] Optionally, the result output module further includes: a reference result output unit for executing the second model service to obtain a reference result and recording the reference result and the request parameters; a parameter update unit for updating the parameters of the first model service according to the reference result and the output result.
[0024] According to another embodiment of the present invention, there is also provided a computer-readable storage medium storing a computer program, where the computer program, when run by a processor, executes the steps in any one of the above method embodiments.
[0025] According to another embodiment of the present invention, there is also provided an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0026] In the present invention, according to the target service request in the request parameters, an interface for invoking a model combination is assembled, where the target service request is used to indicate the target model; the configuration information of the target model combination is determined according to the target service request; the target model file is determined according to the configuration information, and the first model service and the second model service corresponding to the target model are obtained, where the first model service is used to indicate the main model service and the second model service is used to indicate the slave model service; the first model service is executed to obtain an output result, and the output result and the request parameters are recorded. By decoupling the model file and the model service, when using the model, it is necessary to determine the model file to be run through the request parameters. Therefore, it is not necessary to perform complex deployment operations for each model, which reduces the time from model development to completion and improves the resource utilization rate. Description of the Drawings
[0027] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0028] Figure 1 is the hardware structure block diagram of the mobile terminal for the model running method according to the embodiment of the present invention;
[0029] Figure 2 is the flowchart of the model running method according to the embodiment of the present invention;
[0030] Figure 3 is another flowchart of the model running method according to the embodiment of the present invention;
[0031] Figure 4 is the schematic diagram of the model running method according to the embodiment of the present invention;
[0032] Figure 5 is the block diagram of the model running device according to the embodiment of the present invention. Detailed Embodiments
[0033] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0035] Embodiment 1
[0036] The method embodiment provided in the first embodiment of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structural block diagram of the mobile terminal for the model running method according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), and a memory 104 for storing data. Optionally, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0037] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the model running method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.
[0038] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0039] In this embodiment, a model running method running on the above-mentioned mobile terminal or network architecture is provided. Figure 2 is a flowchart of the model running method according to an embodiment of the present invention. As Figure 2As shown, the process includes the following steps:
[0040] Step S202: Assemble the model combination call interface according to the target service request in the request parameters, where the target service request is used to indicate the target model;
[0041] Step S204: Determine the configuration information of the target model combination according to the target service request;
[0042] Step S206: Determine the target model file according to the configuration information, and obtain the first model service and the second model service corresponding to the target model, where the first model service is used to indicate the main model service, and the second model service is used to indicate the slave model service;
[0043] Step S208: Execute the first model service to obtain the output result, and record the output result and the request parameters.
[0044] It should be noted that the target service request is sent by the client to the model operation management platform, which is request data for invoking a specific model or model combination for inference. The request usually contains the identifier of the invoked model, input data, and other necessary parameter information. The configuration information of the target model combination is in the model operation management platform. To achieve flexible invocation and combination of models, multiple models are assembled together in a specific order and manner to form a model combination. The configuration information describes the specific composition of this model combination, including the identifier of the model, version number, location of the model file, service type (main model service or accompanying model service), etc. The target model file is the computational logic carrier of a specific model in the model combination, which can be a file in the PMML, Groovy, etc. format. The model file contains the algorithm implementation, parameter settings, and inference process of the model. The first model service indicates the main model service, which is the service in the model combination responsible for directly processing the request and returning the result. It runs the main model file, processes requests from users, and returns the prediction or inference result. The second model service indicates the slave model service, that is, the accompanying model service. It runs the accompanying model file and executes in parallel with the main model service, but its result is not directly returned to the user, but is used for internal comparison and analysis to optimize the performance of the main model service. Executing the first model service to obtain the output result and recording the output result and the request parameters is the core part of the model inference process. The main model service receives the request parameters, executes the inference algorithm according to the content of the model file, and generates the output result. The result and the request parameters will be recorded for subsequent analysis and auditing.
[0045] In an alternative embodiment, when the model operation management platform receives a request, it will look up and parse the corresponding model combination configuration information based on the parameters in the request (such as the model combination identifier) to determine which models need to be called, their versions, and the running order. Based on the parsed configuration information, the platform further identifies the files of the main model and the shadow model to be executed, as well as their corresponding service instances, and prepares to execute the inference task. After receiving the request, the main model service executes the inference logic according to the content of the model file, generates the prediction result, and records the result together with the original parameters of the request for subsequent analysis and auditing.
[0046] In an alternative embodiment, assume that in a bank credit card approval system, a machine learning model is needed to evaluate the credit risk of applicants. The system has already deployed a main model for daily approval, but in order to continuously optimize the performance of the approval model, the system designs a shadow model mechanism. When the system receives a credit card application (target service request), which contains data such as the applicant's basic information, financial status, and credit history. After receiving this application, the bank's model operation management platform obtains the configuration information of the target model combination from the configuration center (such as Consul) according to the model combination identifier carried in the request. The configuration information indicates that the model combination to be used includes a main model and two shadow models, and the model files are stored in the object storage service. Based on the configuration information, the platform loads the corresponding main model and shadow model files from the object storage, determines the target model files, and calls the corresponding first model service (main model service) and second model service (shadow model service). The main model service is responsible for directly processing the application data, while the shadow model service runs in parallel in the background and does not directly return results. The main model service receives the application parameters, runs the model calculation logic, and obtains the credit risk assessment result of the applicant, such as a credit score. This result will be recorded in the log together with the application parameters. At the same time, the main model service returns this result to the credit card approval system through the API, and the system makes an approval decision based on this result.
[0047] In the present invention, by assembling the model combination call interface according to the target service request in the request parameters, where the target service request is used to indicate the target model; determining the configuration information of the target model combination according to the target service request; determining the target model files according to the configuration information, and obtaining the first model service and the second model service corresponding to the target model, where the first model service is used to indicate the main model service and the second model service is used to indicate the slave model service; executing the first model service to obtain the output result, and recording the output result and the request parameters. Decoupling the model files and the model services, when using the model, it is necessary to determine the model files to be run through the request parameters. Therefore, it is not necessary to perform complex deployment operations for each model, reducing the time from model development to completion and improving the resource utilization rate.
[0048] Optionally, before assembling the model combination call interface according to the target service request in the request parameters, it includes: storing at least one received model file in object storage, where the model file is used to encapsulate the calculation logic; preloading the model file into memory and calling the execution engine corresponding to the model file; updating the configuration information corresponding to the model file.
[0049] It should be noted that object storage is a cloud storage service for storing a large amount of unstructured data (such as model files, pictures, videos, etc.). It stores data as objects and supports access through the network. The calculation logic refers to the algorithm by which the model makes predictions or processes based on input data, including data preprocessing, feature extraction, model invocation, and result postprocessing, etc. Memory refers to the main memory of the computer system, which is used to store running programs and data, and has a much faster access speed than hard disk storage. The execution engine is a software module that runs the model calculation logic according to the type of the model file. For example, the Groovy model uses the ScriptEngine provided by JDK 8, while the PMML model uses the Evaluator under the pmml-evaluator package.
[0050] In an alternative embodiment, when processing a request, according to the service type or target model specified by the user, selectively perform model combination or individual model invocation. Before invoking the model combination service, a series of preparatory work is carried out to ensure the correct storage, loading, and configuration of the model file. After the model file is received, it is first stored in object storage, which is the persistent storage location of the model file, facilitating the management and access of the model. To improve the response speed of the model service, after the model file is stored in object storage, it is preloaded into memory, so that when a request is received, the model logic can be directly executed in memory without reading the hard disk or network again. According to the type of the model file, the corresponding execution engine is called to run the calculation logic of the model. For example, the Groovy script uses the ScriptEngine, and the PMML model uses the Evaluator. After the model file is loaded, the configuration information related to the model file in the configuration center is updated to ensure that the model service can read the latest model configuration, including version numbers, parameter settings, etc.
[0051] In an alternative embodiment, a series of preparatory work before the model service actually processes a request includes storing the model file, preloading the model file into memory, calling the execution engine, and updating the configuration information. Through these steps, it is ensured that the model service can efficiently and stably process user requests. At the same time, preloading the model file into memory and updating the configuration information helps to improve the response speed and flexibility of the model service.
[0052] In an alternative embodiment, assume that we have a model combination service that includes a Groovy model and a PMML model for processing credit scoring requests of bank customers. The configuration information of the model combination is stored in the Consul configuration center. When the models are developed and ready for production, the Groovy model and PMML model files are uploaded to an object storage, such as Alibaba Cloud OSS or AWS S3. At the same time, the configuration information of the Groovy model and PMML model, including the model version, execution parameters, and resource requirements, is registered in Consul. When the model service starts, according to the configuration information in Consul, the Groovy model and PMML model files are read from the object storage and pre-loaded into memory. For the Groovy model, the ScriptEngine of JDK 8 is used for pre-loading; for the PMML model, the Evaluator of pmml-evaluator is used for pre-loading. If the model files are updated, the model service needs to listen for changes in the Consul configuration center. When a model version update is detected, the model version loaded in memory is updated, and the corresponding execution engine is called again to ensure that the model service always runs the latest version of the model. When a bank customer credit scoring request arrives, the model combination service selects the correct model version from the models pre-loaded in memory according to the model encoding and version information in the request parameters for calculation and returns the credit scoring result.
[0053] From the above embodiments, it can be seen that the model service migration solution effectively integrates the storage, loading, execution, and configuration management of model files, providing technical support for the rapid online deployment and efficient operation of models.
[0054] Optionally, pre-loading the model file into memory and calling the execution engine corresponding to the model file includes: determining the model identification information of the model file and determining the path information of the model file in the object storage according to the model identification information; according to the path information of the model file in the object storage, loading the information of the model file from the object storage into the corresponding execution engine.
[0055] It should be noted that the model identification information usually refers to modelVersionId, which is a unique identifier used to identify a specific model version among multiple models. It is an important attribute of the model file in the system to ensure that the correct model file is loaded and run. Object storage is a way to store unstructured data, such as large files like model files, pictures, and videos. Object storage has high scalability and disaster tolerance capabilities and is suitable for storing and managing a large number of model files.
[0056] In an alternative embodiment, the system first determines the modelVersionId of a model file, and then uses this identifier to query the corresponding path information in the database. This path information points to the location of the model file stored in the object storage. Next, the system reads the model file from the object storage according to the queried path information and loads it into the execution engine corresponding to the model format for subsequent model inference.
[0057] In an alternative embodiment, in the model service migration solution, the process of loading the model file into memory and invoking the execution engine. First, the specific location of the model file in the object storage is found through the model identification information, and then the model file information is loaded into the execution engine in the form of a stream. This process is to solve the performance problem during model inference because after the model file is loaded into memory, the latency of reading from disk or network can be significantly reduced, thus accelerating the speed of model inference. In addition, since the model service may run in a cluster environment, ensuring that all nodes can load the correct model file into memory is the key to ensuring service consistency and high availability.
[0058] In an alternative embodiment, assume that there is a model service that needs to run a model file in Groovy format. When the model file is imported into the object storage, the system automatically assigns a modelVersionId to this model file. When the model service starts, it queries the object storage path corresponding to the modelVersionId from the database and caches this path information in the Consul configuration center. When a request arrives, the model service obtains the path of the model file from Consul according to the modelVersionId in the request, and then uses the ScriptEngine provided by JDK8 to load the model file into memory. In this way, the model inference request can be directly executed in memory without the need to frequently read the model file from disk or network, thus greatly improving the response speed of the model service. For example, when a new version of the model file is uploaded to the object storage, the model service will receive a notification of Consul configuration update and then reload the new model file into the ScriptEngine to ensure that the model service uses the latest version of the model. This process summarizes the mechanism of preloading the model file into memory and invoking the execution engine. By using the database and Consul configuration center, dynamic management and efficient loading of the model file are achieved, while also ensuring the stability and efficiency of the model service operation.
[0059] Optionally, update the configuration information corresponding to the model file, including: determining the configuration information of the model file and updating the model encoding, model file name, and version number in the configuration file, where the model encoding is used to indicate the corresponding model service.
[0060] It should be noted that the model encoding is the code that uniquely identifies a model service and is used to quickly locate and manage a specific model service in a complex system environment. The model file name is the unique identifier of the model file and is used to distinguish different model logical implementations. The version number indicates the number of the specific version of the model file and is used for model update and version control to ensure that the model file loaded by the model service is the latest or the specified version.
[0061] In an alternative embodiment, when the model file is updated or a new model file is imported, it is necessary to clarify the detailed parameters on how the model service loads and runs these model files, including information such as the model encoding, model file name, and version number. In the configuration file of the model service, update the model encoding, model file name, and version number to reflect the latest status or changes of the model file, ensuring that the model service can correctly load and run the latest model file.
[0062] In an alternative embodiment, in the model operation management platform, the configuration information of the model file is the key for the model service to correctly load and run the model. The configuration information includes the model encoding, model file name, and version number, and these information together constitute the basis for the model service to identify and load the model file. When the model file is updated or a new version needs to be introduced, it is necessary to synchronously update the configuration information of the model service to ensure that the model file loaded by the model service is the latest or the specified version, so as to achieve the dynamic update and version control of the model service. The model encoding, as the unique identifier of the model service, enables the configuration update to accurately locate the model service that needs to be updated, while the model file name and version number ensure that the correct model file is loaded and run.
[0063] In an alternative embodiment, the configuration information of the model can be saved to Consul. The configuration information can include: serviceCode: the model encoding; imageName: the model file name; zhuVersion: the main service version number, and each main service has one and only one; congVersions: the set of co-running service version numbers, and one main service corresponds to 0 to multiple co-running services.
[0064] In an alternative embodiment, assume that in a bank's credit decision-making system, there is a model service responsible for processing user credit scores. The model service initially runs model version v1. To optimize the scoring model, the development team trains model version v2 and imports the new model file through the model operation management platform. The development team first imports model file v2 into the object storage, and at the same time updates the modelVersionId in the database to associate the information of the new version of the model file. Update the configuration information in the configuration center Consul, including: matching the model encoding (such as serviceCode: CG3WVASN3) with the new model file name and version number (such as imageName: cust_card_model, zhuVersion: v2). The primary / backup service obtains the path of the new version of the model file from the database according to the updated modelVersionId, then loads model file v2 from the object storage into memory and updates its execution engine. When the gateway service receives a request, according to the updated model encoding, model file name and version number in the configuration information, it forwards the request to the updated primary service, and asynchronously calls the backup service for performance comparison.
[0065] Through this process, the model service has successfully migrated to the new version of the model file, while achieving real-time performance comparison and optimization, ensuring the accuracy and efficiency of system decisions. This not only saves server resources, but also improves the flexibility and response speed of model updates, enhancing the intelligence and competitiveness of the bank's credit decision-making system.
[0066] Optionally, determine the configuration information of the model file and update the model encoding, model file name, and version number in the configuration file, including: storing the configuration information of the model file in a configuration cache variable, and when the configuration information changes, reading the changed configuration information; comparing the configuration cache variable with the changed configuration information, and when the configuration cache variable is different from the changed configuration information, storing the changed configuration information in the configuration cache variable and synchronizing the update to the corresponding nodes.
[0067] It should be noted that the configuration cache variable is a variable that stores the latest configuration information. It plays a caching role in the model service, avoiding frequent reading of the configuration source, improving the response speed of the service and reducing the pressure on the configuration center. The model encoding is used to uniquely identify the code of the model service, ensuring that each model service has a globally unique identifier. The version number represents the version of the model file, used to track the update and change history of the model and distinguish different versions of the model file. In a distributed system, a node refers to a server instance that runs the model service, and each node is responsible for executing the inference logic of the model.
[0068] In an alternative embodiment, before the model service runs, it is necessary to clarify the configuration information of the model file, including the model encoding, model file name, and version number, so that the model service knows which model to run and how to load it. After the model service starts, the obtained configuration information will be stored in the configuration cache variable to reduce the number of times of reading the configuration center in subsequent operations and improve the service response speed. When the configuration information in the configuration center changes, the model service needs to be able to recognize this change and read the latest configuration information to ensure that the model service is running the latest model file. The model service will periodically or when triggered by a specific event, check whether the configuration cache variable is consistent with the latest configuration information obtained from the configuration center, which is a key step to ensure the synchronization of the service state with the configuration center state. If it is found that the configuration information has changed, the model service will update the configuration cache variable and synchronize the new configuration information to all running nodes to ensure that all nodes are running the latest version of the model.
[0069] In an alternative embodiment, the model service dynamically obtains and updates configuration information from the configuration center and ensures that all running nodes use the latest version of the model file. Through the comparison and update mechanism of the configuration cache variable, the model service can respond to configuration changes in real time without affecting the continuity and stability of the service. This mechanism is particularly suitable for scenarios that require frequent model updates, such as online learning and optimization of AI models.
[0070] Figure 3 is another flowchart of the model running method according to an embodiment of the present invention; as Figure 3 shown,
[0071] S302, cache Consul configuration. After the project starts, define a variable cacheProp to cache the model file information in Consul. Ensure that the variable cacheProp is correctly initialized when the project starts for subsequent use.
[0072] S304, listen for Consul configuration changes. When the model file information in Consul changes, the Poin platform will publish a corresponding Event event. Ensure that the Poin platform can correctly detect the change of the Consul configuration and publish an event.
[0073] A variable NewProp can also be defined to cache the changed Consul configuration. Listen for Event events and use a retry mechanism to read the latest Consul configuration. Ensure that the event listening mechanism works properly and can capture events in a timely manner. Use a retry mechanism (such as exponential backoff strategy) to handle exceptions such as network jitters, ensuring that each cluster node can read the latest configuration. Avoid differences in the configurations read by different cluster nodes due to network problems. Store the latest Consul configuration in the variable NewProp.
[0074] S306, determine whether the Consul configuration is updated. If not, wait for the next change. The NewProp and cacheProp can be compared.
[0075] S308, in the case of Consul configuration update, load the latest mirror file. If there are modified or newly added Consul configurations, update cacheProp and load the model file information corresponding to the newly added or modified Consul configurations into memory. Ensure that the comparison logic is correct and can accurately identify changes in the configuration. When updating cacheProp, ensure data structure consistency and integrity.
[0076] Through the above steps, caching and synchronization of model file information in Consul can be achieved, ensuring that each node in the cluster environment can pre-load the model file into memory, thereby quickly responding to requests. The specific implementation steps include defining cache variables, listening for configuration changes, reading the latest configuration, and updating the cache and loading the model file. These steps together ensure the real-time nature and consistency of model file information, improving the overall performance and reliability of the system.
[0077] Optionally, determine the configuration information of the target model combination according to the target service request, including: determine the node information corresponding to the target model combination according to the node code in the target service request, where the node is used to indicate a model service in the model combination; determine the configuration information according to the model code in the node information.
[0078] It should be noted that the configuration information of the target model combination refers to the detailed information of the model combination that the model operation management platform needs to load and run to process the target service request, including the path of the model file, the type of model service (main service or shadow service), and its version number, etc. Each model service in the model combination has a unique identifier, called the node code, which is used to correspond to the correct model service in the request data. The node information contains the detailed information of the model service corresponding to the node code, such as service type (main / shadow service), service version number, service address, etc. These information are stored in the database and are used to construct the model combination.
[0079] In an alternative embodiment, after receiving a request from an upstream system, the model operation management platform parses the node encoding in the request and retrieves the detailed information of the specific model service from the database based on this. The platform further analyzes the model encoding in the node information to obtain the complete configuration information of the model service from a configuration center (such as Consul), including which version of the model file to run, the location of the model file in the object storage, etc.
[0080] In an alternative embodiment, first, the platform locates the specific node information, i.e., the metadata of the model service, by parsing the node encoding in the request. Then, based on the model encoding in the node information, the platform loads the configuration information of the model service from the configuration center. This information guides how the model service loads the model file, runs the model logic, and processes the request data. Such a design ensures the flexibility and scalability of the model service, enabling the model operation management platform to efficiently and accurately respond to the requirements of various business scenarios.
[0081] In an alternative embodiment, the model combination information is a set containing multiple nodes, and each node represents an instance of a model service. This information is usually stored in the database in the form of structured data (such as JSON or XML). Parse the model combination information to extract the detailed information of each node, including the node encoding, model encoding, node type, etc. Database storage: Store the parsed node information in the database for subsequent quick query by the node encoding. The node encoding carried in the request parameters is the key information passed when invoking the model service, used to identify the specific node. Query the corresponding node information from the database through the node encoding to ensure the accuracy of the query result. The node information should include the model encoding, which is the key field for obtaining the model service information. Model encoding: The model encoding is the field that uniquely identifies the model service. Through the model encoding, the information of the primary / parallel-running service can be obtained from the Consul configuration. The Consul configuration stores the configuration information corresponding to the model encoding, including the model version number of the primary service and the set of model version numbers of the parallel-running services.
[0082] Through this process, the model operation management platform can dynamically and flexibly determine and load the correct model service according to the request, thereby efficiently processing business requests. At the same time, it can also perform real-time monitoring and optimization of the model performance through the parallel-running service, accelerate the model iteration cycle, reduce the operation cost, and improve the service quality.
[0083] Optionally, after executing the first model service to obtain an output result and recording the output result and the request parameters, it further includes: executing the second model service to obtain a reference result and recording the reference result and the request parameters; updating the parameters of the first model service according to the reference result and the output result.
[0084] It should be noted that the reference result is obtained by the shadow model service, which is used to evaluate the effect of the main model service and provide a basis for model optimization. The parameter update is based on the reference result of the shadow model service to adjust the parameters of the main model service, so as to improve the performance of the main model service, including but not limited to improving the prediction accuracy and optimizing resource consumption.
[0085] In an alternative embodiment, during the operation of the model service, in addition to the normal request processing of the main model service, the execution of the shadow model service is provided to assist in the optimization and iteration of the main model. The shadow model service runs in parallel with the main model service, processes the same requests, generates reference results, and then records the reference results and request parameters for comparative analysis with the output results of the main model service. Based on the analysis results, the parameters of the main model service can be updated automatically or manually to improve its prediction accuracy, response speed, or resource utilization rate. This mechanism can continuously monitor the performance of the main model and make adjustments according to the feedback of the shadow model, which helps in the long-term optimization and improvement of the model.
[0086] In an alternative embodiment, assume that an intelligent credit scoring system is being developed, where the main model service uses version v1 of the model to process users' credit assessment requests in real time. The shadow model service uses version v2 of the model, which may include some experimental algorithm improvements. When a user request arrives, the model operation management platform first executes the main model service v1 to process the user's request data, obtains the output result of the credit score, and records the output result and request parameters. Immediately afterwards, the platform asynchronously executes the shadow model service v2 using the same request data to obtain a reference result of the credit score and also records the reference result and request parameters. Then, the system or model developers will conduct a comparative analysis of the output result and the reference result. If it is found that the scoring result of version v2 is more accurate or the processing speed is faster in certain specific situations (such as when the user has certain risk characteristics), then based on these observations, the parameters of the main model service v1 can be adjusted, such as modifying the feature weights, adjusting the thresholds, or adopting new feature processing techniques.
[0087] Through the above embodiments, the model operation management platform can effectively utilize the reference results of the shadow model service as a feedback mechanism to continuously optimize the performance of the main model service, thereby improving the accuracy and response ability of the entire credit scoring system.
[0088] Figure 4 is a schematic diagram of the model operation method according to an embodiment of the present invention; as Figure 4As shown in the figure, assume there is a financial risk assessment system that uses a model operation management platform to call multiple machine learning models for risk assessment. These models include a main model and a companion model, which are used to evaluate the credit risk of customers. The following steps can be used to demonstrate how to handle a specific request.
[0089] Specifically, the user can execute step S402 in the policy service platform to send a request that contains request parameters and a target service request. Then, in the model operation management platform, step S404 is executed to assemble the model combination call interface, followed by step S406 to obtain configuration information, and then step S408 to obtain node information to determine the main service model. Subsequently, step S410 is executed to send a request to the main service, and step S412 is executed to obtain the information of the main service and the companion service, thereby calling the corresponding models. Then, steps S414-1, adding a circuit breaker and flow limiting mechanism, and step S414-2, synchronously executing the main service model logic, can be performed. After step S414-1, step S418 can be executed to asynchronously execute the companion service model logic. After executing the main service and the companion service, step S420 can be executed to record the model logs. After executing the main service model in step S414-2, step S416 can be executed to return the result.
[0090] In an alternative embodiment, the business system or client sends a request to the model operation management platform. The request contains specific business requirements and data, which will be used for model inference. After the gateway service of the model operation management platform receives the request, it will parse the interface_id in the request. This identifier is used to determine which model combination to call. The gateway service is responsible for forwarding the request to the correct model combination service according to the interface_id. Through the interface_id, the system can read the model combination information in the database. This information describes the structure of the model combination, the included model nodes, and their configurations. The model combination information is parsed to extract the information of each model node, including their types (primary service or accompanying service), version numbers, running logics, etc. This information is stored in the database and used to construct the model combination. Each model node has a node code as a unique identifier. Through the node code in the request parameters, the system can accurately obtain the detailed information of the node from the database, including the path and format of the model file. The node information contains the model code. Through the model code, the system obtains the detailed configurations of the primary service and the accompanying service from the Consul configuration center, including their model version numbers. This step is achieved through the configuration structure saved in Consul, ensuring the unified management and fast access of the model service information. After obtaining the configurations of the primary service and the accompanying service, the system uses the Mocamel logic to execute the model inference of the primary service, and at the same time asynchronously executes the model logic of the accompanying service. In this way, real-time performance comparison and model optimization can be achieved. After each execution of the primary service and the accompanying service is completed, their return results and request parameters will be recorded in the log. These records can be used for subsequent auditing, performance analysis, and model optimization. The execution result of the primary service will be directly returned to the policy platform that initially sent the request, while the result of the accompanying service is mainly used for internal analysis and is not directly provided externally.
[0091] It should be noted that when the co-running service is invoked, the system will add a circuit breaker mechanism and a flow limiting strategy to ensure that the high load of the co-running service will not affect the stability and performance of the main service, and achieve traffic isolation between the main service and the co-running service. When a request arrives at the model operation management platform, the gateway service first parses the interface_id in the request to determine the model combination to be invoked. Then, the system reads the configuration information associated with the serviceCode in Consul to obtain the zhuVersion of the main service and the congVersions of the co-running service, so as to determine the specific model file version to be loaded. The main service loads the model file from the object storage into memory according to the model version number in the configuration information and warms up the execution engine. The co-running service is executed using an asynchronous mechanism to avoid affecting the response time of the main service, and its results are only used for internal analysis and not directly returned to the user. After the model inference is completed, the results of the main service are recorded and directly returned to the policy platform, while the results of the co-running service are recorded for subsequent comparative analysis to provide a basis for model iteration and optimization.
[0092] Specifically, in an alternative embodiment, the real-time method may include the following steps:
[0093] Step 1: The upstream system sends a request
[0094] Operation: A front-end application (upstream system) sends an HTTP request to the model operation management platform to request an assessment of a certain customer's credit risk.
[0095] Request parameters: interface_id: credit_risk_assessment
[0096] node_code: NC12345 (node code)
[0097] Customer data (such as income, credit history, etc.)
[0098] Example request: JSON
[0099]
[0100] Step 2: The gateway service assembles the call interface
[0101] Operation: After receiving the request, the gateway service assembles the corresponding model combination call interface according to the interface_id.
[0102] Assembly result: Generate an internal call interface pointing to the model combination credit_risk_assessment.
[0103] Example Interface: / mocamel / credit_risk_assessment
[0104] Step 3: Obtain Model Combination Information
[0105] Operation: According to the interface_id, obtain the complete information of the credit_risk_assessment model combination from the database.
[0106] Model Combination Information: It contains multiple nodes, and each node represents a model service.
[0107] Example Model Combination Information: JSON
[0108]
[0109]
[0110] Step 4: Parse Node Information
[0111] Operation: Parse the model combination information, extract the detailed information of all nodes, and confirm that this information has been saved in the database.
[0112] Node Information: Node Code: NC12345
[0113] Model Code: M1
[0114] Step 5: Obtain Corresponding Node Information
[0115] Operation: Through the node code NC12345 in the request parameters, obtain the corresponding node information from the database.
[0116] Obtained Result: Node Code: NC12345
[0117] Model Code: M1
[0118] Step 6: Obtain Primary / Backup Service Information
[0119] Operation: According to the model code M1, obtain the information of the primary service and the backup service from the Consul configuration.
[0120] Consul Configuration Information: Primary Service Model Version Number: V1.0
[0121] Backup Service Model Version Number Set: [V1.1, V1.2]
[0122] Step 7: Execute Model Logic
[0123] Operation: Execute the main service model logic (version V1.0) through Mocamel, and simultaneously execute the model logic in the accompanying service asynchronously (versions V1.1 and V1.2).
[0124] Input: Customer data
[0125] Output: Evaluation results of the main service and the accompanying service
[0126] Step 8: Record logs
[0127] Operation: Record the return results and input parameters of the main service and the accompanying service in the log for subsequent analysis and troubleshooting.
[0128] Log content: Request parameters, main service return results, accompanying service return results
[0129] Step 9: Return the main service result
[0130] Operation: Return the execution result of the main service to the policy platform for further decision-making.
[0131] Return result:
[0132] Risk assessment score: 75
[0133] Evaluation result: approved
[0134] Through the above embodiments, it can be clearly seen the entire process from sending a request from the upstream system to returning the main service result. Each step details how to process request parameters, obtain and parse node information, obtain configurations from Consul, execute model logic, record logs, and return results. This form of embodiment helps to understand each link in the actual operation, ensuring the smooth execution of the process and the quick troubleshooting of problems. The calculation logic of the model exists in the model file, and the operation of the model is responsible for the loading engine (i.e., the model service). Which model file the model service runs is determined by the input parameters. This decoupled form of the model service and the model file greatly reduces the time from the completion of model development to production.
[0135] According to another embodiment of the present invention, a model operation device is further provided, Figure 5 which is a block diagram of the model operation device according to the embodiment of the present invention, as Figure 5 shown. The model operation device includes:
[0136] A model combination module 502, configured to assemble a model combination call interface according to the target service request in the request parameters, where the target service request is used to indicate the target model;
[0137] A configuration information determination module 504, configured to determine the configuration information of the target model combination according to the target service request;
[0138] A service determination module 506, configured to determine a target model file according to configuration information, and obtain a first model service and a second model service corresponding to the target model, where the first model service is used to indicate a primary model service, and the second model service is used to indicate a secondary model service;
[0139] A result output module 508, configured to execute the first model service to obtain an output result, and record the output result and request parameters.
[0140] Optionally, the model running device further includes: a storage module, configured to store at least one received model file into an object storage, where the model file is used to encapsulate a computing logic; a preloading module, configured to preload the model file into a memory, and call an execution engine corresponding to the model file; a configuration update module, configured to update configuration information corresponding to the model file.
[0141] Optionally, the preloading module includes: a path determination unit, configured to determine model identification information of the model file, and determine path information of the model file in the object storage according to the model identification information; a loading unit, configured to load information of the model file from the object storage into a corresponding execution engine according to the path information of the model file in the object storage.
[0142] Optionally, the configuration update module is further configured to: determine configuration information of the model file, and update a model encoding, a model file name, and a version number in the configuration file, where the model encoding is used to indicate a corresponding model service.
[0143] Optionally, the configuration update module is further configured to: store the configuration information of the model file into a configuration cache variable, and read the changed configuration information in the case of a change in the configuration information; compare the configuration cache variable with the changed configuration information, and in the case where the configuration cache variable is different from the changed configuration information, store the changed configuration information into the configuration cache variable, and synchronize the update to corresponding nodes.
[0144] Optionally, the configuration information determination module 504 includes: a node determination unit, configured to determine node information corresponding to a target model combination according to a node encoding in a target service request, where the node is used to indicate a model service in the model combination; a configuration information determination unit, configured to determine configuration information according to the model encoding in the node information.
[0145] Optionally, the result output module 508 further includes: a reference result output unit, configured to execute the second model service to obtain a reference result, and record the reference result and request parameters; a parameter update unit, configured to update parameters of the first model service according to the reference result and the output result.
[0146] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: all the above-mentioned modules are located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination.
[0147] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, it executes the steps in any one of the above method embodiments.
[0148] Optionally, in this embodiment, the above storage medium can be set to store a computer program for executing the following steps:
[0149] S1. According to the target service request in the request parameters, assemble the model combination call interface, where the target service request is used to indicate the target model;
[0150] S2. Determine the configuration information of the target model combination according to the target service request;
[0151] S3. Determine the target model file according to the configuration information, and obtain the first model service and the second model service corresponding to the target model, where the first model service is used to indicate the main model service, and the second model service is used to indicate the slave model service;
[0152] S4. Execute the first model service, obtain the output result, and record the output result and the request parameters.
[0153] Optionally, in this embodiment, the above storage medium may include but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks or optical discs that can store computer programs.
[0154] An embodiment of the present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0155] Optionally, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0156] Optionally, in this embodiment, the above processor can be configured to execute the following steps through a computer program:
[0157] S1. Assemble the model combination call interface according to the target service request in the request parameters, where the target service request is used to indicate the target model;
[0158] S2. Determine the configuration information of the target model combination according to the target service request;
[0159] S3. Determine the target model file according to the configuration information, and obtain the first model service and the second model service corresponding to the target model, where the first model service is used to indicate the main model service, and the second model service is used to indicate the slave model service;
[0160] S4. Execute the first model service to obtain the output result, and record the output result and the request parameters.
[0161] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0162] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps of them can be made into a single integrated circuit module to implement. Thus, the present invention is not limited to any specific combination of hardware and software.
[0163] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A model running method, characterized in that, including: assembling a model combination call interface according to a target service request in a request parameter, where the target service request is used to indicate a target model; determining configuration information of a target model combination according to the target service request; determining a target model file according to the configuration information, and obtaining a first model service and a second model service corresponding to the target model, where the first model service is used to indicate a main model service, and the second model service is used to indicate a slave model service; executing the first model service to obtain an output result, and recording the output result and the request parameter; 2. The method according to claim 1, wherein before assembling the model combination call interface according to the target service request in the request parameter, including: storing at least one received model file in an object storage, where the model file is used to encapsulate a calculation logic; preloading the model file into a memory, and calling an execution engine corresponding to the model file; updating configuration information corresponding to the model file; 3. The method according to claim 2, wherein the preloading the model file into the memory and calling the execution engine corresponding to the model file includes: determining model identification information of the model file, and determining path information of the model file in the object storage according to the model identification information; loading information of the model file from the object storage into the corresponding execution engine according to the path information of the model file in the object storage; 4. The method according to claim 2, wherein the updating the configuration information corresponding to the model file includes: determining the configuration information of the model file, and updating a model code, a model file name, and a version number in the configuration file, where the model code is used to indicate a corresponding model service; 5. The method according to claim 4, characterized in that, the determining the configuration information of the model file and updating the model code, the model file name, and the version number in the configuration file includes: storing the configuration information of the model file in a configuration cache variable, and reading the changed configuration information in the case of a change in the configuration information; comparing the configuration cache variable with the changed configuration information, and in the case that the configuration cache variable is different from the changed configuration information, storing the changed configuration information in the configuration cache variable, and synchronizing the update to corresponding nodes; 6. The method according to claim 1, wherein the determining the configuration information of the target model combination according to the target service request includes: determining node information corresponding to the target model combination according to a node code in the target service request, where the node is used to indicate a model service in the model combination; determining the configuration information according to a model code in the node information; 7. The method according to claim 1, wherein after executing the first model service to obtain an output result, and recording the output result and the request parameter, further including: executing the second model service to obtain a reference result, and recording the reference result and the request parameter; updating parameters of the first model service according to the reference result and the output result; 8. A model running device, characterized in that, including: A model combination module, configured to assemble a model combination call interface according to a target service request in request parameters, where the target service request is used to indicate a target model; A configuration information determination module, configured to determine configuration information of a target model combination according to the target service request; A service determination module, configured to determine a target model file according to the configuration information, and obtain a first model service and a second model service corresponding to the target model, where the first model service is used to indicate a main model service, and the second model service is used to indicate a slave model service; A result output module, configured to execute the first model service, obtain an output result, and record the output result and the request parameters.
9. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, where the computer program, when run by a processor, executes the method described in any one of claims 1 to 7.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of claims 1 to 7.